Skip to main content

Random Forest

TYPE 'random_forest' (aliases: rf, forest) trains an ensemble of decision trees with bootstrap sampling and per-split feature subsampling. Compare it with logistic regression and gradient boosting on held-out data; the best choice depends on the dataset and scoring budget.

When to use it​

  • Tabular classification with mixed feature distributions.
  • Datasets where single-tree variance is high — bagging averages it out.
  • Imbalanced labels — class_weight = 'balanced' propagates per-tree.

When NOT to use it​

  • Forecasting tasks (use SARIMA).
  • Regression problems — the native trainer is classification-only.
  • Very wide feature sets where you specifically want sequential corrective learning — use gradient boosting.
  • Latency-critical scoring where num_trees × avg_depth evaluations per row exceeds your budget.

Syntax​

CREATE MODEL <name>(DOUBLE, DOUBLE[, ...]) RETURNS INT
TYPE { 'random_forest' | 'random-forest' | 'rf' | 'forest' }
OPTIONS (...)
AS SELECT <class_label>, <f1>, <f2>, ... FROM <source>;

Options​

OptionDefaultTypeRangeWhat it does
num_trees100int>= 1Number of trees in the forest
max_depth6int>= 1Per-tree maximum depth
num_bins64int>= 2Histogram bins per feature
min_samples_split2int>= 1Minimum rows needed at a node to be eligible to split
feature_subsample_sizefloor(sqrt(num_features))int>= 1Features considered per split (Breiman's recipe)
bootstrap_fraction1.0float(0, 1]Sample fraction per tree, drawn with replacement
num_classes2int>= 2Output class count
ccp_alpha0.0float>= 0Per-tree cost-complexity pruning strength
class_weight(none)string'balanced' onlyAuto-computed per-class weights, propagated to every tree
random_seed0u64anyRNG seed for bootstrap + feature subsampling
partitions1int>= 1Training partitions; see execution and memory limits

Examples​

These fragments assume the named source tables exist. Match training and inference feature order and preprocessing. For a complete dataset and runnable script, follow the linked tutorial.

Minimal:

CREATE MODEL credit_risk(DOUBLE, DOUBLE, DOUBLE) RETURNS INT
TYPE 'random_forest'
AS SELECT
CAST(defaulted AS DOUBLE) AS label,
income, debt_ratio, credit_score
FROM loan_history;

Tuned for a five-feature, imbalanced binary task:

CREATE MODEL credit_risk(DOUBLE, DOUBLE, DOUBLE, DOUBLE, DOUBLE) RETURNS INT
TYPE 'random_forest'
OPTIONS (
num_trees = 200,
max_depth = 12,
feature_subsample_size = 3, -- floor(sqrt(5)) = 2; bump up
bootstrap_fraction = 0.8,
min_samples_split = 20,
num_bins = 128,
num_classes = 2,
class_weight = 'balanced',
random_seed = 42,
partitions = 4
)
AS SELECT
CAST(defaulted AS DOUBLE) AS label,
income, debt_ratio, credit_score, employment_years, num_open_accounts
FROM loan_history;

Output shape​

credit_risk(f1, f2, ...) returns an INT predicted class id in [0, num_classes). The forest votes by majority across trees.

Tuning notes​

  • Start with num_trees = 100 and bump to 200–500 if validation score is still climbing. Diminishing returns set in fast past ~500.
  • Lower bootstrap_fraction (e.g. 0.7) when individual trees are too correlated.
  • Bump feature_subsample_size above sqrt(d) when most features carry signal; lower it when many features are noise.
  • For imbalanced labels set class_weight = 'balanced' — it propagates to every inner tree.
  • See distributed training before choosing the partition count.

Convergence and quality​

The training summary does not establish predictive quality. EVALUATE MODEL emits accuracy, precision, recall, f1 (and macro_* variants for num_classes > 2). Compare multiple seeds on the same held-out split to measure variability; there is no fixed accuracy agreement guarantee.