Random Forest
TYPE 'random_forest' (aliases: rf, forest) trains an ensemble of decision trees with bootstrap sampling and per-split feature subsampling. Compare it with logistic regression and gradient boosting on held-out data; the best choice depends on the dataset and scoring budget.
When to use it
- Tabular classification with mixed feature distributions.
- Datasets where single-tree variance is high — bagging averages it out.
- Imbalanced labels —
class_weight = 'balanced'propagates per-tree.
When NOT to use it
- Forecasting tasks (use SARIMA).
- Regression problems — the native trainer is classification-only.
- Very wide feature sets where you specifically want sequential corrective learning — use gradient boosting.
- Latency-critical scoring where
num_trees × avg_depthevaluations per row exceeds your budget.
Syntax
CREATE MODEL <name>(DOUBLE, DOUBLE[, ...]) RETURNS INT
TYPE { 'random_forest' | 'random-forest' | 'rf' | 'forest' }
OPTIONS (...)
AS SELECT <class_label>, <f1>, <f2>, ... FROM <source>;
Options
| Option | Default | Type | Range | What it does |
|---|---|---|---|---|
num_trees | 100 | int | >= 1 | Number of trees in the forest |
max_depth | 6 | int | >= 1 | Per-tree maximum depth |
num_bins | 64 | int | >= 2 | Histogram bins per feature |
min_samples_split | 2 | int | >= 1 | Minimum rows needed at a node to be eligible to split |
feature_subsample_size | floor(sqrt(num_features)) | int | >= 1 | Features considered per split (Breiman's recipe) |
bootstrap_fraction | 1.0 | float | (0, 1] | Sample fraction per tree, drawn with replacement |
num_classes | 2 | int | >= 2 | Output class count |
ccp_alpha | 0.0 | float | >= 0 | Per-tree cost-complexity pruning strength |
class_weight | (none) | string | 'balanced' only | Auto-computed per-class weights, propagated to every tree |
random_seed | 0 | u64 | any | RNG seed for bootstrap + feature subsampling |
partitions | 1 | int | >= 1 | Training partitions; see execution and memory limits |
Examples
These fragments assume the named source tables exist. Match training and inference feature order and preprocessing. For a complete dataset and runnable script, follow the linked tutorial.
Minimal:
CREATE MODEL credit_risk(DOUBLE, DOUBLE, DOUBLE) RETURNS INT
TYPE 'random_forest'
AS SELECT
CAST(defaulted AS DOUBLE) AS label,
income, debt_ratio, credit_score
FROM loan_history;
Tuned for a five-feature, imbalanced binary task:
CREATE MODEL credit_risk(DOUBLE, DOUBLE, DOUBLE, DOUBLE, DOUBLE) RETURNS INT
TYPE 'random_forest'
OPTIONS (
num_trees = 200,
max_depth = 12,
feature_subsample_size = 3, -- floor(sqrt(5)) = 2; bump up
bootstrap_fraction = 0.8,
min_samples_split = 20,
num_bins = 128,
num_classes = 2,
class_weight = 'balanced',
random_seed = 42,
partitions = 4
)
AS SELECT
CAST(defaulted AS DOUBLE) AS label,
income, debt_ratio, credit_score, employment_years, num_open_accounts
FROM loan_history;
Output shape
credit_risk(f1, f2, ...) returns an INT predicted class id in [0, num_classes). The forest votes by majority across trees.
Tuning notes
- Start with
num_trees = 100and bump to200–500if validation score is still climbing. Diminishing returns set in fast past~500. - Lower
bootstrap_fraction(e.g.0.7) when individual trees are too correlated. - Bump
feature_subsample_sizeabovesqrt(d)when most features carry signal; lower it when many features are noise. - For imbalanced labels set
class_weight = 'balanced'— it propagates to every inner tree. - See distributed training before choosing the partition count.
Convergence and quality
The training summary does not establish predictive quality. EVALUATE MODEL emits accuracy, precision, recall, f1 (and macro_* variants for num_classes > 2). Compare multiple seeds on the same held-out split to measure variability; there is no fixed accuracy agreement guarantee.
Related
-
Decision tree — the single-tree base learner
-
Gradient boosting — sequential boosted alternative