Gradient Boosting
TYPE 'gradient_boosting' (aliases: gbm, gboost) trains a sequential ensemble of shallow regression trees on the gradient of a logistic loss. Each tree corrects the residuals of the running ensemble; a learning_rate shrinkage factor controls how aggressively each tree contributes. Pair a low learning_rate with a higher num_trees and select the combination using held-out metrics.
When to use it
- Tabular classification where you want to push past random-forest accuracy with careful tuning.
- Cases with many weak signal features that compound when learnt sequentially.
- Imbalanced data where you've already tried
class_weight = 'balanced'on a forest.
When NOT to use it
- You don't have validation infrastructure to tune
learning_rate×num_trees×max_depth— random forest is more forgiving out of the box. - Latency-critical scoring with high
num_trees(each row visits every tree). - Regression targets — the native trainer fits a logistic-loss classification head.
Syntax
CREATE MODEL <name>(DOUBLE, DOUBLE[, ...]) RETURNS INT
TYPE { 'gradient_boosting' | 'gradient-boosting' | 'gbm' | 'gboost' }
OPTIONS (...)
AS SELECT <class_label>, <f1>, <f2>, ... FROM <source>;
Options
| Option | Default | Type | Range | What it does |
|---|---|---|---|---|
num_trees | 100 | int | >= 1 | Number of boosting rounds (one tree per round) |
learning_rate | 0.1 | float | (0, 1] | Shrinkage applied to each tree's contribution; smaller updates often need more rounds; validate the tradeoff |
max_depth | 3 | int | >= 1 | Per-tree depth — keep small (2–6); deep trees defeat the boosting premise |
num_bins | 64 | int | >= 2 | Histogram bins per feature for split search |
min_samples_split | 2 | int | >= 1 | Minimum rows needed at a node to split |
min_samples_leaf | 1 | int | >= 1 | Minimum rows required in each child after a split |
min_impurity_decrease | 0.0 | float | >= 0 (finite) | Minimum impurity reduction required to accept a split |
num_classes | 2 | int | >= 2 | Output class count |
partitions | 1 | int | >= 1 | Training partitions; see execution and memory limits |
Examples
These fragments assume the named source tables exist. Match training and inference feature order and preprocessing. For a complete dataset and runnable script, follow the linked tutorial.
Minimal:
CREATE MODEL conversion(DOUBLE, DOUBLE) RETURNS INT
TYPE 'gbm'
AS SELECT
CAST(converted AS DOUBLE) AS label,
time_on_site, page_views
FROM session_features;
Tuned for a slow-and-deep generalization regime:
CREATE MODEL conversion_score(DOUBLE, DOUBLE, DOUBLE, DOUBLE) RETURNS INT
TYPE 'gradient_boosting'
OPTIONS (
num_trees = 500,
learning_rate = 0.03,
max_depth = 4,
num_bins = 128,
min_samples_split = 50,
min_samples_leaf = 20,
min_impurity_decrease = 0.0001,
num_classes = 2
)
AS SELECT
CAST(converted AS DOUBLE) AS label,
time_on_site, page_views, prior_purchases, email_opens
FROM session_features;
Output shape
conversion_score(f1, f2, ...) returns an INT predicted class id in [0, num_classes).
Tuning notes
- Halve
learning_rateand doublenum_treesuntil validation accuracy stops improving — that's the canonical GBM sweep. max_depth = 3–6is the sweet spot. Deeper trees can overfit and increase scoring cost.- Raise
min_samples_leaf(e.g.20–100) before raisingmin_samples_splitto control overfit. - Bump
num_binsto128–255only when feature granularity matters — the cost is linear in bin count. - This trainer does not support
class_weighttoday; for imbalanced data prefer random forest withclass_weight = 'balanced'.
Convergence and quality
Boosting rounds run inside the trainer; the outer convergence flag is not a quality metric. EVALUATE MODEL emits the standard classification metrics (accuracy, precision, recall, f1, plus macro_* for multiclass). Watch for the validation curve flattening or rising — that's the signal to drop num_trees or shrink learning_rate.
Related
-
Random forest — bagged alternative, easier to tune
-
Decision tree — single-tree base learner