Skip to main content

Gradient Boosting

TYPE 'gradient_boosting' (aliases: gbm, gboost) trains a sequential ensemble of shallow regression trees on the gradient of a logistic loss. Each tree corrects the residuals of the running ensemble; a learning_rate shrinkage factor controls how aggressively each tree contributes. Pair a low learning_rate with a higher num_trees and select the combination using held-out metrics.

When to use it​

  • Tabular classification where you want to push past random-forest accuracy with careful tuning.
  • Cases with many weak signal features that compound when learnt sequentially.
  • Imbalanced data where you've already tried class_weight = 'balanced' on a forest.

When NOT to use it​

  • You don't have validation infrastructure to tune learning_rate × num_trees × max_depth — random forest is more forgiving out of the box.
  • Latency-critical scoring with high num_trees (each row visits every tree).
  • Regression targets — the native trainer fits a logistic-loss classification head.

Syntax​

CREATE MODEL <name>(DOUBLE, DOUBLE[, ...]) RETURNS INT
TYPE { 'gradient_boosting' | 'gradient-boosting' | 'gbm' | 'gboost' }
OPTIONS (...)
AS SELECT <class_label>, <f1>, <f2>, ... FROM <source>;

Options​

OptionDefaultTypeRangeWhat it does
num_trees100int>= 1Number of boosting rounds (one tree per round)
learning_rate0.1float(0, 1]Shrinkage applied to each tree's contribution; smaller updates often need more rounds; validate the tradeoff
max_depth3int>= 1Per-tree depth — keep small (2–6); deep trees defeat the boosting premise
num_bins64int>= 2Histogram bins per feature for split search
min_samples_split2int>= 1Minimum rows needed at a node to split
min_samples_leaf1int>= 1Minimum rows required in each child after a split
min_impurity_decrease0.0float>= 0 (finite)Minimum impurity reduction required to accept a split
num_classes2int>= 2Output class count
partitions1int>= 1Training partitions; see execution and memory limits

Examples​

These fragments assume the named source tables exist. Match training and inference feature order and preprocessing. For a complete dataset and runnable script, follow the linked tutorial.

Minimal:

CREATE MODEL conversion(DOUBLE, DOUBLE) RETURNS INT
TYPE 'gbm'
AS SELECT
CAST(converted AS DOUBLE) AS label,
time_on_site, page_views
FROM session_features;

Tuned for a slow-and-deep generalization regime:

CREATE MODEL conversion_score(DOUBLE, DOUBLE, DOUBLE, DOUBLE) RETURNS INT
TYPE 'gradient_boosting'
OPTIONS (
num_trees = 500,
learning_rate = 0.03,
max_depth = 4,
num_bins = 128,
min_samples_split = 50,
min_samples_leaf = 20,
min_impurity_decrease = 0.0001,
num_classes = 2
)
AS SELECT
CAST(converted AS DOUBLE) AS label,
time_on_site, page_views, prior_purchases, email_opens
FROM session_features;

Output shape​

conversion_score(f1, f2, ...) returns an INT predicted class id in [0, num_classes).

Tuning notes​

  • Halve learning_rate and double num_trees until validation accuracy stops improving — that's the canonical GBM sweep.
  • max_depth = 3–6 is the sweet spot. Deeper trees can overfit and increase scoring cost.
  • Raise min_samples_leaf (e.g. 20–100) before raising min_samples_split to control overfit.
  • Bump num_bins to 128–255 only when feature granularity matters — the cost is linear in bin count.
  • This trainer does not support class_weight today; for imbalanced data prefer random forest with class_weight = 'balanced'.

Convergence and quality​

Boosting rounds run inside the trainer; the outer convergence flag is not a quality metric. EVALUATE MODEL emits the standard classification metrics (accuracy, precision, recall, f1, plus macro_* for multiclass). Watch for the validation curve flattening or rising — that's the signal to drop num_trees or shrink learning_rate.