Logistic Regression
Two related trainers share this page. TYPE 'logistic' is the binary classifier — the label must be 0 or 1 and the model emits the probability of class 1. TYPE 'multinomial' (aliases: softmax, multinomial_logistic) is the k-class softmax extension — the label must be a non-negative integer in [0, k) and the model emits the most likely class ID. Both are batch gradient descent-trained with optional L2 regularization.
When to use it
- Binary outcomes (churn, fraud, click) where you need a probability estimate rather than a class id.
- A linear-decision-boundary baseline before reaching for trees / DNNs.
- Multinomial: low-cardinality classification (k of 3 to ~20) where a flat softmax beats a one-vs-rest stack.
When NOT to use it
- Highly non-linear decision boundaries — use random forest or gradient boosting.
- Many imbalanced classes — softmax doesn't reweight; use a tree ensemble with
class_weight = 'balanced'instead. - Streaming labels — use
ALTER MODEL ... ENABLE ONLINE LEARNINGon a logistic model.
Syntax
Binary:
CREATE MODEL <name>(DOUBLE, DOUBLE[, ...]) RETURNS DOUBLE
TYPE { 'logistic' | 'logistic_regression' | 'logistic-regression' }
OPTIONS (...)
AS SELECT <label_0_or_1>, <f1>, <f2>, ... FROM <source>;
Multinomial:
CREATE MODEL <name>(DOUBLE, DOUBLE[, ...]) RETURNS INT
TYPE { 'multinomial' | 'multinomial_logistic' | 'multinomial-logistic' | 'softmax' }
OPTIONS (k = <int>, ...)
AS SELECT <label_in_0_to_k_minus_1>, <f1>, <f2>, ... FROM <source>;
Options
| Option | Default | Type | Applies to | What it does |
|---|---|---|---|---|
k | (required) | int >= 2 | multinomial | Number of classes |
learning_rate | 0.1 | float > 0 | both | batch gradient descent step size |
l2_lambda | 0.0 | float >= 0 | both | L2 regularization strength |
max_iters | 100 | int >= 1 | both | Maximum training iterations |
tolerance | 1e-4 | float >= 0 | both | Convergence threshold on weight delta |
partitions | 1 | int >= 1 | both | Training partitions; see execution and memory limits |
Examples
These fragments assume the named source tables exist. Match training and inference feature order and preprocessing. For a complete dataset and runnable script, follow the linked tutorial.
Binary (minimal):
CREATE MODEL churn(DOUBLE, DOUBLE) RETURNS DOUBLE
TYPE 'logistic'
AS SELECT
CAST(churned AS DOUBLE) AS label,
tenure_days,
monthly_charges
FROM customer_history;
SELECT customer_id, churn(tenure_days, monthly_charges) AS p_churn
FROM active_customers
WHERE churn(tenure_days, monthly_charges) > 0.7;
Multinomial with regularization:
CREATE MODEL article_topic(DOUBLE, DOUBLE, DOUBLE, DOUBLE, DOUBLE) RETURNS INT
TYPE 'softmax'
OPTIONS (
k = 4,
learning_rate = 0.05,
l2_lambda = 0.001,
max_iters = 800,
tolerance = 1e-5
)
AS SELECT
CAST(topic_id AS DOUBLE) AS label,
word_count,
avg_word_length,
sentiment,
keyword_density,
readability_score
FROM labeled_articles;
Output shape
- Binary: a single
DOUBLEin[0, 1]— the probability of class 1. Threshold to your operating point. - Multinomial: one
INTclass ID in[0, k), selected by argmax. The SQL UDF does not expose the full softmax probability vector.
Tuning notes
- Standardise features. Logistic batch gradient descent is sensitive to feature scale.
- If binary predictions stay near
0.5or multiclass predictions collapse to one class, inspect label balance, feature scaling, and training progress before changing the iteration budget. - Add
l2_lambdawhen overfit is visible (training loss drops, validation loss stalls or rises). - For severe class imbalance use a tree ensemble with
class_weight = 'balanced'instead — logistic / softmax don't expose a class-weight knob. - The binary trainer is an online-learning candidate via
ALTER MODEL <name> ENABLE ONLINE LEARNING.
Convergence and quality
converged = true means the weight delta dropped under tolerance. EVALUATE MODEL emits accuracy, precision, recall, f1 for binary; for multinomial it adds macro_precision, macro_recall, macro_f1, and num_classes. Calibrate the binary threshold on a held-out fold rather than reading off the training loss.