Skip to main content

Online learning and regression slices

Connect a base model to labelled feedback and define slices for observing regression.

The practical problem and model​

A model that scores incoming events may eventually receive the true outcome. Online learning uses those delayed labels to update a model as the data changes. This is different from repeatedly scoring new events: without labels, the feedback-driven training step has no supervised signal.

The tutorial starts with a logistic classifier and registers a separate feedback table. It defines US and EU slices so a global improvement cannot hide deterioration in one of those groups.

This script intentionally leaves the feedback table empty. It establishes the base model, feedback contract, and monitoring configuration; it does not pretend that creating the configuration alone produces a new trained version.

Before you run​

Complete the shared setup. This walkthrough uses synthetic data and recreates its tutorial objects. Use a sandbox tenant or schema, and run the steps in order.

Studio: Tutorial: Online Learning + Per-slice Regression Detection. Download the complete SQL.

Background processing requires the tenant service identity, access to training/feedback data, and a running catalog and engine. Provisioning the identity does not grant blanket table permissions.

Step 1: Reset the demo and separate training from feedback​

Disabling the prior demo learner prevents it from consuming the replacement tables while setup runs. Eight labelled rows train the base model; the feedback table also carries region and arrival time.

CREATE CATALOG IF NOT EXISTS tutorial;

CREATE SCHEMA IF NOT EXISTS tutorial.online;

ALTER MODEL tutorial.online.scorer DISABLE ONLINE LEARNING;

CREATE OR REPLACE TABLE tutorial.online.training_history (
feature_a DOUBLE, feature_b DOUBLE, label INT, region VARCHAR
);

INSERT INTO tutorial.online.training_history VALUES
(0.1, 0.9, 1, 'us'), (0.2, 0.8, 1, 'us'), (0.9, 0.1, 0, 'us'),
(0.8, 0.2, 0, 'us'), (0.4, 0.6, 1, 'eu'), (0.6, 0.4, 0, 'eu'),
(0.3, 0.7, 1, 'eu'), (0.7, 0.3, 0, 'eu');

CREATE OR REPLACE TABLE tutorial.online.labeled_outcomes (
feature_a DOUBLE, feature_b DOUBLE, label INT, region VARCHAR,
arrived_at TIMESTAMP
);

Step 2: Train the base model and enable feedback learning​

Training and feedback must agree on the two feature columns and the label. Inspect the enabled learner rather than assuming registration succeeded.

CREATE OR REPLACE MODEL tutorial.online.scorer
(DOUBLE, DOUBLE)
RETURNS DOUBLE
TYPE 'logistic'
OPTIONS (learning_rate = 0.05, max_iters = 500)
AS SELECT CAST(label AS DOUBLE) AS label, feature_a, feature_b
FROM tutorial.online.training_history;

ALTER MODEL tutorial.online.scorer
ENABLE ONLINE LEARNING
WITH FEEDBACK FROM tutorial.online.labeled_outcomes
LEARNING RATE 0.01;

SHOW ONLINE LEARNING;

Step 3: Define slices and inspect runs​

The slice predicates identify regions for monitoring. The run listing can legitimately be empty when no new feedback has arrived.

ALTER MODEL tutorial.online.scorer
ENABLE ONLINE LEARNING SLICE
AS 'eu_customers'
WHERE region = 'eu';

ALTER MODEL tutorial.online.scorer
ENABLE ONLINE LEARNING SLICE
AS 'us_customers'
WHERE region = 'us';

SHOW MODEL SLICES tutorial.online.scorer;

SHOW AUTOML RUNS LIMIT 20;

What to check in the results​

SHOW ONLINE LEARNING should identify the qualified model and its feedback source. SHOW MODEL SLICES should include eu_customers and us_customers with their intended predicates.

An empty feedback table means no update is expected from this script alone. In the separate lifecycle checks, controlled feedback exercised accepted updates, regression handling, and restart recovery. To verify an update in your environment, track feedback consumption, the resulting model version, and slice metrics rather than just seeing an enabled flag.

Adapt it to real data​

Use outcomes that become available after prediction, record event and label-arrival times, and define how corrections and duplicate labels are handled. Keep a stable holdout and a minimum sample requirement for each monitored slice.

Inspect why a candidate was accepted or rejected before allowing unattended updates. Checkpoints, active versions, and feedback offsets need to survive restarts together; a restarted process must not train repeatedly on the same batch.

Streaming and online learning reference

All tutorials · Setup and troubleshooting