Skip to main content

Feature Store

Feature groups give training and inference a shared SQL definition of features such as customer spend, purchase count, or average transaction size. A group has key columns, feature columns, a materialization query, and a refresh interval. The latest materialization supports in-memory point lookups; an offline Iceberg table provides durable data and retained snapshot history.

Feature storage is enabled by default in the engine. Effective availability and capacity are managed by the hosted service. See availability for prerequisites.

Prerequisites​

The catalog provisions the tenant's managed gnok-engine identity at signup or trusted first use for an existing active tenant. An administrator must still grant the background executor access to the group's sources and designated output. Interactive user access alone is insufficient. See background workload setup.

Create an offline target with columns and types matching the materialization output. Keep the catalog, background executor, and lease service available. A definition appearing in SHOW FEATURE GROUPS does not establish that its first refresh succeeded.

Creating Feature Groups​

The parenthesized list contains key column names, not a table-style column definition. The query must return those keys and all feature values. Use one row per key.

After running the feature-group tutorial, its definition is:

CREATE OR REPLACE FEATURE GROUP tutorial.features.customer_30d (customer_id)
MATERIALIZED FROM (
SELECT customer_id,
SUM(amount) AS total_spend_30d,
COUNT(*) AS purchase_count,
CAST(EXTRACT(EPOCH FROM MAX(event_ts)) * 1000 AS BIGINT) AS event_ts
FROM tutorial.features.events
WHERE event_type = 'purchase'
GROUP BY customer_id
)
REFRESH EVERY 300s
OFFLINE TABLE tutorial.features.cust_feat_offline;

The tutorial uses a fixed small event set. The name customer_30d does not itself impose a 30-day filter; an application must include its intended time window in the materialization query.

For composite keys, list the columns, for example (customer_id, region), and group by both in the source query. Keys may be integer or string values; composite lookups use a struct whose fields match the group's key columns.

Refreshing Feature Groups​

Scheduled materialization re-executes the source query. For an immediate refresh:

REFRESH FEATURE GROUP tutorial.features.customer_30d;
SHOW FEATURE GROUPS;
SELECT * FROM tutorial.features.cust_feat_offline ORDER BY customer_id;

An offline-backed refresh uses INSERT OVERWRITE to replace the current output and publishes the resulting data to the online cache. Previous Iceberg snapshots are usable only while retained; configure retention for the historical range your application needs.

Backfill​

Backfill is a bounded historical operation with explicit epoch-millisecond endpoints and time-based chunks. The tutorial's source has an event_ts value in epoch milliseconds:

BACKFILL FEATURE GROUP tutorial.features.customer_30d
FROM 1775001600000 TO 1775433600000
WITH (chunk_size_ms = 432000000, wait = true);

SHOW BACKFILL PROGRESS tutorial.features.customer_30d;

wait = true waits for completion and surfaces failure. Without waiting, inspect progress before consuming the output. Concurrent active runs are rejected with a message directing you to progress; catalog leases coordinate writers. Investigate a failed chunk before retrying.

A backfill time range does not by itself make an aggregate point-in-time correct. Define event-time filters so each training example uses only information that was available at its label timestamp.

Drift detection​

SHOW FEATURE DRIFT;
SHOW FEATURE DRIFT tutorial.features.customer_30d;

A new group may have no prior baseline. Inspect report status and details before interpreting an empty report as absence of drift. Distribution drift is a signal to investigate, not proof that predictive quality changed or that a retraining operation completed.

SQL lookups​

These are scalar functions, not table-valued functions. Each invocation selects one feature for one key.

FunctionSQL result
FEATURE_LOOKUP(group, key, feature)DOUBLE
FEATURE_LOOKUP_INT(group, key, feature)BIGINT
FEATURE_LOOKUP_STR(group, key, feature)VARCHAR
FEATURE_LOOKUP_BOOL(group, key, feature)BOOLEAN
FEATURE_LOOKUP_AS_OF(group, key, feature, as_of_ms)Historical numeric value as DOUBLE
SELECT FEATURE_LOOKUP('tutorial.features.customer_30d', 1, 'total_spend_30d') AS spend,
FEATURE_LOOKUP_INT('tutorial.features.customer_30d', 1, 'purchase_count') AS purchases;

After materializing the tutorial data, customer 1 has spend 130 and purchase count 2. A missing key yields NULL; distinguish that from a legitimate zero. Use the typed lookup matching the stored feature type.

Point-in-time training​

FEATURE_LOOKUP_AS_OF takes an epoch-millisecond snapshot time. It searches retained historical materializations and the offline path where available. This timestamp is distinct from the event timestamps used in the feature query.

A feature column added after the requested historical point is an error by default. The documented session setting SET feature_lookup_missing_column = 'null' opts into a nullable result for that case; it does not recreate missing historical data.

For multi-feature batch training, joining a suitable offline snapshot to labels can be clearer than many scalar lookups. Check both snapshot retention and event-time cutoffs to prevent future information from leaking into training.

Operations​

Use SHOW FEATURE GROUPS to check row counts and refresh timestamps, SHOW BACKFILL PROGRESS for historical jobs, and SHOW FEATURE DRIFT for comparisons. Check executor logs for permission, catalog, source-query, or target-write failures.

Online capacity limits do not automatically turn an oversized group into an offline-only group. Size the materialization for the configured row and group limits. Measure lookup latency and refresh duration with representative key counts.

DROP FEATURE GROUP IF EXISTS tutorial.features.customer_30d;

Dropping a definition is an operational action; manage the associated offline table and retention according to your data lifecycle.

Complete feature-group tutorial · Streaming ML · Configuration