Distributed Training
CREATE MODEL ... TYPE '<algo>' OPTIONS (partitions = N) AS SELECT ... runs the training loop over N partitions of the training input.
Each iteration computes partial results for every partition in
parallel and merges them before the next iteration. Gnok decides where
the partitions run: on the query service itself or, where it is enabled
for your account, spread across additional workers.
When to use partitions
Partitioning helps when each training iteration does enough work to
outweigh the cost of splitting and merging. Benchmark it against the
default (partitions = 1) on representative data before relying on it.
The full training query is still materialized before training starts. Partitioning does not enable out-of-core training or remove the training-row cap (default 1,000,000 rows). Account for the training batches and model state when planning the training dataset, and request capacity changes through Gnok support.
Availability
Spreading partitions across additional workers is a service-managed capability. OPTIONS (partitions = N) requests partitioned training, but it doesn't provision workers. When remote workers aren't available, or a worker fails during a run, Gnok trains the partitions on the query service instead, so the statement still completes. Ask Gnok support about capacity and worker availability for your workload.
Observability
Inspect the training result, row count, iteration count, and convergence status. Gnok support can see where a run's partitions executed; provide the model name and query reference when investigating performance.
Cancellation
ALTER MODEL <name> CANCEL TRAINING works the same way however the
partitions run. Cancellation is checked between iterations, so a
partition that is mid-iteration finishes that iteration first. The
worst-case delay is one iteration of the slowest partition.
Algorithm coverage
These algorithms accept OPTIONS (partitions = N):
- KMeans
- Logistic Regression
- Multinomial Logistic Regression
- Linear Regression
- PCA
- Decision tree
- Random forest
- Gradient boosting
- DNN / MLP
ARIMA / SARIMA is intentionally excluded. Its training reads a single ordered series, and splitting that series into batches would corrupt its lag structure. Train one model per series instead.
Example
-- Request four partitions for a KMeans model.
CREATE OR REPLACE MODEL traffic_clusters(DOUBLE, DOUBLE, DOUBLE, DOUBLE)
RETURNS INT
TYPE 'kmeans'
OPTIONS (k = 8, max_iters = 100, partitions = 4)
AS SELECT f0, f1, f2, f3 FROM weekly_traffic_features;
The example assumes an existing weekly_traffic_features table with four numeric feature columns.
Numerical reproducibility
The training paths use the same trainers, but floating-point addition is not exactly associative. Changing partition count, row order, execution placement, or reduction order can change fitted values. Compare predictions with an appropriate numerical tolerance and task metric instead of requiring bit equality.
For repeatable comparisons, fix the training snapshot, preprocessing, seed where available, and partition count. A fixed random seed controls initialization; it does not eliminate all other sources of variation.