Supported Model Types
Choose a model by its task, output contract, and runtime requirements. The tutorials provide complete datasets and scripts; the algorithm references explain options and tuning. The examples in reference pages use illustrative source tables unless stated otherwise.
Surface summary
| Surface | Registration | Runtime requirements |
|---|---|---|
| Native training | TYPE '<algorithm>' OPTIONS (...) AS SELECT ... | Engine training code; no Python or AI provider |
| ONNX | USING FRAMEWORK 'onnx' AS FROM ... | Built in; supported tensor shape and operators |
| PyTorch | USING FRAMEWORK 'pytorch' AS FROM ... | PyTorch runtime enabled for your account; TorchScript artifact |
| sklearn | USING FRAMEWORK 'sklearn' AS FROM ... | Python runtime enabled for your account; trusted skops artifact |
| Remote endpoint | Remote-model definition | Public HTTPS endpoint approved by Gnok; matching request/response contract |
Registered models can be invoked by qualified name or with ML_PREDICT('catalog.schema.model', feature1, feature2). See imported runtimes, ONNX, and the remote endpoint tutorial.
Native algorithms
For supervised training, put the label first in AS SELECT, followed by the features. The SQL function signature contains features only. KMeans and PCA use every selected column as a feature. ARIMA is the special case: training reads one chronological value series, while inference accepts a forecast horizon.
Native training normalizes supported numeric and Boolean columns to DOUBLE, including decimal inputs. Encode categorical strings explicitly. Preserve feature order, units, and scaling at inference; numeric conversion does not standardize features. Use care with integers or decimals whose precision exceeds DOUBLE.
| Algorithm | Canonical TYPE | SQL result |
|---|---|---|
| KMeans | kmeans | INT cluster ID |
| Linear regression | linear | DOUBLE predicted value |
| Binary logistic regression | logistic | DOUBLE probability of class 1 |
| Multinomial logistic regression | multinomial | INT class ID |
| AR / ARIMA / SARIMA | arima | DOUBLE prediction at a supplied horizon |
| PCA | pca | DOUBLE for one component; ARRAY<DOUBLE> for multiple components |
| Decision tree | decision_tree | INT class ID |
| Random forest | random_forest | INT class ID |
| Gradient boosting | gradient_boosting | INT class ID |
| DNN / MLP | dnn | Regression/binary: DOUBLE; multiclass: INT class ID |
Class IDs are not probabilities. Native tree models are classification-only. Native multiclass models expose the winning class, not the full probability distribution. Declare the return type that matches the output shape.
KMeans clustering
Segment customers by standardized purchase behavior. Cluster IDs have no ranking or meaning until interpreted against their feature profiles. k is required; initialization supports kmeans++ and forgy.
Algorithm reference · Customer segmentation tutorial
Linear regression
Estimate a continuous quantity such as revenue from scaled predictors. Start with a linear baseline and inspect held-out residuals before increasing complexity.
Algorithm reference · Regression tutorial
Logistic regression (binary)
Estimate churn or fraud probability from labelled examples. Choose the decision threshold using the relative cost of false positives and missed cases; probability calibration requires separate evaluation.
Algorithm reference · Churn tutorial
Multinomial logistic regression
Route observations into one of several classes. Labels are integers from 0 to k - 1; predictions are integer class IDs.
Algorithm reference · Multiclass tutorial
AR(p) / ARIMA forecasting
Forecast a regularly sampled series ordered by time. model(1) means the next step after the training series; model(7) means seven steps ahead. Split training and evaluation chronologically.
Algorithm reference · Revenue forecasting tutorial
Principal Component Analysis
Compress correlated measurements into a smaller set of axes. Use num_components = 2 with RETURNS ARRAY<DOUBLE> for a two-dimensional projection. PCA centers inputs; scaling and whitening are separate operations.
Algorithm reference · Dimensionality reduction tutorial
Decision tree
Learn interpretable branching rules for a categorical target. Limit depth and leaf size to control overfitting. The stump alias does not force depth one; set max_depth = 1 explicitly.
Algorithm reference · Tree comparison tutorial
Random forest
Combine bootstrap-sampled trees to reduce single-tree variability. Compare quality, model size, and inference cost on your workload. Balanced class weights are supported; arbitrary weight maps are not.
Algorithm reference · Tree comparison tutorial
Gradient boosting
Build successive trees to correct the ensemble's classification errors. Tune learning rate, round count, and depth together using held-out data.
Algorithm reference · Tree comparison tutorial
Deep neural network
Fit nonlinear interactions with a fully connected network. Use task = 'regression', 'binary', or 'multiclass'; multiclass also requires num_classes.
Algorithm reference · DNN tutorial
Training and evaluation
Iterative native trainers default to max_iters = 100; setting a larger value in an example is a tuning choice. A convergence flag is not evidence of predictive accuracy. Compare held-out predictions and task-appropriate metrics, and keep preprocessing consistent.
Training currently materializes data on the coordinator, including when partitions > 1. The default cap is 1,000,000 rows. Read distributed training before increasing the cap or enabling worker dispatch.
Imported framework models
ONNX
Import an artifact compatible with the engine's numeric tabular adapter. Exportability to ONNX alone does not establish input/output compatibility.
ONNX reference · Import and version activation tutorial
PyTorch (TorchScript)
Use a scripted or traced module; the PyTorch runtime must be enabled for your account. A saved eager-mode state dictionary is not a TorchScript model.
scikit-learn (skops)
Use skops rather than pickle and ensure the worker trusts every serialized estimator type. The numeric prediction adapter does not imply support for arbitrary text pipelines or custom transformers.
Anomaly detection
Durable SQL anomaly streams currently support METHOD 'ema'. Other detector methods, such as isolation forest, are not available as SQL functions or durable stream methods.
Anomaly reference · Streaming tutorial
Vector indexes
Vector indexes accelerate retrieval where the planner and deployment support the access path. Listing an index does not prove that a query used it; inspect EXPLAIN and compare recall against exact search.
Vector reference · Vector search tutorial
Online & streaming variants
Continuous inference applies an existing model to new rows. Online learning additionally updates native linear or logistic model parameters from labelled feedback. These are separate operations with separate background permissions.
Streaming reference · Online learning tutorial · Feature store