Skip to main content

Supported Model Types

Choose a model by its task, output contract, and runtime requirements. The tutorials provide complete datasets and scripts; the algorithm references explain options and tuning. The examples in reference pages use illustrative source tables unless stated otherwise.

Surface summary​

SurfaceRegistrationRuntime requirements
Native trainingTYPE '<algorithm>' OPTIONS (...) AS SELECT ...Engine training code; no Python or AI provider
ONNXUSING FRAMEWORK 'onnx' AS FROM ...Built in; supported tensor shape and operators
PyTorchUSING FRAMEWORK 'pytorch' AS FROM ...PyTorch runtime enabled for your account; TorchScript artifact
sklearnUSING FRAMEWORK 'sklearn' AS FROM ...Python runtime enabled for your account; trusted skops artifact
Remote endpointRemote-model definitionPublic HTTPS endpoint approved by Gnok; matching request/response contract

Registered models can be invoked by qualified name or with ML_PREDICT('catalog.schema.model', feature1, feature2). See imported runtimes, ONNX, and the remote endpoint tutorial.

Native algorithms​

For supervised training, put the label first in AS SELECT, followed by the features. The SQL function signature contains features only. KMeans and PCA use every selected column as a feature. ARIMA is the special case: training reads one chronological value series, while inference accepts a forecast horizon.

Native training normalizes supported numeric and Boolean columns to DOUBLE, including decimal inputs. Encode categorical strings explicitly. Preserve feature order, units, and scaling at inference; numeric conversion does not standardize features. Use care with integers or decimals whose precision exceeds DOUBLE.

AlgorithmCanonical TYPESQL result
KMeanskmeansINT cluster ID
Linear regressionlinearDOUBLE predicted value
Binary logistic regressionlogisticDOUBLE probability of class 1
Multinomial logistic regressionmultinomialINT class ID
AR / ARIMA / SARIMAarimaDOUBLE prediction at a supplied horizon
PCApcaDOUBLE for one component; ARRAY<DOUBLE> for multiple components
Decision treedecision_treeINT class ID
Random forestrandom_forestINT class ID
Gradient boostinggradient_boostingINT class ID
DNN / MLPdnnRegression/binary: DOUBLE; multiclass: INT class ID

Class IDs are not probabilities. Native tree models are classification-only. Native multiclass models expose the winning class, not the full probability distribution. Declare the return type that matches the output shape.

KMeans clustering​

Segment customers by standardized purchase behavior. Cluster IDs have no ranking or meaning until interpreted against their feature profiles. k is required; initialization supports kmeans++ and forgy.

Algorithm reference · Customer segmentation tutorial

Linear regression​

Estimate a continuous quantity such as revenue from scaled predictors. Start with a linear baseline and inspect held-out residuals before increasing complexity.

Algorithm reference · Regression tutorial

Logistic regression (binary)​

Estimate churn or fraud probability from labelled examples. Choose the decision threshold using the relative cost of false positives and missed cases; probability calibration requires separate evaluation.

Algorithm reference · Churn tutorial

Multinomial logistic regression​

Route observations into one of several classes. Labels are integers from 0 to k - 1; predictions are integer class IDs.

Algorithm reference · Multiclass tutorial

AR(p) / ARIMA forecasting​

Forecast a regularly sampled series ordered by time. model(1) means the next step after the training series; model(7) means seven steps ahead. Split training and evaluation chronologically.

Algorithm reference · Revenue forecasting tutorial

Principal Component Analysis​

Compress correlated measurements into a smaller set of axes. Use num_components = 2 with RETURNS ARRAY<DOUBLE> for a two-dimensional projection. PCA centers inputs; scaling and whitening are separate operations.

Algorithm reference · Dimensionality reduction tutorial

Decision tree​

Learn interpretable branching rules for a categorical target. Limit depth and leaf size to control overfitting. The stump alias does not force depth one; set max_depth = 1 explicitly.

Algorithm reference · Tree comparison tutorial

Random forest​

Combine bootstrap-sampled trees to reduce single-tree variability. Compare quality, model size, and inference cost on your workload. Balanced class weights are supported; arbitrary weight maps are not.

Algorithm reference · Tree comparison tutorial

Gradient boosting​

Build successive trees to correct the ensemble's classification errors. Tune learning rate, round count, and depth together using held-out data.

Algorithm reference · Tree comparison tutorial

Deep neural network​

Fit nonlinear interactions with a fully connected network. Use task = 'regression', 'binary', or 'multiclass'; multiclass also requires num_classes.

Algorithm reference · DNN tutorial

Training and evaluation​

Iterative native trainers default to max_iters = 100; setting a larger value in an example is a tuning choice. A convergence flag is not evidence of predictive accuracy. Compare held-out predictions and task-appropriate metrics, and keep preprocessing consistent.

Training currently materializes data on the coordinator, including when partitions > 1. The default cap is 1,000,000 rows. Read distributed training before increasing the cap or enabling worker dispatch.

Imported framework models​

ONNX​

Import an artifact compatible with the engine's numeric tabular adapter. Exportability to ONNX alone does not establish input/output compatibility.

ONNX reference · Import and version activation tutorial

PyTorch (TorchScript)​

Use a scripted or traced module; the PyTorch runtime must be enabled for your account. A saved eager-mode state dictionary is not a TorchScript model.

Runtime reference

scikit-learn (skops)​

Use skops rather than pickle and ensure the worker trusts every serialized estimator type. The numeric prediction adapter does not imply support for arbitrary text pipelines or custom transformers.

Runtime reference

Anomaly detection​

Durable SQL anomaly streams currently support METHOD 'ema'. Other detector methods, such as isolation forest, are not available as SQL functions or durable stream methods.

Anomaly reference · Streaming tutorial

Vector indexes​

Vector indexes accelerate retrieval where the planner and deployment support the access path. Listing an index does not prove that a query used it; inspect EXPLAIN and compare recall against exact search.

Vector reference · Vector search tutorial

Online & streaming variants​

Continuous inference applies an existing model to new rows. Online learning additionally updates native linear or logistic model parameters from labelled feedback. These are separate operations with separate background permissions.

Streaming reference · Online learning tutorial · Feature store