Skip to main content

PyTorch & sklearn Inference

Gnok supports three inference runtime paths, subject to account availability — ONNX, PyTorch (TorchScript), and sklearn — that all register through the same CREATE MODEL DDL with a USING FRAMEWORK clause and surface as scalar UDFs callable from any SELECT / WHERE / JOIN expression.

RuntimeModel formatAvailability
ONNX.onnxBuilt into the hosted service
PyTorchTorchScript .ptRequires the PyTorch runtime to be enabled for your account
sklearnskops .skopsRequires the Python runtime to be enabled for your account

When to pick which​

  • ONNX — the universal export target. Check exporter support and the numeric scalar adapter contract before selecting this path.
  • PyTorch (TorchScript) — when you can't or don't want to go through the ONNX export step (e.g. custom autograd ops, models with a supported TorchScript graph). Requires the corresponding runtime to be enabled.
  • sklearn — when your model is a scikit-learn Pipeline with feature preprocessors that ONNX doesn't round-trip cleanly. Runs in a Python sandbox with a subprocess lifecycle managed by the runtime.

PyTorch​

Runtime availability​

Gnok manages the runtime and its dependencies; there is nothing to install. Ask Gnok support whether this runtime is enabled for your account before exporting an artifact for this framework.

Export → CREATE MODEL​

Trace or script your PyTorch module into TorchScript:

import torch

class FraudClassifier(torch.nn.Module):
def __init__(self):
super().__init__()
self.linear = torch.nn.Linear(3, 1)

def forward(self, x):
return torch.sigmoid(self.linear(x))

model = FraudClassifier()
# Fit model weights on labelled data before exporting for real use.
model.eval()
scripted = torch.jit.script(model)
scripted.save("fraud.torchscript")

Register it:

-- Inline base64 (small models, easy to scope)
CREATE MODEL fraud_pt(FLOAT, FLOAT, INT) RETURNS FLOAT
USING FRAMEWORK 'pytorch'
AS FROM '<base64-torchscript-bytes>';

-- Combined with quantisation
CREATE MODEL fraud_pt(FLOAT, FLOAT, INT) RETURNS FLOAT
QUANTISATION 'fp16' USING FRAMEWORK 'pytorch'
AS FROM '<base64-torchscript-bytes>';

Calling from SQL​

Same surface as ONNX:

SELECT
txn_id,
amount,
fraud_pt(amount, velocity, country_id) AS score
FROM transactions
WHERE fraud_pt(amount, velocity, country_id) > 0.85;

Failure semantics​

The runtime returns an inference error when a batch cannot be scored. This can fail the query; there is no guaranteed automatic per-row NULL fallback. Validate input shapes and preprocessing before scoring a large dataset.

sklearn​

Runtime availability​

Gnok manages the runtime and its dependencies; there is nothing to install. Ask Gnok support whether this runtime is enabled for your account before exporting an artifact for this framework.

Export → CREATE MODEL​

Pickle is rejected. Pickle deserialisation runs arbitrary code — that's a security hole we're deliberately avoiding. Convert your model via skops once at export time:

import skops.io as sio
sio.dump(my_pipeline, "fraud.skops")

Register the skops file:

CREATE MODEL fraud_sk(FLOAT, FLOAT, FLOAT) RETURNS FLOAT
USING FRAMEWORK 'sklearn'
AS FROM '<base64-skops-bytes>';

The worker validates the file's ZIP magic bytes (skops is a ZIP archive) and refuses anything else with a clear error referencing skops.io.dump.

Calling from SQL​

Same surface as PyTorch / ONNX:

SELECT
customer_id,
fraud_sk(amount, velocity, country_id) AS score
FROM live_payments;

Process lifecycle​

  • Gnok starts an isolated worker process for each model on its first inference and stops it after 5 minutes of inactivity or if it exceeds its resource limits.
  • Each model's worker handles one call at a time, so a single sklearn model has limited throughput. Gnok manages this capacity; for high-throughput scoring, convert the model to ONNX or contact Gnok support.
  • Per-call timeout: 20 s.

Trusted types​

The worker loads skops files with no additional trusted types, so only built-in scikit-learn types load. If your pipeline uses third-party estimators (e.g. xgboost wrapped in Pipeline), the load fails with an UntrustedTypesFoundError listing the types. Convert such a model to ONNX instead.

Failure semantics​

Worker prediction failures propagate as inference errors and can fail the query. A predict method alone does not establish compatibility: check supported input types, trusted estimator classes, and the output contract.

ML_PREDICT general-purpose dispatch​

ML_PREDICT is the framework-neutral way to invoke any registered model from SQL. The same function call works against ONNX, PyTorch (TorchScript), sklearn, and trained-in-engine native models — the binder rewrites ML_PREDICT('model_name', col1, col2, …) to a direct UDF call against the named model, so the dispatch path is identical to a plain function call by name.

-- Identical SQL surface regardless of framework
SELECT ML_PREDICT('fraud_v1', amount, velocity, country_id) AS s_onnx,
ML_PREDICT('fraud_pt', amount, velocity, country_id) AS s_pt,
ML_PREDICT('fraud_sk', amount, velocity, country_id) AS s_sk,
ML_PREDICT('fraud_dnn', amount, velocity, country_id) AS s_native
FROM transactions;

Constraints​

The first argument must be a string literal — a column reference or expression is rejected at bind time. The model name has to be resolved at planning time so the registry lookup is constant across the batch:

-- OK — string literal
SELECT ML_PREDICT('fraud_v1', amount) FROM tx;

-- ERROR: ML_PREDICT: first argument must be a literal model name,
-- not a column reference or expression
SELECT ML_PREDICT(model_name_col, amount) FROM tx;

Return type​

ML_PREDICT returns whatever the model's RETURNS clause declared at CREATE MODEL time. For a fraud scorer declared RETURNS FLOAT the SQL type is FLOAT (Float32); RETURNS DOUBLE is Float64. Its interpretation depends on the model, not the type alone. For a clustering model declared RETURNS INT it returns an INT32 cluster id. For a PCA model with k > 1 it returns a LIST<FLOAT64> of projections.

Equivalence with bare-name dispatch​

ML_PREDICT('m', a, b) is exactly equivalent to calling the model by name as m(a, b). The dispatch shape exists so the same SQL is portable across deployments where the model framework might change — registering a new ONNX export of a model under the same name swaps the implementation without touching the call sites.

Versioned model upgrades​

Trained-in-engine native models support the same versioning pipeline as ONNX. CREATE MODEL name VERSION 'vN' (...) stages a new version without flipping the active version; ALTER MODEL <name> ACTIVATE VERSION 'vN' is the explicit cutover. See CREATE MODEL VERSION for the full DDL.

-- Stage v2 of a trained-in-engine logistic model. The previous
-- version stays ACTIVE; SELECT ML_PREDICT('fraud', ...) keeps
-- routing to v1 until you activate v2.
CREATE MODEL fraud VERSION 'v2' (DOUBLE, DOUBLE, DOUBLE)
RETURNS DOUBLE TYPE 'logistic'
OPTIONS (learning_rate = 0.5, max_iters = 200)
AS SELECT is_fraud::DOUBLE, amount::DOUBLE,
merchant_score::DOUBLE, hour_of_day::DOUBLE
FROM transactions
WHERE created_at >= now() - INTERVAL '90 days';

-- Cut over once the new version's offline metrics check out
ALTER MODEL fraud ACTIVATE VERSION 'v2';

OR REPLACE and VERSION are mutually exclusive — the engine rejects CREATE OR REPLACE MODEL … VERSION 'vN' at parse time because the two operations have opposite intent (replace = swap in place; version = additive staging).

Mixing runtimes in one query​

Compare runtimes on the same rows and preprocessing. Use an outer query to filter by derived scores:

WITH scored AS (
SELECT txn_id,
fraud_v1(amount, velocity, country_id) AS onnx_score,
fraud_pt(amount, velocity, country_id) AS pt_score
FROM transactions
)
SELECT txn_id, onnx_score, pt_score,
ABS(onnx_score - pt_score) AS divergence
FROM scored
WHERE ABS(onnx_score - pt_score) > 0.10;

The example assumes existing models with matching scalar score semantics. Comparing a class ID with a probability would not be a meaningful runtime check.

Troubleshooting​

"framework not available" at CREATE MODEL​

The PyTorch or Python runtime isn't enabled for your account, so CREATE MODEL reports that the framework is not supported. Ask Gnok support about availability, or export the model to ONNX, which is built into the hosted service.

"Invalid TorchScript model" at CREATE MODEL​

The artifact doesn't load as a TorchScript module. Common causes:

  • You exported eager-mode .pt instead of TorchScript. Run torch.jit.script(model).save(...) not torch.save(model.state_dict(), ...).
  • PyTorch version mismatch. Export with a PyTorch version that is compatible with Gnok's runtime; Gnok support can tell you which version that is.

"model bytes don't look like skops format" at CREATE MODEL​

You're trying to register a pickle file. Re-export via skops.io.dump. The error message is intentional — if we silently accepted pickle the security model would be broken.

Inference OOM on big batches​

Check the model's tensor shape and per-batch memory requirements. A final SQL LIMIT is not a dependable bound on upstream inference work. Score a bounded input first and filter rows before scoring. Contact Gnok support if a model needs more memory than your account allows.

For a complete ONNX artifact and activation workflow, follow the imported-model tutorial. PyTorch and sklearn require separate runtime and artifact checks.