Skip to main content

AI & ML Overview

Gnok brings predictive models, search, and language operations into SQL. Native training and ONNX inference are built into the hosted service; PyTorch and sklearn models need their runtime enabled for your account. AI functions send authorized inputs to Gnok's AI provider under your organization's AI Governance, and remote models call approved HTTPS endpoints.

Start with a runnable tutorial​

The AI & ML tutorials provide 17 SQL walkthroughs with practical scenarios, model explanations, expected results, and downloadable scripts. Start with setup, then choose a predictive model, search, deployment, or background-processing example.

Choose a capability​

GoalReferencePractical starting point
Predict a number or class from tabular dataSupported models and Algorithm ReferenceChurn, regression, tree ensembles, DNN
Discover segments or compress measurementsKMeans, PCASegmentation, PCA
Forecast an ordered time seriesARIMA / SARIMARevenue forecasting
Run a model trained elsewhereONNX, PyTorch and sklearn, notebooksImport and activate a version, remote inference
Find documents by meaning or wordsVector operations, BM25Vector search and RAG, text search
Summarize or classify text with an LLMAI scalar functionsAISQL cookbook
Ask a question about a datasetNatural language queriesASK and query suggestions
Maintain reusable features and score new eventsFeature store, streaming ML, anomaly detectionFeature groups, online learning, streaming
Analyze distributions or relationshipsStatistical functionsSQL aggregates for baselines and model evaluation
Improve query plans and table layoutsLearned optimization, AutoML advisorsInspect optimizer state and recommendations for your workload

Before running an example​

Use configuration for what Gnok manages and what you control, and setup for tenant permissions and AI Governance. Native training, external inference, and background processing have different requirements.

  • A supervised model signature lists features only; the training query puts the label first.
  • Match the declared return type to the task: a class ID, probability, scalar prediction, and PCA array are different outputs.
  • Apply the same feature order, units, encoding, and scaling during training and inference.
  • Feature storage is enabled by default, but background work still requires a tenant service identity with suitable grants.
  • AI features are off until an organization administrator saves an Available token budget and enables a rollout policy under Operations → AI Governance in Studio.
  • Exact vector results do not establish ANN index use. Check the execution plan before drawing performance conclusions.

From example to application​

Tutorial data makes the behavior easy to inspect. For an application, use representative held-out data, choose metrics that match the business decision, and measure latency and resource use. Inspect missing values and failure behavior as well as successful predictions. See distributed training for training-size limits and worker dispatch.