AI & ML Overview
Gnok brings predictive models, search, and language operations into SQL. Native training and ONNX inference are built into the hosted service; PyTorch and sklearn models need their runtime enabled for your account. AI functions send authorized inputs to Gnok's AI provider under your organization's AI Governance, and remote models call approved HTTPS endpoints.
Start with a runnable tutorial
The AI & ML tutorials provide 17 SQL walkthroughs with practical scenarios, model explanations, expected results, and downloadable scripts. Start with setup, then choose a predictive model, search, deployment, or background-processing example.
Choose a capability
| Goal | Reference | Practical starting point |
|---|---|---|
| Predict a number or class from tabular data | Supported models and Algorithm Reference | Churn, regression, tree ensembles, DNN |
| Discover segments or compress measurements | KMeans, PCA | Segmentation, PCA |
| Forecast an ordered time series | ARIMA / SARIMA | Revenue forecasting |
| Run a model trained elsewhere | ONNX, PyTorch and sklearn, notebooks | Import and activate a version, remote inference |
| Find documents by meaning or words | Vector operations, BM25 | Vector search and RAG, text search |
| Summarize or classify text with an LLM | AI scalar functions | AISQL cookbook |
| Ask a question about a dataset | Natural language queries | ASK and query suggestions |
| Maintain reusable features and score new events | Feature store, streaming ML, anomaly detection | Feature groups, online learning, streaming |
| Analyze distributions or relationships | Statistical functions | SQL aggregates for baselines and model evaluation |
| Improve query plans and table layouts | Learned optimization, AutoML advisors | Inspect optimizer state and recommendations for your workload |
Before running an example
Use configuration for what Gnok manages and what you control, and setup for tenant permissions and AI Governance. Native training, external inference, and background processing have different requirements.
- A supervised model signature lists features only; the training query puts the label first.
- Match the declared return type to the task: a class ID, probability, scalar prediction, and PCA array are different outputs.
- Apply the same feature order, units, encoding, and scaling during training and inference.
- Feature storage is enabled by default, but background work still requires a tenant service identity with suitable grants.
- AI features are off until an organization administrator saves an Available token budget and enables a rollout policy under Operations → AI Governance in Studio.
- Exact vector results do not establish ANN index use. Check the execution plan before drawing performance conclusions.
From example to application
Tutorial data makes the behavior easy to inspect. For an application, use representative held-out data, choose metrics that match the business decision, and measure latency and resource use. Inspect missing values and failure behavior as well as successful predictions. See distributed training for training-size limits and worker dispatch.