AI & ML tutorials
Learn how to train models, score data, search documents, and operate background ML workloads in Gnok. These 17 walkthroughs match the shared tutorials in Studio: in the SQL editor's sidebar, open the Saved tab and expand Shared with me. The tutorials appear there in one list, named Tutorial: … and tagged tutorial and ai-ml, so you can find them by typing tutorial in the search box. Each walkthrough explains the practical problem, the model or technique, the complete SQL, and what a correct result looks like.
Start with setup and troubleshooting. The AISQL, RAG, and natural-language tutorials also need an organization administrator to enable AI Governance first. The examples create synthetic data and intentionally reset their tutorial objects. Every page includes a downloadable SQL file so you can start from the current example instead of an old editor draft.
Choose a starting point
- New to SQL-based ML: start with churn prediction, then linear regression. Learn the distinction between a class probability and a prediction with physical units.
- Working with documents: start with BM25 for literal terms, then vector search and RAG for semantic retrieval and grounded answers.
- Already have a model: use the ONNX import/versioning tutorial or the HTTP remote-model tutorial.
- Need results as data arrives: learn feature groups, then streaming scoring and feedback-driven online learning.
Predictive models and exploration
| Tutorial | Practical question | Technique and result |
|---|---|---|
| Churn prediction | Which customers resemble those who previously cancelled? | Logistic regression; a binary-outcome probability. |
| Linear regression | How can listing characteristics explain price variation? | A numeric baseline with explicit scaling and dollar conversion. |
| Multinomial classification | Which of several categories matches these measurements? | Three-class prediction and a confusion table. |
| K-means segmentation | Which customers have similar engagement patterns? | Unsupervised groups with interpretable feature profiles. |
| Tree ensembles and tuning | How do different tabular classifiers compare? | Trees, forests, boosting, feature importance, and 12 tuning variants. |
| DNN / MLP | Can a model learn an interaction that a simple boundary misses? | A small neural network on an XOR-style fixture. |
| Time-series forecasting | What follows the observations available at forecast time? | An ordered ARIMA example with seven held-out periods. |
| PCA | How can six related measurements be represented in two dimensions? | An unsupervised projection returning a two-element array. |
Search and language
| Tutorial | Practical question | Technique and result |
|---|---|---|
| Vector search and RAG | Which policy passages answer a support question? | Embeddings, metadata filtering, distance ranking, and a grounded answer. |
| AISQL cookbook | What do product reviews say, and which issues need attention? | Classification, summary, extraction, translation, and grouped feedback. |
| BM25 and hybrid retrieval | How do literal terms and semantic similarity complement each other? | Text indexing and reciprocal rank fusion. |
| Natural language to SQL | Which product has the most revenue under a precise definition? | Scoped ASK, scope inspection, conversations, and query suggestions. |
Deploy and operate models
| Tutorial | Practical question | Technique and result |
|---|---|---|
| Import models and activate versions | Can SQL use an externally produced artifact and promote a candidate? | A complete ONNX example with an observable score change. |
| Remote scoring | Can SQL call an existing model-serving API? | A typed HTTP batch contract, called on a demo scorer that Gnok hosts, with a known score. |
| Feature groups and backfill | Can several consumers reuse the same historical customer aggregates? | Offline materialization, bounded backfill, and drift inspection. |
| Online learning and slices | How should delayed labels feed model updates and segment monitoring? | A base model, feedback contract, and US/EU slices. |
| Streaming ML | Can arriving events produce durable scores and anomaly diagnostics? | CDC, derived streams, explicit refresh, and result-count checks. |
Understand the workflow
For supervised models, the training query places the target label before the feature columns. Unsupervised models such as K-means and PCA use features without labels. Forecasting uses an ordered series and a horizon. Imported and remote models follow the input/output contract of the artifact or endpoint. These differences are explained in the individual walkthroughs.