Skip to main content

AI & ML tutorials

Learn how to train models, score data, search documents, and operate background ML workloads in Gnok. These 17 walkthroughs match the shared tutorials in Studio: in the SQL editor's sidebar, open the Saved tab and expand Shared with me. The tutorials appear there in one list, named Tutorial: … and tagged tutorial and ai-ml, so you can find them by typing tutorial in the search box. Each walkthrough explains the practical problem, the model or technique, the complete SQL, and what a correct result looks like.

Start with setup and troubleshooting. The AISQL, RAG, and natural-language tutorials also need an organization administrator to enable AI Governance first. The examples create synthetic data and intentionally reset their tutorial objects. Every page includes a downloadable SQL file so you can start from the current example instead of an old editor draft.

Choose a starting point​

  • New to SQL-based ML: start with churn prediction, then linear regression. Learn the distinction between a class probability and a prediction with physical units.
  • Working with documents: start with BM25 for literal terms, then vector search and RAG for semantic retrieval and grounded answers.
  • Already have a model: use the ONNX import/versioning tutorial or the HTTP remote-model tutorial.
  • Need results as data arrives: learn feature groups, then streaming scoring and feedback-driven online learning.

Predictive models and exploration​

TutorialPractical questionTechnique and result
Churn predictionWhich customers resemble those who previously cancelled?Logistic regression; a binary-outcome probability.
Linear regressionHow can listing characteristics explain price variation?A numeric baseline with explicit scaling and dollar conversion.
Multinomial classificationWhich of several categories matches these measurements?Three-class prediction and a confusion table.
K-means segmentationWhich customers have similar engagement patterns?Unsupervised groups with interpretable feature profiles.
Tree ensembles and tuningHow do different tabular classifiers compare?Trees, forests, boosting, feature importance, and 12 tuning variants.
DNN / MLPCan a model learn an interaction that a simple boundary misses?A small neural network on an XOR-style fixture.
Time-series forecastingWhat follows the observations available at forecast time?An ordered ARIMA example with seven held-out periods.
PCAHow can six related measurements be represented in two dimensions?An unsupervised projection returning a two-element array.

Search and language​

TutorialPractical questionTechnique and result
Vector search and RAGWhich policy passages answer a support question?Embeddings, metadata filtering, distance ranking, and a grounded answer.
AISQL cookbookWhat do product reviews say, and which issues need attention?Classification, summary, extraction, translation, and grouped feedback.
BM25 and hybrid retrievalHow do literal terms and semantic similarity complement each other?Text indexing and reciprocal rank fusion.
Natural language to SQLWhich product has the most revenue under a precise definition?Scoped ASK, scope inspection, conversations, and query suggestions.

Deploy and operate models​

TutorialPractical questionTechnique and result
Import models and activate versionsCan SQL use an externally produced artifact and promote a candidate?A complete ONNX example with an observable score change.
Remote scoringCan SQL call an existing model-serving API?A typed HTTP batch contract, called on a demo scorer that Gnok hosts, with a known score.
Feature groups and backfillCan several consumers reuse the same historical customer aggregates?Offline materialization, bounded backfill, and drift inspection.
Online learning and slicesHow should delayed labels feed model updates and segment monitoring?A base model, feedback contract, and US/EU slices.
Streaming MLCan arriving events produce durable scores and anomaly diagnostics?CDC, derived streams, explicit refresh, and result-count checks.

Understand the workflow​

For supervised models, the training query places the target label before the feature columns. Unsupervised models such as K-means and PCA use features without labels. Forecasting uses an ordered series and a horizon. Imported and remote models follow the input/output contract of the artifact or endpoint. These differences are explained in the individual walkthroughs.