Skip to main content
Apache Iceberg · SQL · Analytics · AI
Gnok

One SQL surface.Analytics, vectors, and AI.

Run distributed analytics, train models, search by meaning, and ask questions in plain English—without stitching together another stack.

Apache Iceberg v2 + v3Flight SQLPostgreSQL wire

Create an organization or sign in to Gnok Studio, open a worksheet, and inspect the result yourself. Follow the quickstart.

homepage-proof.sql Observed run
01SELECT COUNT(*), SUM(ss_net_paid)
02FROM tpcds_catalog.sf1.store_sales;
Observed resultFlight SQL · example
Rows2,880,404
Net paid4.741B
Checked2026-09-13
One engine, four native layers

Stop moving data between systems.

Gnok puts analytical SQL, model execution, semantic search, and Iceberg operations on the same distributed runtime.

AI in SQL01

CREATE MODEL trains, the model serves

Logistic, linear, gradient boosting, random forest, decision tree, K-means, PCA, AR/ARIMA, DNN, and softmax — train and serve from a SQL prompt. ONNX import and remote model endpoints both plug in.

Vector & LLM02

VECTOR(N), HNSW, EMBED, AI_*

Native vector type with HNSW + IVF indexes, hybrid BM25 + cosine ranking, and LLM-backed AI functions: AI_COMPLETE, AI_CLASSIFY_TEXT, AI_AGG, AI_SUMMARIZE.

Talk to your data03

ASK ‘how many orders shipped last week?’

Multi-turn natural-language interface with schema RAG, conversation persistence, and an automatic repair loop on bind errors. Audit-trailed and tenant-scoped.

Iceberg-native04

v2 deletes, v3 types, time travel, WAL streaming

4-level partition pruning, GEOMETRY/VARIANT, snapshot refs, materialized views with staleness budgets, sub-second WAL ingest. Vectorized + distributed.

What it looks like

The shortest path from data to decision.

Keep the workflow declarative. Gnok handles distributed execution, model dispatch, and storage underneath.

Explore the surface
Adapt and run this exampleOpen the guide

CREATE MODEL trains in-engine — logistic, linear, KMeans, GBM, random forest, decision tree, PCA, AR/ARIMA, DNN, or softmax. The trained model becomes a scalar UDF named after itself.

-- Train a churn classifier directly from a SELECT.
-- Numeric labels and features are normalized automatically.
CREATE MODEL churn(DOUBLE, DOUBLE, DOUBLE) RETURNS INT
TYPE 'gradient_boosting'
OPTIONS (num_trees = 200, learning_rate = 0.05, max_depth = 4)
AS SELECT
is_churn AS label,
age,
spend,
tenure_days
FROM customers
WHERE year(signup_date) = 2025;

-- Score every customer. The trained model is invoked as a
-- scalar UDF named after itself; the artifact is hydrated
-- on every worker on first reference.
SELECT customer_id,
churn(age, spend, tenure_days) AS predicted_class
FROM customers;
Built in, not bolted on

What you'd otherwise be gluing together

01
Train without leaving SQL

CREATE MODEL trains in-engine — logistic, GBM, random forest, AR/ARIMA, K-means, PCA, DNN — and the trained model functions as a scalar UDF. No notebook stack, no external model server.

02
First-class vector search

VECTOR(N) with COSINE_DISTANCE, L2_DISTANCE, and INNER_PRODUCT. Persistent HNSW and IVF indexes, hybrid BM25 + cosine fusion in a single ORDER BY.

03
LLM-backed AI in every client

AI_COMPLETE · AI_CLASSIFY_TEXT · AI_SUMMARIZE · AI_TRANSLATE · AI_AGG — scalar and aggregate variants, batched and audit-trailed. Reachable from any tool that speaks SQL.

04
Drop-in for the stack you have

Flight SQL, PostgreSQL wire, and REST on the front; ONNX import and remote model endpoints on the back. Any tool that speaks SQL connects today.

How it runs

Distributed engine, full SQL surface, production ops

Distributed MPP1

Coordinator + workers, vectorised end to end

8192-row Arrow batches, SIMD operators, work-stealing scheduler. Dynamic bloom filters push down to Parquet scans; shuffles travel over Arrow Flight. Reads Iceberg directly from object storage.

Full SQL surface2

Standard SQL, plus the modern shorthand

Broadcast, shuffle-hash, sort-merge, range, and nested-loop joins. Recursive CTEs, window functions, ROLLUP / CUBE / GROUPING SETS, PIVOT / UNPIVOT, QUALIFY, LATERAL FLATTEN, TABLESAMPLE. HLL and t-digest sketches, H3 geospatial, VARIANT / GEOMETRY / GEOGRAPHY.

Workload management3

Auto-classified queries, isolated warehouses

Every query is classified interactive / medium / batch and scheduled on a multi-level feedback queue. Virtual warehouses give each workload its own pool — SIZE, MAX_CONCURRENT, AUTO_SUSPEND_SECS — and suspend when idle.

Built-in pipelines4

Streams, Tasks, and materialized views

Streams capture Iceberg CDC; Tasks run scheduled or chained SQL; materialized views refresh incrementally from snapshot bookmarks. End-to-end ELT without an external orchestrator.

AI in SQL

Train and serve models from a SQL prompt

CREATE MODEL from a SELECT

Logistic, linear, K-means, PCA, decision tree, random forest, gradient boosting, AR/ARIMA, DNN, and softmax — ten native algorithms trained in-engine, plus ONNX import.

Streaming inference & anomaly detection

Score rows in real time with anomaly_score() and anomaly_score_iforest() — Isolation Forest, Z-score, and IQR built in. Same scalar UDF surface as native models.

Inference as a scalar UDF

Trained models are dispatched as scalar UDFs over Arrow batches, with version pinning. In-engine native runtimes and remote model endpoints are both supported.

Feature store

FEATURE_LOOKUP and FEATURE_LOOKUP_AS_OF for point-in-time training data. Catalog-registered feature lineage; multiple value types.

Vector & natural language

Embeddings, hybrid search, and English-to-SQL

VECTOR(N) + HNSW / IVF

First-class fixed-dimension vector type with persistent HNSW and IVF indexes, sharded across workers. L2_DISTANCE, COSINE_DISTANCE, INNER_PRODUCT for nearest-neighbor search.

EMBED() at write or query time

EMBED(text) generates embeddings via any registered remote model provider. Batched, deduped, async-dispatched.

BM25 + vector hybrid ranking

Persistent BM25 text indexes (Okapi, configurable k1/b/tokenizer). Combine lexical + semantic scores in a single ORDER BY.

ASK — natural language to SQL

Schema-aware text-to-SQL with session-scoped multi-turn context and an automatic repair loop on bind errors. Single statement: ASK '…' WITH SCOPE <catalog>.<schema>.

Iceberg & streaming

Iceberg-native runtime, sub-second writes

Iceberg v2 + v3 first class

Equality / position deletes, GEOMETRY, GEOGRAPHY, VARIANT, branches, tags. 4-level partition pruning end to end.

Sub-second WAL streaming

Vectorized + distributed (8192-row batches, SIMD, work-stealing) with WAL-durable writes that become queryable in under a second.

Time travel by snapshot or timestamp

FOR VERSION AS OF and FOR TIMESTAMP AS OF query any prior snapshot — plus an AT(SNAPSHOT => …) / AT(TIMESTAMP => …) function-call form.

Learned cardinality optimizer

Execution feedback trains lightweight cardinality and cost models in the background. Future plans pick up the corrected estimates automatically — no user action required.

Connect from anywhere

Pick a wire protocol, get to work

ClientProtocolUse case
gnok-studioFlight SQLSQL editor, history, AI assistant, dashboards
Any Postgres-wire clientPostgreSQL wireDrop-in for tools and CLIs that already speak Postgres
Java / Python / Rust SDKFlight SQLVectorised Arrow streams to your application
JDBC driversFlight SQLBI dashboards and analytics tools
curl / cron / scriptsRESTQuick scripted reads, CI checks, lightweight integrations
Streaming producersgRPC ingestWAL-backed streaming writes
Try it yourself

Start with a working example.

Explore your workspace

Sign in to Gnok Studio and work through the SQL, vector, model, and time-travel examples shown above. The guides include prerequisites, executable queries, and expected results.

Start the quickstart →