Skip to main content

One ALTER TABLE for Millisecond Writes

· 5 min read
Gnok Team

A single-row INSERT into an Iceberg table is a strange thing to benchmark. The row is a few dozen bytes; the commit that lands it rewrites metadata, swaps a pointer, and waits for the catalog to say yes. On our own cluster that costs about 712 ms, and no amount of tuning changes the shape of it — you are paying for a catalog transaction, once per statement, whatever the statement contains.

That is fine for the workload Iceberg was built for. It is not fine for an application that edits one order, appends one comment, or increments one counter and then wants to read it back.

Gnok now takes a different path for those writes, and turning it on is one line of DDL:

ALTER TABLE orders SET TBLPROPERTIES ('gnok.write.mode' = 'command');

That write now acknowledges in 5–7 ms.

Build an Incremental Inventory Change Feed with Gnok CDC Streams

· 6 min read
Gnok Team

An inventory table may contain millions of products, but a downstream stock monitor usually needs only the rows that changed. Repeatedly scanning the full table wastes compute and makes deletes especially difficult to detect.

Gnok CDC streams turn the Iceberg snapshots behind a table into a bounded change feed. In this tutorial, we will create an inventory stream, inspect inserts, updates, and deletes, and acknowledge a batch after processing it.

From JSON to Iceberg: Bulk-Loading a Document Dataset into Gnok

· 7 min read
Gnok Team

We recently needed to land a large pile of JSON documents — roughly 50 GB across ~70 datasets, including a couple of 15+ GB monsters — into Gnok so it could be queried with SQL alongside everything else. The documents were schemaless and deeply nested; Gnok tables are columnar Iceberg. This post walks through the pipeline we landed on, and the handful of gotchas that shaped it.

From Data to SQL-Callable Model: An End-to-End Tour of Gnok Notebooks

· 4 min read
Gnok Team

Gnok Studio now has an in-product Python notebook. You write Python against governed Gnok data, train a model, and register it so it's callable from SQL with ML_PREDICT — the whole train→serve loop, in one place, with no data movement and no separate notebook server to babysit.

This post builds a complete example end to end: read data, explore it, engineer features, train a scikit-learn pipeline, register it, and score new rows from SQL — then let Gnok AI write a cell for us.

AI in the WHERE Clause: AI_FILTER and AI_FILTER_AGG

· 5 min read
Gnok Team

SELECT * FROM reviews WHERE AI_FILTER('is a complaint about shipping', body).

If that line of SQL works, a whole class of "I just need to find the rows where X holds, and X isn't a regex" problems disappears. Gnok's AI_FILTER family lands the natural-language predicate in the place where you'd actually write it — alongside =, LIKE, and BETWEEN.

This post walks through the four new functions, the patterns they unlock, and how to keep cost predictable when LLMs are sitting in your hot path.

Real Iceberg Rollback in Pure SQL

· 4 min read
Gnok Team

Most engines say they support time travel. Watch what happens when you actually try to roll back. Plenty of catalogs have shipped a "snapshot rollback" that quietly only edits the table's current-snapshot-id property — subsequent reads still hit the latest data. Until recently, gnok was one of them.

This post walks through how that path is now wired end-to-end: ALTER TABLE … SET SNAPSHOT actually flips the read pointer, the new is_current / parent_snapshot_id / sequence_number columns surface state honestly, and the $snapshots system view makes Iceberg metadata first-class SQL.

Learned Optimization: Inspecting What AutoML Has Picked Up

· 4 min read
Gnok Team

After watching ten thousand queries, Gnok has opinions about your workload. Learned scorers now feed the planner's partition and index advisors, so those opinions reflect the cost savings actually observed across your queries — not just the original heuristics.

The advisors run continuously as part of the service. This post walks through the surface that lets you see what they've learned: three SHOW … RECOMMENDATIONS commands that expose the live state of the join-order, materialized-view, and overall AutoML caches.

AI/ML in Gnok Goes Production-Grade: Filtered ANN, Point-in-Time Features, and Schema-Aware Autocomplete

· 6 min read
Gnok Team

The AI/ML stack in Gnok has been usable for a while — register an ONNX model, build an HNSW index, materialize a feature group, run inference inline. What's new this month is that every track now holds up at production scale: vector search with WHERE clauses no longer falls back to brute-force, point-in-time training joins read durably from Iceberg time-travel, and the SQL editor knows which models you've actually registered.

This post walks through the three biggest changes — filtered ANN, offline-backed feature lookups, and the schema-aware ML autocomplete — with the SQL you can run today.

Building a Real-Time Fraud Detection Pipeline with Gnok

· 10 min read
Gnok Team

This tutorial builds a complete fraud detection system inside Gnok -- from data ingestion to ML scoring to alerting -- without any external services. We'll use COPY INTO for bulk loading, vector similarity for merchant profiling, statistical anomaly detection for flagging outliers, and natural language queries for ad-hoc investigation.

Iceberg v3 in Gnok: Deletion Vectors, VARIANT, Spatial Types, and Row Lineage

· 8 min read
Gnok Team

Apache Iceberg format version 3 brings four major capabilities: deletion vectors for more efficient row-level mutations, the VARIANT type for semi-structured data, GEOMETRY/GEOGRAPHY types for spatial analytics, and row lineage for stable row identity across compaction. Gnok now supports all of them.

This post explains what each feature does, when to use it, and how to get started.