Iceberg Compatibility
Gnok implements a native Apache Iceberg runtime supporting format versions 1, 2, and 3. This page details which features are available at each format version level.
Format Version Overview
| Version | Focus | Default |
|---|---|---|
| v1 | Foundational table format — immutable data files, schema evolution, partition specs, sort orders | No |
| v2 | Row-level mutations — position deletes, equality deletes, sequence numbers, snapshot references | Yes |
| v3 | Extended types & lineage — deletion vectors, VARIANT type, spatial types, row lineage, default values | No |
Gnok defaults to format version 2 when creating new tables. You can specify the version explicitly, and the requested version is honored at create time (the metadata is written at that format version, not just recorded as a property):
CREATE TABLE events (
id BIGINT,
payload VARCHAR
)
WITH ('format-version' = '3');
A table that declares a V3-only column type (VARIANT, GEOMETRY, GEOGRAPHY) is automatically
created at format version 3 even if a lower version is requested, since those types are invalid at
v1/v2 and would be rejected by other engines.
v1 Features
All v1 features are fully supported and form the foundation of every Iceberg table.
Table Metadata
| Feature | Status | Notes |
|---|---|---|
| Manifest lists | Supported | Metadata tree: manifest list → manifests → data files |
| Column statistics (min/max/null counts) | Supported | Per-field in manifest entries; used for file-level pruning |
| Partition specs | Supported | Identity, year, month, day, hour, bucket(N), truncate(W) |
| Sort orders | Supported | Per-table sort specification with field, transform, direction, null ordering |
| Schema evolution | Supported | Add, drop, rename, widen columns with field ID preservation |
| Table properties | Supported | Key-value metadata (compression, file size, bloom filter config) |
| Snapshot management | Supported | Append and overwrite operations; parent snapshot tracking |
| Multiple schemas | Supported | Schema history with unique schema_id per revision |
Scan & Pruning
Gnok implements 4-level pruning that applies to all format versions:
- Manifest list pruning — skip manifests by partition summary bounds
- File pruning — skip data files by per-column min/max statistics
- Row group pruning — skip Parquet row groups by page-level statistics
- Dynamic filter pruning — runtime bloom/min-max filters from join build sides
Catalogs
| Catalog | Status |
|---|---|
| REST Catalog (Iceberg REST spec) | Supported |
| AWS Glue Data Catalog | Supported |
| Nessie | Supported (external) |
| Hive Metastore | Supported (external) |
All catalog types support JWT token forwarding for multi-tenant isolation.
v2 Features
Format version 2 adds row-level delete capabilities and sequence-based ordering.
Row-Level Deletes
| Feature | Status | Notes |
|---|---|---|
| Position deletes | Supported | (file_path, pos) pairs in Parquet delete files |
| Equality deletes | Supported | Column-value predicates stored as delete files |
| Merge-on-Read (MOR) | Supported | Default strategy for UPDATE and DELETE |
| Copy-on-Write (COW) | Supported | Full file rewrite for small, targeted changes |
Position deletes use an adaptive index: Vec<i64> with binary search for fewer than 50,000 positions (service-managed threshold), and RoaringBitmap for larger sets.
Delete Strategies
| Strategy | When Selected |
|---|---|
| Broadcast | Small delete sets kept in memory |
| SortMerge | High delete-to-data ratio |
| BloomFilter | Large, sparse delete patterns |
Sequence Numbers
Sequence numbers provide a total ordering of commits:
sequence_numberon snapshots — monotonically increasing per commitmin_sequence_numberon manifest files — enables efficient delete file matchingdata_sequence_numberon data files — identifies which commit produced each file
Snapshot References (Branches & Tags)
-- View snapshots
SHOW SNAPSHOTS FROM orders;
-- Time travel by snapshot id (these forms are all accepted)
SELECT * FROM orders FOR SYSTEM_VERSION AS OF 1234567890;
SELECT * FROM orders FOR VERSION AS OF 1234567890;
SELECT * FROM orders AT(SNAPSHOT => 1234567890);
-- Time travel by timestamp
SELECT * FROM orders FOR SYSTEM_TIME AS OF TIMESTAMP '2024-01-15 10:00:00';
| Feature | Status |
|---|---|
| Branch references | Supported |
| Tag references | Supported |
Branch retention (min_snapshots_to_keep, max_snapshot_age_ms) | Supported |
Reference expiration (max_ref_age_ms) | Supported |
DML Operations
All DML operations produce v2-compliant snapshots with position delete files:
| Operation | Strategy | Notes |
|---|---|---|
| INSERT | Append data files | Partition-aware routing, idempotent file naming |
| INSERT OVERWRITE | Overwrite affected partitions | Metadata-level partition replacement |
| UPDATE | MOR (position deletes + new data files) | Atomic snapshot with both delete and data files |
| DELETE | MOR (position delete files) | Partition pruning applied to scan |
| MERGE | MOR (hash table probe + position deletes) | Supports MATCHED, NOT MATCHED, NOT MATCHED BY SOURCE |
| TRUNCATE | Metadata-only empty snapshot | No scan required |
| COPY INTO | Bulk append from staged files | SIMD-accelerated CSV/Parquet/JSON parsing |
Statistics (Puffin Format)
| Feature | Status |
|---|---|
| Statistics files | Supported |
| Theta sketch blobs (NDV estimation) | Supported |
| Blob metadata (snapshot ID, sequence number, field IDs) | Supported |
Partition statistics files (Iceberg partition-statistics, 12-field schema; per-partition row/file/size + bounds, registered via SetPartitionStatistics) | Supported |
NaN Value Counts
nan_value_counts per column in manifest entries — used for accurate statistics on floating-point columns. Supported in v2+ tables.
Partition Evolution
Tables can change their partition spec over time without rewriting existing data:
-- Original partitioning
CREATE TABLE events (...) PARTITIONED BY (day(event_time));
-- Evolve to hourly partitioning (new data uses new spec; old data retains old spec)
ALTER TABLE events SET PARTITION SPEC (hour(event_time));
| Feature | Status |
|---|---|
| Multiple partition specs in history | Supported |
| Mixed-spec scans (old + new data) | Supported |
| Partition field source tracking | Supported |
Schema Evolution
| Operation | Status |
|---|---|
| Add column | Supported |
| Drop column | Supported |
| Rename column | Supported |
| Widen type (e.g., INT → BIGINT) | Supported |
| Reorder columns | Supported |
| Set/drop NOT NULL | Supported |
| Set column comment | Supported |
Field IDs are preserved across schema changes — Gnok never reuses a retired field ID.
Encryption
| Feature | Status |
|---|---|
| Key hierarchy (Master → KEK → File Key) | Supported |
| Parquet AES-GCM | Supported |
| Per-file key tracking | Supported |
| Key metadata in table properties | Supported |
v3 Features
Format version 3 introduces extended type support, deletion vectors, and row lineage.
Deletion Vectors
Deletion vectors replace position delete files with a more compact, Puffin-based representation using 64-bit Roaring bitmaps.
| Feature | Status | Notes |
|---|---|---|
DV binary format (magic 0xD1D33964 + Roaring bitmap + IEEE CRC-32) | Supported | Spec-compliant Puffin framing (4-byte footer flags, blob length = magic+bitmap); the checksum is IEEE CRC-32 (== java.util.zip.CRC32 / zlib), not CRC-32C |
| Puffin blob storage with LZ4/Zstd compression | Supported | |
DV manifest classification (external content=1 + file_format=puffin) | Supported | Manifest reader reclassifies spec-shaped DVs (Spark/Trino/Doris/PyIceberg) to deletion vectors |
DV-to-data-file matching via referenced_data_file | Supported | |
| Scan-path execution (read Puffin blob + merge with position deletes) | Supported | Live-validated: scan preloads the DV and applies it during read |
| DV write path (auto-gated on v3) | Supported | UPDATE / DELETE / MERGE on a format-version-3 table automatically write Puffin DVs (content=1+file_format=puffin+content_offset) — gated on format_version >= 3, managed automatically by the service. Sequential edits to the same data file union into a single DV and retire the superseded manifest entry (one DV per data file), with snapshot added-rows recorded |
-- format-version 3 tables both read and write deletion vectors
CREATE TABLE orders (...) WITH ('format-version' = '3');
gnok's V3 deletion vectors are read back by reference engines. Apache Spark 4.0.1 / Iceberg 1.10 reads the full V3 DML lifecycle written by gnok (INSERT → UPDATE → UPDATE → DELETE, and a 3-clause MERGE), including the merged deletion vectors and row lineage. PyIceberg and Doris read the spec-shaped DVs. Trino ≤ 479 predates Iceberg's DV reader, so it reads gnok V3 tables but not their deletion vectors (use Spark, or a Trino build with DV support).
VARIANT Type
Semi-structured data stored natively in Iceberg's VARIANT format, with shredding support for columnar extraction.
-- Create a table with VARIANT column
CREATE TABLE events (
id BIGINT,
payload VARIANT
) WITH ('format-version' = '3');
-- Insert JSON data as VARIANT
INSERT INTO events VALUES (1, PARSE_JSON('{"user": "alice", "action": "click"}'));
-- Query VARIANT data
SELECT
id,
VARIANT_GET(payload, 'user', 'VARCHAR') AS user_name,
VARIANT_TYPE(payload) AS vtype
FROM events;
| Function | Description |
|---|---|
PARSE_JSON(string) | Parse JSON string into VARIANT |
TO_VARIANT(expr) | Convert a scalar value to VARIANT |
VARIANT_GET(v, path, type) | Extract typed value at dot-path |
VARIANT_EXTRACT_STRING(v, path) | Extract value as string |
VARIANT_TYPE(v) | Return the VARIANT's runtime type name |
VARIANT_AS_BOOL(v) | Cast to BOOLEAN |
VARIANT_AS_INT(v) | Cast to BIGINT |
VARIANT_AS_DOUBLE(v) | Cast to DOUBLE |
VARIANT_AS_STRING(v) | Cast to VARCHAR |
IS_VARIANT_NULL(v) | True if VARIANT value is null |
Arrow representation: LargeBinary.
gnok stores VARIANT values on disk as the Iceberg V3 Variant binary encoding — a Parquet group
of two binaries {metadata, value} (the unshredded variant layout) — produced at the Parquet write
boundary and decoded back on read. Within the engine the column is handled as LargeBinary (JSON
text) so all VARIANT functions work unchanged. Legacy columns previously written as JSON text are
read back transparently (no migration needed). A variant binary written by another engine
(Spark/Trino/PyIceberg) is decoded by gnok's size-bit-aware reader. The Parquet group is stamped
with the Parquet LogicalType::Variant annotation (via the arrow.parquet.variant Arrow extension),
so pure-Parquet readers resolve it as variant; Iceberg-aware readers resolve it from the Iceberg
schema regardless. Shredded typed_value is not produced (unshredded {metadata, value} layout).
GEOMETRY / GEOGRAPHY Types
Spatial types for geometric and geographic data, following PostGIS-style function conventions.
-- Create a table with spatial columns
CREATE TABLE stores (
id BIGINT,
name VARCHAR,
location GEOMETRY
) WITH ('format-version' = '3');
-- Insert a point
INSERT INTO stores VALUES (1, 'Downtown', ST_POINT(-73.9857, 40.7484));
-- Spatial query
SELECT name
FROM stores
WHERE ST_DWITHIN(location, ST_POINT(-73.98, 40.75), 1000);
| Function | Description |
|---|---|
ST_POINT(x, y) | Construct a point |
ST_DISTANCE(g1, g2) | Euclidean or haversine distance |
ST_WITHIN(g1, g2) | True if g1 is within g2 |
ST_CONTAINS(g1, g2) | True if g1 contains g2 |
ST_INTERSECTS(g1, g2) | True if geometries intersect |
ST_AREA(g) | Area of a polygon |
ST_LENGTH(g) | Length of a linestring |
ST_ASTEXT(g) | Convert to WKT |
ST_GEOMFROMTEXT(wkt) | Parse WKT to geometry |
ST_GEOGFROMTEXT(wkt) | Parse WKT to geography |
ST_X(g) / ST_Y(g) | Extract X/Y coordinates |
ST_SRID(g) | Get the spatial reference ID |
ST_TRANSFORM(g, srid) | Reproject to a different CRS |
ST_BUFFER(g, distance) | Buffer around geometry |
ST_DWITHIN(g1, g2, dist) | True if distance ≤ threshold |
ST_ENVELOPE(g) | Bounding box as polygon |
Supported CRS projections: WGS84 (EPSG:4326), Web Mercator (EPSG:3857), UTM zones (EPSG:326xx/327xx).
Spatial pruning is applied at the scan level using bounding box statistics.
Row Lineage
Row lineage provides stable, unique row identifiers that persist across compaction and schema evolution.
| Feature | Status | Notes |
|---|---|---|
_row_id metadata column (field ID 2147483540) | Supported | first_row_id + file_position |
_last_updated_sequence_number metadata column (field ID 2147483539) | Supported | Sequence number of the commit that last modified the row |
first_row_id on DataFile (manifest field 142) | Supported | Base row ID assigned during commit |
first_row_id on Snapshot | Supported | |
next_row_id on TableMetadata | Supported | Monotonically increasing row ID counter (catalog-derived on commit) |
added-rows on Snapshot | Supported | Required by readers when first-row-id is set |
Row lineage is maintained across writes, single-node and distributed: an UPDATE or a
MERGE-matched row preserves its original _row_id and bumps _last_updated_sequence_number;
INSERT and MERGE-inserted rows receive a fresh, unique _row_id; next_row_id advances
monotonically (and INSERT OVERWRITE commits without reusing row IDs). A pure DELETE does not advance
the counter.
-- Query row lineage columns (v3 tables only)
SELECT _row_id, _last_updated_sequence_number, *
FROM orders;
Default Values
Columns can specify default values that apply when existing data predates the column addition.
| Feature | Status | Notes |
|---|---|---|
initial_default | Supported | Default for existing rows when reading data written before the column was added |
write_default | Supported | Default for new rows when the column value is omitted on INSERT |
ALTER TABLE orders ADD COLUMN priority INT DEFAULT 0;
Metadata Columns
All Iceberg metadata columns are queryable (all format versions, but some are v3-specific):
| Column | Type | Version | Description |
|---|---|---|---|
_file | VARCHAR | v1+ | Source data file path |
_pos | BIGINT | v1+ | Row position within the file |
_spec_id | INT | v1+ | Partition spec ID |
_partition | STRUCT | v1+ | Partition values |
_file_size | BIGINT | v1+ | Data file size in bytes |
_row_id | BIGINT | v3 | Globally unique row identifier |
_last_updated_sequence_number | BIGINT | v3 | Commit sequence of last row modification |
-- Metadata columns are excluded from SELECT * but can be requested explicitly
SELECT _file, _pos, order_id, amount FROM orders LIMIT 10;
Feature Matrix
| Feature | v1 | v2 | v3 | Gnok Status |
|---|---|---|---|---|
| Core | ||||
| Manifest lists & data files | ✓ | ✓ | ✓ | Full |
| Schema evolution (add/drop/rename/widen) | ✓ | ✓ | ✓ | Full |
| Partition specs & transforms | ✓ | ✓ | ✓ | Full |
| Sort orders | ✓ | ✓ | ✓ | Full |
| Column statistics (min/max/null) | ✓ | ✓ | ✓ | Full |
| Table properties | ✓ | ✓ | ✓ | Full |
| Snapshot management | ✓ | ✓ | ✓ | Full |
| Time travel (snapshot ID, timestamp) | ✓ | ✓ | ✓ | Full |
| 4-level scan pruning | ✓ | ✓ | ✓ | Full |
| v2 Additions | ||||
| Position deletes | ✓ | ✓ | Full | |
| Equality deletes | ✓ | ✓ | Full | |
| Merge-on-Read (MOR) | ✓ | ✓ | Full | |
| Sequence numbers | ✓ | ✓ | Full | |
| Partition evolution | ✓ | ✓ | Full | |
| Snapshot references (branches & tags) | ✓ | ✓ | Full | |
| Statistics files (Puffin) | ✓ | ✓ | Full | |
| NaN value counts | ✓ | ✓ | Full | |
| DML (INSERT/UPDATE/DELETE/MERGE) | ✓ | ✓ | Full | |
| Encryption (AES-GCM) | ✓ | ✓ | Full | |
| Default values | ✓ | ✓ | Full | |
| v3 Additions | ||||
| Format version 3 create + commit | ✓ | Full (honored at CREATE; INSERT/UPDATE/DELETE commit) | ||
| Deletion vectors (Roaring bitmap / Puffin) | ✓ | Read + Write: Full — auto-gated on v3 (UPDATE/DELETE/MERGE), one DV per file, IEEE CRC-32; read by Spark 4.0.1 / Iceberg 1.10 | ||
| VARIANT type | ✓ | Full — stored as the Iceberg variant binary group {metadata,value}, stamped with the Parquet LogicalType::Variant annotation (shredded typed_value not produced) | ||
| GEOMETRY / GEOGRAPHY types | ✓ | Full (numeric args coerced; stored as WKB; CRS round-trips) | ||
Row lineage (_row_id, _last_updated_sequence_number) | ✓ | Read + Write: Full — preserved across UPDATE/MERGE, fresh on INSERT, next_row_id advances (single + distributed) | ||
Default values (initial-default / write-default) | ✓ | ✓ | Full (persisted to Iceberg schema JSON) | |
| Spatial type support | ✓ | Full |
Format Version Upgrade
Tables can be upgraded from v1 → v2 or v2 → v3 via table properties:
ALTER TABLE orders SET TBLPROPERTIES ('format-version' = '2');
ALTER TABLE orders SET TBLPROPERTIES ('format-version' = '3');
Format version upgrades are one-way — downgrading is not supported. Ensure all consumers of the table support the target version before upgrading.
Change Data Capture
A dedicated changelog SQL surface (a READ CHANGES / table_changes()-style statement returning
per-row _change_type) is not yet implemented. The internal incremental-scan engine exists but
is currently wired only to materialized-view / dynamic-table refresh (append-only snapshot ranges).
For now, capture changes via V3 row lineage — query rows whose
_last_updated_sequence_number exceeds a stored watermark.