Skip to main content

Iceberg Compatibility

Gnok implements a native Apache Iceberg runtime supporting format versions 1, 2, and 3. This page details which features are available at each format version level.

Format Version Overview​

VersionFocusDefault
v1Foundational table format — immutable data files, schema evolution, partition specs, sort ordersNo
v2Row-level mutations — position deletes, equality deletes, sequence numbers, snapshot referencesYes
v3Extended types & lineage — deletion vectors, VARIANT type, spatial types, row lineage, default valuesNo

Gnok defaults to format version 2 when creating new tables. You can specify the version explicitly, and the requested version is honored at create time (the metadata is written at that format version, not just recorded as a property):

CREATE TABLE events (
id BIGINT,
payload VARCHAR
)
WITH ('format-version' = '3');

A table that declares a V3-only column type (VARIANT, GEOMETRY, GEOGRAPHY) is automatically created at format version 3 even if a lower version is requested, since those types are invalid at v1/v2 and would be rejected by other engines.


v1 Features​

All v1 features are fully supported and form the foundation of every Iceberg table.

Table Metadata​

FeatureStatusNotes
Manifest listsSupportedMetadata tree: manifest list → manifests → data files
Column statistics (min/max/null counts)SupportedPer-field in manifest entries; used for file-level pruning
Partition specsSupportedIdentity, year, month, day, hour, bucket(N), truncate(W)
Sort ordersSupportedPer-table sort specification with field, transform, direction, null ordering
Schema evolutionSupportedAdd, drop, rename, widen columns with field ID preservation
Table propertiesSupportedKey-value metadata (compression, file size, bloom filter config)
Snapshot managementSupportedAppend and overwrite operations; parent snapshot tracking
Multiple schemasSupportedSchema history with unique schema_id per revision

Scan & Pruning​

Gnok implements 4-level pruning that applies to all format versions:

  1. Manifest list pruning — skip manifests by partition summary bounds
  2. File pruning — skip data files by per-column min/max statistics
  3. Row group pruning — skip Parquet row groups by page-level statistics
  4. Dynamic filter pruning — runtime bloom/min-max filters from join build sides

Catalogs​

CatalogStatus
REST Catalog (Iceberg REST spec)Supported
AWS Glue Data CatalogSupported
NessieSupported (external)
Hive MetastoreSupported (external)

All catalog types support JWT token forwarding for multi-tenant isolation.


v2 Features​

Format version 2 adds row-level delete capabilities and sequence-based ordering.

Row-Level Deletes​

FeatureStatusNotes
Position deletesSupported(file_path, pos) pairs in Parquet delete files
Equality deletesSupportedColumn-value predicates stored as delete files
Merge-on-Read (MOR)SupportedDefault strategy for UPDATE and DELETE
Copy-on-Write (COW)SupportedFull file rewrite for small, targeted changes

Position deletes use an adaptive index: Vec<i64> with binary search for fewer than 50,000 positions (service-managed threshold), and RoaringBitmap for larger sets.

Delete Strategies​

StrategyWhen Selected
BroadcastSmall delete sets kept in memory
SortMergeHigh delete-to-data ratio
BloomFilterLarge, sparse delete patterns

Sequence Numbers​

Sequence numbers provide a total ordering of commits:

  • sequence_number on snapshots — monotonically increasing per commit
  • min_sequence_number on manifest files — enables efficient delete file matching
  • data_sequence_number on data files — identifies which commit produced each file

Snapshot References (Branches & Tags)​

-- View snapshots
SHOW SNAPSHOTS FROM orders;

-- Time travel by snapshot id (these forms are all accepted)
SELECT * FROM orders FOR SYSTEM_VERSION AS OF 1234567890;
SELECT * FROM orders FOR VERSION AS OF 1234567890;
SELECT * FROM orders AT(SNAPSHOT => 1234567890);

-- Time travel by timestamp
SELECT * FROM orders FOR SYSTEM_TIME AS OF TIMESTAMP '2024-01-15 10:00:00';
FeatureStatus
Branch referencesSupported
Tag referencesSupported
Branch retention (min_snapshots_to_keep, max_snapshot_age_ms)Supported
Reference expiration (max_ref_age_ms)Supported

DML Operations​

All DML operations produce v2-compliant snapshots with position delete files:

OperationStrategyNotes
INSERTAppend data filesPartition-aware routing, idempotent file naming
INSERT OVERWRITEOverwrite affected partitionsMetadata-level partition replacement
UPDATEMOR (position deletes + new data files)Atomic snapshot with both delete and data files
DELETEMOR (position delete files)Partition pruning applied to scan
MERGEMOR (hash table probe + position deletes)Supports MATCHED, NOT MATCHED, NOT MATCHED BY SOURCE
TRUNCATEMetadata-only empty snapshotNo scan required
COPY INTOBulk append from staged filesSIMD-accelerated CSV/Parquet/JSON parsing

Statistics (Puffin Format)​

FeatureStatus
Statistics filesSupported
Theta sketch blobs (NDV estimation)Supported
Blob metadata (snapshot ID, sequence number, field IDs)Supported
Partition statistics files (Iceberg partition-statistics, 12-field schema; per-partition row/file/size + bounds, registered via SetPartitionStatistics)Supported

NaN Value Counts​

nan_value_counts per column in manifest entries — used for accurate statistics on floating-point columns. Supported in v2+ tables.

Partition Evolution​

Tables can change their partition spec over time without rewriting existing data:

-- Original partitioning
CREATE TABLE events (...) PARTITIONED BY (day(event_time));

-- Evolve to hourly partitioning (new data uses new spec; old data retains old spec)
ALTER TABLE events SET PARTITION SPEC (hour(event_time));
FeatureStatus
Multiple partition specs in historySupported
Mixed-spec scans (old + new data)Supported
Partition field source trackingSupported

Schema Evolution​

OperationStatus
Add columnSupported
Drop columnSupported
Rename columnSupported
Widen type (e.g., INT → BIGINT)Supported
Reorder columnsSupported
Set/drop NOT NULLSupported
Set column commentSupported

Field IDs are preserved across schema changes — Gnok never reuses a retired field ID.

Encryption​

FeatureStatus
Key hierarchy (Master → KEK → File Key)Supported
Parquet AES-GCMSupported
Per-file key trackingSupported
Key metadata in table propertiesSupported

v3 Features​

Format version 3 introduces extended type support, deletion vectors, and row lineage.

Deletion Vectors​

Deletion vectors replace position delete files with a more compact, Puffin-based representation using 64-bit Roaring bitmaps.

FeatureStatusNotes
DV binary format (magic 0xD1D33964 + Roaring bitmap + IEEE CRC-32)SupportedSpec-compliant Puffin framing (4-byte footer flags, blob length = magic+bitmap); the checksum is IEEE CRC-32 (== java.util.zip.CRC32 / zlib), not CRC-32C
Puffin blob storage with LZ4/Zstd compressionSupported
DV manifest classification (external content=1 + file_format=puffin)SupportedManifest reader reclassifies spec-shaped DVs (Spark/Trino/Doris/PyIceberg) to deletion vectors
DV-to-data-file matching via referenced_data_fileSupported
Scan-path execution (read Puffin blob + merge with position deletes)SupportedLive-validated: scan preloads the DV and applies it during read
DV write path (auto-gated on v3)SupportedUPDATE / DELETE / MERGE on a format-version-3 table automatically write Puffin DVs (content=1+file_format=puffin+content_offset) — gated on format_version >= 3, managed automatically by the service. Sequential edits to the same data file union into a single DV and retire the superseded manifest entry (one DV per data file), with snapshot added-rows recorded
-- format-version 3 tables both read and write deletion vectors
CREATE TABLE orders (...) WITH ('format-version' = '3');
Cross-engine interoperability

gnok's V3 deletion vectors are read back by reference engines. Apache Spark 4.0.1 / Iceberg 1.10 reads the full V3 DML lifecycle written by gnok (INSERT → UPDATE → UPDATE → DELETE, and a 3-clause MERGE), including the merged deletion vectors and row lineage. PyIceberg and Doris read the spec-shaped DVs. Trino ≤ 479 predates Iceberg's DV reader, so it reads gnok V3 tables but not their deletion vectors (use Spark, or a Trino build with DV support).

VARIANT Type​

Semi-structured data stored natively in Iceberg's VARIANT format, with shredding support for columnar extraction.

-- Create a table with VARIANT column
CREATE TABLE events (
id BIGINT,
payload VARIANT
) WITH ('format-version' = '3');

-- Insert JSON data as VARIANT
INSERT INTO events VALUES (1, PARSE_JSON('{"user": "alice", "action": "click"}'));

-- Query VARIANT data
SELECT
id,
VARIANT_GET(payload, 'user', 'VARCHAR') AS user_name,
VARIANT_TYPE(payload) AS vtype
FROM events;
FunctionDescription
PARSE_JSON(string)Parse JSON string into VARIANT
TO_VARIANT(expr)Convert a scalar value to VARIANT
VARIANT_GET(v, path, type)Extract typed value at dot-path
VARIANT_EXTRACT_STRING(v, path)Extract value as string
VARIANT_TYPE(v)Return the VARIANT's runtime type name
VARIANT_AS_BOOL(v)Cast to BOOLEAN
VARIANT_AS_INT(v)Cast to BIGINT
VARIANT_AS_DOUBLE(v)Cast to DOUBLE
VARIANT_AS_STRING(v)Cast to VARCHAR
IS_VARIANT_NULL(v)True if VARIANT value is null

Arrow representation: LargeBinary.

note

gnok stores VARIANT values on disk as the Iceberg V3 Variant binary encoding — a Parquet group of two binaries {metadata, value} (the unshredded variant layout) — produced at the Parquet write boundary and decoded back on read. Within the engine the column is handled as LargeBinary (JSON text) so all VARIANT functions work unchanged. Legacy columns previously written as JSON text are read back transparently (no migration needed). A variant binary written by another engine (Spark/Trino/PyIceberg) is decoded by gnok's size-bit-aware reader. The Parquet group is stamped with the Parquet LogicalType::Variant annotation (via the arrow.parquet.variant Arrow extension), so pure-Parquet readers resolve it as variant; Iceberg-aware readers resolve it from the Iceberg schema regardless. Shredded typed_value is not produced (unshredded {metadata, value} layout).

GEOMETRY / GEOGRAPHY Types​

Spatial types for geometric and geographic data, following PostGIS-style function conventions.

-- Create a table with spatial columns
CREATE TABLE stores (
id BIGINT,
name VARCHAR,
location GEOMETRY
) WITH ('format-version' = '3');

-- Insert a point
INSERT INTO stores VALUES (1, 'Downtown', ST_POINT(-73.9857, 40.7484));

-- Spatial query
SELECT name
FROM stores
WHERE ST_DWITHIN(location, ST_POINT(-73.98, 40.75), 1000);
FunctionDescription
ST_POINT(x, y)Construct a point
ST_DISTANCE(g1, g2)Euclidean or haversine distance
ST_WITHIN(g1, g2)True if g1 is within g2
ST_CONTAINS(g1, g2)True if g1 contains g2
ST_INTERSECTS(g1, g2)True if geometries intersect
ST_AREA(g)Area of a polygon
ST_LENGTH(g)Length of a linestring
ST_ASTEXT(g)Convert to WKT
ST_GEOMFROMTEXT(wkt)Parse WKT to geometry
ST_GEOGFROMTEXT(wkt)Parse WKT to geography
ST_X(g) / ST_Y(g)Extract X/Y coordinates
ST_SRID(g)Get the spatial reference ID
ST_TRANSFORM(g, srid)Reproject to a different CRS
ST_BUFFER(g, distance)Buffer around geometry
ST_DWITHIN(g1, g2, dist)True if distance ≤ threshold
ST_ENVELOPE(g)Bounding box as polygon

Supported CRS projections: WGS84 (EPSG:4326), Web Mercator (EPSG:3857), UTM zones (EPSG:326xx/327xx).

Spatial pruning is applied at the scan level using bounding box statistics.

Row Lineage​

Row lineage provides stable, unique row identifiers that persist across compaction and schema evolution.

FeatureStatusNotes
_row_id metadata column (field ID 2147483540)Supportedfirst_row_id + file_position
_last_updated_sequence_number metadata column (field ID 2147483539)SupportedSequence number of the commit that last modified the row
first_row_id on DataFile (manifest field 142)SupportedBase row ID assigned during commit
first_row_id on SnapshotSupported
next_row_id on TableMetadataSupportedMonotonically increasing row ID counter (catalog-derived on commit)
added-rows on SnapshotSupportedRequired by readers when first-row-id is set

Row lineage is maintained across writes, single-node and distributed: an UPDATE or a MERGE-matched row preserves its original _row_id and bumps _last_updated_sequence_number; INSERT and MERGE-inserted rows receive a fresh, unique _row_id; next_row_id advances monotonically (and INSERT OVERWRITE commits without reusing row IDs). A pure DELETE does not advance the counter.

-- Query row lineage columns (v3 tables only)
SELECT _row_id, _last_updated_sequence_number, *
FROM orders;

Default Values​

Columns can specify default values that apply when existing data predates the column addition.

FeatureStatusNotes
initial_defaultSupportedDefault for existing rows when reading data written before the column was added
write_defaultSupportedDefault for new rows when the column value is omitted on INSERT
ALTER TABLE orders ADD COLUMN priority INT DEFAULT 0;

Metadata Columns​

All Iceberg metadata columns are queryable (all format versions, but some are v3-specific):

ColumnTypeVersionDescription
_fileVARCHARv1+Source data file path
_posBIGINTv1+Row position within the file
_spec_idINTv1+Partition spec ID
_partitionSTRUCTv1+Partition values
_file_sizeBIGINTv1+Data file size in bytes
_row_idBIGINTv3Globally unique row identifier
_last_updated_sequence_numberBIGINTv3Commit sequence of last row modification
-- Metadata columns are excluded from SELECT * but can be requested explicitly
SELECT _file, _pos, order_id, amount FROM orders LIMIT 10;

Feature Matrix​

Featurev1v2v3Gnok Status
Core
Manifest lists & data files✓✓✓Full
Schema evolution (add/drop/rename/widen)✓✓✓Full
Partition specs & transforms✓✓✓Full
Sort orders✓✓✓Full
Column statistics (min/max/null)✓✓✓Full
Table properties✓✓✓Full
Snapshot management✓✓✓Full
Time travel (snapshot ID, timestamp)✓✓✓Full
4-level scan pruning✓✓✓Full
v2 Additions
Position deletes✓✓Full
Equality deletes✓✓Full
Merge-on-Read (MOR)✓✓Full
Sequence numbers✓✓Full
Partition evolution✓✓Full
Snapshot references (branches & tags)✓✓Full
Statistics files (Puffin)✓✓Full
NaN value counts✓✓Full
DML (INSERT/UPDATE/DELETE/MERGE)✓✓Full
Encryption (AES-GCM)✓✓Full
Default values✓✓Full
v3 Additions
Format version 3 create + commit✓Full (honored at CREATE; INSERT/UPDATE/DELETE commit)
Deletion vectors (Roaring bitmap / Puffin)✓Read + Write: Full — auto-gated on v3 (UPDATE/DELETE/MERGE), one DV per file, IEEE CRC-32; read by Spark 4.0.1 / Iceberg 1.10
VARIANT type✓Full — stored as the Iceberg variant binary group {metadata,value}, stamped with the Parquet LogicalType::Variant annotation (shredded typed_value not produced)
GEOMETRY / GEOGRAPHY types✓Full (numeric args coerced; stored as WKB; CRS round-trips)
Row lineage (_row_id, _last_updated_sequence_number)✓Read + Write: Full — preserved across UPDATE/MERGE, fresh on INSERT, next_row_id advances (single + distributed)
Default values (initial-default / write-default)✓✓Full (persisted to Iceberg schema JSON)
Spatial type support✓Full

Format Version Upgrade​

Tables can be upgraded from v1 → v2 or v2 → v3 via table properties:

ALTER TABLE orders SET TBLPROPERTIES ('format-version' = '2');
ALTER TABLE orders SET TBLPROPERTIES ('format-version' = '3');
caution

Format version upgrades are one-way — downgrading is not supported. Ensure all consumers of the table support the target version before upgrading.

Change Data Capture​

note

A dedicated changelog SQL surface (a READ CHANGES / table_changes()-style statement returning per-row _change_type) is not yet implemented. The internal incremental-scan engine exists but is currently wired only to materialized-view / dynamic-table refresh (append-only snapshot ranges). For now, capture changes via V3 row lineage — query rows whose _last_updated_sequence_number exceeds a stored watermark.