From JSON to Iceberg: Bulk-Loading a Document Dataset into Gnok
· 7 min read
We recently needed to land a large pile of JSON documents — roughly 50 GB across ~70 datasets, including a couple of 15+ GB monsters — into Gnok so it could be queried with SQL alongside everything else. The documents were schemaless and deeply nested; Gnok tables are columnar Iceberg. This post walks through the pipeline we landed on, and the handful of gotchas that shaped it.