Python Notebooks
Gnok Studio ships an in-product Python notebook: cells run in a sandboxed
CPython kernel (numpy, pandas, scikit-learn pre-installed) with a pre-wired
gnok helper. You write Python against governed Gnok data, train a model, and
register it so it's callable from SQL via ML_PREDICT — the train→serve
loop, without leaving the product.
gnok vs pygnokgnok(this page) is the helper pre-installed inside the notebook kernel. You don'tpip installit and you don't pass connection details — it auto-connects as the signed-in user (so RBAC, row masking, and vended credentials all apply) and adds anmlnamespace for the model round-trip.pygnokis the standalone Python SDK youpip installto connect to Gnok from your own scripts.gnokis a thin wrapper around it;gnokhasgnok.ml,pygnokdoes not.
Quick start
import gnok
from gnok import ml
# Query — returns a DB-API cursor; .fetch_df() gives a pandas DataFrame.
df = gnok.sql("SELECT region, SUM(amount) AS total FROM sales.orders GROUP BY 1").fetch_df()
# Or read a whole table straight into pandas.
orders = gnok.read_table("sales.orders", limit=10_000)
No connect(), no token — the kernel is wired to the engine as you.
Data API
gnok.sql(query, params=None)
Run a query on the shared connection; returns a DB-API cursor.
cur = gnok.sql("SELECT * FROM sales.orders WHERE region = ?", ["EMEA"])
rows = cur.fetchall() # list of tuples
df = gnok.sql("SELECT ...").fetch_df() # pandas DataFrame
gnok.read_sql(query, params=None)
The pandas-idiomatic shortcut — run a query and get a DataFrame directly.
df = gnok.read_sql("SELECT region, SUM(amount) FROM sales.orders GROUP BY 1")
gnok.read_table(name, limit=None)
Read a table into a pandas DataFrame (optionally LIMIT-ed).
df = gnok.read_table("sales.orders", limit=1000)
gnok.write(df, table, mode="append")
Write a pandas DataFrame to a table; returns the row count.
mode | Behavior |
|---|---|
"append" (default) | INSERT into an existing table. |
"create" | CREATE TABLE (schema inferred from dtypes), then load. |
"replace" | CREATE OR REPLACE TABLE, then load. |
For create/replace, use a fully-qualified catalog.schema.table. For large
volumes prefer CREATE TABLE AS / COPY INTO.
gnok.write(scored_df, "sales.scored") # append
gnok.write(features_df, "sales.features", mode="replace") # (re)create + load
Discovery
Explore the catalog without leaving the notebook — each returns a DataFrame.
gnok.catalogs() # SHOW CATALOGS
gnok.schemas("sales") # SHOW SCHEMAS IN sales
gnok.tables("sales", "public") # SHOW TABLES IN sales.public
gnok.describe("sales.public.orders") # column names + types
| Method | Returns |
|---|---|
gnok.catalogs() | All catalogs. |
gnok.schemas(catalog=None) | Schemas (scope with catalog). |
gnok.tables(catalog=None, schema=None) | Tables (scope with catalog/schema). |
gnok.describe(table) | column_name, data_type, is_nullable for a table. |
gnok.connect(**overrides) / gnok.reset()
connect() opens a fresh connection with the notebook's own credentials
(override catalog or schema if needed). reset() drops the cached
shared connection so the next call reconnects. The helper does this for you
when Studio refreshes your sign-in; you rarely need it.
ML round-trip: gnok.ml
The helper exports and registers an artifact; it does not establish that every sklearn estimator is supported by the serving adapter. Check the ONNX input/output contract and score representative rows after registration.
Train in the notebook, serve in SQL. gnok.ml exports a fitted scikit-learn
estimator to ONNX, uploads it to an internal stage, and registers it as a
SQL-callable model.
gnok.ml.register_model()
The one call that does it all — export → upload → CREATE MODEL.
gnok.ml.register_model(
estimator, # a fitted scikit-learn estimator
name: str, # fully-qualified: "catalog.schema.model"
x_sample, # sample inputs; column count -> SQL signature
source: str | None = None, # skip the upload; register an existing artifact
returns: str = "FLOAT", # SQL output type the model is called by
stage: str = "models", # internal stage to upload the ONNX to
) -> str
| Parameter | Type | Default | Description |
|---|---|---|---|
estimator | sklearn estimator | — | A fitted estimator whose ONNX export matches the numeric scalar adapter. |
name | str | — | Model name; fully-qualify as catalog.schema.model. |
x_sample | array / DataFrame | — | Sample inputs — its column count sets the (FLOAT, …) signature. |
source | str | None | If set, skip export/upload and register this stage ref / URI instead of the trained estimator. |
returns | str | "FLOAT" | SQL output type the model is called by ("BIGINT" for an integer label, etc.). |
stage | str | "models" | Internal stage the ONNX is uploaded to. |
A full example:
import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.ensemble import GradientBoostingRegressor
from gnok import ml
X = orders[["feature_a", "feature_b", "feature_c", "feature_d"]].to_numpy("float32")
y = orders["amount"].to_numpy("float32")
pipe = Pipeline([
("scale", StandardScaler()),
("gbr", GradientBoostingRegressor(n_estimators=100, max_depth=3)),
]).fit(X, y)
ml.register_model(pipe, "sales.imported.amount_v1", X)
This example assumes orders contains the four named feature columns and a continuous amount target. Fit preprocessing on training rows and evaluate on separate data before registration. The input signature ((FLOAT, FLOAT, FLOAT, FLOAT)) is inferred from
x_sample — one FLOAT per feature. Then, from anywhere that speaks SQL:
SELECT ML_PREDICT(
'sales.imported.amount_v1',
CAST(feature_a AS REAL), CAST(feature_b AS REAL), CAST(feature_c AS REAL), CAST(feature_d AS REAL)
) AS predicted_amount
FROM sales.candidates;
See ONNX Inference for how registered models execute.
- Use a fully-qualified model name (
catalog.schema.model). The notebook kernel has no default catalog/schema, so an unqualified name lands somewhereML_PREDICTcan't find. returnsmust match the model's output. A regressor outputs a float →returns="FLOAT"(the default). The ONNX adapter currently requires a scalar Float32 first output. A classifier export with an integer label or probability map needs an adapted artifact; changingreturnsalone does not fix it.- Use explicit numeric casts when needed in
ML_PREDICT(e.g.CAST(3.0 AS REAL)) — bare decimal literals are typed asDECIMAL, which the ONNX input rejects.
gnok.ml.predict(model, df, features=None)
Score a pandas DataFrame with a registered model and get predictions back as a
pandas Series (aligned to df.index) — the Python-side complement to
register_model (it builds the ML_PREDICT query, casts, and preserves row
order for you). features defaults to all of df's columns, in order.
preds = gnok.ml.predict("sales.imported.amount_v1", candidates[FEATURES])
candidates["predicted_amount"] = preds
Best for modest frames; for million-row batches call ML_PREDICT in SQL (e.g.
CREATE TABLE AS SELECT ML_PREDICT(...)).
gnok.ml.models()
List registered models as a DataFrame (SHOW MODELS) — qualified name, input
types, return type.
gnok.ml.models()
gnok.ml.log_run(table, model, metrics, params=None) / gnok.ml.runs(table, model=None)
Lightweight experiment tracking that lives in your own Gnok table (no
external MLOps service). log_run creates the table on first use and appends a
run (metrics + hyperparameters, stored as JSON); runs reads them back as a
DataFrame (newest first, metrics expanded into metric.* columns).
ml.log_run("ml.experiments.churn", "churn_v1",
{"auc": 0.91, "rmse": 0.12}, {"n_estimators": 200})
ml.runs("ml.experiments.churn", model="churn_v1") # -> DataFrame, metric.auc, metric.rmse, …
gnok.ml.to_onnx() / gnok.ml.register_onnx()
The lower-level building blocks, if you want to manage the artifact yourself:
gnok.ml.to_onnx(estimator, x_sample) -> bytes
gnok.ml.register_onnx(
name: str, # fully-qualified model name
source: str, # stage ref / URI the engine can read
inputs: Sequence[str], # input column types, e.g. ["FLOAT", "FLOAT"]
returns: str = "FLOAT", # SQL output type
framework: str = "onnx", # artifact format
) -> str
onnx_bytes = ml.to_onnx(pipe, X) # -> bytes
# ...upload onnx_bytes somewhere the engine can read, then:
ml.register_onnx("sales.imported.amount_v1",
"s3://my-bucket/models/amount_v1.onnx",
inputs=["FLOAT", "FLOAT", "FLOAT", "FLOAT"])
source is a stage ref (@stage/models/x.onnx) or an object-storage URI
(s3://…) that your account is authorized to read. to_onnx serializes a fitted
estimator to ONNX bytes via skl2onnx; register_onnx runs the CREATE OR REPLACE MODEL with the given signature.
Generative AI & embeddings
Gnok's AI and embedding providers are callable from Python too. gnok.ai
calls are governed like SQL AI_* calls: an organization administrator must
first enable AI Governance. gnok.embed
uses your organization's monthly embedding allowance.
gnok.ai.generate(prompt)
Generate text via SQL AI_GENERATE, an alias of AI_COMPLETE — summaries, classification, extraction, and the like. The rollout policy must allow ai_complete.
from gnok import ai
label = ai.generate("Classify the sentiment of 'love it' in one word.")
gnok.embed(text)
Embed text into a vector (engine EMBED) for similarity search / RAG. Returns
a list of floats.
vec = gnok.embed("wireless noise-cancelling headphones")
Gnok AI in the notebook
The notebook has a ✨ Gnok AI panel that writes and fixes cells for you,
grounded in the gnok helper. Ask it to "read 1,000 rows of sales.orders and
plot amount by region" or "train a model to predict amount" and it returns a
ready-to-run cell. When a cell errors, ✨ Fix sends the traceback to the
model and proposes a corrected cell.
Persistence, history & lifecycle
Notebooks (cells + outputs) are saved per user. Every save also snapshots a version — the 🕘 History panel lists prior versions and Restore loads one back into the editor (then Save to make it current).
Kernels are ephemeral — your saved cells are the source of truth — and are reaped when idle, so a long-lived notebook never holds resources it isn't using. Re-running a cell after the kernel is reaped transparently starts a fresh one.
Compute availability: notebook packages and accelerator access are determined by the hosted environment. Confirm the capabilities enabled for your account before depending on them.
Headless runs & schedules
Every notebook has two run paths, side by side:
- ▶ Run (per cell) — interactive, against your open kernel session. Fast, state is shared across cells. The default.
- ▶▶ Headless (per cell) and ▶ Run all (toolbar) — the headless executor spawns a transient kernel via the control plane, runs the cell(s), persists the outputs back, and tears the kernel down. Doesn't disturb your open kernel; useful for re-running against a clean state or for the unattended scheduled path.
The ⏰ Schedule panel lets you set a notebook to run on a fixed cadence (seconds, minimum 60). The scheduler executes a snapshot of the cells captured when you save the schedule — re-save the schedule to refresh the snapshot. Scheduled runs land in the notebook's 📜 Runs panel (status, trigger, duration; click a row to expand for full error + the persisted cell outputs); they never overwrite the live notebook, so a notebook you have open is never clobbered by a 3 a.m. run.
The cross-notebook view at ⏰ All schedules (sidebar link →
/notebooks/schedules) lists every schedule you own — interval, enabled/
paused, next-run timestamp — with Pause/Resume (flips enabled only,
preserves the snapshot) and Delete.
Architecture (where the compute actually runs)
Gnok manages notebook kernels, the scheduler, and execution infrastructure. Scheduled runs require a valid execution identity and access to their inputs. Use the Runs panel to inspect failures and contact Gnok support for runtime or scheduling-service problems. Customers do not need internal service tokens or Kubernetes configuration.