Skip to main content

Python Notebooks

Gnok Studio ships an in-product Python notebook: cells run in a sandboxed CPython kernel (numpy, pandas, scikit-learn pre-installed) with a pre-wired gnok helper. You write Python against governed Gnok data, train a model, and register it so it's callable from SQL via ML_PREDICT — the train→serve loop, without leaving the product.

gnok vs pygnok
  • gnok (this page) is the helper pre-installed inside the notebook kernel. You don't pip install it and you don't pass connection details — it auto-connects as the signed-in user (so RBAC, row masking, and vended credentials all apply) and adds an ml namespace for the model round-trip.
  • pygnok is the standalone Python SDK you pip install to connect to Gnok from your own scripts. gnok is a thin wrapper around it; gnok has gnok.ml, pygnok does not.

Quick start​

import gnok
from gnok import ml

# Query — returns a DB-API cursor; .fetch_df() gives a pandas DataFrame.
df = gnok.sql("SELECT region, SUM(amount) AS total FROM sales.orders GROUP BY 1").fetch_df()

# Or read a whole table straight into pandas.
orders = gnok.read_table("sales.orders", limit=10_000)

No connect(), no token — the kernel is wired to the engine as you.

Data API​

gnok.sql(query, params=None)​

Run a query on the shared connection; returns a DB-API cursor.

cur = gnok.sql("SELECT * FROM sales.orders WHERE region = ?", ["EMEA"])
rows = cur.fetchall() # list of tuples
df = gnok.sql("SELECT ...").fetch_df() # pandas DataFrame

gnok.read_sql(query, params=None)​

The pandas-idiomatic shortcut — run a query and get a DataFrame directly.

df = gnok.read_sql("SELECT region, SUM(amount) FROM sales.orders GROUP BY 1")

gnok.read_table(name, limit=None)​

Read a table into a pandas DataFrame (optionally LIMIT-ed).

df = gnok.read_table("sales.orders", limit=1000)

gnok.write(df, table, mode="append")​

Write a pandas DataFrame to a table; returns the row count.

modeBehavior
"append" (default)INSERT into an existing table.
"create"CREATE TABLE (schema inferred from dtypes), then load.
"replace"CREATE OR REPLACE TABLE, then load.

For create/replace, use a fully-qualified catalog.schema.table. For large volumes prefer CREATE TABLE AS / COPY INTO.

gnok.write(scored_df, "sales.scored")                      # append
gnok.write(features_df, "sales.features", mode="replace") # (re)create + load

Discovery​

Explore the catalog without leaving the notebook — each returns a DataFrame.

gnok.catalogs()                          # SHOW CATALOGS
gnok.schemas("sales") # SHOW SCHEMAS IN sales
gnok.tables("sales", "public") # SHOW TABLES IN sales.public
gnok.describe("sales.public.orders") # column names + types
MethodReturns
gnok.catalogs()All catalogs.
gnok.schemas(catalog=None)Schemas (scope with catalog).
gnok.tables(catalog=None, schema=None)Tables (scope with catalog/schema).
gnok.describe(table)column_name, data_type, is_nullable for a table.

gnok.connect(**overrides) / gnok.reset()​

connect() opens a fresh connection with the notebook's own credentials (override catalog or schema if needed). reset() drops the cached shared connection so the next call reconnects. The helper does this for you when Studio refreshes your sign-in; you rarely need it.

ML round-trip: gnok.ml​

The helper exports and registers an artifact; it does not establish that every sklearn estimator is supported by the serving adapter. Check the ONNX input/output contract and score representative rows after registration.

Train in the notebook, serve in SQL. gnok.ml exports a fitted scikit-learn estimator to ONNX, uploads it to an internal stage, and registers it as a SQL-callable model.

gnok.ml.register_model()​

The one call that does it all — export → upload → CREATE MODEL.

gnok.ml.register_model(
estimator, # a fitted scikit-learn estimator
name: str, # fully-qualified: "catalog.schema.model"
x_sample, # sample inputs; column count -> SQL signature
source: str | None = None, # skip the upload; register an existing artifact
returns: str = "FLOAT", # SQL output type the model is called by
stage: str = "models", # internal stage to upload the ONNX to
) -> str
ParameterTypeDefaultDescription
estimatorsklearn estimator—A fitted estimator whose ONNX export matches the numeric scalar adapter.
namestr—Model name; fully-qualify as catalog.schema.model.
x_samplearray / DataFrame—Sample inputs — its column count sets the (FLOAT, …) signature.
sourcestrNoneIf set, skip export/upload and register this stage ref / URI instead of the trained estimator.
returnsstr"FLOAT"SQL output type the model is called by ("BIGINT" for an integer label, etc.).
stagestr"models"Internal stage the ONNX is uploaded to.

A full example:

import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.ensemble import GradientBoostingRegressor
from gnok import ml

X = orders[["feature_a", "feature_b", "feature_c", "feature_d"]].to_numpy("float32")
y = orders["amount"].to_numpy("float32")

pipe = Pipeline([
("scale", StandardScaler()),
("gbr", GradientBoostingRegressor(n_estimators=100, max_depth=3)),
]).fit(X, y)

ml.register_model(pipe, "sales.imported.amount_v1", X)

This example assumes orders contains the four named feature columns and a continuous amount target. Fit preprocessing on training rows and evaluate on separate data before registration. The input signature ((FLOAT, FLOAT, FLOAT, FLOAT)) is inferred from x_sample — one FLOAT per feature. Then, from anywhere that speaks SQL:

SELECT ML_PREDICT(
'sales.imported.amount_v1',
CAST(feature_a AS REAL), CAST(feature_b AS REAL), CAST(feature_c AS REAL), CAST(feature_d AS REAL)
) AS predicted_amount
FROM sales.candidates;

See ONNX Inference for how registered models execute.

Three things to get right
  1. Use a fully-qualified model name (catalog.schema.model). The notebook kernel has no default catalog/schema, so an unqualified name lands somewhere ML_PREDICT can't find.
  2. returns must match the model's output. A regressor outputs a float → returns="FLOAT" (the default). The ONNX adapter currently requires a scalar Float32 first output. A classifier export with an integer label or probability map needs an adapted artifact; changing returns alone does not fix it.
  3. Use explicit numeric casts when needed in ML_PREDICT (e.g. CAST(3.0 AS REAL)) — bare decimal literals are typed as DECIMAL, which the ONNX input rejects.

gnok.ml.predict(model, df, features=None)​

Score a pandas DataFrame with a registered model and get predictions back as a pandas Series (aligned to df.index) — the Python-side complement to register_model (it builds the ML_PREDICT query, casts, and preserves row order for you). features defaults to all of df's columns, in order.

preds = gnok.ml.predict("sales.imported.amount_v1", candidates[FEATURES])
candidates["predicted_amount"] = preds

Best for modest frames; for million-row batches call ML_PREDICT in SQL (e.g. CREATE TABLE AS SELECT ML_PREDICT(...)).

gnok.ml.models()​

List registered models as a DataFrame (SHOW MODELS) — qualified name, input types, return type.

gnok.ml.models()

gnok.ml.log_run(table, model, metrics, params=None) / gnok.ml.runs(table, model=None)​

Lightweight experiment tracking that lives in your own Gnok table (no external MLOps service). log_run creates the table on first use and appends a run (metrics + hyperparameters, stored as JSON); runs reads them back as a DataFrame (newest first, metrics expanded into metric.* columns).

ml.log_run("ml.experiments.churn", "churn_v1",
{"auc": 0.91, "rmse": 0.12}, {"n_estimators": 200})

ml.runs("ml.experiments.churn", model="churn_v1") # -> DataFrame, metric.auc, metric.rmse, …

gnok.ml.to_onnx() / gnok.ml.register_onnx()​

The lower-level building blocks, if you want to manage the artifact yourself:

gnok.ml.to_onnx(estimator, x_sample) -> bytes

gnok.ml.register_onnx(
name: str, # fully-qualified model name
source: str, # stage ref / URI the engine can read
inputs: Sequence[str], # input column types, e.g. ["FLOAT", "FLOAT"]
returns: str = "FLOAT", # SQL output type
framework: str = "onnx", # artifact format
) -> str
onnx_bytes = ml.to_onnx(pipe, X)                       # -> bytes
# ...upload onnx_bytes somewhere the engine can read, then:
ml.register_onnx("sales.imported.amount_v1",
"s3://my-bucket/models/amount_v1.onnx",
inputs=["FLOAT", "FLOAT", "FLOAT", "FLOAT"])

source is a stage ref (@stage/models/x.onnx) or an object-storage URI (s3://…) that your account is authorized to read. to_onnx serializes a fitted estimator to ONNX bytes via skl2onnx; register_onnx runs the CREATE OR REPLACE MODEL with the given signature.

Generative AI & embeddings​

Gnok's AI and embedding providers are callable from Python too. gnok.ai calls are governed like SQL AI_* calls: an organization administrator must first enable AI Governance. gnok.embed uses your organization's monthly embedding allowance.

gnok.ai.generate(prompt)​

Generate text via SQL AI_GENERATE, an alias of AI_COMPLETE — summaries, classification, extraction, and the like. The rollout policy must allow ai_complete.

from gnok import ai
label = ai.generate("Classify the sentiment of 'love it' in one word.")

gnok.embed(text)​

Embed text into a vector (engine EMBED) for similarity search / RAG. Returns a list of floats.

vec = gnok.embed("wireless noise-cancelling headphones")

Gnok AI in the notebook​

The notebook has a ✨ Gnok AI panel that writes and fixes cells for you, grounded in the gnok helper. Ask it to "read 1,000 rows of sales.orders and plot amount by region" or "train a model to predict amount" and it returns a ready-to-run cell. When a cell errors, ✨ Fix sends the traceback to the model and proposes a corrected cell.

Persistence, history & lifecycle​

Notebooks (cells + outputs) are saved per user. Every save also snapshots a version — the 🕘 History panel lists prior versions and Restore loads one back into the editor (then Save to make it current).

Kernels are ephemeral — your saved cells are the source of truth — and are reaped when idle, so a long-lived notebook never holds resources it isn't using. Re-running a cell after the kernel is reaped transparently starts a fresh one.

Compute availability: notebook packages and accelerator access are determined by the hosted environment. Confirm the capabilities enabled for your account before depending on them.

Headless runs & schedules​

Every notebook has two run paths, side by side:

  • ▶ Run (per cell) — interactive, against your open kernel session. Fast, state is shared across cells. The default.
  • ▶▶ Headless (per cell) and ▶ Run all (toolbar) — the headless executor spawns a transient kernel via the control plane, runs the cell(s), persists the outputs back, and tears the kernel down. Doesn't disturb your open kernel; useful for re-running against a clean state or for the unattended scheduled path.

The ⏰ Schedule panel lets you set a notebook to run on a fixed cadence (seconds, minimum 60). The scheduler executes a snapshot of the cells captured when you save the schedule — re-save the schedule to refresh the snapshot. Scheduled runs land in the notebook's 📜 Runs panel (status, trigger, duration; click a row to expand for full error + the persisted cell outputs); they never overwrite the live notebook, so a notebook you have open is never clobbered by a 3 a.m. run.

The cross-notebook view at ⏰ All schedules (sidebar link → /notebooks/schedules) lists every schedule you own — interval, enabled/ paused, next-run timestamp — with Pause/Resume (flips enabled only, preserves the snapshot) and Delete.

Architecture (where the compute actually runs)​

Gnok manages notebook kernels, the scheduler, and execution infrastructure. Scheduled runs require a valid execution identity and access to their inputs. Use the Runs panel to inspect failures and contact Gnok support for runtime or scheduling-service problems. Customers do not need internal service tokens or Kubernetes configuration.