Skip to main content

Python SDK (pygnok)

pygnok is a DB-API 2.0 compliant Python client for Gnok, built on Arrow Flight SQL.

Inside a Gnok Studio notebook, use gnok

Notebook kernels come with a pre-installed gnok helper that wraps pygnok, auto-connects as the signed-in user, and adds gnok.ml for the train→serve round-trip. Reach for pygnok (below) when connecting from your own scripts.

Installation​

Client connectivity is set up by request

Connecting with pygnok from your own scripts isn't self-service for new organizations yet; use Gnok Studio in the meantime. Email support@gnok.io to request client connectivity. Gnok support provides the endpoint, client package, and credentials for your organization.

Gnok support provides the pygnok package and installation instructions when your client connectivity is set up. In a Studio notebook, use the pre-installed gnok helper instead. See client packages.

Connect to the Flight SQL host and port provided for your organization with TLS enabled. flight.example.com:443 below is a placeholder. The examples read your token from a variable named GNOK_TOKEN in your own shell or secret manager, so the token stays out of your code.

Quick Start​

import os
import pygnok

# Connect
conn = pygnok.connect(host="flight.example.com", port=443, tls=True, token=os.environ["GNOK_TOKEN"])
cursor = conn.cursor()

# Execute a query
cursor.execute("SELECT * FROM my_table WHERE date > '2025-01-01'")

# Fetch as rows
rows = cursor.fetchall()

# Or fetch as an Arrow table (zero-copy)
table = cursor.fetch_arrow_table()

# Or convert to pandas
df = table.to_pandas()

Connection Options​

conn = pygnok.connect(
host="flight.example.com",
tls=True,
port=443,
token=os.environ["GNOK_TOKEN"], # JWT token for authentication
query_timeout=30, # Query timeout in seconds
)

DB-API 2.0 Compliance​

pygnok implements the standard PEP 249 interface:

  • connect() — create a connection
  • cursor() — create a cursor
  • execute() / executemany() — run queries
  • fetchone() / fetchmany() / fetchall() — retrieve results
  • close() — close connections and cursors

PyArrow Integration​

For high-performance data pipelines, use the Arrow-native interface:

cursor.execute("SELECT * FROM large_table")
arrow_table = cursor.fetch_arrow_table()

# Direct to Parquet
import pyarrow.parquet as pq
pq.write_table(arrow_table, "output.parquet")