Skip to main content

Workload Management

Gnok automatically classifies queries by complexity, routes them to resource groups with configurable limits, and schedules execution using a multi-level feedback queue. This helps prioritize interactive work alongside batch workloads; response time still depends on query cost and available resources.

Workload Classification​

Every query is automatically classified based on analysis of its physical plan. Classification determines default resource limits and scheduling priority.

ClassCriteriaDefault Priority
InteractiveNo joins, no window functions, estimated rows < 100KHigh
Medium1-2 joins, or has window function, or 3+ scans, or estimated rows 100K-10MNormal
Batch3+ joins, or window + aggregate combo, or estimated rows > 10MLow

Classification is based on structural plan metrics:

  • Join count -- number of join operators (hash, sort-merge, broadcast, nested-loop, range, as-of)
  • Scan count -- number of table scan operators
  • Window/aggregate count -- presence of window functions or aggregates
  • Estimated rows -- sum of estimated rows across all scans (from Iceberg metadata)

Manual Override​

Override automatic classification at the session level:

SET gnok.workload_class = 'batch';
SET gnok.workload_class = 'medium';
SET gnok.workload_class = 'interactive';

Each class has configurable timeout and memory limits:

ClassDefault TimeoutDefault Memory Cap
Interactive30s1 GB
Medium3 min4 GB
Batch10 min8 GB

Service-managed limits determine the effective timeout, memory, and execution settings for each workload class.

Resource Groups​

Resource groups provide hierarchical admission control for query execution. They are tree-structured (inspired by Trino's resource groups) with per-group concurrency, queuing, and memory limits.

Configuration​

Resource-group provisioning and service routing settings are managed by Gnok. Use the authorized workload controls available to your account; a local configuration file does not change the hosted service.

Group Properties​

PropertyDefaultDescription
hard_concurrency_limit100Maximum concurrent running queries. New queries are queued when this limit is reached.
soft_concurrency_limit100Target concurrency. Above this, CPU penalty scaling reduces effective concurrency.
max_queued_queries1000Maximum queued queries. New submissions are rejected when the queue is full.
soft_memory_limit_bytes(none)Memory threshold. New queries are blocked when the group's aggregate memory exceeds this.
soft_cpu_limit_ms(none)Cumulative CPU time threshold. When exceeded, effective concurrency decreases linearly toward 1.
hard_cpu_limit_ms(none)Cumulative CPU time hard limit. Effective concurrency reduced to 1. Defaults to 2x soft if unset.
scheduling_policyfairHow queries are dequeued within the group.
scheduling_weight1Weight for parent's weighted scheduling. Higher weight = more share.

Scheduling Policies​

PolicyDescription
fairFIFO ordering -- queries are dequeued in arrival order.
weightedWeighted fair sharing across sibling groups. Groups with a lower running/weight ratio get priority.
query_priorityDequeue the highest-priority query first (based on the query's execution priority).

Query Routing​

Queries are routed to resource groups using routing rules evaluated top-to-bottom. The first matching rule wins.

Rule FieldDescription
match_tenantMatch organization (tenant) ID (exact). Omit to match any organization.
match_userMatch user ID (exact). Omit to match any user.
match_sourceMatch query source (e.g., dashboard, etl).
match_resource_poolMatch resource pool name from session config.
match_query_typeMatch statement type: select, insert, update, delete.
target_groupTarget group path (e.g., root.interactive.dashboard).

If no rule matches, the query is routed to the root group.

Priority Scheduling (MLFQ)​

Within each resource group, fragment execution on workers is scheduled using a Multi-Level Feedback Queue (MLFQ). The MLFQ gives short-running queries natural priority over long-running ones without requiring explicit priority assignment.

How It Works​

The MLFQ has 5 levels defined by cumulative scheduled time:

LevelCPU Time ThresholdTypical Queries
00 - 1sSimple lookups, point queries
11s - 10sSmall aggregations, filtered scans
210s - 60sMulti-table joins, moderate analytics
360s - 300sLarge joins, complex aggregations
4300s+Long-running ETL, full-table scans

New queries start at level 0. As a query accumulates CPU time across its fragments, its subsequent fragments are assigned to higher levels. Level 0 gets 2x more CPU time than level 1, which gets 2x more than level 2, and so on.

High-priority queries (priority >= High) bypass the MLFQ and always execute at level 0, regardless of accumulated CPU time.

MLFQ Configuration​

Execution concurrency and scheduler thresholds are managed by Gnok. Request workload capacity changes through your administrator or support@gnok.io.

Query Killing​

Queries that exceed their resource limits are automatically cancelled:

  • Timeout -- queries exceeding their workload class timeout are cancelled
  • Memory -- queries exceeding their memory limit trigger spill-to-disk first; if spilling is not possible, the query is cancelled
  • Queue overflow -- when a resource group's queue is full, new submissions are rejected immediately

Cancelled queries are recorded in query history with status cancelled and an error message indicating the reason.

Monitoring​

SHOW RESOURCE GROUPS​

View the current state of all resource groups:

SHOW RESOURCE GROUPS;

Returns: group path, running queries, queued queries, hard/soft concurrency limits, scheduling policy, and cumulative CPU time.

QoS Statistics API​

Use SHOW RESOURCE GROUPS in Studio. Calling the HTTP API from outside Studio isn't self-service yet; email support@gnok.io to set up client connectivity. query.example.com below is a placeholder.

GET /api/qos/stats

Returns quality-of-service metrics for the resource groups and scheduler that serve your queries:

curl -s 'https://query.example.com/api/qos/stats' \
-H "Authorization: Bearer $TOKEN" | jq

Response:

{
"resource_groups": [
{
"path": "/root/interactive",
"running": 12,
"queued": 0,
"hard_concurrency_limit": 50,
"scheduling_policy": "fair"
},
{
"path": "/root/batch",
"running": 8,
"queued": 3,
"hard_concurrency_limit": 30,
"scheduling_policy": "query_priority"
}
],
"mlfq": {
"running": 20,
"max_concurrent": 64,
"level_queue_depths": [0, 0, 2, 1, 0],
"level_scheduled_ms": [45000, 120000, 340000, 890000, 1200000],
"tracked_queries": 15
}
}