Workload Management
Gnok automatically classifies queries by complexity, routes them to resource groups with configurable limits, and schedules execution using a multi-level feedback queue. This helps prioritize interactive work alongside batch workloads; response time still depends on query cost and available resources.
Workload Classification
Every query is automatically classified based on analysis of its physical plan. Classification determines default resource limits and scheduling priority.
| Class | Criteria | Default Priority |
|---|---|---|
| Interactive | No joins, no window functions, estimated rows < 100K | High |
| Medium | 1-2 joins, or has window function, or 3+ scans, or estimated rows 100K-10M | Normal |
| Batch | 3+ joins, or window + aggregate combo, or estimated rows > 10M | Low |
Classification is based on structural plan metrics:
- Join count -- number of join operators (hash, sort-merge, broadcast, nested-loop, range, as-of)
- Scan count -- number of table scan operators
- Window/aggregate count -- presence of window functions or aggregates
- Estimated rows -- sum of estimated rows across all scans (from Iceberg metadata)
Manual Override
Override automatic classification at the session level:
SET gnok.workload_class = 'batch';
SET gnok.workload_class = 'medium';
SET gnok.workload_class = 'interactive';
Each class has configurable timeout and memory limits:
| Class | Default Timeout | Default Memory Cap |
|---|---|---|
| Interactive | 30s | 1 GB |
| Medium | 3 min | 4 GB |
| Batch | 10 min | 8 GB |
Service-managed limits determine the effective timeout, memory, and execution settings for each workload class.
Resource Groups
Resource groups provide hierarchical admission control for query execution. They are tree-structured (inspired by Trino's resource groups) with per-group concurrency, queuing, and memory limits.
Configuration
Resource-group provisioning and service routing settings are managed by Gnok. Use the authorized workload controls available to your account; a local configuration file does not change the hosted service.
Group Properties
| Property | Default | Description |
|---|---|---|
hard_concurrency_limit | 100 | Maximum concurrent running queries. New queries are queued when this limit is reached. |
soft_concurrency_limit | 100 | Target concurrency. Above this, CPU penalty scaling reduces effective concurrency. |
max_queued_queries | 1000 | Maximum queued queries. New submissions are rejected when the queue is full. |
soft_memory_limit_bytes | (none) | Memory threshold. New queries are blocked when the group's aggregate memory exceeds this. |
soft_cpu_limit_ms | (none) | Cumulative CPU time threshold. When exceeded, effective concurrency decreases linearly toward 1. |
hard_cpu_limit_ms | (none) | Cumulative CPU time hard limit. Effective concurrency reduced to 1. Defaults to 2x soft if unset. |
scheduling_policy | fair | How queries are dequeued within the group. |
scheduling_weight | 1 | Weight for parent's weighted scheduling. Higher weight = more share. |
Scheduling Policies
| Policy | Description |
|---|---|
fair | FIFO ordering -- queries are dequeued in arrival order. |
weighted | Weighted fair sharing across sibling groups. Groups with a lower running/weight ratio get priority. |
query_priority | Dequeue the highest-priority query first (based on the query's execution priority). |
Query Routing
Queries are routed to resource groups using routing rules evaluated top-to-bottom. The first matching rule wins.
| Rule Field | Description |
|---|---|
match_tenant | Match organization (tenant) ID (exact). Omit to match any organization. |
match_user | Match user ID (exact). Omit to match any user. |
match_source | Match query source (e.g., dashboard, etl). |
match_resource_pool | Match resource pool name from session config. |
match_query_type | Match statement type: select, insert, update, delete. |
target_group | Target group path (e.g., root.interactive.dashboard). |
If no rule matches, the query is routed to the root group.
Priority Scheduling (MLFQ)
Within each resource group, fragment execution on workers is scheduled using a Multi-Level Feedback Queue (MLFQ). The MLFQ gives short-running queries natural priority over long-running ones without requiring explicit priority assignment.
How It Works
The MLFQ has 5 levels defined by cumulative scheduled time:
| Level | CPU Time Threshold | Typical Queries |
|---|---|---|
| 0 | 0 - 1s | Simple lookups, point queries |
| 1 | 1s - 10s | Small aggregations, filtered scans |
| 2 | 10s - 60s | Multi-table joins, moderate analytics |
| 3 | 60s - 300s | Large joins, complex aggregations |
| 4 | 300s+ | Long-running ETL, full-table scans |
New queries start at level 0. As a query accumulates CPU time across its fragments, its subsequent fragments are assigned to higher levels. Level 0 gets 2x more CPU time than level 1, which gets 2x more than level 2, and so on.
High-priority queries (priority >= High) bypass the MLFQ and always execute at level 0, regardless of accumulated CPU time.
MLFQ Configuration
Execution concurrency and scheduler thresholds are managed by Gnok. Request workload capacity changes through your administrator or support@gnok.io.
Query Killing
Queries that exceed their resource limits are automatically cancelled:
- Timeout -- queries exceeding their workload class timeout are cancelled
- Memory -- queries exceeding their memory limit trigger spill-to-disk first; if spilling is not possible, the query is cancelled
- Queue overflow -- when a resource group's queue is full, new submissions are rejected immediately
Cancelled queries are recorded in query history with status cancelled and an error message indicating the reason.
Monitoring
SHOW RESOURCE GROUPS
View the current state of all resource groups:
SHOW RESOURCE GROUPS;
Returns: group path, running queries, queued queries, hard/soft concurrency limits, scheduling policy, and cumulative CPU time.
QoS Statistics API
Use SHOW RESOURCE GROUPS in Studio. Calling the HTTP API from outside Studio isn't self-service yet; email support@gnok.io to set up client connectivity. query.example.com below is a placeholder.
GET /api/qos/stats
Returns quality-of-service metrics for the resource groups and scheduler that serve your queries:
curl -s 'https://query.example.com/api/qos/stats' \
-H "Authorization: Bearer $TOKEN" | jq
Response:
{
"resource_groups": [
{
"path": "/root/interactive",
"running": 12,
"queued": 0,
"hard_concurrency_limit": 50,
"scheduling_policy": "fair"
},
{
"path": "/root/batch",
"running": 8,
"queued": 3,
"hard_concurrency_limit": 30,
"scheduling_policy": "query_priority"
}
],
"mlfq": {
"running": 20,
"max_concurrent": 64,
"level_queue_depths": [0, 0, 2, 1, 0],
"level_scheduled_ms": [45000, 120000, 340000, 890000, 1200000],
"tracked_queries": 15
}
}