Resource Governance
Gnok's resource governance controls how compute resources are allocated across organizations and workloads.
Virtual Warehouses
Virtual warehouses are isolated compute pools. Each warehouse has a dedicated worker pool and resource limits:
-- Create a warehouse with size and auto-suspend
CREATE WAREHOUSE analytics_wh
WITH SIZE = 'medium', MAX_CONCURRENT = 8, AUTO_SUSPEND_SECS = 300;
-- Route queries to a specific warehouse
USE WAREHOUSE analytics_wh;
SELECT * FROM large_table;
-- Suspend when idle, resume when needed
ALTER WAREHOUSE analytics_wh SUSPEND;
ALTER WAREHOUSE analytics_wh RESUME;
See DDL: Warehouses for the full SQL reference.
Warehouse Sizes
| Size | Workers | Description |
|---|---|---|
xsmall | 1 | Development and testing |
small | 2 | Light workloads |
medium | 4 | Standard workloads |
large | 8 | Heavy analytics |
xlarge | 16 | Large-scale processing |
x2large | 32 | Enterprise workloads |
x3large | 64 | Very large-scale processing |
x4large | 128 | Maximum scale |
custom | configurable | Custom worker count |
Warehouses transition through states: running → suspended (via auto-suspend or ALTER WAREHOUSE ... SUSPEND), and back to running (via ALTER WAREHOUSE ... RESUME). Other transient states include resizing and stopping. Auto-suspend pauses the warehouse after a configurable idle timeout (default: 300 seconds / 5 minutes).
Resource Groups
Resource groups provide hierarchical workload isolation within a warehouse. Groups form a tree structure where each group has its own concurrency limits, memory budget, and scheduling policy. Child groups share the resource budget of their parent.
-- Create a top-level resource group
CREATE RESOURCE GROUP interactive
WITH (
hard_concurrency_limit = 50,
soft_concurrency_limit = 30,
max_queued_queries = 200,
soft_memory_limit_bytes = '32GB',
scheduling_policy = 'fair'
);
-- Create a child group for a specific team
CREATE RESOURCE GROUP interactive.data_science
WITH (
hard_concurrency_limit = 20,
soft_concurrency_limit = 10,
max_queued_queries = 50,
soft_memory_limit_bytes = '16GB',
scheduling_policy = 'weighted',
scheduling_weight = 3
);
-- Assign a role to a resource group
ALTER ROLE analyst SET RESOURCE_GROUP = 'interactive';
ALTER ROLE data_scientist SET RESOURCE_GROUP = 'interactive.data_science';
Resource Group Settings
| Setting | Type | Description |
|---|---|---|
hard_concurrency_limit | Integer | Maximum concurrent queries. Queries beyond this limit are rejected. |
soft_concurrency_limit | Integer | Target concurrency. Queries beyond this limit are queued rather than rejected. |
max_queued_queries | Integer | Maximum queries waiting in the queue. Excess queries are rejected. |
soft_memory_limit_bytes | String | Memory budget for the group. Queries may exceed this transiently but triggers spill. |
scheduling_policy | Enum | Fair (round-robin), Weighted (proportional to weight), or QueryPriority (by priority tier). |
scheduling_weight | Integer | Relative weight when using Weighted scheduling. Higher weight means more resources. |
Scheduling Policies
- Fair: Each resource group at the same level receives an equal share of scheduling slots. Queries within a group are scheduled in FIFO order.
- Weighted: Resource groups receive scheduling slots proportional to their
scheduling_weight. A group with weight 3 receives three times the slots of a group with weight 1. - QueryPriority: Queries are scheduled based on their priority tier (critical, high, normal, low, background). Within the same priority, queries are scheduled in FIFO order.
Query Admission Control
Before a query begins execution on a worker, Gnok evaluates whether the worker has sufficient resources to accept the query fragment. Admission control uses jemalloc memory statistics to make accurate decisions based on actual process memory usage rather than estimated reservations.
Memory Zones
The admission controller evaluates the worker's memory utilization and classifies it into one of four zones:
| Zone | Utilization | Behavior |
|---|---|---|
| Normal | Below 70% | All fragments admitted without restriction. |
| Cautious | 70% -- 80% | Fragments admitted. A log message is emitted at debug level. |
| Pressured | 80% -- 90% | Fragments admitted with spill pressure. Operators are signaled to spill intermediate state to disk proactively. |
| Critical | 90% and above | New fragments are rejected. The coordinator retries on other workers or queues the query. |
When a worker has zero active fragments but reports Critical utilization (possible due to jemalloc arena retention), the system purges jemalloc arenas and re-evaluates before rejecting.
Fragment-Level Limits
The service manages worker concurrency and fragment limits. Customer workload limits and available warehouse capacity are determined by the account configuration.
Memory Management
Gnok tracks memory at the process level using jemalloc statistics (stats.allocated) and at the query level using per-query memory accounting.
Dynamic Memory Budget
Memory allocation is managed by Gnok. Reduce unnecessary input volume and join expansion, and use appropriate available compute for larger workloads.
Spill-to-Disk
Supported operators can spill intermediate data to service-managed storage. Spilling can increase latency and does not remove query limits. Customers do not configure worker filesystems.
Query-Level Memory Tracking
Each query tracks its own memory consumption. If a query exceeds the per-query memory limit (max_memory_per_query), the query is cancelled with a descriptive error message. This prevents a single runaway query from consuming all available memory.
SET max_memory_per_query = '8GB';
Query Killing
Queries that exceed resource limits are automatically cancelled by the query killer:
- Memory limit: If a query's tracked memory exceeds
max_memory_per_query, the query is cancelled. - Time limit: If a query's wall-clock execution time exceeds
query_timeout, the query is cancelled. - Queue timeout: If a query waits in the admission queue longer than
max_queue_wait_time(default: 5 minutes), the query is cancelled.
Cancelled queries return an error message indicating the reason for cancellation. The cancellation event is recorded in the audit log.
-- Set per-session limits
SET query_timeout = '120s';
SET max_memory_per_query = '8GB';
SET max_rows_returned = 1000000;
-- Per-query hint
SELECT /*+ MEMORY_LIMIT('2GB') */ * FROM large_table;
Query Limits
Set per-query resource limits:
SET query_timeout = '120s';
SET max_memory_per_query = '8GB';
SET max_rows_returned = 1000000;
Organization Quotas
Organization concurrency, scan, and resource quotas are managed by Gnok. Ask your administrator or email support@gnok.io about the applicable limits and authorized changes.
Priority Queues
Queries can be assigned priority levels. Higher-priority queries get preferential scheduling when resources are contended:
SET query_priority = 'high';
SLO Tiers
| Tier | p50 Target | p99 Target | Description |
|---|---|---|---|
critical | 100ms | 500ms | Highest priority, strictest latency targets |
high | 500ms | 2s | Important queries |
normal | 2s | 10s | Default tier |
low | 10s | 60s | Batch-style workloads |
background | best-effort | best-effort | Lowest priority |
When resources are contended, the engine uses adaptive throttling to maintain SLO targets for higher-priority queries. Low-priority queries are queued while high-priority queries are scheduled first.
Monitoring
Gnok monitors service-side capacity, memory, and admission metrics. To see how your own workloads are doing:
SHOW RESOURCE GROUPS: running and queued queries, concurrency limits, and scheduling policy per resource group. See workload management.- Query history: per-query duration, bytes scanned, status, and compute units. See query history.
- Warehouse usage: consumption per warehouse. See billing.
Best Practices
- Separate interactive and batch workloads into distinct resource groups. Give interactive groups a higher
scheduling_weightand a lowerhard_concurrency_limitto maintain low latency. Give batch groups a higher concurrency limit with lower priority. - Set per-query memory limits to prevent runaway queries from consuming the worker's entire memory budget. A reasonable default is 25% of the worker's total memory.
- Monitor admission rejection rates. Frequent rejections in the Critical zone indicate that workers are undersized or that the workload exceeds the cluster's capacity.
- Use auto-suspend on warehouses that have idle periods. This reduces cost without requiring manual intervention.