ZAT-SCS
Zero-Overhead Active Telemetry & Speculative Context Streaming
01 · why this exists
Every prompt was paying for a cold start
Run a local llama-server on the same GPU as a background workload like Folding@Home and every prompt you send arrives cold. The card has to shed the background job, load weights, and restore context before it can produce a first token. That overhead is hundreds of milliseconds, and it is paid on every single prompt.
This is the same problem Peridot describes as treating idle GPU memory as a design defect. The answer on both sides is the same idea: the hardware should be working while you are not, and it should already be yours by the time you are. ZAT-SCS is the half of that which decides when.
It removes the overhead by predicting the prompt instead of reacting to it. By the time you press Enter, the GPU has already been reclaimed and the context is already loaded.
02 · how it works
Two sensors, one probability, four states
- A keyboard tracker and a microphone tracker feed raw signals into a probability engine at 10Hz.
- The engine fuses them into a single score, P(I_t), using temporal decay and a weighted contribution from each sensor.
- A finite state machine watches that score. Crossing the threshold moves it from
COLD_IDLEtoSPECULATIVE_PREPARED, which throttles the background process, flags Unified Memory, and fires an async context-restore request atllama-server. - If a prompt does arrive, time to first token is near zero. If none arrives within 15 seconds, that counts as a false positive: the system rolls back, and may raise its own threshold for next time.
The fusion
keyboard.py captures keystrokes through pynput and keeps a sliding window of the last ten timestamps, computing a typing acceleration coefficient, f_c, from the average inter-key interval. audio.py opens a mono 44.1kHz input stream through sounddevice and computes an RMS envelope, g_a, normalised to [0.0, 1.0]. processor.py combines them each tick:
P(I_t) = P(I_t-1) * e^(-lambda * dt) + w_key * f_c + w_aud * g_aThe result is clamped to [0.0, 1.0]. The decay factor is pre-compiled at init, because the tick interval is fixed at 100ms and recomputing it ten times a second buys nothing.
03 · the state machine
Four states
COLD_IDLE
background at 100%
Nothing speculative is running. Background jobs have the card.
SPECULATIVE_PREPARED
background at 10% SM
P(I_t) crossed the threshold. Background compute is throttled, Unified Memory is flagged, and an async context restore is fired at the inference server.
ACTIVE_INFERENCE
background at 0%
A prompt actually arrived. The whole card goes to the model.
DEGRADED
background at 100%
The inference server is offline. Everything is rolled back and speculative operations pause until it returns.
It tunes its own threshold
Prepare speculatively and get no prompt within 15 seconds: false positive. After three consecutive ones, the threshold rises by 0.05. A prompt that arrives while the system is still in COLD_IDLE is a false negative, and the threshold drops by 0.05. The threshold is clamped between 0.40 and 0.90, so neither direction can run away.
04 · the pieces
Telemetry, orchestration, client
Telemetry
keyboard.py for keystroke acceleration, audio.py for the mic RMS envelope, processor.py for the fusion that outputs P(I_t).
Orchestration
fsm.py holds the state machine and the self-tuning. mps.py handles GPU resource control: on Linux with nvidia-cuda-mps-control present it sets SM thread capacity directly, and on Windows or without MPS it falls back to psutil, either lowering the background process priority or suspending it, configurably. It caches the current allocation so redundant system calls are skipped. uvm.py sets GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 so the runtime can stage weight tensors into Unified Virtual Memory.
Client
api.py wraps a persistent requests.Session with two methods, a health check and a slot restore, so connection overhead stays low. prefetcher.py runs the restore on a daemon thread so it never blocks the 10Hz loop, retrying with exponential backoff from 1s up to 16s if the server goes unreachable mid-restore.
05 · configuration
Every tunable, and its default
All of these live in one file, config/settings.py. The values below are the shipped defaults.
| Parameter | Default | Description |
|---|---|---|
INITIAL_THRESHOLD | 0.65 | Starting probability threshold for speculative preemption |
MIN_THRESHOLD | 0.40 | Floor after self-tuning lowers the threshold |
MAX_THRESHOLD | 0.90 | Ceiling after self-tuning raises the threshold |
ADJUSTMENT_STEP | 0.05 | How much the threshold moves per calibration event |
CONSECUTIVE_FALSE_POSITIVES_LIMIT | 3 | False positives needed before raising the threshold |
INACTIVITY_TIMER | 15.0s | Seconds before a speculative cycle counts as a false positive |
LAMBDA_DECAY | 0.1 | Temporal decay rate for the probability model |
WEIGHT_KEY | 0.45 | Keystroke acceleration weight in probability fusion |
WEIGHT_AUD | 0.35 | Audio amplitude weight in probability fusion |
TELEMETRY_HZ | 10 | Sampling and evaluation frequency |
BACKGROUND_PROCESS_NAME | fahclient | Name of the process to throttle |
FALLBACK_TO_SUSPEND | False | Suspend the process entirely instead of lowering its priority |
06 · proving it
Diagnostics ship with it
mock_server.py is a Flask stand-in for llama-server that accepts a slot restore and sleeps 200ms to simulate NVMe read latency, so the whole system can be exercised with no real inference server running. benchmark.py measures cold-start time to first token against speculative time to first token and prints the ratio.
stress_test.py is five phases of deliberate abuse. It floods the keyboard tracker at 100Hz to check the main loop still holds 10Hz. It flaps the server health endpoint on and off to test DEGRADED recovery. It forces 15 false negatives and 50 false positives to verify the threshold clamps at 0.40 and 0.90. It disables MPS and runs the psutil software fallback. Then it fires SIGINT mid-loop and checks teardown completes with no zombie threads.
Clean shutdown matters more here than it sounds. The daemon holds GPU resources away from a background job while it runs, so an unclean exit would leave the card throttled with nothing left to claim it. Ctrl+C restores 100% before exiting.
Part of Peridot
Same principle, one layer down: your hardware, your schedule.