projects peridot zat-scs

ZAT-SCS

Zero-Overhead Active Telemetry & Speculative Context Streaming

A predictive preemption engine for local LLM inference, part of Peridot from uncoalesced. It watches keystroke acceleration and microphone amplitude at 10Hz, estimates the probability that a prompt is about to arrive, and reclaims the GPU before it does.

Part of Peridot, not a separate product. Since Peridot 1.6.0 it is opt-in and off by default.

01 · why this exists

Every prompt was paying for a cold start

Run a local llama-server on the same GPU as a background workload like Folding@Home and every prompt you send arrives cold. The card has to shed the background job, load weights, and restore context before it can produce a first token. That overhead is hundreds of milliseconds, and it is paid on every single prompt.

This is the same problem Peridot describes as treating idle GPU memory as a design defect. The answer on both sides is the same idea: the hardware should be working while you are not, and it should already be yours by the time you are. ZAT-SCS is the half of that which decides when.

It removes the overhead by predicting the prompt instead of reacting to it. By the time you press Enter, the GPU has already been reclaimed and the context is already loaded.

  • predictive preemption
  • local inference
  • GPU scheduling
  • Python

02 · how it works

Two sensors, one probability, four states

  1. A keyboard tracker and a microphone tracker feed raw signals into a probability engine at 10Hz.
  2. The engine fuses them into a single score, P(I_t), using temporal decay and a weighted contribution from each sensor.
  3. A finite state machine watches that score. Crossing the threshold moves it from COLD_IDLE to SPECULATIVE_PREPARED, which throttles the background process, flags Unified Memory, and fires an async context-restore request at llama-server.
  4. If a prompt does arrive, time to first token is near zero. If none arrives within 15 seconds, that counts as a false positive: the system rolls back, and may raise its own threshold for next time.

The fusion

keyboard.py captures keystrokes through pynput and keeps a sliding window of the last ten timestamps, computing a typing acceleration coefficient, f_c, from the average inter-key interval. audio.py opens a mono 44.1kHz input stream through sounddevice and computes an RMS envelope, g_a, normalised to [0.0, 1.0]. processor.py combines them each tick:

P(I_t) = P(I_t-1) * e^(-lambda * dt) + w_key * f_c + w_aud * g_a

The result is clamped to [0.0, 1.0]. The decay factor is pre-compiled at init, because the tick interval is fixed at 100ms and recomputing it ten times a second buys nothing.

03 · the state machine

Four states

COLD_IDLE

background at 100%

Nothing speculative is running. Background jobs have the card.

SPECULATIVE_PREPARED

background at 10% SM

P(I_t) crossed the threshold. Background compute is throttled, Unified Memory is flagged, and an async context restore is fired at the inference server.

ACTIVE_INFERENCE

background at 0%

A prompt actually arrived. The whole card goes to the model.

DEGRADED

background at 100%

The inference server is offline. Everything is rolled back and speculative operations pause until it returns.

It tunes its own threshold

Prepare speculatively and get no prompt within 15 seconds: false positive. After three consecutive ones, the threshold rises by 0.05. A prompt that arrives while the system is still in COLD_IDLE is a false negative, and the threshold drops by 0.05. The threshold is clamped between 0.40 and 0.90, so neither direction can run away.

04 · the pieces

Telemetry, orchestration, client

Telemetry

keyboard.py for keystroke acceleration, audio.py for the mic RMS envelope, processor.py for the fusion that outputs P(I_t).

Orchestration

fsm.py holds the state machine and the self-tuning. mps.py handles GPU resource control: on Linux with nvidia-cuda-mps-control present it sets SM thread capacity directly, and on Windows or without MPS it falls back to psutil, either lowering the background process priority or suspending it, configurably. It caches the current allocation so redundant system calls are skipped. uvm.py sets GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 so the runtime can stage weight tensors into Unified Virtual Memory.

Client

api.py wraps a persistent requests.Session with two methods, a health check and a slot restore, so connection overhead stays low. prefetcher.py runs the restore on a daemon thread so it never blocks the 10Hz loop, retrying with exponential backoff from 1s up to 16s if the server goes unreachable mid-restore.

05 · configuration

Every tunable, and its default

All of these live in one file, config/settings.py. The values below are the shipped defaults.

ZAT-SCS configuration parameters and defaults
ParameterDefaultDescription
INITIAL_THRESHOLD0.65Starting probability threshold for speculative preemption
MIN_THRESHOLD0.40Floor after self-tuning lowers the threshold
MAX_THRESHOLD0.90Ceiling after self-tuning raises the threshold
ADJUSTMENT_STEP0.05How much the threshold moves per calibration event
CONSECUTIVE_FALSE_POSITIVES_LIMIT3False positives needed before raising the threshold
INACTIVITY_TIMER15.0sSeconds before a speculative cycle counts as a false positive
LAMBDA_DECAY0.1Temporal decay rate for the probability model
WEIGHT_KEY0.45Keystroke acceleration weight in probability fusion
WEIGHT_AUD0.35Audio amplitude weight in probability fusion
TELEMETRY_HZ10Sampling and evaluation frequency
BACKGROUND_PROCESS_NAMEfahclientName of the process to throttle
FALLBACK_TO_SUSPENDFalseSuspend the process entirely instead of lowering its priority

06 · proving it

Diagnostics ship with it

mock_server.py is a Flask stand-in for llama-server that accepts a slot restore and sleeps 200ms to simulate NVMe read latency, so the whole system can be exercised with no real inference server running. benchmark.py measures cold-start time to first token against speculative time to first token and prints the ratio.

stress_test.py is five phases of deliberate abuse. It floods the keyboard tracker at 100Hz to check the main loop still holds 10Hz. It flaps the server health endpoint on and off to test DEGRADED recovery. It forces 15 false negatives and 50 false positives to verify the threshold clamps at 0.40 and 0.90. It disables MPS and runs the psutil software fallback. Then it fires SIGINT mid-loop and checks teardown completes with no zombie threads.

Clean shutdown matters more here than it sounds. The daemon holds GPU resources away from a background job while it runs, so an unclean exit would leave the card throttled with nothing left to claim it. Ctrl+C restores 100% before exiting.

Part of Peridot

Same principle, one layer down: your hardware, your schedule.