01 · overview
The model runs on your hardware. Full stop.
Peridot is a local AI kernel built around the same principle as the rest of the uncoalesced stack: the model runs on your hardware, under your control, with nothing phoning home. Configuration and system behavior are governed by an explicit, versioned “constitution” file rather than opaque defaults, and the project maintains a running changelog across releases.
The current stable release is v1.5.4. It added native Linux support (Debian 12, Ubuntu 22.04+ and Arch) next to the Windows runtime, and closed a gap where a stale .env file could quietly switch the offline flags back on.
02 · how it works
Air-gapped by architecture, not by policy
A sovereign, local kernel
Everything runs inside the boundary of your machine. Local hardware feeds a sovereign kernel, which drives the user interface and the audio module directly. There is no network edge in the loop, and nothing to exfiltrate data through.

Idle compute, instant handoff
Idle GPU memory is treated as a design defect. While you're away, the kernel puts the hardware to work on background compute like protein folding, and when you return it hands the GPU back to active inference in roughly 21 milliseconds. You never wait for your own machine.
Deciding when to hand it back is its own problem, and it has its own project inside Peridot: ZAT-SCS reads keystroke acceleration and microphone amplitude at 10Hz to predict that a prompt is coming, and reclaims the card before you press Enter rather than after.

A versioned constitution
No opaque defaults. The kernel's configuration and system behavior are governed by an explicit “constitution” file that is versioned, diffable and yours to edit. Every release ships against a running changelog, so behavior changes are visible history, not surprises.
Tested against RAG workloads
Retrieval-augmented generation is part of the kernel's testing scope. The “RAG” tag on this project names something it is tested against.

03 · stack
Tech stack
| Language | Python 3.11 |
|---|---|
| Inference | llama-cpp-python with cuBLAS, running GGUF models |
| Default model | Qwen2.5-14B-Instruct, Q4_K_M, 8192-token sliding context |
| Memory | Split tensor allocation across GPU VRAM and system RAM |
| Retrieval | TurboVec, a Rust-backed vector index with 4-bit TurboQuant compression |
| History | A local SQLite conversation ledger, no accounts or cloud sync |
| Platforms | Windows, and native Linux on Debian 12, Ubuntu 22.04+ and Arch |
| License | MIT |
The benchmark figures above were measured on an NVIDIA RTX 5050 Laptop GPU.
04 · on github
Development in the open
05 · from the lab
Illustrations
More Peridot illustrations are being produced and will drop into this gallery as they land.
Runs local. Stays local.
Built, shipped, and getting regular updates.

