projects peridot

PeridotPERIDOT, a wordmark in which the O is a faceted green hexagonal gem with an amber dot at its center

Shipped

A sovereign, local-first AI inference kernel from uncoalesced. No cloud dependency, no telemetry, and you decide how the model runs.

Built, shipped, and getting regular updates.

01 · overview

The model runs on your hardware. Full stop.

Peridot is a local AI kernel built around the same principle as the rest of the uncoalesced stack: the model runs on your hardware, under your control, with nothing phoning home. Configuration and system behavior are governed by an explicit, versioned “constitution” file rather than opaque defaults, and the project maintains a running changelog across releases.

The current stable release is v1.5.4. It added native Linux support (Debian 12, Ubuntu 22.04+ and Arch) next to the Windows runtime, and closed a gap where a stale .env file could quietly switch the offline flags back on.

  • local inference
  • privacy-first
  • AI kernel
  • RAG
  • Windows + Linux

02 · how it works

Air-gapped by architecture, not by policy

A sovereign, local kernel

Everything runs inside the boundary of your machine. Local hardware feeds a sovereign kernel, which drives the user interface and the audio module directly. There is no network edge in the loop, and nothing to exfiltrate data through.

Air-gapped system diagram: inside a bounded frame labeled 'AIR-GAPPED SYSTEM' and 'ZERO DATA EXFILTRATION', local hardware connects to the sovereign kernel, which branches to the user interface and the audio module.
The air-gapped architecture: zero data exfiltration by design.

Idle compute, instant handoff

Idle GPU memory is treated as a design defect. While you're away, the kernel puts the hardware to work on background compute like protein folding, and when you return it hands the GPU back to active inference in roughly 21 milliseconds. You never wait for your own machine.

Deciding when to hand it back is its own problem, and it has its own project inside Peridot: ZAT-SCS reads keystroke acceleration and microphone amplitude at 10Hz to predict that a prompt is coming, and reclaims the card before you press Enter rather than after.

State diagram: a box labeled 'IDLE: PROTEIN FOLDING' connects to a box labeled 'ACTIVE: AI INFERENCE' along a green arrow marked '21ms handoff'.
VRAM handoff between background compute and active inference.

A versioned constitution

No opaque defaults. The kernel's configuration and system behavior are governed by an explicit “constitution” file that is versioned, diffable and yours to edit. Every release ships against a running changelog, so behavior changes are visible history, not surprises.

Tested against RAG workloads

Retrieval-augmented generation is part of the kernel's testing scope. The “RAG” tag on this project names something it is tested against.

Software specification sheet benchmarked on an NVIDIA RTX 5050 laptop GPU: an inference speed bar chart alongside figures reading short response 58.76 tokens per second, medium response 59.77, long response 57.71, and handoff latency 21.10 milliseconds.
Real benchmark data from an NVIDIA RTX 5050 Laptop GPU, numbers unaltered.

03 · stack

Tech stack

Peridot tech stack
LanguagePython 3.11
Inferencellama-cpp-python with cuBLAS, running GGUF models
Default modelQwen2.5-14B-Instruct, Q4_K_M, 8192-token sliding context
MemorySplit tensor allocation across GPU VRAM and system RAM
RetrievalTurboVec, a Rust-backed vector index with 4-bit TurboQuant compression
HistoryA local SQLite conversation ledger, no accounts or cloud sync
PlatformsWindows, and native Linux on Debian 12, Ubuntu 22.04+ and Arch
LicenseMIT

The benchmark figures above were measured on an NVIDIA RTX 5050 Laptop GPU.

04 · on github

Development in the open

stars
last commit
loading live stats…Loading GitHub stats

05 · from the lab

More Peridot illustrations are being produced and will drop into this gallery as they land.

Runs local. Stays local.

Built, shipped, and getting regular updates.