projects cordon

Cordon

Shipped

Resource control for AI coding agents from uncoalesced, one tool call at a time. To most resource controllers a pytest run and a git status both look like “a subprocess”. Cordon tells them apart, because one needs 500MB and the other needs 13MB.

Python 3.11+ · Linux cgroup v2 · MIT · v1.0.0

01 · what it is

Measure first, then act on what you measured

Cordon hooks into Claude Code, Codex CLI, Hermes Agent, Cursor CLI and Gemini CLI (Aider too, at a coarser grain) and tracks memory and CPU for every tool call the agent makes. One container-wide limit can’t fit both a test suite and a status check, so Cordon doesn’t try.

Enforcement runs each guarded call in its own short-lived cgroup, sized from what the agent says it’s about to do, and tells the agent when a limit actually bites. The design builds on two papers, AgentCgroup and AgentSight, and its reports compare what it measures against their numbers.

  • AI coding agents
  • cgroup v2
  • resource limits
  • Claude Code
  • Python

02 · measurement

One sampler per session, not one per call

Every supported agent fires a PreToolUse and PostToolUse event around each tool call, with a JSON payload on stdin. Cordon registers one command against all of them and irons out the differences in a single alias table instead of shipping five integrations. Hooks are a stable boundary the agent can’t route around. Patching each framework’s tool loop would break on every release.

The first hook starts one background sampler for the whole session, and after that the hooks only write cheap timestamped markers. Spawning a process inside PreToolUse costs about 100ms on Windows, landing right inside the window being measured. The quiet time between calls matters too: it’s where the agent’s baseline memory shows up.

The sampler polls memory (RSS summed over the agent’s whole process tree) and CPU every 250ms, because the bursts it’s after last a second or two. On the reference machine one tick has a 6.82ms median against a live Claude Code tree, about 2.73% of one core. The README asks you to re-measure that on your own machine before trusting a batch.

cordon reduce joins markers to samples, one record per tool call. cordon analyze then runs five passes over a batch (execution-time split, peak-to-average memory, per-tool breakdown, retry loops, CPU and memory correlation) and writes a report with a measured-versus-paper verdict for each.

It never breaks the agent it watches

Every hook path exits 0 whatever happens inside it, and every sampling or analysis failure is logged and skipped rather than raised.

03 · enforcement

Throttle, don’t kill

cordon control run guards one command inside its own cgroup, created just before the subprocess starts and torn down after it exits. The limit comes from a hint: set AGENT_RESOURCE_HINT=memory:high before a call and that call alone gets a memory.high soft limit and a cpu.weight.

Cordon resource hint tiers
TierShare of RAMOn 16GBcpu.weight
low2.5%410 MB25
medium (default)10%1.6 GB100
high35%5.7 GB400
maxunlimitednone1000

Hints are advice, not trust. Crossing memory.high slows a call down under pressure; it doesn’t end it. Cordon never sets memory.max, because an OOM kill throws away whatever context the agent had built up.

When a call spends more than max(200ms, 5% of its runtime) stalled on memory, read from PSI, Cordon adds a note to its stderr after it exits:

[cordon] This tool call was resource-limited. It peaked at 1842.0 MB against a memory:medium limit of 1638.4 MB. It stalled 1.50s (54% of its 2.8s runtime) waiting on memory. Consider narrowing the scope of this command. If it genuinely needs more, set AGENT_RESOURCE_HINT=memory:high before retrying.

A freeze or an OOM kill is always reported. From the third warning on the same command, the note adds that retrying it unchanged probably won’t help.

04 · not built yet

The kernel-speed half is waiting on the kernel

Everything above runs on ordinary Linux with cgroup v2, no patches. What’s missing is moving the throttle decision into the kernel itself, where it would take microseconds instead of a userspace loop’s tens of milliseconds. Against a burst that lasts a second or two, that gap matters.

It needs sched_ext (Linux 6.12+) for CPU policy and memcg_bpf_ops, which is still an RFC and not upstream, for memory. Neither is stubbed or faked. cordon control probe reports both as absent on machines that don’t have them.

There’s also no automatic freeze loop above the throttle, on purpose. A userspace loop that polls pressure and decides when to freeze would just be a slower oomd.

05 · supported agents

Five hook dialects, one command

The five hook systems are the same idea under different names: matcher and command groups, JSON on stdin, an exit code to block. cordon install-hooks writes the right config for whichever agents you actually use.

Agents Cordon supports and at what grain
Claude Codeper tool call, the default target
Codex CLIper tool call, hooks are opt-in there
Hermes Agentper tool call, user-global config
Cursor CLIper tool call, or reuses Claude Code’s hooks
Gemini CLIper tool call, hooks on by default from v0.26.0
Aiderwhole session only, via cordon wrap (no hook system)

06 · status

Shipped

v1.0.0, MIT licensed

Measurement works on Windows and Linux. Enforcement needs Linux with cgroup v2; on a machine without it, guarded commands still run, and Cordon logs what it would have applied instead of enforcing anything.