Governed compute for unified-memory AI hardware — the machines where CPU and
GPU share one pool and there is no separate VRAM allocation to bounce off.
Over-commit that pool and the box thrashes and wedges, SSH and ping included,
before the OOM killer gets a turn.
Compute does not run inference. It supervises the servers that do:
- Admission control. A model starts only if committed + requested + margin
fits the budget. The refusal is the feature.
- A 1 Hz watchdog on MemAvailable that stops the newest model before thrash.
- Scenes: named sets of models activated as one transactional unit, with
pre-flight validation and rollback to the previously active Scene on
failure. Scenes reference model ids, never weight paths or commands, so a
Scene obtained from elsewhere cannot introduce code.
- Process ownership bound to (boot_id, pid, start_time_ticks, pgid == pid),
so a reused PID can never be group-killed.
- A protocol-transparent TCP gateway, so clients keep one address while model
runtimes move behind it.
- An MCP server, so agents drive the node as tools rather than as a CLI.
One binary, six direct dependencies, no async runtime outside the MCP surface.
Published from the internal monorepo with a fresh history. The private
development tree keeps its own history; nothing here carries it.