Lumbridge Compute
ci / rust (push) Successful in 2m26s

Governed compute for unified-memory AI hardware — the machines where CPU and
GPU share one pool and there is no separate VRAM allocation to bounce off.
Over-commit that pool and the box thrashes and wedges, SSH and ping included,
before the OOM killer gets a turn.

Compute does not run inference. It supervises the servers that do:

- Admission control. A model starts only if committed + requested + margin
  fits the budget. The refusal is the feature.
- A 1 Hz watchdog on MemAvailable that stops the newest model before thrash.
- Scenes: named sets of models activated as one transactional unit, with
  pre-flight validation and rollback to the previously active Scene on
  failure. Scenes reference model ids, never weight paths or commands, so a
  Scene obtained from elsewhere cannot introduce code.
- Process ownership bound to (boot_id, pid, start_time_ticks, pgid == pid),
  so a reused PID can never be group-killed.
- A protocol-transparent TCP gateway, so clients keep one address while model
  runtimes move behind it.
- An MCP server, so agents drive the node as tools rather than as a CLI.

One binary, six direct dependencies, no async runtime outside the MCP surface.

Published from the internal monorepo with a fresh history. The private
development tree keeps its own history; nothing here carries it.
This commit is contained in:
Karti Tripathi
2026-08-03 16:22:21 -07:00
commit 91a47fb42c
38 changed files with 5086 additions and 0 deletions
+22
View File
@@ -0,0 +1,22 @@
apiVersion: lumbridge/v1
kind: Scene
metadata:
name: studio
version: 2
description: "Live voice assistant — ears, brain, mouth, and music."
author: karti
tags: [assistant, voice, always-on]
# ears 5 + brain 66 + voice 4 + music (XL 4B sft) 24 = 99 GB declared.
# Empirically validated on the reference node: all four load with ~20 GB MemAvailable free.
# budget bumped to 108 (box is 121 GB) so the Governor admits the measured-safe set.
models:
- ears
- brain
- voice
- music
budget_gb: 108
activation:
order: footprint-asc # small first, so brain's big load spike lands last
wait_healthy: true