Files
compute/docs/scene-spec.md
T
Karti Tripathi 75aa63d737
ci / rust (push) Successful in 2m26s
Lumbridge Compute
Governed compute for unified-memory AI hardware — the machines where CPU and
GPU share one pool and there is no separate VRAM allocation to bounce off.
Over-commit that pool and the box thrashes and wedges, SSH and ping included,
before the OOM killer gets a turn.

Compute does not run inference. It supervises the servers that do:

- Admission control against two ceilings: a declared budget, and what the
  machine actually has free. The refusal is the feature.
- A 1 Hz watchdog on MemAvailable that stops the newest model before thrash,
  and defers to a Scene transition rather than racing it.
- Scenes: named sets of models activated as one transactional unit, with
  pre-flight validation and rollback to the previously active Scene on
  failure. Scenes reference model ids, never weight paths or commands, so a
  Scene obtained from elsewhere cannot introduce code.
- Process ownership bound to (boot_id, pid, start_time_ticks, pgid == pid),
  so a reused PID can never be group-killed.
- A protocol-transparent TCP gateway, so clients keep one address while model
  runtimes move behind it.
- An MCP server, so agents drive the node as tools rather than as a CLI.

One binary, no async runtime outside the MCP surface. Apache-2.0.

Generated from the internal monorepo by scripts/publish-compute.sh, which
refuses to publish a tree it cannot prove clean.
2026-08-03 16:47:06 -07:00

179 lines
6.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Lumbridge Compute Scene Spec (`lumbridge/v1`)
This is the public contract. Everything else — the Governor, the CLI, the
scheduler — can be refactored freely. This format cannot, once people publish
scenes against it. So it is deliberately small.
## The core idea: scenes reference ids, not weights
A **scene** is a manifest listing **model ids**. A **model registry** resolves each
id to actual weights + a launch command. The registry churns as models evolve (new
quants, new backends, bigger context); the scene stays stable.
```
scene (stable, shareable) registry (local, evolves)
───────────────────────── ──────────────────────────────
models: [ears, brain, voice] ──▶ brain → ~/models/Qwen3.6-35B-A3B-NVFP4-Fast
→ vllm serve, gmu 0.55, flashinfer...
```
This is the same decoupling the original harness used: it kept the id `brain` and a
stable `served_name` while the underlying weights swapped dense-27B → 35B-A3B MoE,
and downstream agents never noticed. The spec formalizes that as the mechanism that
lets scenes "update over time as models evolve" without breaking anyone.
### Why this also solves scene-sharing security
A published scene contains **only ids and parameters — never shell commands.** The
launch commands live in your *local, vetted* registry. So activating a downloaded
scene can only ever start models your own registry already trusts. If a scene
references an id you don't have, Lumbridge Compute asks you to add it to your registry, showing
the launch command for review — an explicit opt-in, not silent remote code
execution. Declarative-by-construction; there is no field in which a scene can smuggle
a command.
---
## Scene manifest
`scenes/studio.scene.yaml`
```yaml
apiVersion: lumbridge/v1
kind: Scene
metadata:
name: studio
version: 3 # bump on any change; a published scene is reproducible
description: "Live voice assistant — ears, brain, mouth, and music."
author: karti
tags: [assistant, voice, always-on]
models: # stable ids; the registry resolves each
- ears # ASR
- brain # MoE LLM
- voice # TTS
- music # ACE-Step
budget_gb: 100 # optional; overrides the Governor's global budget
activation:
order: footprint-asc # small models first so the big load spike lands last
wait_healthy: true # block until each model's health check passes
```
`scenes/darkroom.scene.yaml`
```yaml
apiVersion: lumbridge/v1
kind: Scene
metadata:
name: darkroom
version: 1
description: "Overnight image farm — drops the brain to make room for FLUX.2-dev."
tags: [image, overnight, unattended]
models:
- ears
- voice
- music
- image # FLUX.2-dev — only fits because `brain` is not in this scene
activation:
order: footprint-asc
wait_healthy: true
```
### Fields
| Field | Required | Meaning |
|---|---|---|
| `apiVersion` | ✔ | `lumbridge/v1`. |
| `kind` | ✔ | `Scene`. |
| `metadata.name` | ✔ | Unique scene name; the CLI handle. |
| `metadata.version` | ✔ | Integer, bumped on any change. Reproducibility. |
| `metadata.description` | ✔ | One line, shown in `lumbridge-compute scene ls`. |
| `metadata.tags` | | For the (future) registry search. |
| `models` | ✔ | Ordered list of model ids resolved via the registry. |
| `budget_gb` | | Per-scene budget override; defaults to the global Governor budget. |
| `activation.order` | | `footprint-asc` (default) \| `listed`. |
| `activation.wait_healthy` | | Default `true`. Block until health checks pass. |
A scene **never** contains: weight paths, shell commands, or GPU flags. Those live in
the registry. This is load-bearing for both stability and security.
---
## Model registry
`registry/models.yaml` — local to each box, evolves freely. Ids are the stable
contract; everything under `serve` can change.
```yaml
apiVersion: lumbridge/v1
kind: Registry
version: 1
models:
brain:
name: "Qwen3.6 35B-A3B MoE (NVFP4)"
footprint_gb: 66 # worst-case unified memory once serving (weights + KV + encoder)
channel: stable # stable | latest — how aggressively to track new weights
health: "http://localhost:8001/v1/models"
serve:
kind: vllm
port: 8001
weights: "~/models/Qwen3.6-35B-A3B-NVFP4-Fast"
served_name: # stable aliases so downstream clients survive a weight swap
- brain
- "local-moe"
- "unsloth/Qwen3.6-35B-A3B-NVFP4-Fast"
args:
max-model-len: 65536
kv-cache-dtype: fp8
gpu-memory-utilization: 0.55
enforce-eager: true
moe-backend: flashinfer_b12x # Unsloth DGX Spark recipe; critical for speed
limit-mm-per-prompt: '{"image": 0, "video": 0}'
env:
CUTE_DSL_ARCH: sm_121a
image:
name: "FLUX.2-dev (FP8)"
footprint_gb: 32
channel: stable
health: "http://localhost:8007/health"
serve:
kind: diffusers # not vllm — a separate runtime (ComfyUI/diffusers)
port: 8007
weights: "~/models/FLUX.2-dev"
args: { dtype: fp8 }
# ears / voice / music elaborated the same way (ASR, Chatterbox, ACE-Step).
```
### `channel`: how scenes track evolving weights
- `stable` — pin the exact `weights` path. Reproducible; you update deliberately.
- `latest` — the registry may resolve to a newer quant of the same model family
(e.g. a fresh NVFP4 build) on activation. Bleeding edge; use for your own box, not
for scenes you publish for others.
The scene picks the *id*; the registry's `channel` decides how much the weights are
allowed to drift underneath it. That's the whole "scenes evolve as models evolve"
story, made explicit and controllable.
---
## Governor interaction
On `activate`, Lumbridge Compute computes the diff between the running set and the target
scene's `models`, then:
1. **Stops** running models not in the scene (frees their footprint first).
2. **Starts** the scene's models in `activation.order`, each passing **admission
control** against `budget_gb` before launch.
3. Waits for health if `wait_healthy`.
The watchdog runs throughout, unchanged — the safety net if any `footprint_gb` is
wrong. A scene can never talk the Governor into over-committing; admission control is
not bypassable by a scene.