Governed compute for unified-memory AI hardware — machines where CPU and GPU share one pool and there is no separate VRAM allocation to bounce off. Over-commit that pool and the box thrashes and wedges, SSH and ping included, before the OOM killer gets a turn. Compute does not run inference. It supervises the servers that do: - Admission control against two ceilings — a declared budget, and what the machine actually has free. The refusal is the feature. - A 1 Hz watchdog on MemAvailable that stops the newest model before thrash, and defers to a Scene transition rather than racing it. - Scenes: named sets of models activated as one transactional unit, with pre-flight validation and rollback to the previously active Scene on failure. Scenes reference model ids, never weight paths or commands, so a Scene obtained from elsewhere cannot introduce code. - Process ownership bound to (boot_id, pid, start_time_ticks, pgid == pid), so a reused PID can never be group-killed. - A protocol-transparent TCP gateway, so clients keep one address while model runtimes move behind it. - An MCP server, so agents drive the node as tools. - A read-only HTTP API for dashboards: loopback by default, CORS off unless an origin is named, and it never mutates. One binary, six direct dependencies, no async runtime outside the MCP and API surfaces. Apache-2.0. Generated from the internal monorepo by scripts/publish-compute.sh, which refuses to publish a tree it cannot prove clean.
This commit is contained in:
@@ -0,0 +1,27 @@
|
||||
apiVersion: lumbridge/v1
|
||||
kind: Scene
|
||||
metadata:
|
||||
name: darkroom
|
||||
version: 1
|
||||
description: "Overnight image farm — drops the brain to make room for FLUX.2-klein-9B."
|
||||
author: karti
|
||||
tags: [image, overnight, unattended]
|
||||
|
||||
# Drops `brain` so `image` (66 GB, on-hardware measured) fits. ears+voice+music+
|
||||
# image commits ~99 GB — tight against the 8 GB admission margin, so this scene
|
||||
# needs the full 108 GB ceiling (matching studio/voice-laguna), not the old 100.
|
||||
models:
|
||||
- ears
|
||||
- voice
|
||||
- music
|
||||
- image
|
||||
|
||||
budget_gb: 108
|
||||
activation:
|
||||
order: footprint-asc
|
||||
wait_healthy: true
|
||||
|
||||
# Scheduled scenes (planned): hand the box to darkroom overnight, back at 08:00.
|
||||
# schedule:
|
||||
# - activate: "03:00"
|
||||
# - handoff: "04:00" -> studio
|
||||
@@ -0,0 +1,20 @@
|
||||
apiVersion: lumbridge/v1
|
||||
kind: Scene
|
||||
metadata:
|
||||
name: music
|
||||
version: 1
|
||||
description: "Music studio — ACE-Step XL 4B sft, with ears and voice, no brain."
|
||||
author: karti
|
||||
tags: [music, creative]
|
||||
|
||||
# ears 5 + voice 4 + music 24 = 33 GB. Brain-free, so tons of headroom — the clean
|
||||
# scene for a music-generation demo when the MoE is not needed.
|
||||
models:
|
||||
- ears
|
||||
- voice
|
||||
- music
|
||||
|
||||
budget_gb: 100
|
||||
activation:
|
||||
order: footprint-asc
|
||||
wait_healthy: true
|
||||
@@ -0,0 +1,22 @@
|
||||
apiVersion: lumbridge/v1
|
||||
kind: Scene
|
||||
metadata:
|
||||
name: studio
|
||||
version: 2
|
||||
description: "Live voice assistant — ears, brain, mouth, and music."
|
||||
author: karti
|
||||
tags: [assistant, voice, always-on]
|
||||
|
||||
# ears 5 + brain 66 + voice 4 + music (XL 4B sft) 24 = 99 GB declared.
|
||||
# Empirically validated on the reference node: all four load with ~20 GB MemAvailable free.
|
||||
# budget bumped to 108 (box is 121 GB) so the Governor admits the measured-safe set.
|
||||
models:
|
||||
- ears
|
||||
- brain
|
||||
- voice
|
||||
- music
|
||||
|
||||
budget_gb: 108
|
||||
activation:
|
||||
order: footprint-asc # small first, so brain's big load spike lands last
|
||||
wait_healthy: true
|
||||
@@ -0,0 +1,22 @@
|
||||
apiVersion: lumbridge/v1
|
||||
kind: Scene
|
||||
metadata:
|
||||
name: voice-gemma
|
||||
version: 1
|
||||
description: "Efficient voice assistant — Nemotron ASR, Gemma 4 31B, and Chatterbox."
|
||||
author: karti
|
||||
tags: [assistant, voice, efficient, multimodal]
|
||||
|
||||
# Gemma's mostly-sliding-window attention keeps its footprint well under brain
|
||||
# and brain-laguna's, leaving real headroom in the budget beyond Lumbridge Compute's 8 GB
|
||||
# admission margin — useful until brain-gemma's footprint is confirmed on
|
||||
# real hardware.
|
||||
models:
|
||||
- ears
|
||||
- brain-gemma
|
||||
- voice
|
||||
|
||||
budget_gb: 80
|
||||
activation:
|
||||
order: footprint-asc
|
||||
wait_healthy: true
|
||||
@@ -0,0 +1,21 @@
|
||||
apiVersion: lumbridge/v1
|
||||
kind: Scene
|
||||
metadata:
|
||||
name: voice-laguna
|
||||
version: 1
|
||||
description: "High-capability voice assistant — Nemotron ASR, Laguna S 2.1, and Chatterbox."
|
||||
author: karti
|
||||
tags: [assistant, voice, coding, reasoning]
|
||||
|
||||
# Laguna replaces Qwen and intentionally excludes ACE-Step. The 108 GB scene
|
||||
# budget matches the empirically safe studio ceiling on the reference node while retaining
|
||||
# Lumbridge Compute's separate 8 GB admission margin.
|
||||
models:
|
||||
- ears
|
||||
- brain-laguna
|
||||
- voice
|
||||
|
||||
budget_gb: 108
|
||||
activation:
|
||||
order: footprint-asc
|
||||
wait_healthy: true
|
||||
@@ -0,0 +1,18 @@
|
||||
apiVersion: lumbridge/v1
|
||||
kind: Scene
|
||||
metadata:
|
||||
name: voice-qwen
|
||||
version: 1
|
||||
description: "Low-latency voice assistant — Nemotron ASR, Qwen MoE, and Chatterbox."
|
||||
author: karti
|
||||
tags: [assistant, voice, low-latency, always-on]
|
||||
|
||||
models:
|
||||
- ears
|
||||
- brain
|
||||
- voice
|
||||
|
||||
budget_gb: 100
|
||||
activation:
|
||||
order: footprint-asc
|
||||
wait_healthy: true
|
||||
Reference in New Issue
Block a user