Lumbridge Compute
ci / rust (push) Successful in 2m26s

Governed compute for unified-memory AI hardware — machines where CPU and GPU
share one pool and there is no separate VRAM allocation to bounce off.
Over-commit that pool and the box thrashes and wedges, SSH and ping included,
before the OOM killer gets a turn.

Compute does not run inference. It supervises the servers that do: admission
control against both a declared budget and what the machine actually has free,
a 1 Hz watchdog that stops the newest model before thrash, Scenes activated as
one transactional unit with rollback, process ownership bound to
(boot_id, pid, start_time_ticks, pgid) so a reused PID can never be
group-killed, a protocol-transparent gateway, an MCP server, and a read-only
HTTP API for dashboards.

Registry footprints in this release are measured on a live node rather than
estimated.

One binary, six direct dependencies. Apache-2.0.

Generated by scripts/publish-compute.sh, which refuses to publish a tree it
cannot prove clean.
This commit is contained in:
Karti Tripathi
2026-08-03 22:23:56 -07:00
commit a4490ec80e
40 changed files with 6030 additions and 0 deletions
+26
View File
@@ -0,0 +1,26 @@
apiVersion: lumbridge/v1
kind: Scene
metadata:
name: voice-qwen
version: 2
description: "Superseded by a newer scene on this node. Kept only because it is still the recorded fallback."
author: karti
tags: [assistant, voice, low-latency, deprecated]
# DEPRECATED as of 2026-08-02, superseded by a newer scene. Retained solely
# because last_known_good still points at it; delete once the fallback rolls forward.
#
# Its ordering was made brain-first to match its successor. `brain` now runs MTP, whose
# KV-cache profiling spike is not bounded by gpu-memory-utilization; the old
# footprint-asc order would start voice+ears first and leave brain short enough
# to trip the 3GB watchdog floor. A fallback that cannot come up is worse than
# no fallback.
models:
- brain
- ears
- voice
budget_gb: 100
activation:
order: listed
wait_healthy: true