Governed compute for unified-memory AI hardware — machines where CPU and GPU share one pool and there is no separate VRAM allocation to bounce off. Over-commit that pool and the box thrashes and wedges, SSH and ping included, before the OOM killer gets a turn. Compute does not run inference. It supervises the servers that do: admission control against both a declared budget and what the machine actually has free, a 1 Hz watchdog that stops the newest model before thrash, Scenes activated as one transactional unit with rollback, process ownership bound to (boot_id, pid, start_time_ticks, pgid) so a reused PID can never be group-killed, a protocol-transparent gateway, an MCP server, and a read-only HTTP API for dashboards. Registry footprints in this release are measured on a live node rather than estimated. One binary, six direct dependencies. Apache-2.0. Generated by scripts/publish-compute.sh, which refuses to publish a tree it cannot prove clean.
This commit is contained in:
@@ -0,0 +1,21 @@
|
||||
apiVersion: lumbridge/v1
|
||||
kind: Scene
|
||||
metadata:
|
||||
name: voice-laguna
|
||||
version: 1
|
||||
description: "High-capability voice assistant — Nemotron ASR, Laguna S 2.1, and Chatterbox."
|
||||
author: karti
|
||||
tags: [assistant, voice, coding, reasoning]
|
||||
|
||||
# Laguna replaces Qwen and intentionally excludes ACE-Step. The 108 GB scene
|
||||
# budget matches the empirically safe studio ceiling on the reference node while retaining
|
||||
# Lumbridge Compute's separate 8 GB admission margin.
|
||||
models:
|
||||
- ears
|
||||
- brain-laguna
|
||||
- voice
|
||||
|
||||
budget_gb: 108
|
||||
activation:
|
||||
order: footprint-asc
|
||||
wait_healthy: true
|
||||
Reference in New Issue
Block a user