a4490ec80e
ci / rust (push) Successful in 2m26s
Governed compute for unified-memory AI hardware — machines where CPU and GPU share one pool and there is no separate VRAM allocation to bounce off. Over-commit that pool and the box thrashes and wedges, SSH and ping included, before the OOM killer gets a turn. Compute does not run inference. It supervises the servers that do: admission control against both a declared budget and what the machine actually has free, a 1 Hz watchdog that stops the newest model before thrash, Scenes activated as one transactional unit with rollback, process ownership bound to (boot_id, pid, start_time_ticks, pgid) so a reused PID can never be group-killed, a protocol-transparent gateway, an MCP server, and a read-only HTTP API for dashboards. Registry footprints in this release are measured on a live node rather than estimated. One binary, six direct dependencies. Apache-2.0. Generated by scripts/publish-compute.sh, which refuses to publish a tree it cannot prove clean.
23 lines
627 B
YAML
23 lines
627 B
YAML
apiVersion: lumbridge/v1
|
|
kind: Scene
|
|
metadata:
|
|
name: studio
|
|
version: 2
|
|
description: "Live voice assistant — ears, brain, mouth, and music."
|
|
author: karti
|
|
tags: [assistant, voice, always-on]
|
|
|
|
# ears 5 + brain 66 + voice 4 + music (XL 4B sft) 24 = 99 GB declared.
|
|
# Empirically validated on the reference node: all four load with ~20 GB MemAvailable free.
|
|
# budget bumped to 108 (box is 121 GB) so the Governor admits the measured-safe set.
|
|
models:
|
|
- ears
|
|
- brain
|
|
- voice
|
|
- music
|
|
|
|
budget_gb: 108
|
|
activation:
|
|
order: footprint-asc # small first, so brain's big load spike lands last
|
|
wait_healthy: true
|