75aa63d737
ci / rust (push) Successful in 2m26s
Governed compute for unified-memory AI hardware — the machines where CPU and GPU share one pool and there is no separate VRAM allocation to bounce off. Over-commit that pool and the box thrashes and wedges, SSH and ping included, before the OOM killer gets a turn. Compute does not run inference. It supervises the servers that do: - Admission control against two ceilings: a declared budget, and what the machine actually has free. The refusal is the feature. - A 1 Hz watchdog on MemAvailable that stops the newest model before thrash, and defers to a Scene transition rather than racing it. - Scenes: named sets of models activated as one transactional unit, with pre-flight validation and rollback to the previously active Scene on failure. Scenes reference model ids, never weight paths or commands, so a Scene obtained from elsewhere cannot introduce code. - Process ownership bound to (boot_id, pid, start_time_ticks, pgid == pid), so a reused PID can never be group-killed. - A protocol-transparent TCP gateway, so clients keep one address while model runtimes move behind it. - An MCP server, so agents drive the node as tools rather than as a CLI. One binary, no async runtime outside the MCP surface. Apache-2.0. Generated from the internal monorepo by scripts/publish-compute.sh, which refuses to publish a tree it cannot prove clean.
1.6 KiB
1.6 KiB
Lumbridge Compute roadmap
Lumbridge Compute is the resident AI workload control plane for unified-memory machines.
Core contracts
- Models — local vetted launch recipes with immutable revisions and measured footprints.
- Scenes — named workload sets activated through admission control.
- Evals — versioned datasets, prompts, harnesses, assertions, and portable results.
- Artifacts — resumable weights, adapters, checkpoints, datasets, and result bundles.
- Jobs — inference, evaluation, download, conversion, SFT, and RL lifecycles.
Near term
- Resumable, checksum-verified
lumbridge-compute model pullwith leases and progress state. - Eval comparison, regression thresholds, warmup policy, concurrency/load tests, server-native Prometheus metrics, and hardware/config provenance.
- Scene hooks for preflight, post-activation eval gates, rollback, and schedules.
- Gateway aliases so clients follow the active scene without configuration changes.
- Training scenes and a Prime Intellect adapter for Qwen 0.6B/1.7B experiments.
- Checkpoint discovery, pause/resume, retention, and promotion into the model registry.
Open-source readiness
- Keep public scene/eval manifests command-free; commands remain in the trusted local registry.
- Add CI across x86_64 and aarch64, unit/integration tests, security policy, contribution guide, code of conduct, changelog, and versioned JSON schemas.
- Remove private hostnames, voices, paths, and finance datasets from public fixtures; ship generic examples and keep personal overlays outside the repository.