Files
compute/docs/roadmap.md
T
Karti Tripathi 75aa63d737
ci / rust (push) Successful in 2m26s
Lumbridge Compute
Governed compute for unified-memory AI hardware — the machines where CPU and
GPU share one pool and there is no separate VRAM allocation to bounce off.
Over-commit that pool and the box thrashes and wedges, SSH and ping included,
before the OOM killer gets a turn.

Compute does not run inference. It supervises the servers that do:

- Admission control against two ceilings: a declared budget, and what the
  machine actually has free. The refusal is the feature.
- A 1 Hz watchdog on MemAvailable that stops the newest model before thrash,
  and defers to a Scene transition rather than racing it.
- Scenes: named sets of models activated as one transactional unit, with
  pre-flight validation and rollback to the previously active Scene on
  failure. Scenes reference model ids, never weight paths or commands, so a
  Scene obtained from elsewhere cannot introduce code.
- Process ownership bound to (boot_id, pid, start_time_ticks, pgid == pid),
  so a reused PID can never be group-killed.
- A protocol-transparent TCP gateway, so clients keep one address while model
  runtimes move behind it.
- An MCP server, so agents drive the node as tools rather than as a CLI.

One binary, no async runtime outside the MCP surface. Apache-2.0.

Generated from the internal monorepo by scripts/publish-compute.sh, which
refuses to publish a tree it cannot prove clean.
2026-08-03 16:47:06 -07:00

30 lines
1.6 KiB
Markdown

# Lumbridge Compute roadmap
Lumbridge Compute is the resident AI workload control plane for unified-memory machines.
## Core contracts
1. **Models** — local vetted launch recipes with immutable revisions and measured footprints.
2. **Scenes** — named workload sets activated through admission control.
3. **Evals** — versioned datasets, prompts, harnesses, assertions, and portable results.
4. **Artifacts** — resumable weights, adapters, checkpoints, datasets, and result bundles.
5. **Jobs** — inference, evaluation, download, conversion, SFT, and RL lifecycles.
## Near term
- Resumable, checksum-verified `lumbridge-compute model pull` with leases and progress state.
- Eval comparison, regression thresholds, warmup policy, concurrency/load tests,
server-native Prometheus metrics, and hardware/config provenance.
- Scene hooks for preflight, post-activation eval gates, rollback, and schedules.
- Gateway aliases so clients follow the active scene without configuration changes.
- Training scenes and a Prime Intellect adapter for Qwen 0.6B/1.7B experiments.
- Checkpoint discovery, pause/resume, retention, and promotion into the model registry.
## Open-source readiness
- Keep public scene/eval manifests command-free; commands remain in the trusted local registry.
- Add CI across x86_64 and aarch64, unit/integration tests, security policy,
contribution guide, code of conduct, changelog, and versioned JSON schemas.
- Remove private hostnames, voices, paths, and finance datasets from public fixtures;
ship generic examples and keep personal overlays outside the repository.