Files
compute/docs/roadmap.md
T
Karti Tripathi 75aa63d737
ci / rust (push) Successful in 2m26s
Lumbridge Compute
Governed compute for unified-memory AI hardware — the machines where CPU and
GPU share one pool and there is no separate VRAM allocation to bounce off.
Over-commit that pool and the box thrashes and wedges, SSH and ping included,
before the OOM killer gets a turn.

Compute does not run inference. It supervises the servers that do:

- Admission control against two ceilings: a declared budget, and what the
  machine actually has free. The refusal is the feature.
- A 1 Hz watchdog on MemAvailable that stops the newest model before thrash,
  and defers to a Scene transition rather than racing it.
- Scenes: named sets of models activated as one transactional unit, with
  pre-flight validation and rollback to the previously active Scene on
  failure. Scenes reference model ids, never weight paths or commands, so a
  Scene obtained from elsewhere cannot introduce code.
- Process ownership bound to (boot_id, pid, start_time_ticks, pgid == pid),
  so a reused PID can never be group-killed.
- A protocol-transparent TCP gateway, so clients keep one address while model
  runtimes move behind it.
- An MCP server, so agents drive the node as tools rather than as a CLI.

One binary, no async runtime outside the MCP surface. Apache-2.0.

Generated from the internal monorepo by scripts/publish-compute.sh, which
refuses to publish a tree it cannot prove clean.
2026-08-03 16:47:06 -07:00

1.6 KiB

Lumbridge Compute roadmap

Lumbridge Compute is the resident AI workload control plane for unified-memory machines.

Core contracts

  1. Models — local vetted launch recipes with immutable revisions and measured footprints.
  2. Scenes — named workload sets activated through admission control.
  3. Evals — versioned datasets, prompts, harnesses, assertions, and portable results.
  4. Artifacts — resumable weights, adapters, checkpoints, datasets, and result bundles.
  5. Jobs — inference, evaluation, download, conversion, SFT, and RL lifecycles.

Near term

  • Resumable, checksum-verified lumbridge-compute model pull with leases and progress state.
  • Eval comparison, regression thresholds, warmup policy, concurrency/load tests, server-native Prometheus metrics, and hardware/config provenance.
  • Scene hooks for preflight, post-activation eval gates, rollback, and schedules.
  • Gateway aliases so clients follow the active scene without configuration changes.
  • Training scenes and a Prime Intellect adapter for Qwen 0.6B/1.7B experiments.
  • Checkpoint discovery, pause/resume, retention, and promotion into the model registry.

Open-source readiness

  • Keep public scene/eval manifests command-free; commands remain in the trusted local registry.
  • Add CI across x86_64 and aarch64, unit/integration tests, security policy, contribution guide, code of conduct, changelog, and versioned JSON schemas.
  • Remove private hostnames, voices, paths, and finance datasets from public fixtures; ship generic examples and keep personal overlays outside the repository.