Files
compute/docs/evals.md
T
Karti Tripathi a4490ec80e
ci / rust (push) Successful in 2m26s
Lumbridge Compute
Governed compute for unified-memory AI hardware — machines where CPU and GPU
share one pool and there is no separate VRAM allocation to bounce off.
Over-commit that pool and the box thrashes and wedges, SSH and ping included,
before the OOM killer gets a turn.

Compute does not run inference. It supervises the servers that do: admission
control against both a declared budget and what the machine actually has free,
a 1 Hz watchdog that stops the newest model before thrash, Scenes activated as
one transactional unit with rollback, process ownership bound to
(boot_id, pid, start_time_ticks, pgid) so a reused PID can never be
group-killed, a protocol-transparent gateway, an MCP server, and a read-only
HTTP API for dashboards.

Registry footprints in this release are measured on a live node rather than
estimated.

One binary, six direct dependencies. Apache-2.0.

Generated by scripts/publish-compute.sh, which refuses to publish a tree it
cannot prove clean.
2026-08-03 22:23:56 -07:00

1.8 KiB

Lumbridge Compute evaluations (lumbridge/v1)

Lumbridge Compute evaluates the model configuration that is actually serving: weights, quantization, context, runtime, parsers, and speculative decoder. A model name without its serving configuration is not a reproducible benchmark target.

lumbridge-compute eval ls
lumbridge-compute eval run smoke
lumbridge-compute eval run performance --model brain --repeat 5
lumbridge-compute eval run finance-core --base-url http://your-node:8001/v1

Suites live in evals/*.eval.yaml. Each case has a stable id, prompt, category, generation limit, and deterministic assertions. Runs produce append-only JSON in eval-results/ with raw outputs and per-sample metrics.

Metrics

  • TTFT: wall time until the first streamed content or reasoning token.
  • Prefill tok/s (approximate): API-reported prompt tokens divided by TTFT. This is client-observed and includes queueing/scheduling; server-native prefill metrics should be added as a separate source rather than conflated with it.
  • Decode tok/s: completion tokens divided by time after the first token.
  • Score: share of samples satisfying every declared assertion.

Performance runs should include warmups in automation and record hardware, Lumbridge Compute scene, runtime version, model revision, and cold/warm cache state. The v1 artifact is deliberately local and portable; a future registry can ingest the same JSON.

Eval-to-training bridge

Capability suites should graduate into environment packages containing a dataset, harness, and reward function. That common contract can be adapted to Prime Intellect verifiers for evaluation, synthetic-data generation, SFT, or RL with prime-rl. Lumbridge Compute owns scene scheduling, memory admission, checkpoints, and process lifecycle; the training framework owns optimization and distributed execution.