Files
compute/docs/evals.md
T
2026-08-31 16:09:08 -07:00

2.1 KiB

Lumbridge Compute evaluations (lumbridge/v1)

Lumbridge Compute evaluates the model configuration that is actually serving: weights, quantization, context, runtime, parsers, and speculative decoder. A model name without its serving configuration is not a reproducible benchmark target.

lumbridge-compute eval ls
lumbridge-compute eval run smoke
lumbridge-compute eval run performance --model brain --repeat 5
lumbridge-compute eval run finance-core --base-url http://your-node:8001/v1

Suites live in evals/*.eval.yaml. Each case has a stable id, prompt, category, generation limit, and deterministic assertions. Runs produce append-only JSON in eval-results/ with raw outputs and per-sample metrics.

Metrics

  • TTFT: wall time until the first streamed content or reasoning token.
  • Prefill tok/s (approximate): API-reported prompt tokens divided by TTFT. This is client-observed and includes queueing/scheduling; server-native prefill metrics should be added as a separate source rather than conflated with it.
  • Decode tok/s: completion tokens divided by time after the first token.
  • Score: share of samples satisfying every declared assertion.

Performance runs should include warmups in automation and record hardware, Lumbridge Compute scene, runtime version, model revision, and cold/warm cache state. The v1 artifact is deliberately local and portable; a future registry can ingest the same JSON.

Boundary with Bench, Arena, and Forge

Compute's native suites are post-activation health and performance checks for the exact configuration serving on a node. They do not grow into another training harness. Bench owns longitudinal model evidence, Arena owns task distributions and rewards on upstream Prime Intellect Verifiers, and Forge records any Prime-RL handoff and returned training artifact.

Compute resumes ownership only after an approved checkpoint has been returned and verified: local transfer, footprint admission, registry promotion, scene scheduling, and serving-process lifecycle. The training framework owns optimization and distributed execution.