# Lumbridge Compute evaluations (`lumbridge/v1`) Lumbridge Compute evaluates the model configuration that is actually serving: weights, quantization, context, runtime, parsers, and speculative decoder. A model name without its serving configuration is not a reproducible benchmark target. ```bash lumbridge-compute eval ls lumbridge-compute eval run smoke lumbridge-compute eval run performance --model brain --repeat 5 lumbridge-compute eval run finance-core --base-url http://your-node:8001/v1 ``` Suites live in `evals/*.eval.yaml`. Each case has a stable id, prompt, category, generation limit, and deterministic assertions. Runs produce append-only JSON in `eval-results/` with raw outputs and per-sample metrics. ## Metrics - **TTFT**: wall time until the first streamed content or reasoning token. - **Prefill tok/s (approximate)**: API-reported prompt tokens divided by TTFT. This is client-observed and includes queueing/scheduling; server-native prefill metrics should be added as a separate source rather than conflated with it. - **Decode tok/s**: completion tokens divided by time after the first token. - **Score**: share of samples satisfying every declared assertion. Performance runs should include warmups in automation and record hardware, Lumbridge Compute scene, runtime version, model revision, and cold/warm cache state. The v1 artifact is deliberately local and portable; a future registry can ingest the same JSON. ## Eval-to-training bridge Capability suites should graduate into environment packages containing a dataset, harness, and reward function. That common contract can be adapted to Prime Intellect `verifiers` for evaluation, synthetic-data generation, SFT, or RL with `prime-rl`. Lumbridge Compute owns scene scheduling, memory admission, checkpoints, and process lifecycle; the training framework owns optimization and distributed execution.