This commit is contained in:
@@ -0,0 +1,40 @@
|
||||
# Lumbridge Compute evaluations (`lumbridge/v1`)
|
||||
|
||||
Lumbridge Compute evaluates the model configuration that is actually serving: weights,
|
||||
quantization, context, runtime, parsers, and speculative decoder. A model name
|
||||
without its serving configuration is not a reproducible benchmark target.
|
||||
|
||||
```bash
|
||||
lumbridge-compute eval ls
|
||||
lumbridge-compute eval run smoke
|
||||
lumbridge-compute eval run performance --model brain --repeat 5
|
||||
lumbridge-compute eval run finance-core --base-url http://your-node:8001/v1
|
||||
```
|
||||
|
||||
Suites live in `evals/*.eval.yaml`. Each case has a stable id, prompt, category,
|
||||
generation limit, and deterministic assertions. Runs produce append-only JSON in
|
||||
`eval-results/` with raw outputs and per-sample metrics.
|
||||
|
||||
## Metrics
|
||||
|
||||
- **TTFT**: wall time until the first streamed content or reasoning token.
|
||||
- **Prefill tok/s (approximate)**: API-reported prompt tokens divided by TTFT.
|
||||
This is client-observed and includes queueing/scheduling; server-native prefill
|
||||
metrics should be added as a separate source rather than conflated with it.
|
||||
- **Decode tok/s**: completion tokens divided by time after the first token.
|
||||
- **Score**: share of samples satisfying every declared assertion.
|
||||
|
||||
Performance runs should include warmups in automation and record hardware, Lumbridge Compute
|
||||
scene, runtime version, model revision, and cold/warm cache state. The v1 artifact
|
||||
is deliberately local and portable; a future registry can ingest the same JSON.
|
||||
|
||||
## Boundary with Bench, Arena, and Forge
|
||||
|
||||
Compute's native suites are post-activation health and performance checks for the exact
|
||||
configuration serving on a node. They do not grow into another training harness. Bench owns
|
||||
longitudinal model evidence, Arena owns task distributions and rewards on upstream Prime
|
||||
Intellect Verifiers, and Forge records any Prime-RL handoff and returned training artifact.
|
||||
|
||||
Compute resumes ownership only after an approved checkpoint has been returned and verified:
|
||||
local transfer, footprint admission, registry promotion, scene scheduling, and serving-process
|
||||
lifecycle. The training framework owns optimization and distributed execution.
|
||||
Reference in New Issue
Block a user