4.4 KiB
Lumbridge Bench
Private evaluation and serving-performance evidence for models running on Lumbridge Compute.
Two questions, one score card:
- Is this model good enough on the work we actually do?
- Is it fast enough on our own hardware to put in front of that work?
Most eval frameworks answer only the first. A self-hosting decision needs
both, so every run produces quality scores and a serving perf profile for
the same (model × host × quant × serving config) target.
Place in Lumbridge
Lumbridge Compute owns nodes, model artifacts, Scenes, and workload state. Lumbridge Bench owns measurement truth: versioned Evals, exact target snapshots, immutable runs, comparisons, and selected public score cards. The kbench command remains the compatible runner name.
Why this exists
your-node currently serves Qwen3.6-35B-A3B-NVFP4 behind the brain and
local-moe aliases — in production, with nothing measuring whether
that was the right call. This bench exists to answer that, and to be the eval
half of the loop when we fine-tune our own model.
Install
uv venv --python 3.12
uv pip install -e ".[reference]"
Use
kbench targets # what can be benchmarked
kbench tasks # what it gets benchmarked on
kbench snapshot your-node # live serving flags, for models.yaml
kbench perf brain # perf sweep only (no task authoring needed)
kbench run brain --limit 50 # quality + perf -> results/*.json
kbench results # every score card recorded so far
kbench card results/<file> # re-render one
Raw Inspect logs land in logs/ (gitignored). View them with:
.venv/bin/inspect view --log-dir logs
Layout
kbench/
registry/ targets: (model × host × quant × serving config × checkpoint)
tasks/ task catalog — reference tier and signal tier
scorers/ custom + model-graded scorers
perf.py async load generator, any OpenAI-compatible endpoint
results.py the committed result schema
run.py orchestration: target + tasks + perf -> score card
data/
public/ a few example samples, shown on the site
private/ the real holdout — never leaves this repo
canary/ GUID-tagged samples for proving contamination later
results/ one JSON per run, committed. the git history IS the record.
docs/ DECISIONS.md — why any of this is shaped the way it is
Two tiers of task
| tier | source | purpose |
|---|---|---|
| reference | inspect_evals, unmodified |
calibration + smoke test. Public, likely contaminated. Not a ranking. |
| signal | ours, private | the actual measurement |
If a reference score lands far from published values, the harness is broken, not the model. That is what it is for.
Rules that are not negotiable
- The private eval set never becomes public. Not a secrets concern
primarily — a public test set gets scraped into the next training corpus and
stops measuring anything. See
docs/DECISIONS.md#d3. - Scrub samples at authoring time. No keys, tokens, tailscale addresses, or customer data. Repo privacy is defence in depth, not the control.
- Never publish a capped run as a full run.
--limitis recorded in the result and surfaced on the card. - Results are only comparable across identical serving flags. Use
kbench snapshotrather than trusting memory.
Manual model suggestions
The public site accepts suggestions, not executable submissions. A suggestion records a Hugging Face reference and optional notes in the Lumbridge control plane for Karti to review. Nothing downloads or runs automatically; Karti verifies the model card, license, files, compatibility, footprint, and serving path before manually scheduling any run.
Built on Inspect AI as a
dependency, not a fork — see docs/DECISIONS.md#d1.
License
Apache-2.0 — see LICENSE and NOTICE.
The private signal set, the contamination canary, and the sample seeder are not
published. Publishing an eval destroys it: the signal set is worthless once it has been
crawled, and a published canary GUID is reproduced by every model that ever read the
repo, so the detector would report contamination for everyone forever. kbench ships
with the public example family; you author your own private families locally.