Lumbridge Bench
CI / verify (push) Successful in 24s
CI / deploy (push) Failing after 1m14s

This commit is contained in:
Karti Tripathi
2026-08-04 00:44:07 -07:00
commit 006feee0f7
65 changed files with 13516 additions and 0 deletions
+106
View File
@@ -0,0 +1,106 @@
# Lumbridge Bench
Private evaluation and serving-performance evidence for models running on Lumbridge Compute.
Two questions, one score card:
1. **Is this model good enough** on the work we actually do?
2. **Is it fast enough** on our own hardware to put in front of that work?
Most eval frameworks answer only the first. A self-hosting decision needs
both, so every run produces quality scores *and* a serving perf profile for
the same `(model × host × quant × serving config)` target.
## Place in Lumbridge
Lumbridge Compute owns nodes, model artifacts, Scenes, and workload state. Lumbridge Bench owns measurement truth: versioned Evals, exact target snapshots, immutable runs, comparisons, and selected public score cards. The `kbench` command remains the compatible runner name.
## Why this exists
`your-node` currently serves `Qwen3.6-35B-A3B-NVFP4` behind the `brain` and
`local-moe` aliases — in production, with nothing measuring whether
that was the right call. This bench exists to answer that, and to be the eval
half of the loop when we fine-tune our own model.
## Install
```bash
uv venv --python 3.12
uv pip install -e ".[reference]"
```
## Use
```bash
kbench targets # what can be benchmarked
kbench tasks # what it gets benchmarked on
kbench snapshot your-node # live serving flags, for models.yaml
kbench perf brain # perf sweep only (no task authoring needed)
kbench run brain --limit 50 # quality + perf -> results/*.json
kbench results # every score card recorded so far
kbench card results/<file> # re-render one
```
Raw Inspect logs land in `logs/` (gitignored). View them with:
```bash
.venv/bin/inspect view --log-dir logs
```
## Layout
```
kbench/
registry/ targets: (model × host × quant × serving config × checkpoint)
tasks/ task catalog — reference tier and signal tier
scorers/ custom + model-graded scorers
perf.py async load generator, any OpenAI-compatible endpoint
results.py the committed result schema
run.py orchestration: target + tasks + perf -> score card
data/
public/ a few example samples, shown on the site
private/ the real holdout — never leaves this repo
canary/ GUID-tagged samples for proving contamination later
results/ one JSON per run, committed. the git history IS the record.
docs/ DECISIONS.md — why any of this is shaped the way it is
```
## Two tiers of task
| tier | source | purpose |
|---|---|---|
| **reference** | `inspect_evals`, unmodified | calibration + smoke test. Public, likely contaminated. Not a ranking. |
| **signal** | ours, private | the actual measurement |
If a reference score lands far from published values, the harness is broken,
not the model. That is what it is for.
## Rules that are not negotiable
- **The private eval set never becomes public.** Not a secrets concern
primarily — a public test set gets scraped into the next training corpus and
stops measuring anything. See `docs/DECISIONS.md#d3`.
- **Scrub samples at authoring time.** No keys, tokens, tailscale addresses,
or customer data. Repo privacy is defence in depth, not the control.
- **Never publish a capped run as a full run.** `--limit` is recorded in the
result and surfaced on the card.
- **Results are only comparable across identical serving flags.** Use
`kbench snapshot` rather than trusting memory.
## Manual model suggestions
The public site accepts suggestions, not executable submissions. A suggestion records a Hugging Face reference and optional notes in the Lumbridge control plane for Karti to review. Nothing downloads or runs automatically; Karti verifies the model card, license, files, compatibility, footprint, and serving path before manually scheduling any run.
Built on [Inspect AI](https://github.com/UKGovernmentBEIS/inspect_ai) as a
dependency, not a fork — see `docs/DECISIONS.md#d1`.
## License
Apache-2.0 — see [LICENSE](LICENSE) and [NOTICE](NOTICE).
The private signal set, the contamination canary, and the sample seeder are **not**
published. Publishing an eval destroys it: the signal set is worthless once it has been
crawled, and a published canary GUID is reproduced by every model that ever read the
repo, so the detector would report contamination for everyone forever. `kbench` ships
with the public example family; you author your own private families locally.