Files
bench/README.md
T
Karti Tripathi 006feee0f7
CI / verify (push) Successful in 24s
CI / deploy (push) Failing after 1m14s
Lumbridge Bench
2026-08-04 00:44:07 -07:00

4.4 KiB
Raw Blame History

Lumbridge Bench

Private evaluation and serving-performance evidence for models running on Lumbridge Compute.

Two questions, one score card:

  1. Is this model good enough on the work we actually do?
  2. Is it fast enough on our own hardware to put in front of that work?

Most eval frameworks answer only the first. A self-hosting decision needs both, so every run produces quality scores and a serving perf profile for the same (model × host × quant × serving config) target.

Place in Lumbridge

Lumbridge Compute owns nodes, model artifacts, Scenes, and workload state. Lumbridge Bench owns measurement truth: versioned Evals, exact target snapshots, immutable runs, comparisons, and selected public score cards. The kbench command remains the compatible runner name.

Why this exists

your-node currently serves Qwen3.6-35B-A3B-NVFP4 behind the brain and local-moe aliases — in production, with nothing measuring whether that was the right call. This bench exists to answer that, and to be the eval half of the loop when we fine-tune our own model.

Install

uv venv --python 3.12
uv pip install -e ".[reference]"

Use

kbench targets                 # what can be benchmarked
kbench tasks                   # what it gets benchmarked on
kbench snapshot your-node        # live serving flags, for models.yaml

kbench perf brain              # perf sweep only (no task authoring needed)
kbench run brain --limit 50    # quality + perf -> results/*.json
kbench results                 # every score card recorded so far
kbench card results/<file>     # re-render one

Raw Inspect logs land in logs/ (gitignored). View them with:

.venv/bin/inspect view --log-dir logs

Layout

kbench/
  registry/     targets: (model × host × quant × serving config × checkpoint)
  tasks/        task catalog — reference tier and signal tier
  scorers/      custom + model-graded scorers
  perf.py       async load generator, any OpenAI-compatible endpoint
  results.py    the committed result schema
  run.py        orchestration: target + tasks + perf -> score card
data/
  public/       a few example samples, shown on the site
  private/      the real holdout — never leaves this repo
  canary/       GUID-tagged samples for proving contamination later
results/        one JSON per run, committed. the git history IS the record.
docs/           DECISIONS.md — why any of this is shaped the way it is

Two tiers of task

tier source purpose
reference inspect_evals, unmodified calibration + smoke test. Public, likely contaminated. Not a ranking.
signal ours, private the actual measurement

If a reference score lands far from published values, the harness is broken, not the model. That is what it is for.

Rules that are not negotiable

  • The private eval set never becomes public. Not a secrets concern primarily — a public test set gets scraped into the next training corpus and stops measuring anything. See docs/DECISIONS.md#d3.
  • Scrub samples at authoring time. No keys, tokens, tailscale addresses, or customer data. Repo privacy is defence in depth, not the control.
  • Never publish a capped run as a full run. --limit is recorded in the result and surfaced on the card.
  • Results are only comparable across identical serving flags. Use kbench snapshot rather than trusting memory.

Manual model suggestions

The public site accepts suggestions, not executable submissions. A suggestion records a Hugging Face reference and optional notes in the Lumbridge control plane for Karti to review. Nothing downloads or runs automatically; Karti verifies the model card, license, files, compatibility, footprint, and serving path before manually scheduling any run.

Built on Inspect AI as a dependency, not a fork — see docs/DECISIONS.md#d1.

License

Apache-2.0 — see LICENSE and NOTICE.

The private signal set, the contamination canary, and the sample seeder are not published. Publishing an eval destroys it: the signal set is worthless once it has been crawled, and a published canary GUID is reproduced by every model that ever read the repo, so the detector would report contamination for everyone forever. kbench ships with the public example family; you author your own private families locally.