3.8 KiB
Working guide — Lumbridge Bench
Read docs/DECISIONS.md before changing anything structural. It records why
the shape is the shape; the code does not.
What this repo is
A Lumbridge evidence lab that answers "should we self-host this model?" — quality on our real
work, plus serving performance on our own boxes, in one score card per
(model × host × quant × serving config × checkpoint) target.
Built on Inspect AI as a dependency. It is not a fork and must not become one (D1). Adding a benchmark means adding a task, not patching the harness.
Ground rules
-
Never commit anything to
data/private/that was not scrubbed. No API keys, tokens, tailscale hostnames/IPs, customer data, or personal information. Samples derived from real tool calls get sanitised at authoring time. The content will be sent to third-party APIs (judges, baselines) regardless of how private this repo is. -
Never publish or export
data/private/. Not in the site build, not in an artifact, not in a share card, not in a bug report. A public test set is a dead test set (D3). -
Never present a capped run as a full run.
--limitis recorded in the result JSON and rendered on the card. Keep it that way. -
Do not change
results/schema fields in place. Old result files must stay readable. BumpSCHEMA_VERSIONand add fields; never repurpose one. -
Results are only comparable across identical serving flags. Before recording a run, confirm the target's flags with
kbench snapshot <host>. If they drifted, updatemodels.yamlfirst — a result carries its own snapshot, so a stale registry silently produces incomparable numbers.
Adding a reference task
Reference tasks come from inspect_evals unmodified. Add an entry to
CATALOG in kbench/tasks/__init__.py with its primary_metric (the headline
key in the scorer's metrics) and, if sample scores are dict-valued,
primary_sample_key. Do not modify the upstream task.
Adding a signal task
Signal tasks are ours. Samples live in data/private/<family>.jsonl, one JSON
object per line. See data/README.md for the format. Use kbench add rather
than hand-editing so the canary tag and split assignment are applied.
Assign split at authoring time and never move a sample between splits —
test samples must never be seen during fine-tuning, or every number after
that point is a lie you cannot detect.
Perf measurement
kbench/perf.py deliberately defeats prefix caching (unique nonce per
request) and fixes output length (ignore_eos). If you touch it, preserve
both — without them the numbers look better and mean nothing (D6).
Concurrency-1 numbers describe what a single user feels. Aggregate throughput at higher concurrency describes what the box can serve. Report both; they can differ by more than 20x on the Spark.
Environment
- Runner: EQ (this box). It only needs HTTP to the serving host.
your-nodeserves vLLM on:8001, OpenAI-compatible.- Inspect addresses it as
openai-api/spark/<model>withSPARK_BASE_URL/SPARK_API_KEY, both derived from the registry — do not hardcode them.
Deploy
The standalone site is retired (2026-08-01): bench.karti.ai permanently redirects to
lumbridgecorp.com/bench, and bench-karti on web-host :8909 is stopped and disabled. Bench ships as
part of Lumbridge, not as its own destination — see docs/DEPLOY.md before standing
any UI back up. Supabase stores auth and the private manual-suggestion inbox in schema bench; it is
not an execution queue. The kbench CLI runner is unaffected.
Public model suggestions never execute code. The suggestion endpoint may record a reference for owner review, but it must never download, import, schedule, or run a model. D10 supersedes D8 for product behavior; every run begins with explicit owner approval.