Files
Karti Tripathi 006feee0f7
CI / verify (push) Successful in 24s
CI / deploy (push) Failing after 1m14s
Lumbridge Bench
2026-08-04 00:44:07 -07:00

3.8 KiB
Raw Permalink Blame History

Working guide — Lumbridge Bench

Read docs/DECISIONS.md before changing anything structural. It records why the shape is the shape; the code does not.

What this repo is

A Lumbridge evidence lab that answers "should we self-host this model?" — quality on our real work, plus serving performance on our own boxes, in one score card per (model × host × quant × serving config × checkpoint) target.

Built on Inspect AI as a dependency. It is not a fork and must not become one (D1). Adding a benchmark means adding a task, not patching the harness.

Ground rules

  1. Never commit anything to data/private/ that was not scrubbed. No API keys, tokens, tailscale hostnames/IPs, customer data, or personal information. Samples derived from real tool calls get sanitised at authoring time. The content will be sent to third-party APIs (judges, baselines) regardless of how private this repo is.

  2. Never publish or export data/private/. Not in the site build, not in an artifact, not in a share card, not in a bug report. A public test set is a dead test set (D3).

  3. Never present a capped run as a full run. --limit is recorded in the result JSON and rendered on the card. Keep it that way.

  4. Do not change results/ schema fields in place. Old result files must stay readable. Bump SCHEMA_VERSION and add fields; never repurpose one.

  5. Results are only comparable across identical serving flags. Before recording a run, confirm the target's flags with kbench snapshot <host>. If they drifted, update models.yaml first — a result carries its own snapshot, so a stale registry silently produces incomparable numbers.

Adding a reference task

Reference tasks come from inspect_evals unmodified. Add an entry to CATALOG in kbench/tasks/__init__.py with its primary_metric (the headline key in the scorer's metrics) and, if sample scores are dict-valued, primary_sample_key. Do not modify the upstream task.

Adding a signal task

Signal tasks are ours. Samples live in data/private/<family>.jsonl, one JSON object per line. See data/README.md for the format. Use kbench add rather than hand-editing so the canary tag and split assignment are applied.

Assign split at authoring time and never move a sample between splits — test samples must never be seen during fine-tuning, or every number after that point is a lie you cannot detect.

Perf measurement

kbench/perf.py deliberately defeats prefix caching (unique nonce per request) and fixes output length (ignore_eos). If you touch it, preserve both — without them the numbers look better and mean nothing (D6).

Concurrency-1 numbers describe what a single user feels. Aggregate throughput at higher concurrency describes what the box can serve. Report both; they can differ by more than 20x on the Spark.

Environment

  • Runner: EQ (this box). It only needs HTTP to the serving host.
  • your-node serves vLLM on :8001, OpenAI-compatible.
  • Inspect addresses it as openai-api/spark/<model> with SPARK_BASE_URL / SPARK_API_KEY, both derived from the registry — do not hardcode them.

Deploy

The standalone site is retired (2026-08-01): bench.karti.ai permanently redirects to lumbridgecorp.com/bench, and bench-karti on web-host :8909 is stopped and disabled. Bench ships as part of Lumbridge, not as its own destination — see docs/DEPLOY.md before standing any UI back up. Supabase stores auth and the private manual-suggestion inbox in schema bench; it is not an execution queue. The kbench CLI runner is unaffected.

Public model suggestions never execute code. The suggestion endpoint may record a reference for owner review, but it must never download, import, schedule, or run a model. D10 supersedes D8 for product behavior; every run begins with explicit owner approval.