Throughput scales{" "}
{peak && single
? `${(peak.output_tps_total / single.output_tps_per_stream).toFixed(1)}×`
: "—"}{" "}
from one stream to {peak?.concurrency ?? "—"}, but per-stream
decode falls to {peak?.output_tps_per_stream.toFixed(1) ?? "—"}{" "}
tok/s. Batch work and interactive work want opposite settings on
this box.
Lumbridge Bench is the evidence layer for Lumbridge Compute. Two
numbers decide a self-hosting call: whether a model is good enough at
the work we actually do, and whether it is fast enough on the box we
would run it on. Most leaderboards report only the first. Every card
here reports both, measured together, on the same machine.
Signal scores come from private
tasks built from our own workloads. They are the measurement. The
set is not published — a public test set gets scraped into the next
training run and stops measuring anything.
Reference scores come from
public benchmarks, unmodified. They are calibration, not a ranking:
if one lands far from its published value, our harness is wrong, not
the model.