MULTI_TURN.md scopes the multi-turn gap and records the finding that made it cheap: multi-turn needs no container and no real harness. vf.Env.run programs the control flow host-side and the simulator runs in-process as the user, with the harness left null. PLAN.md carries the settled decisions. EXECUTION.md is the post-adversarial-review plan — every position in it survived a reviewer, and where a design spec was overturned the corrected position is what is written down. Two corrections are marked in place rather than silently rewritten, because both were believed and acted on: - The provenance hole is src/arena/index.ts, not offices/sites.ts. Mutating offices/sites.ts produces a bit-identical episode; index.ts re-exports the scripted baselines that ARE the reward denominator, and a halved office-nav baseline moves return 4.744 -> 4.098 while every tera hash gate prints ok. - office-jobs-v1 is not phantom. It was absent at 13:00 and present at 15:50, because another session fast-forward-merged it into the shared tera tree mid-session. The counts are twelve, five and twenty. docs/FIRST_EVAL.md is the runbook for the first evaluation Arena ever completed: seven environments, 224 episodes, 223 scored, against brain-qwen38-dspark. It warns that the mean is sum(score x weight) — an unweighted mean over components is a different number, and it is how six of seven means were first misreported. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
9.3 KiB
Running an eval — the runbook
Written 2026-08-21, the day Arena first measured anything end to end. Before this,
outputs/ held one config.toml and zero episodes: every number downstream of an
eval — the dashboard, eval-results.json, the leaderboard, the before/after
training table — was waiting on a run that had never completed.
This is that run, written down so the next one is a copy-paste.
1. The command that works
cd ~/repos/gitea/arena
export SPARK_API_KEY=dummy
uv run --project environments/canary_trap eval @ configs/canary_trap.toml \
--model brain-qwen38-dspark \
--client.base-url http://100.127.247.67:8001/v1 \
--client.api-key-var SPARK_API_KEY \
--no-push --no-rich \
-c 8 -o outputs/run-<stamp>/canary-trap
Swap environments/<pkg> and configs/<pkg>.toml for any of the seven. The
sweep script that ran all of them is outputs/run_all.sh (gitignored); the
counter is outputs/count_rewards.py.
Why each flag is there
| flag | why |
|---|---|
--model brain-qwen38-dspark |
The served id, pinned. --model brain does not 404. SGLang does not validate the model field on the chat route: it returns 200, Qwen answers, and the trace records brain. The old outputs/canary-trap--brain--null/ dir is exactly that mistake. Pin the id GET /v1/models reports. |
--client.base-url |
The default is https://api.pinference.ai/api/v1, which needs a PRIME_API_KEY this box does not have. |
--client.api-key-var SPARK_API_KEY |
spark-1 requires no key — a bare POST /v1/chat/completions returns 200. But the client still resolves some env var, so point it at a dummy one rather than leaving it on PRIME_API_KEY (empty → the failure mode of the one earlier attempt). |
--no-push |
Hub push is parked (PLAN.md). Push is on by default and needs $PRIME_API_KEY. |
--no-rich |
The live dashboard is in-process only and unreadable in a log. Without it you get one rollout done: … reward=… line per rollout. |
-c 8 |
spark-1 is shared with Chatterbox, vox and ASR. The config's own max_concurrent = 128 will hammer it. |
-o <dir> |
Otherwise the run lands in outputs/<env>--<model>--<harness>/<uuid>/ and you have to go hunting for the uuid. |
Flags our own docs get wrong
README.md's "Running one" line omits--client.base-url/--client.api-key-varentirely, so copy-pasting it points at Prime's inference API with an empty key. That is the one attempt that was ever made, and whyoutputs/was empty.MULTI_TURN.md's preconditions say the spark-1 model id isbrain. The served id isbrain-qwen38-dspark;brainsilently mislabels the trace.EXECUTION.mdis right that--harness.id nullis not valid —EvalConfighas no top-levelharness. The configs already carry[env.agent.harness] id = "null", so nothing needs passing on the command line. The CLI form, if you ever need it, is--env.agent.harness.id null.- The
@ config.tomlform is a positional argument, not a flag:eval @ file.toml. Inside a shell script quote it as"@"so it is not glob-expanded.
2. The output layout verifiers 0.3.0 actually produces
verifiers/v1/cli/output.py:38 — output_path():
outputs/<env>--<model>--<harness>/<uuid>/
├── config.toml # the RESOLVED config, re-runnable as `eval @ config.toml`
├── traces.jsonl # one JSON *Episode* per line, appended as rollouts land
└── eval.log # only when you redirect it there yourself
--output-dir replaces the whole <name>/<uuid> pair, so an explicit -o gives
you a flat dir with no uuid leaf. Slashes in the model id become --.
One Episode per line, not one Trace. Shape:
{"id": "...", "env": {...}, "ok": true, "errors": [],
"traces": [ { "id": "...", "agent": "...", "rewards": {...},
"metrics": {...}, "stop_condition": "agent_completed",
"errors": [], "timing": {...} } ] }
Two traps in that file, both of which bite an ingester:
Reward.valueis a@propertyand never appears in the JSON. Each entry inrewardsis{"score": float, "weight": float}. The episode total issum(score * weight). Thereward=0.525in the log line is computed, not stored.write_episodedumps withexclude_none=True. ANonereward is dropped from the file entirely, so "not measured" and "key absent" are indistinguishable on the wire. Carry the run's declared reward-name set fromconfig.tomlif you need to tell them apart.
⚠️ 0.4.0 moves this
Upstream HEAD (~/vendor/prime-intellect/verifiers) produces one flat
<env>--<model>--<harness>--<short-id> directory with no nested uuid, and the
resolved config moves to configs/<cli>.json (JSON, because JSON keeps nulls).
RunConfig.id is a PrivateAttr and is absent from that file — derive the run id
from the directory basename, not from the config.
3. Counting what actually scored
eval exits 0 even when every rollout errors. An errored trace is written with
rewards: {}. Never report a run on its exit code.
uv run python outputs/count_rewards.py outputs/run-<stamp>/<env>
It prints traces_scored (non-empty rewards), traces_errored (rewards: {}),
provider_errors, the stop-condition histogram, and two means:
reward_mean_scored— the mean over what survived. This is Arena's own "score what survived" trap in reporting form: a run where 30 of 32 rollouts errored and the 2 survivors scored 0.9 reads as 0.900.reward_mean_attempted— errors counted as 0. Rank on this one, and refuse to rank at all below a scored-fraction floor.
A quick eyeball without the script:
jq -r '.traces[] | if (.rewards|length)>0 then "scored" else "errored" end' \
outputs/run-<stamp>/<env>/traces.jsonl | sort | uniq -c
4. What it cost, and what it scored
The run: outputs/run-20260821-1401/. Seven data environments, 32 rollouts each,
224 episodes, 223 scored. Model brain-qwen38-dspark (Qwen3.8-27B) on spark-1.
Wall-clock 14:03 → 14:58, 55 minutes, -c 8.
⚠️ Compute the mean as sum(score × weight) per trace. An unweighted mean over the
reward components is a different number and it is wrong — it disagrees with what the
harness itself prints. Cross-check against the reward= field on each env's 32
rollout done lines in eval.log; the two must agree exactly. The first transcription
of this run got six of seven means wrong by taking the unweighted mean.
| environment | traces | scored | errored | reward_mean_attempted | reward_mean_scored |
|---|---|---|---|---|---|
fault-localisation |
32 | 32 | 0 | 0.9531 | 0.9531 |
canary-trap |
32 | 32 | 0 | 0.6289 | 0.6289 |
schema-migration |
32 | 32 | 0 | 0.4432 | 0.4432 |
drop-table-inference |
32 | 31 | 1 | 0.4174 | 0.4309 |
redaction-pressure |
32 | 32 | 0 | 0.4040 | 0.4040 |
bot-detection |
32 | 32 | 0 | 0.3488 | 0.3488 |
grand-exchange |
32 | 32 | 0 | 0.0055 | 0.0055 |
The single errored episode is a genuine spark-1 gateway timeout (ProviderError, 504) in
drop-table-inference — not a scoring bug. Rank on _attempted (errors as 0); the
_scored column is Arena's own "score what survived" trap in reporting form.
Two numbers not to publish as capability scores
⚠️ grand-exchange 0.0055 measures a sampling configuration, not a capability.
In 31 of 32 rollouts the model returned a valid, parseable, degenerate answer —
{"expected_profit": 0, "orders": []} in 18 completion tokens, finish_reason=stop,
never entering a thinking block. The one rollout that did reason (4,954 completion tokens)
placed 4 real orders and scored 0.176 weighted, with filled=1233 against planned=2100.
The parser, the execution engine and the reward all work end to end — the environment
is not broken. But probe.py rates the crude baseline here at 0.082, so 0.0055 sits
below crude and barely above inaction. Re-run with thinking forced on before this
number goes anywhere near /evals.
⚠️ fault-localisation 0.9531 has two of four components pegged at maximum.
evidence mean 1.000 (max 1.000) and fault mean 1.000 (max 1.000) across all 32
rollouts; only gate (0.9062) and service (0.9062) still move. Half the reward mass
no longer discriminates at this capability level. Nemotron scored 0.352 here on 19 August —
a 2.7× swing that needs separating into a stronger model, a prompting difference, or an
environment that has become too easy, before it appears on a scoreboard.
The gate component carried no gradient in four of seven environments
Across all 32 rollouts each, gate scored exactly 0.000, max 0.000, in
bot-detection (weight 0.25), grand-exchange (0.25), redaction-pressure (0.30) and
schema-migration (0.25). It fired elsewhere — canary-trap 0.2188, drop-table 0.2258,
fault-localisation 0.9062.
This is not a house-rule-3 violation: probe.py's oracle row is 1.000 for every one of
those environments, which requires gate = 1, so it is reachable. But a quarter to a third
of the reward mass produced no gradient in this run, and any ranking or before/after
training table over those four environments is really a ranking over their two soft
components. Understand that before training against them.