# Running an eval — the runbook Written 2026-08-21, the day Arena first measured anything end to end. Before this, `outputs/` held one `config.toml` and zero episodes: every number downstream of an eval — the dashboard, `eval-results.json`, the leaderboard, the before/after training table — was waiting on a run that had never completed. This is that run, written down so the next one is a copy-paste. ⚠️ **Updated 2026-08-25.** The pin is now the stable `verifiers==0.3.1` release published 2026-08-24. It removes the 600-second model-call ceiling §5 blames for every truncation here, and moves the output layout. §1 and §2 are the post-bump instructions; §4 and §5 are kept as written because they are the record of what the thinking-**off** run measured, and §6 is the thinking-**on** re-measurement the bump made possible. **Every reward number in this file now carries its sampling configuration.** --- ## 1. The command that works ```bash cd ~/repos/gitea/arena export SPARK_API_KEY=dummy uv run --project environments/canary_trap eval @ configs/canary_trap.toml \ --model brain-qwen38-dspark \ --client.base-url http://100.127.247.67:8001/v1 \ --client.api-key-var SPARK_API_KEY \ --env.agent.runtime.type subprocess \ --no-push --no-rich \ -c 8 -o outputs/run- --run.dir canary-trap ``` Swap `environments/` and `configs/.toml` for any of the nine in `configs/`. The thinking-**on** sweep is `outputs/thinking-n32/run_all.sh` (gitignored, and it takes a `flock` — see below); the counters are `outputs/count_rewards.py` and `outputs/summarise_run.py`. **Thinking is off unless you ask for it**, and asking goes in a config file, not a flag — see §5. The seven data environments have thinking-on copies in `outputs/thinking-n32/*.toml`; §6 is what they measured. ### Why each flag is there | flag | why | |---|---| | `--model brain-qwen38-dspark` | **The served id, pinned.** `--model brain` does *not* 404. SGLang does not validate the model field on the chat route: it returns 200, Qwen answers, and the trace records `brain`. The old `outputs/canary-trap--brain--null/` dir is exactly that mistake. Pin the id `GET /v1/models` reports. | | `--client.base-url` | The default is `https://api.pinference.ai/api/v1`, which needs a `PRIME_API_KEY` this box does not have. | | `--client.api-key-var SPARK_API_KEY` | spark-1 requires **no** key — a bare `POST /v1/chat/completions` returns 200. But the client still resolves *some* env var, so point it at a dummy one rather than leaving it on `PRIME_API_KEY` (empty → the failure mode of the one earlier attempt). | | `--no-push` | Hub push is parked (PLAN.md). Push is **on by default** and needs `$PRIME_API_KEY`. | | `--no-rich` | The live dashboard is in-process only and unreadable in a log. Without it you get one `rollout done: … reward=…` line per rollout. | | `-c 8` | spark-1 is shared with Chatterbox, vox and ASR. The config's own `max_concurrent = 128` will hammer it. | | `--env.agent.runtime.type subprocess` | ⚠️ **New, and not optional.** 0.3.0 defaulted the agent runtime to `subprocess`; `0.3.1` defaults it to `prime` — a *remote* sandbox. A remote role makes `Env._runs_local()` false, which makes the interception pool mint a Prime tunnel, which calls `ensure_prime_auth()` and raises `SystemExit: not authenticated with prime`. The run dies before a single rollout, having already created its output directory and an empty `traces.jsonl`. The null harness needs no runtime at all; naming a local one is the whole fix. | | `-o ` + `--run.dir ` | `-o` is `output_dir` and the run writes to `output_dir / run.dir`. On 0.3.0 a bare `-o` gave you a flat directory; on 0.3.1 it gives you `outputs//------/`. Pass `--run.dir` to name the leaf yourself. | ### Flags our own docs get wrong - `README.md`'s "Running one" line omits `--client.base-url` / `--client.api-key-var` entirely, so copy-pasting it points at Prime's inference API with an empty key. That is the one attempt that was ever made, and why `outputs/` was empty. - `MULTI_TURN.md`'s preconditions say the spark-1 model id is `brain`. The served id is `brain-qwen38-dspark`; `brain` silently mislabels the trace. - `EXECUTION.md` is right that `--harness.id null` is not valid — `EvalConfig` has no top-level `harness`. The configs already carry `[env.agent.harness] id = "null"`, so nothing needs passing on the command line. The CLI form, if you ever need it, is `--env.agent.harness.id null`. - The `@ config.toml` form is a *positional* argument, not a flag: `eval @ file.toml`. Inside a shell script quote it as `"@"` so it is not glob-expanded. --- ## 2. The pin, and the output layout it produces ### There is no `verifiers` 0.4.0 `~/vendor/prime-intellect/verifiers` is **not** a 0.4.0 release, and nothing on any index is. `git describe --tags` there says `v0.3.0-59-g4bcb48e5`: fifty-nine commits past the 0.3.0 tag, version derived from git tags by `hatch-vcs`. The `version = "0.4.0"` that earlier notes read off its `pyproject.toml:136` belongs to a `[[tool.uv.dependency-metadata]]` block for **nemo-gym**, not to verifiers. Anyone who goes looking for `verifiers==0.4.0` will not find it. The stable `0.3.1` wheel was published to PyPI on 2026-08-24 after the dev stream. It contains the unbounded model-call timeout and output-layout changes this runbook requires; `verifiers/v1/clients/base.py` reads: ```python DEFAULT_TIMEOUT = httpx.Timeout(connect=5.0, read=None, write=None, pool=None) ``` **So the pin is an ordinary index dependency, not a path or git source.** Every `environments/*/pyproject.toml` now reads `verifiers==0.3.1`. It is an ordinary stable index dependency — no `--prerelease=allow`, no `[tool.uv.sources]`, no `tool.uv` block, and nothing that resolves differently in CI than it does here. A path dependency on `~/vendor` would have been reproducible only on this box; a git URL would have pinned a commit no wheel matches. Neither was necessary. ``` Resolved 114 packages in 620ms - verifiers==0.3.0 + verifiers==0.3.1 ``` `uv sync` each environment after pulling. Verify the thing you bumped *for*, not the version string: ```bash for d in environments/*/; do printf "%-34s " "$d" grep -h "DEFAULT_TIMEOUT = " $d/.venv/lib/python*/site-packages/verifiers/v1/clients/base.py done # environments/bot_detection/ DEFAULT_TIMEOUT = httpx.Timeout(connect=5.0, read=None, write=None, pool=None) # ... nine of nine ``` `uv sync` rewrites each `environments/*/uv.lock`; none of them are tracked any more (`.gitignore:5` is `environments/*/uv.lock`, and `git ls-files environments | grep uv.lock` is empty), so the bump does not put a lock file into a commit. CI's lock-policy step is `test -z "$(find environments -name uv.lock -print -quit)"` — it looks on **disk**, not in git, so it passes on a fresh checkout and would fail on a working tree that has been synced. Do not run it locally and conclude the repo is broken. ### The layout 0.3.1 writes `verifiers/v1/cli/output.py` — `output_path()` is `config.output_dir / config.run.dir`, and `run.dir` defaults to `run.name`, which is `------`, lowercased, with slashes in the model id becoming `--`. Measured, not read off the source: ``` outputs//------/ ├── configs/eval.json # the RESOLVED config, re-runnable as `eval @ configs/eval.json` ├── traces.jsonl # one JSON *Episode* per line, appended as rollouts land └── logs/eval.log # written for you now — no redirect needed ``` Three differences from 0.3.0 that an ingester or a script has to handle: 1. **The uuid leaf is gone.** 0.3.0 wrote `----//`; 0.3.1 writes one flat directory with an 8-hex suffix on the name. `-o` no longer flattens anything — it sets `output_dir`, and the run still makes its own leaf underneath. `--run.dir ` is what pins the leaf, and it is what `outputs/thinking-n32/run_all.sh` uses. 2. **`config.toml` → `configs/eval.json`.** JSON because JSON keeps nulls, so an explicit `None` round-trips on re-parse. `RunConfig._id` is a `PrivateAttr` and is **not** in it — derive the run id from the directory basename. 3. **`logs/eval.log` is written unconditionally.** On 0.3.0 you only got it if you redirected stdout yourself, which is why `outputs/run-20260821-1401/*.log` sit beside their run dirs rather than inside them. Anything that globs for a log should look in both places. 4. **`--resume` exists now**, and it is the answer to §5's "a timed-out run leaves no evidence": `uv run eval @ /configs/eval.json --resume` re-runs only the missing and errored rollouts, in place. (Its own `--help` still says `@ /config.toml` — the 0.3.0 filename. The file is `configs/eval.json`.) The episode schema did **not** move, despite `4bcb48e5 feat(v1): make episodes training-native` landing in this range. One Episode per line, not one Trace: ```jsonc {"id": "...", "env": {...}, "ok": true, "errors": [], "traces": [ { "id": "...", "agent": "...", "rewards": {...}, "metrics": {...}, "stop_condition": "agent_completed", "errors": [], "timing": {...} } ] } ``` `outputs/count_rewards.py`, written against 0.3.0, reproduces every published number from `run-20260821-1401` unchanged against a 0.3.1 run. The two traps in that file are also unchanged, and both still bite an ingester: 1. **`Reward.value` is a `@property` and never appears in the JSON.** Each entry in `rewards` is `{"score": float, "weight": float}`. The episode total is `sum(score * weight)`. The `reward=0.525` in the log line is computed, not stored. 2. **`write_episode` dumps with `exclude_none=True`.** A `None` reward is dropped from the file entirely, so "not measured" and "key absent" are indistinguishable on the wire. Carry the run's declared reward-name set from the resolved config if you need to tell them apart. ### What the bump does not fix `sampling` in `configs/eval.json` now carries `chat_template_kwargs` through resolution, so a run's own config records whether thinking was requested. **The trace still does not.** Every `calls[].sampling` is `{}` — `ChatDialect.parse_sampling` keeps an allow-list of known keys — and `usage.reasoning_tokens` is still null, because SGLang reports it at the top level of `usage` while `Usage.from_openai` reads it out of `completion_tokens_details`. The only in-trace evidence that a model thought is a `reasoning_content` field on the assistant message. `outputs/summarise_run.py` counts exactly that, and prints the run's resolved `sampling` beside every mean it reports. ## 3. Counting what actually scored **`eval` exits 0 even when every rollout errors.** An errored trace is written with `rewards: {}`. Never report a run on its exit code. ```bash uv run python outputs/count_rewards.py outputs/run-/ ``` It prints `traces_scored` (non-empty `rewards`), `traces_errored` (`rewards: {}`), `provider_errors`, the stop-condition histogram, and two means: - `reward_mean_scored` — the mean over what survived. This is Arena's own "score what survived" trap in reporting form: a run where 30 of 32 rollouts errored and the 2 survivors scored 0.9 reads as 0.900. - `reward_mean_attempted` — errors counted as 0. **Rank on this one**, and refuse to rank at all below a scored-fraction floor. A quick eyeball without the script: ```bash jq -r '.traces[] | if (.rewards|length)>0 then "scored" else "errored" end' \ outputs/run-//traces.jsonl | sort | uniq -c ``` ### And under what sampling configuration `count_rewards.py` answers *did anything score*. It does not answer the question that made `run-20260821-1401` unpublishable: *how was it sampled*. `outputs/summarise_run.py` adds that, and nothing else you cannot get from the first script: ```bash uv run --project environments/ python outputs/summarise_run.py outputs/ ``` It reads the run's own resolved config (`configs/eval.json` on dev59, `config.toml` on 0.3.0) and prints, beside the means: - `sampling` — the run's own record of `max_tokens` / `chat_template_kwargs`. A run whose `sampling` is `{}` was sampled at spark-1's server-side default, which is `enable_thinking: false`. - `traces_with_reasoning` — assistant messages carrying `reasoning_content` (or a literal `` block). **This is the only in-trace evidence that the model was allowed to think.** Verify the flag from here, never from the fact that you passed it. - `finish_reasons` and the completion-token quartiles — `length` in that histogram means truncation, and a truncated answer scores whatever an empty answer scores. - `reward_mean_from_eval_log` — the mean of the `reward=` field on the harness's own `rollout done` lines. It must equal `reward_mean_scored` exactly. An unweighted mean over reward *components* does not, which is how six of seven means were misreported once. Its arithmetic is checked against the published run: it reproduces `fault-localisation` 0.9531, `canary-trap` 0.6289, `grand-exchange` 0.0055 and `drop-table-inference` 0.4174 from `run-20260821-1401`, with `reward_mean_from_eval_log` agreeing in every case, and `traces_with_reasoning` 0/32 in every case. --- ## 4. What it cost, and what it scored **The run: `outputs/run-20260821-1401/`.** Seven data environments, 32 rollouts each, 224 episodes, **223 scored**. Model `brain-qwen38-dspark` (Qwen3.8-27B) on spark-1. Wall-clock 14:03 → 14:58, 55 minutes, `-c 8`. > ⚠️ **Sampling configuration for every number in this section — read it before quoting > one.** `verifiers==0.3.0`; `sampling = {}`, i.e. no `max_tokens` and no > `chat_template_kwargs`, so spark-1's server-side default applied and > **`enable_thinking` was `false`**; `temperature`, `top_p` and `reasoning_effort` all > unset; `-c 8`; harness `null`; agent runtime `subprocess`. `traces_with_reasoning` is > **0 of 224**. These are thinking-**off** numbers. §6 re-measures the same eight > environments with thinking on, and the two columns are not interchangeable — > `grand-exchange` moves by two orders of magnitude between them. ⚠️ **Compute the mean as `sum(score × weight)` per trace.** An unweighted mean over the reward components is a different number and it is wrong — it disagrees with what the harness itself prints. Cross-check against the `reward=` field on each env's 32 `rollout done` lines in `eval.log`; the two must agree exactly. The first transcription of this run got six of seven means wrong by taking the unweighted mean. | environment | traces | scored | errored | reward_mean_attempted | reward_mean_scored | |---|---|---|---|---|---| | `fault-localisation` | 32 | 32 | 0 | **0.9531** | 0.9531 | | `canary-trap` | 32 | 32 | 0 | **0.6289** | 0.6289 | | `schema-migration` | 32 | 32 | 0 | **0.4432** | 0.4432 | | `drop-table-inference` | 32 | 31 | 1 | **0.4174** | 0.4309 | | `redaction-pressure` | 32 | 32 | 0 | **0.4040** | 0.4040 | | `bot-detection` | 32 | 32 | 0 | **0.3488** | 0.3488 | | `grand-exchange` | 32 | 32 | 0 | **0.0055** | 0.0055 | The single errored episode is a genuine spark-1 gateway timeout (`ProviderError`, 504) in `drop-table-inference` — not a scoring bug. **Rank on `_attempted`** (errors as 0); the `_scored` column is Arena's own "score what survived" trap in reporting form. ### Two numbers not to publish as capability scores ⚠️ **`grand-exchange` 0.0055 measures a sampling configuration, not a capability.** In **31 of 32 rollouts** the model returned a valid, parseable, degenerate answer (27 of them in exactly 18 completion tokens; the full distribution is `{18: 27, 19: 1, 24: 3, 4954: 1}`) — `{"expected_profit": 0, "orders": []}` in 18 completion tokens, `finish_reason=stop`, never entering a thinking block. The one rollout that did reason (4,954 completion tokens) placed 4 real orders and scored 0.176 weighted, with `filled=1233` against `planned=2100`. **The parser, the execution engine and the reward all work end to end** — the environment is not broken. But `probe.py` rates the *crude* baseline here at 0.082, so 0.0055 sits **below crude** and barely above inaction. Re-run with thinking forced on before this number goes anywhere near `/evals`. **That re-run happened — see §5.** With thinking on the same model scored **1.000 on every component** of the one rollout that finished inside its token budget. ⚠️ **`drop-table-inference` 0.4174 is two populations, not one.** Twelve of its 31 scored rollouts answered in 116–174 tokens and averaged 0.0894; the other nineteen wrote 4,585–11,034 tokens and averaged 0.6465. §5 has the distribution. ⚠️ **`fault-localisation` 0.9531 has two of four components pegged at maximum.** `evidence` mean 1.000 (max 1.000) and `fault` mean 1.000 (max 1.000) across all 32 rollouts; only `gate` (0.9062) and `service` (0.9062) still move. Half the reward mass no longer discriminates at this capability level. Nemotron scored 0.352 here on 19 August — a 2.7× swing that needs separating into a stronger model, a prompting difference, or an environment that has become too easy, before it appears on a scoreboard. ### The `gate` component carried no gradient in four of seven environments Across all 32 rollouts each, `gate` scored **exactly 0.000, max 0.000**, in `bot-detection` (weight 0.25), `grand-exchange` (0.25), `redaction-pressure` (0.30) and `schema-migration` (0.25). It fired elsewhere — `canary-trap` 0.2188, `drop-table` 0.2258, `fault-localisation` 0.9062. This is **not** a house-rule-3 violation: `probe.py`'s oracle row is 1.000 for every one of those environments, which requires `gate = 1`, so it is reachable. But a quarter to a third of the reward mass produced no gradient in this run, and any ranking or before/after training table over those four environments is really a ranking over their two soft components. Understand that before training against them. --- ## 5. The run was measured with thinking OFF, and nothing in the run says so Appended 2026-08-21, after re-measuring `grand-exchange`. ### The cause spark-1's SGLang is started with a default that overrides the chat template: ```bash curl -s http://100.127.247.67:8001/get_server_info | jq '.default_chat_template_kwargs, .reasoning_parser' # {"enable_thinking": false} # "qwen3" ``` **Every one of the 224 episodes in `run-20260821-1401` was sampled with Qwen3's thinking mode disabled**, and not one of them requested otherwise. Confirmed from the traces: no assistant message in any of the seven environments carries a `reasoning_content` field, and none contains a `` block — 0 of 223 scored traces. Nothing in the run's own artefacts records this. The resolved `config.toml` says `sampling = {}`, and `ChatDialect.parse_sampling` keeps only an allow-list of known keys, so even a run that *does* pass `chat_template_kwargs` writes a `calls[].sampling` that does not mention it. `usage.reasoning_tokens` is null too: SGLang reports `reasoning_tokens` at the top level of `usage`, while `Usage.from_openai` reads it out of `completion_tokens_details`. **A trace in this repository cannot tell you whether the model was allowed to think.** ### Turning it on `SamplingConfig` is `extra="allow"` and `ChatDialect.apply_overrides` merges the whole dump into the request body, so an untyped key rides through to the provider untouched. There is no CLI flag — `--sampling.*` is typed — so it goes in a config file: ```toml [sampling] chat_template_kwargs = { enable_thinking = true } max_tokens = 16384 ``` `outputs/thinking/grand_exchange_thinking.toml` is that file. **Verify it worked from the trace, not from the flag**: the assistant message gains a `reasoning_content` field. ### What thinking on actually did to `grand-exchange` **`outputs/thinking/ge-think-n6/`** — 6 tasks × 1 rollout, `-c 2`, `max_tokens = 16384`, thinking on. All 6 traces carry `reasoning_content`, so the flag reached the model. | tokens | `finish_reason` | content | weighted reward | |---|---|---|---| | 16384 | length | 0 chars | 0.000 | | 16384 | length | 0 chars | 0.000 | | 16384 | length | 0 chars | 0.000 | | 16384 | length | 0 chars | 0.000 | | 16384 | length | 0 chars | 0.000 | | **8335** | **stop** | 1089 chars | **1.000** | **Mean 0.1667** against 0.0055 with thinking off — a 30× move, and above `probe.py`'s crude baseline of 0.082 for the first time. But read the column, not the mean: it is one perfect run and five truncations. The run that finished scored **1.000 on every component** — `profit` 1.000, `discipline` 1.000, **`gate` 1.000** — with `realised` 40,081 gp against a `reference_realised` of 36,157, `committed` 0.994 of the purse and `conversion` 0.835. Four orders, a written rationale, `orders_dropped` 0: ```json {"expected_profit": 15000, "orders": [ {"item": "Chipped bone charm", "quantity": 400, "buy": 120, "sell": 134}, {"item": "Coarse fletching feather", "quantity": 700, "buy": 97, "sell": 114}, {"item": "Emberglass shard", "quantity": 80, "buy": 1420, "sell": 1640}, {"item": "Heart of the sunken cairn", "quantity": 2, "buy": 9500, "sell": 13500}]} ``` ⚠️ **So the environment is not merely unbroken — it is fully solvable by this model, gate included, and the thinking-off number measured none of that.** `arith_ok` was still 0.0 on that rollout (`expected_profit` 15,000 against an implied 40,184, `arith_error` 0.627), which is exactly the line the taskset docstring says it records for this reason. The five truncations are not junk either. The reasoning is the analysis the environment asks for, item by item — counting crossings, applying the 25%-of-volume cap, sizing against the buy limit — it simply does not converge inside 16k: ``` **Heart of the sunken cairn**: Buy ticks (<=10000): 9267, 8846, 9656, 9541, 9703, 8853, 9945 = 7 ticks Volumes: 2, 2, 3, 6, 2, 3, 3 25%: 0.5, 0.5, 0.75, 1.5, 0.5, 0.75, 0.75 Total: 5.25 in 56 ticks. In 30 ticks: ~2.8. So I can only buy about 2-3 units! ``` A single rollout at `-c 1` (`outputs/thinking/smoke/`) does the same: 486 seconds, `completion_tokens` 16384, `reasoning_content` 30,121 characters, `content` empty, reward 0.000. So does the same prompt sent straight at spark-1 with `curl`, no verifiers in the path — `reasoning_tokens` 16390, `finish_reason` length, content empty. **This is not a harness bug, and it is not fixed by raising `max_tokens`** — see the next section for why. ⚠️ **`grand-exchange` 0.0055 was a sampling artefact and 0.1667 is a truncation artefact.** Neither is the environment's number. The number is somewhere at or above 0.1667, taken with a budget large enough that the model finishes, and this box cannot currently run that. ### ⚠️ The harness has a hard 600-second ceiling on one model call `verifiers/v1/clients/base.py:11`: ```python DEFAULT_TIMEOUT = httpx.Timeout(connect=5.0, read=600.0, write=600.0, pool=600.0) ``` `BaseClientConfig` exposes `base_url`, `api_key_var` and `headers` and nothing else. **There is no config key and no CLI flag for this timeout.** At spark-1's ~34 tok/s single-stream it caps one call at roughly 20,000 completion tokens, and far fewer under concurrency. That is not theoretical. `outputs/thinking/ge-think-n8.log` is 8 tasks at `-c 8` with `max_tokens = 32768`: ``` 16:09:27 INFO rollout start: ... (×8) 16:19:30 WARNING model call failed: id=... ProviderError: (×8) ``` Eight failures, all at 16:19:30, exactly 603 seconds after the starts, with an empty `ProviderError` message and `traces.jsonl` left at **zero lines**. `eval` still exited 0. The `-c 1` smoke run survived at 486 seconds; the same work at `-c 8` did not. `-c 2` is not safe either. `ge-think-n6.log` at `max_tokens = 16384` still logged **4 `ProviderError`s in 6 rollouts**, each exactly 601–603 seconds after its rollout started — `16:36:12 → 16:46:15`, `16:46:15 → 16:56:16`. Those rollouts survived only because the null harness's program retries the call itself (tenacity), which is why their traces carry two `calls` entries, the first with a null `finish_reason`. A retry that also lands at 16k tokens costs another ten minutes and scores zero anyway. **Raising `max_tokens` therefore does not fix the truncations.** The five rollouts that ran out at 16,384 tokens want more budget; the wall says they cannot have it, because 16,384 tokens is already 486–603 seconds of wall clock on this box. The two limits point in opposite directions and there is no setting that satisfies both. **And the thinking block cannot be bounded from the request either.** SGLang's `thinking_budget` is accepted (HTTP 200) and ignored by this build: the same prompt with `{"max_tokens": 8192, "thinking_budget": 3000}` came back with `reasoning_tokens` **8195**, `finish_reason` length, content empty. The server also reports `enable_strict_thinking: false`. The only lever is `max_tokens`, and `max_tokens` cuts the answer, not the reasoning. Three consequences for anyone re-running this suite against a reasoning model: 1. **Keep `-c` low and `max_tokens` under the wall**, or budget the run knowing calls will be cut at ten minutes. spark-1 is shared — another lane's eval was running against it during this measurement, and contention alone moves a rollout across the line. 2. **A timed-out run leaves no evidence.** Zero traces, exit code 0. `count_rewards.py` reports on a file that does not exist. This is the same trap as §3 with a bigger blast radius: check the log for `ProviderError` and check `wc -l traces.jsonl` before believing any run. 3. ⚠️ **CORRECTED: the 600-second ceiling is the INSTALLED WHEEL, not upstream.** The remedy stated here first — "the 600-second default gets a config key upstream" — is wrong, and it is wrong in the expensive direction, because it reads as blocked on someone else. Upstream **already removed it**: `verifiers/v1/clients/base.py:12` in `~/vendor/prime-intellect/verifiers` reads `httpx.Timeout(connect=5.0, read=None, ...)`, commit `a298bcfe fix(v1): restore the unbounded model-call timeout (#2304)`, 2026-08-08. Every `environments/*/.venv` here pins `read=600.0` because it resolves verifiers 0.3.0. **The fix is a dependency bump**, and an n=32 thinking-on run is achievable today. That makes the "grand-exchange is solvable" claim — currently resting on six rollouts — cheap to settle properly. ✅ **RETIRED 2026-08-21 evening.** The bump happened — `verifiers==0.3.1`, §2 — and every venv on this box now reads `read=None`. One correction to the paragraph above, which matters because it is the version anyone will type: the fix is **not** `verifiers==0.4.0`. There is no such release. `a298bcfe` is `v0.3.0-8-ga298bcfe`; the vendored HEAD is `v0.3.0-59-g4bcb48e5`; the wheel that carries them is `0.3.1`. Everything below this line about 600 seconds describes the old pin and is kept as the record of why the bump was worth doing, not as current behaviour. The `max_tokens` / wall-clock contradiction it describes — "the two limits point in opposite directions and there is no setting that satisfies both" — no longer holds: with `read=None` the only limit left is `max_tokens`, and §6 spends it. ### Did the other six collapse too? Only one of them, and only partly. Completion tokens per trace, and the mean weighted reward in each half of the length distribution: | environment | min | p25 | median | p75 | max | shape | |---|---|---|---|---|---|---| | `grand-exchange` | 18 | 18 | 18 | 18 | 4954 | **collapsed**; distribution `{18: 27, 19: 1, 24: 3, 4954: 1}` — 27/32 at 18 tokens, and **31/32 planned no orders at all** | | `drop-table-inference` | 0 | 159 | 5347 | 7205 | 11034 | **bimodal** | | `fault-localisation` | 34 | 37 | 37 | 39 | 191 | short by design | | `schema-migration` | 105 | 145 | 168 | 271 | 1687 | unimodal | | `redaction-pressure` | 153 | 240 | 310 | 369 | 2014 | unimodal | | `canary-trap` | 56 | 139 | 148 | 626 | 3128 | unimodal | | `bot-detection` | 1156 | 2508 | 3457 | 4603 | 8146 | unimodal | ⚠️ **`drop-table-inference` 0.4174 is a mixture of two regimes, not a capability estimate.** Twelve of its 31 scored rollouts answered in 116–174 tokens and averaged **0.0894**; the other nineteen wrote 4,585–11,034 tokens and averaged **0.6465**. The same prompt, the same sampling settings, a 7× gap in outcome depending on whether the model chose to work. Report it with that split or not at all. The other five are unimodal and their reward does not track length — `fault-localisation` answers three fields in 37 tokens and scores 0.9531, `canary-trap`'s shortest half scores 0.6141 against its longest half's 0.6437. Those five numbers stand as measured. They are still **thinking-off numbers**, which is a property of the run nobody chose and nothing recorded. ### What to do before the next run - Put the sampling configuration in the run's own name or its notes. A number that changes by two orders of magnitude with one server-side default is not a capability score unless the configuration travels with it. - Ask spark-1 what it defaults to (`/get_server_info`) rather than assuming a model's own default applies. - If `reasoning_tokens` matters to you, read it off the provider — the trace drops it. --- ## 6. The thinking-ON re-measurement, n=32 Run 2026-08-21 evening, after the `verifiers==0.3.1` bump of §2. This section exists because two of §4's seven numbers were artefacts of a sampling configuration nobody chose and nothing recorded, and because §5's remedy — "an n=32 thinking-on run is achievable today" — was owed a run. > **Sampling configuration for every number in this section.** `verifiers==0.3.1` > (`read=None`); `chat_template_kwargs = {enable_thinking = true}`; `max_tokens = 49152`; > `temperature`, `top_p`, `reasoning_effort` unset; `-c 8`; harness `null`; agent runtime > `subprocess`; model `brain-qwen38-dspark` on spark-1. The configs are > `outputs/thinking-n32/*.toml`, the sweep is `outputs/thinking-n32/run_all.sh`, the runs > are `outputs/thinking-n32/run//`. **Thinking is verified from the traces, not from the flag.** `traces_with_reasoning` is **32 of 32** for `grand-exchange`, mean 56,560 characters of `reasoning_content` per trace, against 0 of 224 for `run-20260821-1401`. ### `grand-exchange` — 0.0055 → **0.4338** | | thinking off (§4) | 6 rollouts, 16k (§5) | **thinking on, n=32** | |---|---|---|---| | `reward_mean_attempted` | 0.0055 | 0.1667 | **0.4338** | | traces scored | 32/32 | 6/6 | **32/32** | | traces with reasoning | 0/32 | 6/6 | **32/32** | | `finish_reason=length` | 0 | 5/6 | **2/32** | | median completion tokens | 18 | 16,384 | **29,376** | | `gate` mean (max) | 0.0000 (0.000) | — | **0.2188 (1.000)** | `probe.py` rates the crude maximiser at 0.082 and the plausible strategy at 0.608. **0.4338 sits between them**, which is the first time this environment has produced a number that means anything: 0.0055 was below crude, and 0.1667 was one perfect run and five truncations. Wall clock 18:28 → 20:33, 2 h 5 min, 965,217 completion tokens for 32 rollouts. Zero `ProviderError`s. **Under the old pin every one of these rollouts would have died**: the median call ran ~20 minutes against a 600-second ceiling, and the median completion is 29,376 tokens against the 16,384 that section 5 could afford. The distribution is properly graded — not the two-spike shape thinking-off produced: ``` total gate profit disc tokens finish 0.0 0.0 0.0 0.0 20305 stop <- 7 genuine zeros: finished, planned nothing usable ... 0.0 0.0 0.0 0.0 49152 length <- 2 truncations, both scored 0 0.102 0.0 0.203 0.038 38913 stop 0.162 0.0 0.302 0.087 30042 stop 0.413 0.0 0.653 0.397 37978 stop 0.56 0.0 0.746 0.746 22000 stop 0.683 0.0 0.91 0.91 20886 stop 0.747 0.0 0.996 0.996 32545 stop 0.993 1.0 1.0 0.977 23531 stop 1.0 1.0 1.0 1.0 20579 stop <- 6 perfect rollouts ``` Nine zeros, six 1.000s, seventeen strictly between. Median 0.4165. Dropping the two truncations gives 0.4627 over the thirty that finished — quote 0.4338 and say why, rather than quoting the higher number. ⚠️ **The dead `gate` in `grand-exchange` was partly a sampling artefact.** `docs/GATE_DIAGNOSIS.md` records it at exactly 0.000, max 0.000, across all 32 thinking-off rollouts. With thinking on it fires on **7 of 32**, mean 0.2188, max 1.000. That does not repair the shape — it is still exact equality, still all-or-nothing, and it still only opens for a rollout that is already at `profit` ≈ 1.0, so it adds no gradient anywhere below the top. But the diagnosis's evidence for *this* environment was measured under a configuration in which the model answered in 18 tokens, and the three other dead gates were measured under the same one. **Re-read GATE_DIAGNOSIS against this column before acting on it.** ### `drop-table-inference` — thinking on does not fit in 49,152 tokens The question §5 left was whether thinking on collapses this environment's two regimes (twelve short rollouts at 0.0894, nineteen long ones at 0.6465). **It does not collapse them; it pushes the whole distribution off the end of the token budget.** The first eight rollouts of the n=32 run, all eight carrying `reasoning_content`: ``` weighted fit gate rare tokens finish 0.4082 0.692 0.0 0.277 31649 stop 0.4082 0.692 0.0 0.277 33803 stop 0.0000 0.0 0.0 0.0 49152 length 0.0000 0.0 0.0 0.0 49152 length 0.0000 0.0 0.0 0.0 49152 length 0.0000 0.0 0.0 0.0 49152 length 0.0000 0.0 0.0 0.0 49152 length 0.0000 0.0 0.0 0.0 49152 length ``` **Six of eight hit `finish_reason=length` at 49,152 completion tokens** — three times the budget §5 could afford, on a model whose reasoning here runs 74–86 thousand characters. The two that finished scored 0.408 each, above §4's short regime (0.0894) and below its long one (0.6465). ⚠️ **Whatever mean this run finishes with is a truncation artefact, exactly like `grand-exchange`'s 0.1667 was.** Do not publish it as a capability score, and do not read it as evidence about the mixture. The comparison `grand-exchange` earned — a distribution where 30 of 32 finish — this environment has not: it needs a larger `max_tokens` (spark-1's context is 262,144, so there is room) and a correspondingly larger wall-clock budget. One 49,152-token rollout costs roughly 25 minutes at `-c 8` on this box. ### What is still running, and how to pick it up `outputs/thinking-n32/run_all.sh` was launched detached (`setsid`) and works down the priority order in §6: grand-exchange (done), drop-table-inference, bot-detection, redaction-pressure, schema-migration, canary-trap, fault-localisation. It holds a `flock` on `outputs/thinking-n32/.lock`, so a second launch refuses rather than writing two evals into one `traces.jsonl`. ```bash tail -f outputs/thinking-n32/sweep.log # stage transitions + per-rollout grep 'exit=' outputs/thinking-n32/sweep.log # which stages finished, with counts uv run --project environments/ python outputs/summarise_run.py \ outputs/thinking-n32/run/ # the numbers, with the sampling config pkill -f 'thinking-n32/run_all.sh' # stop it after the current stage, pkill -f 'bin/eval @ outputs/thinking-n32' # then kill the eval in flight ``` ⚠️ **Read every stage's `finish_reasons` before its mean.** `length` in that histogram means the model was cut off mid-answer, and a cut-off answer scores what an empty one scores. A stage whose histogram is mostly `length` has measured `max_tokens`, not the model — that is the single mistake this section exists to stop repeating. ### Every number here is comparable only to the thinking-on column `grand-exchange` moves by 79× between the two sampling configurations. Nothing in §4 and nothing in §6 may be tabulated beside the other without the configuration attached, and no number from either belongs on `/evals` or in `eval-results.json` without it. That is the whole lesson of the last two phases, stated as a rule: **A reward mean is not a number. It is a number plus a sampling configuration, and the pair travels or neither does.**