docs: the scope, the plan, the post-review execution plan, and the eval runbook

MULTI_TURN.md scopes the multi-turn gap and records the finding that made it
cheap: multi-turn needs no container and no real harness. vf.Env.run programs the
control flow host-side and the simulator runs in-process as the user, with the
harness left null.

PLAN.md carries the settled decisions. EXECUTION.md is the post-adversarial-review
plan — every position in it survived a reviewer, and where a design spec was
overturned the corrected position is what is written down.

Two corrections are marked in place rather than silently rewritten, because both
were believed and acted on:

- The provenance hole is src/arena/index.ts, not offices/sites.ts. Mutating
  offices/sites.ts produces a bit-identical episode; index.ts re-exports the
  scripted baselines that ARE the reward denominator, and a halved office-nav
  baseline moves return 4.744 -> 4.098 while every tera hash gate prints ok.
- office-jobs-v1 is not phantom. It was absent at 13:00 and present at 15:50,
  because another session fast-forward-merged it into the shared tera tree
  mid-session. The counts are twelve, five and twenty.

docs/FIRST_EVAL.md is the runbook for the first evaluation Arena ever completed:
seven environments, 224 episodes, 223 scored, against brain-qwen38-dspark. It
warns that the mean is sum(score x weight) — an unweighted mean over components
is a different number, and it is how six of seven means were first misreported.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-21 15:55:43 -07:00
co-authored by Claude Opus 5
parent 162d67a83c
commit e4cc2bf1c3
4 changed files with 917 additions and 0 deletions
+208
View File
@@ -0,0 +1,208 @@
# Scope: the multi-turn gap
Status: scoping document, 2026-08-21. Nothing here is built yet.
Every environment in this repository sets `[env.agent.harness] id = "null"` — one
completion, no tools, no runtime. That is the right choice for a pure-trace data
environment and it is also the whole of what Arena currently measures. Nothing here
scores an agent *acting*: choosing what to look at next, paying for a turn, or
revising after seeing a consequence.
This document scopes closing that gap. The headline finding is that it is three
separable tiers, not one, and the cheapest tier needs no new infrastructure at all.
## What upstream actually requires
Read from `~/vendor/prime-intellect/verifiers` (v1, installed version 0.3.0), not
from blog posts.
**Multi-turn does not imply a container.** `vf.Env.run(task, agents)` programs the
control flow host-side. `agents.<role>.interaction(task)` opens a rollout and holds
it open; `await interaction.turn(observation)` sends one user turn, runs one harness
segment, and returns a `Segment` with `.last_reply` and `.terminated`. The simulator
runs in-process, as the user. The harness can stay `null` throughout.
The reference implementation is `verifiers/v1/tasksets/textarena/taskset.py` — 90
lines, drives a real game engine, scores off the engine directly. Read it before
writing anything.
**An Env ships inside the taskset's own wheel.** `EnvConfig.id` empty means "the
taskset's own env, else `SingleAgentEnv`". Export the `Env` subclass from the
package `__init__.py`, exactly as `environments/wordle` exports `TextArenaEnv`, and
`[env.taskset] id = "..."` picks it up with no extra config.
**Available in the installed version** — confirmed, not assumed:
- harnesses: `null`, `bash`, `browser_use`, `claude_code`, `codex`, `hermes_agent`,
`kimi_code`, `mini_swe_agent`, `openclaw`, `pi`, `pool`, `rlm`, `terminus_2`
- runtimes: `subprocess` (debug only — rollouts share side effects), `docker`,
`modal`, `prime`
- prebuilt envs: `single_agent`, `best_of_n`, `agentic_judge`,
`shared_agentic_judge`, `user_sim`
**The interception server** sits between harness and provider. It speaks each
harness's native dialect (Anthropic Messages for Claude Code, OpenAI Responses for
Codex), builds the trace live, forces sampling params the harness will not expose,
and can rewrite tool results and web-search results to block reward hacks. It is
mandatory for tiers 2 and 3 and free for tier 1.
## The three tiers
| | Mechanism | Harness | Runtime | New infra |
|---|---|---|---|---|
| **1. Interactive** | `vf.Env.run` + `interaction.turn`, simulator in-process | `null` | none | **none** |
| **2. Tool-using** | Taskset exports a `Toolset`, served as MCP | `bash` | `subprocess` / `docker` | image per env |
| **3. Agentic** | Full coding agent in a sandbox | `claude_code`, `codex`, `openclaw`, `pi` | `docker` / `prime` | image + interception + budget |
Tier 1 is the one to build first, and it is where our existing simulators already
sit. Tiers 2 and 3 are real work with real cost and should follow evidence from
tier 1, not precede it.
## What we already have that plugs in
**Six of the seven environments here own a simulator that is already steppable.**
They are one-shot only because the taskset hands the model the whole observation up
front. Each has an obvious interactive form:
| environment | one-shot today | interactive form | tier |
|---|---|---|---|
| `grand-exchange` | read 30 ticks of history, place one basket of orders | quote, see fills and the next tick, re-quote, over a turn budget — inventory and unfilled orders carry | 1 |
| `drop-table-inference` | read a fixed kill log, estimate the table | **buy** observations: each turn spends budget on more kills, then answers — measures when to stop looking | 1 |
| `fault-localisation` | read the whole incident log, name the cause | the log is *not* given: query it by service, time window and level under a query budget | 1 → 2 |
| `redaction-pressure` | write rules, scored by a scanner | two agents: a redactor and an extractor scored on what it recovers — the counterweight becomes an opponent | 1 (multi-agent) |
| `schema-migration` | emit a migration, executed against held-out rows | a real SQLite in a container the agent runs SQL against, with rollback | 3 |
| `bot-detection` | classify from six labelled examples | request labels one at a time under a budget (active learning) | 1 |
| `canary-trap` | write probes | no natural interactive form — leave one-shot | — |
**Tera's four spatial environments are already multi-turn — in the wrong language.**
`~/repos/gitea/tera/src/arena/` implements `tera.arena/v1`: Gym-style
`reset(seed, scenario)` / `step(action)` / `snapshot` / `restore` / `trace` /
`replay`, FNV-1a-64 checksums, SHA-256 source pins, train/dev splits, counterweighted
rewards, and a CI gate that proves inaction finishes below zero and a scripted
baseline finishes positive. That is a stricter contract than most of what is on the
Hub — and none of it is reachable from `uv run eval`, because it is TypeScript with
its own contract rather than a verifiers taskset.
The work there is a **bridge**, not a simulator: a Python `vf.Env` whose `run()`
drives the TS package over a subprocess or a small local HTTP loop, turning
`step()` observations into `interaction.turn()` messages and the existing component
rewards into `trace.record_reward` calls. The replay checksum makes the bridge
verifiable — a bridged episode must replay identically in the TS package or the
bridge is wrong. That is a rare property and worth building the bridge around.
⚠️ **RESOLVED 2026-08-21, 15:50 — `office-jobs-v1` EXISTS.** It was absent at 13:00 and
present at 15:50, because another session fast-forward-merged `feat/tera-ultracode-integration`
into `~/repos/gitea/tera` mid-session (reflog `8738367`). It is committed (`e378a03`,
`120eac8`, dated 2026-08-19), is in `ARENA_MANIFESTS`, is documented in `tera/ARENA.md`, and
passes `npm run arena:source-hashes`. **The true counts are twelve environments, five spatial,
twenty scenarios** — measured, not inferred. The site's original copy was right; the
correction to eleven/four/sixteen was right when made and wrong an hour later, and has been
reverted. Lesson: `tera` is a shared tree with live concurrent sessions — re-check its HEAD
before acting on a claim about what it contains.
~~⚠️ **`office-jobs-v1` does not exist.** `lumbridgecorp.com/arena` and
`services/api/src/seo.ts` both advertise five spatial environments; `ARENA_MANIFESTS`
in `src/arena/index.ts` holds four (`drive-101-v1`, `office-nav-v1`, `crow-nav-v1`,
`california-flight-v1`) and there is no `officeJobs.ts`. The real total is eleven,
not twelve. Decide whether to build it or correct the page before publishing
anything that cites the count.~~ (superseded — see above)
## The house rules under multi-turn
The three rules were written for one-shot scoring. Two survive unchanged and one
does not, and there is a fourth the interactive tier forces.
1. **Graded on what the model did not see** — unchanged and easier: a stepped
simulator generates its held-out continuation by construction.
2. **No single-sided reward***harder*. Multi-turn adds degenerate strategies that
have no one-shot analogue: stalling to burn the clock when the score is
time-averaged, spamming cheap turns when observations are free, refusing to act
when acting risks a penalty. Every new reward needs a counterweight in the *turn*
dimension, not just the answer dimension.
3. **A zero floor and a reachable ceiling, demonstrated** — unchanged, but the floor
now has to hold across a whole trajectory, not one reply.
4. **New: the budget must bind.** An interactive environment where the optimal
policy is "use every turn" is measuring nothing the one-shot version did not.
Observations must cost something, and the oracle must be shown to stop early.
**`probe.py` must grow degenerate *policies*, not just degenerate answers.** Today
it scores four fixed strings per environment. The interactive tier needs at minimum:
the staller (never acts), the spammer (acts every turn at random), the impatient
(answers on turn 1 with no observation), and the exhaustive (spends the entire
budget then answers). The gate stays the same shape — inaction 0.000, oracle 1.000,
exit 1 otherwise — plus a new assertion that the exhaustive policy scores *below*
the oracle. If it does not, rule 4 is violated and the environment does not ship.
This is the differentiator. Nothing in `prime-envs` (88 environments) or
`community-environments` (109) ships a degenerate-policy gate.
## Dual publishing
Both targets take the same wheel; the repo layout already mirrors verifiers' own.
- **Prime Intellect Hub**, our own namespace: `prime env push <name> -v PUBLIC`.
We keep authorship and versioning, and the listing can link back to
lumbridgecorp.com/arena.
- **`community-environments`**: a PR to that repo, published under the
`primeintellect` org by their CI on a version bump. Extra distribution, but it
hands the namespace to them. Secondary, not primary.
- **lumbridgecorp.com/arena**: unchanged — the page documents the same wheel.
⚠️ Its README still documents the **legacy** flow (`load_environment`, `vf-eval`,
`uv pip install -e .`). Ignore it; follow `verifiers/docs/v1/`.
⚠️ **Blocker: `PRIME_API_KEY` is not set on amd-server.** `prime config view` shows
`API Key: Not set`. The `pit_` key lives in karti's memory note on cloud-1
(`.claude/projects/-home-ubuntu/memory/prime-intellect.md`) and is the same
credential for `api.pinference.ai` and `api.primeintellect.ai`. Nothing pushes until
it is configured here.
## Preconditions, verified on this box
- `uv run eval @ configs/canary_trap.toml --dry-run` → exit 0, config written ✓
- verifiers 0.3.0 with all 13 harnesses and 4 runtimes present ✓
- Docker 29.1.3 running (tier 2/3 runtime) ✓
- spark-1 `brain` on `http://100.127.247.67:8001/v1`, model id `brain`
- `prime` CLI installed (0.6.25; 0.6.26 available) ✓
- `PRIME_API_KEY` ✗ — see above
## Proposed sequence
**Phase 1 — one interactive environment, end to end.** `grand-exchange` is the
candidate: the market simulation already steps, the counterweight (liquidity, the
random-walk trap) already exists, and turning one basket into a quote/fill loop is
the smallest change that produces a genuinely new measurement. Deliverable: a
`vf.Env` in the existing wheel, a second config, and probe coverage for the four
degenerate policies. This is where rule 4 gets written for real.
**Phase 2 — extend `probe.py` to policies and re-gate everything.** Whatever
phase 1 teaches about degenerate trajectories applies to every later environment.
Do this before building a second one, not after. Note that `tests/test_probe.py`
and the wheel-build loop in `.gitea/workflows/ci.yml` hardcode the environment
list — both need editing or CI goes red.
**Phase 3 — the Tera bridge.** One Python `vf.Env` that drives the existing TS
package, validated by replaying every bridged episode through `replay()` and
requiring an identical checksum. Lands four environments at once and makes the
strongest existing asset publishable. Larger and riskier than phase 1; the
process boundary and the action encoding are the unknowns.
**Phase 4 — first Hub push**, once phases 13 give three to four interactive
environments alongside the seven one-shot ones. Needs the API key.
**Phase 5 — tier 2, then tier 3.** `fault-localisation` as a query-budgeted
tool-using environment, then `schema-migration` against a real database in a
container. Only after tier 1 has produced measured scores worth extending, and
with the concurrency cost of a container per rollout understood first.
## Open questions
- Does a bridged Tera episode replay to an identical checksum through a process
boundary? If not, the bridge cannot be verified and phase 3 gets much more
expensive.
- What is the per-rollout wall-clock cost of the `docker` runtime here, at the
concurrency an eval sweep needs? Unmeasured. It gates tier 3 entirely.
- Should the interactive form of an environment be a **separate** taskset id
(`grand-exchange-live`) or a config flag on the existing one? Separate ids keep
the one-shot scores comparable across time; a flag keeps the wheel count down.
Leaning separate.