docs: the scope, the plan, the post-review execution plan, and the eval runbook
MULTI_TURN.md scopes the multi-turn gap and records the finding that made it cheap: multi-turn needs no container and no real harness. vf.Env.run programs the control flow host-side and the simulator runs in-process as the user, with the harness left null. PLAN.md carries the settled decisions. EXECUTION.md is the post-adversarial-review plan — every position in it survived a reviewer, and where a design spec was overturned the corrected position is what is written down. Two corrections are marked in place rather than silently rewritten, because both were believed and acted on: - The provenance hole is src/arena/index.ts, not offices/sites.ts. Mutating offices/sites.ts produces a bit-identical episode; index.ts re-exports the scripted baselines that ARE the reward denominator, and a halved office-nav baseline moves return 4.744 -> 4.098 while every tera hash gate prints ok. - office-jobs-v1 is not phantom. It was absent at 13:00 and present at 15:50, because another session fast-forward-merged it into the shared tera tree mid-session. The counts are twelve, five and twenty. docs/FIRST_EVAL.md is the runbook for the first evaluation Arena ever completed: seven environments, 224 episodes, 223 scored, against brain-qwen38-dspark. It warns that the mean is sum(score x weight) — an unweighted mean over components is a different number, and it is how six of seven means were first misreported. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+208
@@ -0,0 +1,208 @@
|
||||
# Scope: the multi-turn gap
|
||||
|
||||
Status: scoping document, 2026-08-21. Nothing here is built yet.
|
||||
|
||||
Every environment in this repository sets `[env.agent.harness] id = "null"` — one
|
||||
completion, no tools, no runtime. That is the right choice for a pure-trace data
|
||||
environment and it is also the whole of what Arena currently measures. Nothing here
|
||||
scores an agent *acting*: choosing what to look at next, paying for a turn, or
|
||||
revising after seeing a consequence.
|
||||
|
||||
This document scopes closing that gap. The headline finding is that it is three
|
||||
separable tiers, not one, and the cheapest tier needs no new infrastructure at all.
|
||||
|
||||
## What upstream actually requires
|
||||
|
||||
Read from `~/vendor/prime-intellect/verifiers` (v1, installed version 0.3.0), not
|
||||
from blog posts.
|
||||
|
||||
**Multi-turn does not imply a container.** `vf.Env.run(task, agents)` programs the
|
||||
control flow host-side. `agents.<role>.interaction(task)` opens a rollout and holds
|
||||
it open; `await interaction.turn(observation)` sends one user turn, runs one harness
|
||||
segment, and returns a `Segment` with `.last_reply` and `.terminated`. The simulator
|
||||
runs in-process, as the user. The harness can stay `null` throughout.
|
||||
|
||||
The reference implementation is `verifiers/v1/tasksets/textarena/taskset.py` — 90
|
||||
lines, drives a real game engine, scores off the engine directly. Read it before
|
||||
writing anything.
|
||||
|
||||
**An Env ships inside the taskset's own wheel.** `EnvConfig.id` empty means "the
|
||||
taskset's own env, else `SingleAgentEnv`". Export the `Env` subclass from the
|
||||
package `__init__.py`, exactly as `environments/wordle` exports `TextArenaEnv`, and
|
||||
`[env.taskset] id = "..."` picks it up with no extra config.
|
||||
|
||||
**Available in the installed version** — confirmed, not assumed:
|
||||
|
||||
- harnesses: `null`, `bash`, `browser_use`, `claude_code`, `codex`, `hermes_agent`,
|
||||
`kimi_code`, `mini_swe_agent`, `openclaw`, `pi`, `pool`, `rlm`, `terminus_2`
|
||||
- runtimes: `subprocess` (debug only — rollouts share side effects), `docker`,
|
||||
`modal`, `prime`
|
||||
- prebuilt envs: `single_agent`, `best_of_n`, `agentic_judge`,
|
||||
`shared_agentic_judge`, `user_sim`
|
||||
|
||||
**The interception server** sits between harness and provider. It speaks each
|
||||
harness's native dialect (Anthropic Messages for Claude Code, OpenAI Responses for
|
||||
Codex), builds the trace live, forces sampling params the harness will not expose,
|
||||
and can rewrite tool results and web-search results to block reward hacks. It is
|
||||
mandatory for tiers 2 and 3 and free for tier 1.
|
||||
|
||||
## The three tiers
|
||||
|
||||
| | Mechanism | Harness | Runtime | New infra |
|
||||
|---|---|---|---|---|
|
||||
| **1. Interactive** | `vf.Env.run` + `interaction.turn`, simulator in-process | `null` | none | **none** |
|
||||
| **2. Tool-using** | Taskset exports a `Toolset`, served as MCP | `bash` | `subprocess` / `docker` | image per env |
|
||||
| **3. Agentic** | Full coding agent in a sandbox | `claude_code`, `codex`, `openclaw`, `pi` | `docker` / `prime` | image + interception + budget |
|
||||
|
||||
Tier 1 is the one to build first, and it is where our existing simulators already
|
||||
sit. Tiers 2 and 3 are real work with real cost and should follow evidence from
|
||||
tier 1, not precede it.
|
||||
|
||||
## What we already have that plugs in
|
||||
|
||||
**Six of the seven environments here own a simulator that is already steppable.**
|
||||
They are one-shot only because the taskset hands the model the whole observation up
|
||||
front. Each has an obvious interactive form:
|
||||
|
||||
| environment | one-shot today | interactive form | tier |
|
||||
|---|---|---|---|
|
||||
| `grand-exchange` | read 30 ticks of history, place one basket of orders | quote, see fills and the next tick, re-quote, over a turn budget — inventory and unfilled orders carry | 1 |
|
||||
| `drop-table-inference` | read a fixed kill log, estimate the table | **buy** observations: each turn spends budget on more kills, then answers — measures when to stop looking | 1 |
|
||||
| `fault-localisation` | read the whole incident log, name the cause | the log is *not* given: query it by service, time window and level under a query budget | 1 → 2 |
|
||||
| `redaction-pressure` | write rules, scored by a scanner | two agents: a redactor and an extractor scored on what it recovers — the counterweight becomes an opponent | 1 (multi-agent) |
|
||||
| `schema-migration` | emit a migration, executed against held-out rows | a real SQLite in a container the agent runs SQL against, with rollback | 3 |
|
||||
| `bot-detection` | classify from six labelled examples | request labels one at a time under a budget (active learning) | 1 |
|
||||
| `canary-trap` | write probes | no natural interactive form — leave one-shot | — |
|
||||
|
||||
**Tera's four spatial environments are already multi-turn — in the wrong language.**
|
||||
`~/repos/gitea/tera/src/arena/` implements `tera.arena/v1`: Gym-style
|
||||
`reset(seed, scenario)` / `step(action)` / `snapshot` / `restore` / `trace` /
|
||||
`replay`, FNV-1a-64 checksums, SHA-256 source pins, train/dev splits, counterweighted
|
||||
rewards, and a CI gate that proves inaction finishes below zero and a scripted
|
||||
baseline finishes positive. That is a stricter contract than most of what is on the
|
||||
Hub — and none of it is reachable from `uv run eval`, because it is TypeScript with
|
||||
its own contract rather than a verifiers taskset.
|
||||
|
||||
The work there is a **bridge**, not a simulator: a Python `vf.Env` whose `run()`
|
||||
drives the TS package over a subprocess or a small local HTTP loop, turning
|
||||
`step()` observations into `interaction.turn()` messages and the existing component
|
||||
rewards into `trace.record_reward` calls. The replay checksum makes the bridge
|
||||
verifiable — a bridged episode must replay identically in the TS package or the
|
||||
bridge is wrong. That is a rare property and worth building the bridge around.
|
||||
|
||||
⚠️ **RESOLVED 2026-08-21, 15:50 — `office-jobs-v1` EXISTS.** It was absent at 13:00 and
|
||||
present at 15:50, because another session fast-forward-merged `feat/tera-ultracode-integration`
|
||||
into `~/repos/gitea/tera` mid-session (reflog `8738367`). It is committed (`e378a03`,
|
||||
`120eac8`, dated 2026-08-19), is in `ARENA_MANIFESTS`, is documented in `tera/ARENA.md`, and
|
||||
passes `npm run arena:source-hashes`. **The true counts are twelve environments, five spatial,
|
||||
twenty scenarios** — measured, not inferred. The site's original copy was right; the
|
||||
correction to eleven/four/sixteen was right when made and wrong an hour later, and has been
|
||||
reverted. Lesson: `tera` is a shared tree with live concurrent sessions — re-check its HEAD
|
||||
before acting on a claim about what it contains.
|
||||
|
||||
~~⚠️ **`office-jobs-v1` does not exist.** `lumbridgecorp.com/arena` and
|
||||
`services/api/src/seo.ts` both advertise five spatial environments; `ARENA_MANIFESTS`
|
||||
in `src/arena/index.ts` holds four (`drive-101-v1`, `office-nav-v1`, `crow-nav-v1`,
|
||||
`california-flight-v1`) and there is no `officeJobs.ts`. The real total is eleven,
|
||||
not twelve. Decide whether to build it or correct the page before publishing
|
||||
anything that cites the count.~~ (superseded — see above)
|
||||
|
||||
## The house rules under multi-turn
|
||||
|
||||
The three rules were written for one-shot scoring. Two survive unchanged and one
|
||||
does not, and there is a fourth the interactive tier forces.
|
||||
|
||||
1. **Graded on what the model did not see** — unchanged and easier: a stepped
|
||||
simulator generates its held-out continuation by construction.
|
||||
2. **No single-sided reward** — *harder*. Multi-turn adds degenerate strategies that
|
||||
have no one-shot analogue: stalling to burn the clock when the score is
|
||||
time-averaged, spamming cheap turns when observations are free, refusing to act
|
||||
when acting risks a penalty. Every new reward needs a counterweight in the *turn*
|
||||
dimension, not just the answer dimension.
|
||||
3. **A zero floor and a reachable ceiling, demonstrated** — unchanged, but the floor
|
||||
now has to hold across a whole trajectory, not one reply.
|
||||
4. **New: the budget must bind.** An interactive environment where the optimal
|
||||
policy is "use every turn" is measuring nothing the one-shot version did not.
|
||||
Observations must cost something, and the oracle must be shown to stop early.
|
||||
|
||||
**`probe.py` must grow degenerate *policies*, not just degenerate answers.** Today
|
||||
it scores four fixed strings per environment. The interactive tier needs at minimum:
|
||||
the staller (never acts), the spammer (acts every turn at random), the impatient
|
||||
(answers on turn 1 with no observation), and the exhaustive (spends the entire
|
||||
budget then answers). The gate stays the same shape — inaction 0.000, oracle 1.000,
|
||||
exit 1 otherwise — plus a new assertion that the exhaustive policy scores *below*
|
||||
the oracle. If it does not, rule 4 is violated and the environment does not ship.
|
||||
|
||||
This is the differentiator. Nothing in `prime-envs` (88 environments) or
|
||||
`community-environments` (109) ships a degenerate-policy gate.
|
||||
|
||||
## Dual publishing
|
||||
|
||||
Both targets take the same wheel; the repo layout already mirrors verifiers' own.
|
||||
|
||||
- **Prime Intellect Hub**, our own namespace: `prime env push <name> -v PUBLIC`.
|
||||
We keep authorship and versioning, and the listing can link back to
|
||||
lumbridgecorp.com/arena.
|
||||
- **`community-environments`**: a PR to that repo, published under the
|
||||
`primeintellect` org by their CI on a version bump. Extra distribution, but it
|
||||
hands the namespace to them. Secondary, not primary.
|
||||
- **lumbridgecorp.com/arena**: unchanged — the page documents the same wheel.
|
||||
|
||||
⚠️ Its README still documents the **legacy** flow (`load_environment`, `vf-eval`,
|
||||
`uv pip install -e .`). Ignore it; follow `verifiers/docs/v1/`.
|
||||
|
||||
⚠️ **Blocker: `PRIME_API_KEY` is not set on amd-server.** `prime config view` shows
|
||||
`API Key: Not set`. The `pit_` key lives in karti's memory note on cloud-1
|
||||
(`.claude/projects/-home-ubuntu/memory/prime-intellect.md`) and is the same
|
||||
credential for `api.pinference.ai` and `api.primeintellect.ai`. Nothing pushes until
|
||||
it is configured here.
|
||||
|
||||
## Preconditions, verified on this box
|
||||
|
||||
- `uv run eval @ configs/canary_trap.toml --dry-run` → exit 0, config written ✓
|
||||
- verifiers 0.3.0 with all 13 harnesses and 4 runtimes present ✓
|
||||
- Docker 29.1.3 running (tier 2/3 runtime) ✓
|
||||
- spark-1 `brain` on `http://100.127.247.67:8001/v1`, model id `brain` ✓
|
||||
- `prime` CLI installed (0.6.25; 0.6.26 available) ✓
|
||||
- `PRIME_API_KEY` ✗ — see above
|
||||
|
||||
## Proposed sequence
|
||||
|
||||
**Phase 1 — one interactive environment, end to end.** `grand-exchange` is the
|
||||
candidate: the market simulation already steps, the counterweight (liquidity, the
|
||||
random-walk trap) already exists, and turning one basket into a quote/fill loop is
|
||||
the smallest change that produces a genuinely new measurement. Deliverable: a
|
||||
`vf.Env` in the existing wheel, a second config, and probe coverage for the four
|
||||
degenerate policies. This is where rule 4 gets written for real.
|
||||
|
||||
**Phase 2 — extend `probe.py` to policies and re-gate everything.** Whatever
|
||||
phase 1 teaches about degenerate trajectories applies to every later environment.
|
||||
Do this before building a second one, not after. Note that `tests/test_probe.py`
|
||||
and the wheel-build loop in `.gitea/workflows/ci.yml` hardcode the environment
|
||||
list — both need editing or CI goes red.
|
||||
|
||||
**Phase 3 — the Tera bridge.** One Python `vf.Env` that drives the existing TS
|
||||
package, validated by replaying every bridged episode through `replay()` and
|
||||
requiring an identical checksum. Lands four environments at once and makes the
|
||||
strongest existing asset publishable. Larger and riskier than phase 1; the
|
||||
process boundary and the action encoding are the unknowns.
|
||||
|
||||
**Phase 4 — first Hub push**, once phases 1–3 give three to four interactive
|
||||
environments alongside the seven one-shot ones. Needs the API key.
|
||||
|
||||
**Phase 5 — tier 2, then tier 3.** `fault-localisation` as a query-budgeted
|
||||
tool-using environment, then `schema-migration` against a real database in a
|
||||
container. Only after tier 1 has produced measured scores worth extending, and
|
||||
with the concurrency cost of a container per rollout understood first.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Does a bridged Tera episode replay to an identical checksum through a process
|
||||
boundary? If not, the bridge cannot be verified and phase 3 gets much more
|
||||
expensive.
|
||||
- What is the per-rollout wall-clock cost of the `docker` runtime here, at the
|
||||
concurrency an eval sweep needs? Unmeasured. It gates tier 3 entirely.
|
||||
- Should the interactive form of an environment be a **separate** taskset id
|
||||
(`grand-exchange-live`) or a config flag on the existing one? Separate ids keep
|
||||
the one-shot scores comparable across time; a flag keeps the wheel count down.
|
||||
Leaning separate.
|
||||
Reference in New Issue
Block a user