12 KiB
Scope: the multi-turn gap
Status: scoping document, 2026-08-21. Nothing here is built yet.
Every environment in this repository sets [env.agent.harness] id = "null" — one
completion, no tools, no runtime. That is the right choice for a pure-trace data
environment and it is also the whole of what Arena currently measures. Nothing here
scores an agent acting: choosing what to look at next, paying for a turn, or
revising after seeing a consequence.
This document scopes closing that gap. The headline finding is that it is three separable tiers, not one, and the cheapest tier needs no new infrastructure at all.
What upstream actually requires
Read from stable verifiers==0.3.1 and the matching checkout under
~/vendor/prime-intellect/verifiers, not
from blog posts.
Multi-turn does not imply a container. vf.Env.run(task, agents) programs the
control flow host-side. agents.<role>.interaction(task) opens a rollout and holds
it open; await interaction.turn(observation) sends one user turn, runs one harness
segment, and returns a Segment with .last_reply and .terminated. The simulator
runs in-process, as the user. The harness can stay null throughout.
The reference implementation is verifiers/v1/tasksets/textarena/taskset.py — 90
lines, drives a real game engine, scores off the engine directly. Read it before
writing anything.
An Env ships inside the taskset's own wheel. EnvConfig.id empty means "the
taskset's own env, else SingleAgentEnv". Export the Env subclass from the
package __init__.py, exactly as environments/wordle exports TextArenaEnv, and
[env.taskset] id = "..." picks it up with no extra config.
Available in the installed version — confirmed, not assumed:
- harnesses:
null,bash,browser_use,claude_code,codex,hermes_agent,kimi_code,mini_swe_agent,openclaw,pi,pool,rlm,terminus_2 - runtimes:
subprocess(debug only — rollouts share side effects),docker,modal,prime - prebuilt envs:
single_agent,best_of_n,agentic_judge,shared_agentic_judge,user_sim
The interception server sits between harness and provider. It speaks each harness's native dialect (Anthropic Messages for Claude Code, OpenAI Responses for Codex), builds the trace live, forces sampling params the harness will not expose, and can rewrite tool results and web-search results to block reward hacks. It is mandatory for tiers 2 and 3 and free for tier 1.
The three tiers
| Mechanism | Harness | Runtime | New infra | |
|---|---|---|---|---|
| 1. Interactive | vf.Env.run + interaction.turn, simulator in-process |
null |
none | none |
| 2. Tool-using | Taskset exports a Toolset, served as MCP |
bash |
subprocess / docker |
image per env |
| 3. Agentic | Full coding agent in a sandbox | claude_code, codex, openclaw, pi |
docker / prime |
image + interception + budget |
Tier 1 is the one to build first, and it is where our existing simulators already sit. Tiers 2 and 3 are real work with real cost and should follow evidence from tier 1, not precede it.
What we already have that plugs in
Six of the seven environments here own a simulator that is already steppable. They are one-shot only because the taskset hands the model the whole observation up front. Each has an obvious interactive form:
| environment | one-shot today | interactive form | tier |
|---|---|---|---|
grand-exchange |
read 30 ticks of history, place one basket of orders | quote, see fills and the next tick, re-quote, over a turn budget — inventory and unfilled orders carry | 1 |
drop-table-inference |
read a fixed kill log, estimate the table | buy observations: each turn spends budget on more kills, then answers — measures when to stop looking | 1 |
fault-localisation |
read the whole incident log, name the cause | the log is not given: query it by service, time window and level under a query budget | 1 → 2 |
redaction-pressure |
write rules, scored by a scanner | two agents: a redactor and an extractor scored on what it recovers — the counterweight becomes an opponent | 1 (multi-agent) |
schema-migration |
emit a migration, executed against held-out rows | a real SQLite in a container the agent runs SQL against, with rollback | 3 |
bot-detection |
classify from six labelled examples | request labels one at a time under a budget (active learning) | 1 |
canary-trap |
write probes | no natural interactive form — leave one-shot | — |
Tera's four spatial environments are already multi-turn — in the wrong language.
~/repos/gitea/tera/src/arena/ implements tera.arena/v1: Gym-style
reset(seed, scenario) / step(action) / snapshot / restore / trace /
replay, FNV-1a-64 checksums, SHA-256 source pins, train/dev splits, counterweighted
rewards, and a CI gate that proves inaction finishes below zero and a scripted
baseline finishes positive. That is a stricter contract than most of what is on the
Hub — and none of it is reachable from uv run eval, because it is TypeScript with
its own contract rather than a verifiers taskset.
The work there is a bridge, not a simulator: a Python vf.Env whose run()
drives the TS package over a subprocess or a small local HTTP loop, turning
step() observations into interaction.turn() messages and the existing component
rewards into trace.record_reward calls. The replay checksum makes the bridge
verifiable — a bridged episode must replay identically in the TS package or the
bridge is wrong. That is a rare property and worth building the bridge around.
⚠️ RESOLVED 2026-08-21, 15:50 — office-jobs-v1 EXISTS. It was absent at 13:00 and
present at 15:50, because another session fast-forward-merged feat/tera-ultracode-integration
into ~/repos/gitea/tera mid-session (reflog 8738367). It is committed (e378a03,
120eac8, dated 2026-08-19), is in ARENA_MANIFESTS, is documented in tera/ARENA.md, and
passes npm run arena:source-hashes. The true counts are twelve environments, five spatial,
twenty scenarios — measured, not inferred. The site's original copy was right; the
correction to eleven/four/sixteen was right when made and wrong an hour later, and has been
reverted. Lesson: tera is a shared tree with live concurrent sessions — re-check its HEAD
before acting on a claim about what it contains.
⚠️ (superseded — see above)office-jobs-v1 does not exist. lumbridgecorp.com/arena and
services/api/src/seo.ts both advertise five spatial environments; ARENA_MANIFESTS
in src/arena/index.ts holds four (drive-101-v1, office-nav-v1, crow-nav-v1,
california-flight-v1) and there is no officeJobs.ts. The real total is eleven,
not twelve. Decide whether to build it or correct the page before publishing
anything that cites the count.
The house rules under multi-turn
The three rules were written for one-shot scoring. Two survive unchanged and one does not, and there is a fourth the interactive tier forces.
- Graded on what the model did not see — unchanged and easier: a stepped simulator generates its held-out continuation by construction.
- No single-sided reward — harder. Multi-turn adds degenerate strategies that have no one-shot analogue: stalling to burn the clock when the score is time-averaged, spamming cheap turns when observations are free, refusing to act when acting risks a penalty. Every new reward needs a counterweight in the turn dimension, not just the answer dimension.
- A zero floor and a reachable ceiling, demonstrated — unchanged, but the floor now has to hold across a whole trajectory, not one reply.
- New: the budget must bind. An interactive environment where the optimal policy is "use every turn" is measuring nothing the one-shot version did not. Observations must cost something, and the oracle must be shown to stop early.
probe.py must grow degenerate policies, not just degenerate answers. Today
it scores four fixed strings per environment. The interactive tier needs at minimum:
the staller (never acts), the spammer (acts every turn at random), the impatient
(answers on turn 1 with no observation), and the exhaustive (spends the entire
budget then answers). The gate stays the same shape — inaction 0.000, oracle 1.000,
exit 1 otherwise — plus a new assertion that the exhaustive policy scores below
the oracle. If it does not, rule 4 is violated and the environment does not ship.
This is the differentiator. Nothing in prime-envs (88 environments) or
community-environments (109) ships a degenerate-policy gate.
Dual publishing
Both targets take the same wheel; the repo layout already mirrors verifiers' own.
- Prime Intellect Hub, our own namespace:
prime env push <name> -v PUBLIC. We keep authorship and versioning, and the listing can link back to lumbridgecorp.com/arena. community-environments: a PR to that repo, published under theprimeintellectorg by their CI on a version bump. Extra distribution, but it hands the namespace to them. Secondary, not primary.- lumbridgecorp.com/arena: unchanged — the page documents the same wheel.
⚠️ Its README still documents the legacy flow (load_environment, vf-eval,
uv pip install -e .). Ignore it; follow verifiers/docs/v1/.
⚠️ Blocker: PRIME_API_KEY is not set on amd-server. prime config view shows
API Key: Not set. The pit_ key lives in karti's memory note on cloud-1
(.claude/projects/-home-ubuntu/memory/prime-intellect.md) and is the same
credential for api.pinference.ai and api.primeintellect.ai. Nothing pushes until
it is configured here.
Preconditions, verified on this box
uv run eval @ configs/canary_trap.toml --dry-run→ exit 0, config written ✓- stable verifiers 0.3.1 with upstream harnesses and runtimes present ✓
- Docker 29.1.3 running (tier 2/3 runtime) ✓
- spark-1
brainonhttp://100.127.247.67:8001/v1, model idbrain✓ primeCLI installed (0.6.25; 0.6.26 available) ✓PRIME_API_KEY✗ — see above
Proposed sequence
Phase 1 — one interactive environment, end to end. grand-exchange is the
candidate: the market simulation already steps, the counterweight (liquidity, the
random-walk trap) already exists, and turning one basket into a quote/fill loop is
the smallest change that produces a genuinely new measurement. Deliverable: a
vf.Env in the existing wheel, a second config, and probe coverage for the four
degenerate policies. This is where rule 4 gets written for real.
Phase 2 — extend probe.py to policies and re-gate everything. Whatever
phase 1 teaches about degenerate trajectories applies to every later environment.
Do this before building a second one, not after. Note that tests/test_probe.py
and the wheel-build loop in .gitea/workflows/ci.yml hardcode the environment
list — both need editing or CI goes red.
Phase 3 — the Tera bridge. One Python vf.Env that drives the existing TS
package, validated by replaying every bridged episode through replay() and
requiring an identical checksum. Lands four environments at once and makes the
strongest existing asset publishable. Larger and riskier than phase 1; the
process boundary and the action encoding are the unknowns.
Phase 4 — first Hub push, once phases 1–3 give three to four interactive environments alongside the seven one-shot ones. Needs the API key.
Phase 5 — tier 2, then tier 3. fault-localisation as a query-budgeted
tool-using environment, then schema-migration against a real database in a
container. Only after tier 1 has produced measured scores worth extending, and
with the concurrency cost of a container per rollout understood first.
Open questions
- Does a bridged Tera episode replay to an identical checksum through a process boundary? If not, the bridge cannot be verified and phase 3 gets much more expensive.
- What is the per-rollout wall-clock cost of the
dockerruntime here, at the concurrency an eval sweep needs? Unmeasured. It gates tier 3 entirely. - Should the interactive form of an environment be a separate taskset id
(
grand-exchange-live) or a config flag on the existing one? Separate ids keep the one-shot scores comparable across time; a flag keeps the wheel count down. Leaning separate.