Two measurement findings, both larger than the lanes that found them. FIRST_EVAL.md: every number in the first run was sampled with thinking OFF — 0 of 224 traces carry reasoning_content or a <think> block, and spark-1 serves with enable_thinking False. No artefact of the run recorded that. Five of the seven numbers still stand; two do not. grand-exchange 0.0055 measured nothing about the environment — with thinking on it scores 0.1667 on six rollouts and one of them scored a clean 1.000, so it is fully solvable. drop-table-inference 0.4174 is a mixture of two regimes (0.0894 short, 0.6465 long) and should be reported split or not at all. The defect is the missing sampling footnote, not the values. The 600-second per-call ceiling that blocked an n=32 thinking-on run is the INSTALLED wheel, not upstream: verifiers fixed it in a298bcfe on 2026-08-08 and 0.3.0 predates it. The remedy is a dependency bump, not a wait. GATE_DIAGNOSIS.md: the `gate` component scored exactly 0.000, max 0.000, across all 32 rollouts in four environments. One shape explains three of them — exact equality. Nothing below the oracle earns a fraction: not the plausible strategy, not six of redaction's seven rules, not the oracle bot list minus one account. bot-detection and schema-migration are a threshold set at the ceiling; redaction-pressure is that and genuinely hard; grand-exchange's zero was the sampling artefact above. The repo already contains the fix and already uses it twice — drop-table's GATE_MARGIN and grand-exchange's TARGET_SHARE are margins, not equalities. Recommended, not yet measured. Consequence worth stating plainly: 25-30% of the reward mass on three environments carries identically zero gradient, so they train against a 0.70-0.75 objective while being scored out of 1.00. The mirror failure exists too — fault-localisation's evidence and fault components are pinned at exactly 1.0 on all 32 rollouts, so half its headline 0.9531 is constant. tools/gate_probe.py is the reproduction, moved out of the gitignored outputs/ so the finding survives. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Arena
RL environments from Lumbridge. Apache-2.0, built to the verifiers
v1 spec so the same wheel installs from the Prime Intellect Environments Hub and runs here
unchanged.
An environment is an eval you take the gradient of. That is not a rename — it changes what the artifact has to carry. A benchmark's score is its output, so a benchmark can be a number in a table. An environment's score is the input to the next weight update, so a defect in one is inherited by the model rather than printed beside it.
Most of what is on the Hub today is a benchmark that was ported into the spec. A port cannot generate a held-out slice, because its dataset is fixed — so it is contaminated the moment anyone trains on it. And an exact-match reward is single-sided by construction, so it teaches whatever maxes the string comparison. Nothing here is a port.
The house rules
Every environment in this repository is held to three, and ships the probe that proves it. An interactive one is held to a fourth.
- Graded on what the model did not see. If a held-out slice cannot be generated, the environment does not measure generalisation and does not ship.
- No single-sided reward. Every reward has a counterweight in the same episode. A reward you can max by doing the crude thing is a reward that teaches the crude thing.
- A zero floor and a reachable ceiling, demonstrated. Inaction scores 0.0 and an oracle
scores 1.0, measured, in
probe.py. An unreachable component is dead weight in the gradient; a floor above zero pays for doing nothing. - A turn budget that binds. In an environment the model acts in over several turns, a
policy that spends every turn must score measurably below one that spends the turns worth
spending. Every reward here saturates, so without a cost for looking the budget is free
and the multi-turn form measures nothing the one-shot form did not. In
grand-exchange-livethe gap is 0.369, stdev 0.054 over five blocks of twenty-four, worst block 0.316.
Rules 3 and 4 are not decoration. Every environment here was wrong the first time and the probe is what caught it — see the table below, and the git history.
The environments
| What it measures | The trap | |
|---|---|---|
redaction-pressure |
Write provenance-safe redaction rules for a ticket corpus, run against the next records the generator would have produced. | Partial edits remain residual secrets, every non-secret source character is protected, and hostile regexes share one episode budget. |
canary-trap |
Write contamination probes that catch a model trained on your eval set. | Probes are run against a model that memorised the corpus and one that learned the same facts from a reworded copy. GUID-recall probes are sound and see only half of what is there. |
fault-localisation |
Read an incident log and name the root cause, its class, and the line proving it. | The service that broke emits one line and goes quiet; its dependents emit a dozen timeouts. Ranking by error volume answers the victim, every time. |
schema-migration |
Split a text column into a number and a unit, executed against held-out rows. | The visible rows are tidy. The graded ones carry thousands separators, negatives, a unit containing a slash, and a value with no unit at all. |
bot-detection |
Classify accounts as bot or human from activity logs, learning the rule from six labelled examples. | Click latency, session length and route count are drawn before the generator decides who is a bot, so every timing rule sits at chance. The bots that cloak leak on one behavioural channel each, and there are three. |
drop-table-inference |
Estimate a monster's drop rates from a kill log, graded on the next forty thousand kills. | Copying the observed frequencies is right about the common drops and asserts that the item which appeared zero times is impossible. Nothing in the counts distinguishes one silent item from another, so the only evidence about them is the rungs the observed items did not take and what the silent ones are worth. |
grand-exchange |
Read price and volume history for a basket of items and place limit orders, executed against the next thirty ticks. | One or two items per basket are a random walk, not a mean-reverting one — the widest-swinging lines on the board and the ones with no anchor to revert to — and you are not told how many there are, so counting is not a substitute for the shape statistic. And liquidity in gp per tick is uncorrelated with price, so the fattest visible margins sit where the purse cannot go. |
grand-exchange-live |
The same basket, traded while the window is running: eight looks over sixty ticks, with stock, offers standing between looks, and a purse that sale proceeds come back to. | Looking is not free. Re-quoting an item that already carries an offer freezes the whole book for one to four ticks, an amended offer restarts at the back of the fill queue, and the per-item buy limit is cumulative over the window rather than per offer — so the policy that re-quotes on every look scores 0.631 against a reference that looks four times and scores 1.000. |
What the probes measure
Reward for the degenerate strategies and for an oracle, from uv run python probe.py:
| environment | inaction | crude maximiser | plausible attempt | oracle |
|---|---|---|---|---|
redaction-pressure |
0.000 | 0.387 (redact everything) | 0.549 | 1.000 |
canary-trap |
0.000 | 0.000 (public-knowledge probes) | 0.525 (GUID recall) | 1.000 |
fault-localisation |
0.000 | 0.250 (blame the loudest) | 0.500 | 1.000 |
schema-migration |
0.000 | 0.000 (add columns, touch nothing) | 0.630 | 1.000 |
bot-detection |
0.000 | 0.000 (ban everyone) | 0.400 (one tell of three) | 1.000 |
drop-table-inference |
0.000 | 0.179 (copy the frequencies, floor the rest at a memorised constant) | 0.476 (per-item posterior over the grid) | 1.000 |
grand-exchange |
0.000 | 0.082 (buy and sell at market) | 0.608 (anchor on the mean and size to the volume, trap included) | 1.000 |
grand-exchange-live |
0.000 | 0.126 (buy and sell at market, on one look) | 0.406 (anchor on the mean, no trap filter, one look and no re-quote) | 1.000 |
Running one
uv run --project environments/redaction_pressure eval @ configs/redaction_pressure.toml --model <model-id>
Run its release gate from the repository root:
uv sync --project environments/redaction_pressure
uv run --project environments/redaction_pressure \
python -m unittest discover -s environments/redaction_pressure/tests -v
uv run --with regex python probe.py
Arena contains independently published environment libraries, so per-environment
uv.lock files are intentionally not committed. Redaction v0.2 pins its runtime contract
in pyproject.toml; CI resolves it on Python 3.11 and 3.12, builds every environment, and
runs the shared four-environment probe.
Publishing
The layout mirrors verifiers' own, so an environment goes to the Hub without a fork:
prime env push redaction-pressure -v PUBLIC
What a model actually scores
First evaluation, 2026-08-19: Nemotron 3.5 Lightning 30B-A3B (NVFP4, thinking off) served on one DGX Spark. 16 tasks x 2 rollouts per environment, 32 rollouts each.
| environment | reward | where it fails |
|---|---|---|
canary-trap |
0.570 | specificity 1.00 — it never accuses a clean model. Detection 0.55: it writes GUID-recall probes and is blind to the paraphrase, which is the failure the environment was built to expose |
fault-localisation |
0.352 | fault class 0.56, evidence 0.53, service 0.16. In 12 of the 17 cases where it cited the correct line it still named the wrong service — the service is printed on that line |
schema-migration |
0.350 | schema 0.59, integrity 0.38. It reaches the target shape and loses the awkward rows |
redaction-pressure |
0.343 | recall 0.45, precision 0.53 |
grand-exchange |
0.219 | profit 0.33, discipline 0.18 |
bot-detection |
0.142 | caught 0.19, spared 0.19 — it accuses humans, which is the expensive error |
drop-table-inference |
0.059 | fit 0.01. It answers in fractions (48/128, 1/3072), so it has the structural idea, and the rates are still wrong |
Every gate is at or near zero — 0.00 on four of the seven. None of these is solved, and
probe.py shows none is unsolvable. That gap is the whole point: an environment a model
already passes has no gradient left in it, and one nothing can pass has none yet.
The replies parse: 32/32 rollouts produced non-zero metrics in every environment, so these are model scores rather than a harness failing to read its own output.