kartiandClaude Opus 5 00d21025ff grand-exchange-live: the stepped market engine, and the turn budget binds
Arena's first interactive environment, as its own wheel beside grand-exchange so
the one-shot scores stay comparable rather than conflated. One basket of orders
becomes a loop: quote, see fills and the next tick, re-quote, with inventory,
open orders and cash carrying between turns.

The engine and its measurement only. No taskset, no config, no probe row — the
spec's constants were derived from a throwaway engine and did not reconcile
(impatient computed to 0.612 against a reported 0.623), so every constant here
is chosen from a ladder re-measured in the code that ships.

House rule 4 asks whether the turn budget binds. It does: exhaustive 0.631 vs
oracle 1.000, gap 0.369, stdev 0.054 over five blocks, worst block 0.316, and
0.354-0.405 at four further seed bases. That is 18x the proposed margin and it
clears the 0.25 quantum of the binary gate outright, so the comparison is not a
cliff. The saturation worry does not fire: per-basket clipping is not clipping
the mean, so exhaustive averages 0.849 on profit_ratio rather than 1.000.

The first draft's gap was FAKE — a clean-looking 0.152 produced by an under-tuned
denominator whose every-look sibling earned 22% more gp while scoring below it.
It was caught only by sweeping the reference's own family. Nothing that scores
below the oracle may earn more than it; that check belongs in every interactive
environment that follows this one.

freeze is derived from market.seed inside build_market so viability and scoring
see one value, roc is on peak rather than average capital employed, and the
reference is cached per seed so scoring never re-simulates the played episode.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 15:55:12 -07:00

Arena

RL environments from Lumbridge. Apache-2.0, built to the verifiers v1 spec so the same wheel installs from the Prime Intellect Environments Hub and runs here unchanged.

An environment is an eval you take the gradient of. That is not a rename — it changes what the artifact has to carry. A benchmark's score is its output, so a benchmark can be a number in a table. An environment's score is the input to the next weight update, so a defect in one is inherited by the model rather than printed beside it.

Most of what is on the Hub today is a benchmark that was ported into the spec. A port cannot generate a held-out slice, because its dataset is fixed — so it is contaminated the moment anyone trains on it. And an exact-match reward is single-sided by construction, so it teaches whatever maxes the string comparison. Nothing here is a port.

The house rules

Every environment in this repository is held to three, and ships the probe that proves it.

  1. Graded on what the model did not see. If a held-out slice cannot be generated, the environment does not measure generalisation and does not ship.
  2. No single-sided reward. Every reward has a counterweight in the same episode. A reward you can max by doing the crude thing is a reward that teaches the crude thing.
  3. A zero floor and a reachable ceiling, demonstrated. Inaction scores 0.0 and an oracle scores 1.0, measured, in probe.py. An unreachable component is dead weight in the gradient; a floor above zero pays for doing nothing.

Rule 3 is not decoration. Every environment here was wrong the first time and the probe is what caught it — see the table below, and the git history.

The environments

What it measures The trap
redaction-pressure Write provenance-safe redaction rules for a ticket corpus, run against the next records the generator would have produced. Partial edits remain residual secrets, every non-secret source character is protected, and hostile regexes share one episode budget.
canary-trap Write contamination probes that catch a model trained on your eval set. Probes are run against a model that memorised the corpus and one that learned the same facts from a reworded copy. GUID-recall probes are sound and see only half of what is there.
fault-localisation Read an incident log and name the root cause, its class, and the line proving it. The service that broke emits one line and goes quiet; its dependents emit a dozen timeouts. Ranking by error volume answers the victim, every time.
schema-migration Split a text column into a number and a unit, executed against held-out rows. The visible rows are tidy. The graded ones carry thousands separators, negatives, a unit containing a slash, and a value with no unit at all.
bot-detection Classify accounts as bot or human from activity logs, learning the rule from six labelled examples. Click latency, session length and route count are drawn before the generator decides who is a bot, so every timing rule sits at chance. The bots that cloak leak on one behavioural channel each, and there are three.
drop-table-inference Estimate a monster's drop rates from a kill log, graded on the next forty thousand kills. Copying the observed frequencies is right about the common drops and asserts that the item which appeared zero times is impossible. Nothing in the counts distinguishes one silent item from another, so the only evidence about them is the rungs the observed items did not take and what the silent ones are worth.
grand-exchange Read price and volume history for a basket of items and place limit orders, executed against the next thirty ticks. One or two items per basket are a random walk, not a mean-reverting one — the widest-swinging lines on the board and the ones with no anchor to revert to — and you are not told how many there are, so counting is not a substitute for the shape statistic. And liquidity in gp per tick is uncorrelated with price, so the fattest visible margins sit where the purse cannot go.

What the probes measure

Reward for the degenerate strategies and for an oracle, from uv run python probe.py:

environment inaction crude maximiser plausible attempt oracle
redaction-pressure 0.000 0.387 (redact everything) 0.549 1.000
canary-trap 0.000 0.000 (public-knowledge probes) 0.525 (GUID recall) 1.000
fault-localisation 0.000 0.250 (blame the loudest) 0.500 1.000
schema-migration 0.000 0.000 (add columns, touch nothing) 0.630 1.000
bot-detection 0.000 0.000 (ban everyone) 0.400 (one tell of three) 1.000
drop-table-inference 0.000 0.179 (copy the frequencies, floor the rest at a memorised constant) 0.476 (per-item posterior over the grid) 1.000
grand-exchange 0.000 0.082 (buy and sell at market) 0.608 (anchor on the mean and size to the volume, trap included) 1.000

Running one

uv run --project environments/redaction_pressure eval @ configs/redaction_pressure.toml --model <model-id>

Run its release gate from the repository root:

uv sync --project environments/redaction_pressure
uv run --project environments/redaction_pressure \
  python -m unittest discover -s environments/redaction_pressure/tests -v
uv run --with regex python probe.py

Arena contains independently published environment libraries, so per-environment uv.lock files are intentionally not committed. Redaction v0.2 pins its runtime contract in pyproject.toml; CI resolves it on Python 3.11 and 3.12, builds every environment, and runs the shared four-environment probe.

Publishing

The layout mirrors verifiers' own, so an environment goes to the Hub without a fork:

prime env push redaction-pressure -v PUBLIC

What a model actually scores

First evaluation, 2026-08-19: Nemotron 3.5 Lightning 30B-A3B (NVFP4, thinking off) served on one DGX Spark. 16 tasks x 2 rollouts per environment, 32 rollouts each.

environment reward where it fails
canary-trap 0.570 specificity 1.00 — it never accuses a clean model. Detection 0.55: it writes GUID-recall probes and is blind to the paraphrase, which is the failure the environment was built to expose
fault-localisation 0.352 fault class 0.56, evidence 0.53, service 0.16. In 12 of the 17 cases where it cited the correct line it still named the wrong service — the service is printed on that line
schema-migration 0.350 schema 0.59, integrity 0.38. It reaches the target shape and loses the awkward rows
redaction-pressure 0.343 recall 0.45, precision 0.53
grand-exchange 0.219 profit 0.33, discipline 0.18
bot-detection 0.142 caught 0.19, spared 0.19 — it accuses humans, which is the expensive error
drop-table-inference 0.059 fit 0.01. It answers in fractions (48/128, 1/3072), so it has the structural idea, and the rates are still wrong

Every gate is at or near zero — 0.00 on four of the seven. None of these is solved, and probe.py shows none is unsolvable. That gap is the whole point: an environment a model already passes has no gradient left in it, and one nothing can pass has none yet.

The replies parse: 32/32 rollouts produced non-zero metrics in every environment, so these are model scores rather than a harness failing to read its own output.

S
Description
RL environments from Lumbridge, built to the verifiers v1 spec.
Readme Apache-2.0
2 MiB
Languages
Python 100%