probe: gate grand-exchange-live, the first interactive environment

Eight environments gated, one warning left (tera-crow-nav). The probe calls
measure_ladder.ladder() rather than re-deriving the policies, and seeds from the
skip-aware baskets() helper — a contiguous range() grades a different set of
baskets than the model is served.

Beyond the four-rung ladder it asserts what house rule 4 actually claims: the
turn budget binds in every block of twenty-four, sweeps all 64 cells of the
reference family, scans the rendered turns for held-out ticks, and checks that
hostile replies parse to a hold and score zero.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-21 17:44:20 -07:00
co-authored by Claude Opus 5
parent f953c03bd1
commit 94ab513eda
3 changed files with 184 additions and 6 deletions
+11 -2
View File
@@ -17,6 +17,7 @@ whatever maxes the string comparison. Nothing here is a port.
## The house rules ## The house rules
Every environment in this repository is held to three, and ships the probe that proves it. Every environment in this repository is held to three, and ships the probe that proves it.
An interactive one is held to a fourth.
1. **Graded on what the model did not see.** If a held-out slice cannot be generated, the 1. **Graded on what the model did not see.** If a held-out slice cannot be generated, the
environment does not measure generalisation and does not ship. environment does not measure generalisation and does not ship.
@@ -25,9 +26,15 @@ Every environment in this repository is held to three, and ships the probe that
3. **A zero floor and a reachable ceiling, demonstrated.** Inaction scores 0.0 and an oracle 3. **A zero floor and a reachable ceiling, demonstrated.** Inaction scores 0.0 and an oracle
scores 1.0, measured, in `probe.py`. An unreachable component is dead weight in the scores 1.0, measured, in `probe.py`. An unreachable component is dead weight in the
gradient; a floor above zero pays for doing nothing. gradient; a floor above zero pays for doing nothing.
4. **A turn budget that binds.** In an environment the model acts in over several turns, a
policy that spends every turn must score measurably below one that spends the turns worth
spending. Every reward here saturates, so without a cost for looking the budget is free
and the multi-turn form measures nothing the one-shot form did not. In
`grand-exchange-live` the gap is 0.369, stdev 0.054 over five blocks of twenty-four,
worst block 0.316.
Rule 3 is not decoration. Every environment here was wrong the first time and the probe is Rules 3 and 4 are not decoration. Every environment here was wrong the first time and the
what caught it — see the table below, and the git history. probe is what caught it — see the table below, and the git history.
## The environments ## The environments
@@ -40,6 +47,7 @@ what caught it — see the table below, and the git history.
| [`bot-detection`](environments/bot_detection) | Classify accounts as bot or human from activity logs, learning the rule from six labelled examples. | Click latency, session length and route count are drawn before the generator decides who is a bot, so every timing rule sits at chance. The bots that cloak leak on one behavioural channel each, and there are three. | | [`bot-detection`](environments/bot_detection) | Classify accounts as bot or human from activity logs, learning the rule from six labelled examples. | Click latency, session length and route count are drawn before the generator decides who is a bot, so every timing rule sits at chance. The bots that cloak leak on one behavioural channel each, and there are three. |
| [`drop-table-inference`](environments/drop_table_inference) | Estimate a monster's drop rates from a kill log, graded on the next forty thousand kills. | Copying the observed frequencies is right about the common drops and asserts that the item which appeared zero times is impossible. Nothing in the counts distinguishes one silent item from another, so the only evidence about them is the rungs the observed items did not take and what the silent ones are worth. | | [`drop-table-inference`](environments/drop_table_inference) | Estimate a monster's drop rates from a kill log, graded on the next forty thousand kills. | Copying the observed frequencies is right about the common drops and asserts that the item which appeared zero times is impossible. Nothing in the counts distinguishes one silent item from another, so the only evidence about them is the rungs the observed items did not take and what the silent ones are worth. |
| [`grand-exchange`](environments/grand_exchange) | Read price and volume history for a basket of items and place limit orders, executed against the next thirty ticks. | One or two items per basket are a random walk, not a mean-reverting one — the widest-swinging lines on the board and the ones with no anchor to revert to — and you are not told how many there are, so counting is not a substitute for the shape statistic. And liquidity in gp per tick is uncorrelated with price, so the fattest visible margins sit where the purse cannot go. | | [`grand-exchange`](environments/grand_exchange) | Read price and volume history for a basket of items and place limit orders, executed against the next thirty ticks. | One or two items per basket are a random walk, not a mean-reverting one — the widest-swinging lines on the board and the ones with no anchor to revert to — and you are not told how many there are, so counting is not a substitute for the shape statistic. And liquidity in gp per tick is uncorrelated with price, so the fattest visible margins sit where the purse cannot go. |
| [`grand-exchange-live`](environments/grand_exchange_live) | The same basket, traded while the window is running: eight looks over sixty ticks, with stock, offers standing between looks, and a purse that sale proceeds come back to. | Looking is not free. Re-quoting an item that already carries an offer freezes the **whole** book for one to four ticks, an amended offer restarts at the back of the fill queue, and the per-item buy limit is cumulative over the window rather than per offer — so the policy that re-quotes on every look scores 0.631 against a reference that looks four times and scores 1.000. |
## What the probes measure ## What the probes measure
@@ -54,6 +62,7 @@ Reward for the degenerate strategies and for an oracle, from `uv run python prob
| `bot-detection` | 0.000 | 0.000 *(ban everyone)* | 0.400 *(one tell of three)* | **1.000** | | `bot-detection` | 0.000 | 0.000 *(ban everyone)* | 0.400 *(one tell of three)* | **1.000** |
| `drop-table-inference` | 0.000 | 0.179 *(copy the frequencies, floor the rest at a memorised constant)* | 0.476 *(per-item posterior over the grid)* | **1.000** | | `drop-table-inference` | 0.000 | 0.179 *(copy the frequencies, floor the rest at a memorised constant)* | 0.476 *(per-item posterior over the grid)* | **1.000** |
| `grand-exchange` | 0.000 | 0.082 *(buy and sell at market)* | 0.608 *(anchor on the mean and size to the volume, trap included)* | **1.000** | | `grand-exchange` | 0.000 | 0.082 *(buy and sell at market)* | 0.608 *(anchor on the mean and size to the volume, trap included)* | **1.000** |
| `grand-exchange-live` | 0.000 | 0.126 *(buy and sell at market, on one look)* | 0.406 *(anchor on the mean, no trap filter, one look and no re-quote)* | **1.000** |
## Running one ## Running one
+173 -3
View File
@@ -60,6 +60,17 @@ the first time and this is what caught it:
this file asserts all three — over five blocks of twenty-four rather this file asserts all three — over five blocks of twenty-four rather
than one, because thirty ticks of P&L is the noisiest number here and than one, because thirty ticks of P&L is the noisiest number here and
the first block was the shortcut's luckiest. the first block was the shortcut's luckiest.
grand-exchange-live reported a rule-4 gap of 0.152 that was not a fact about the turn
budget at all. Its reference never redeployed the coins its own sales
earned, so a policy that simply re-quoted on every look earned 22% MORE
gp and 12% more return while scoring below it — the same imitation-score
defect the one-shot form paid for once already, wearing a different
number. The reference redeploys idle cash now, the gap is 0.369 with a
worst block of 0.316, and this file sweeps the WHOLE reference family
rather than four named rungs. Separately, `TurnView` carries the live
market, so the protocol layer is the only thing standing between the
held-out ticks and the screen: a policy that read them earned 9.4% over
the oracle, and the sentinel scan below is what holds that shut.
Exit code is the API: 1 if any environment's floor is above zero or its ceiling is Exit code is the API: 1 if any environment's floor is above zero or its ceiling is
unreachable. Nothing here is listed by hand — the environments are `environments/*/pyproject.toml` unreachable. Nothing here is listed by hand — the environments are `environments/*/pyproject.toml`
@@ -82,6 +93,7 @@ import sys
import tomllib import tomllib
import types import types
from collections.abc import Callable from collections.abc import Callable
from dataclasses import replace as dataclass_replace
from pathlib import Path from pathlib import Path
ENVS = Path(__file__).parent / "environments" ENVS = Path(__file__).parent / "environments"
@@ -812,6 +824,160 @@ def drops() -> dict[str, float]:
} }
@probes("grand-exchange-live")
def exchange_live() -> dict[str, float]:
"""The stepped form. Rules 3 and 4, and the two ways this one could be quietly wrong.
The ladder is not re-derived here. `environments/grand_exchange_live/measure_ladder.py`
is where every constant in this environment was measured, it plays the shipped engine
through the shipped reward, and re-implementing its nine policies in this file would
give the probe a second opinion about what a strategy scores. It is imported instead —
120 baskets in about a second — and the seeds come from `measure_ladder.baskets`, which
is SKIP-AWARE: `viable_live_market` passes over a basket the live reference loses money
in, so a contiguous `range` would grade a different set of baskets from the one a run
serves.
"""
sys.path.insert(0, str(ENVS / "grand_exchange_live"))
import measure_ladder as ladder_module
from grand_exchange_live.live import IDLE_SHARE, REQUOTE_BAND, Reference, live_reference, play
from grand_exchange_live.protocol import parse_sheet, render
from grand_exchange_live.reward import score
BLOCK = ladder_module.BLOCK
LADDER = ladder_module.baskets(BLOCK * ladder_module.BLOCKS)
BLOCKS = [LADDER[i:i + BLOCK] for i in range(0, len(LADDER), BLOCK)]
rows = ladder_module.ladder(LADDER)
# --- rule 4, the reason this environment exists in a second form ---------------------
# Every reward in Arena saturates, so a policy that spends the whole turn budget scores
# the ceiling unless spending it COSTS something. Block by block, not on the 120-basket
# mean: this generator swings 0.48 against 0.18 between blocks of twenty-four, and a
# mean that holds while one block of the size a run actually grades goes the other way
# is a mean measured off the seed.
gaps = []
for block in BLOCKS:
gap = 0.0
for market in block:
reference = live_reference(market)
exhaustive = play(market, Reference(requote_band=0.0))
gap += score(reference, reference)["total"] - score(exhaustive, reference)["total"]
gaps.append(gap / len(block))
assert gaps[-1] > 0.10, \
f"baskets {block[0].seed}-{block[-1].seed}: spending every look scores " \
f"{gaps[-1]:.3f} below the reference — the turn budget is a formality"
assert statistics.fmean(gaps) > 0.25, \
f"the turn budget binds by {statistics.fmean(gaps):.3f} on average — too thin to " \
"survive the noise in sixty ticks of P&L"
# --- the denominator has to be the profit-maximising point of its own family ---------
# Same assertion the one-shot form ships, over the two triggers that make the LIVE
# reference act rather than over its bands. If some other member of the family earns
# more gp and scores less, the reward's argmax is a pair of constants rather than a
# strategy — which is exactly what the first cut of this engine did, at 63,974 gp
# against the reference's 51,520.
#
# The shipped reference is not the family's argmax and the tolerances are the measured
# distance to it: (1.0, 0.50) earns 1.0111x and every member that out-earns the
# reference scores at least 0.972. Moving the denominator there was measured and not
# taken — a requote band of 1.0 never fires, so the denominator could not express "the
# anchor moved" at all, and the freeze exists to price exactly that judgement. It buys
# no gap either: 0.372 mean against 0.369, and a worse worst block.
shipped = statistics.fmean(live_reference(m).realised for m in LADDER)
for band in (0.0, 0.005, 0.010, 0.015, 0.025, 0.040, 0.060, 1.0):
for idle in (0.0, 0.05, 0.10, 0.15, 0.20, 0.30, 0.50, 1.01):
if (band, idle) == (REQUOTE_BAND, IDLE_SHARE):
continue
plays = [play(m, Reference(band, idle)) for m in LADDER]
earned = statistics.fmean(p.realised for p in plays)
assert earned <= 1.02 * shipped, \
f"reference family member ({band}, {idle}) earns {earned:,.0f} gp against " \
f"the shipped reference's {shipped:,.0f} — the denominator is leaving money " \
"on the table, so the reward's argmax is its constants rather than the profit"
if earned <= shipped:
continue
got = statistics.fmean(
score(p, live_reference(m))["total"] for p, m in zip(plays, LADDER)
)
assert got >= 0.90, \
f"reference family member ({band}, {idle}) earns {earned:,.0f} gp against " \
f"{shipped:,.0f} and scores {got:.3f} — the reward is an imitation score, " \
"not a profit metric"
# --- house rule 1, which in this environment lives in the protocol layer -------------
# `TurnView` carries the live `Market`, so `view.market.items[i].prices` is the whole
# held-out series and a policy that reads it earns 9.4% more than the oracle. Nothing
# but `protocol.render` stands between that and the model. Every graded tick of this
# basket is replaced by a value no real price or volume can take, encoding its own
# index; any sentinel that reaches the screen must be a tick the view had already
# observed.
PRICE, VOLUME = 9_000_000, 8_000_000
leaky = LADDER[0]
items = [
dataclass_replace(i, prices=i.prices[:leaky.visible]
+ [PRICE + t for t in range(leaky.visible, len(i.prices))],
volumes=i.volumes[:leaky.visible]
+ [VOLUME + t for t in range(leaky.visible, len(i.volumes))])
for i in leaky.items
]
sentinels = 0
def watch(view):
nonlocal sentinels
for token in re.findall(r"\d+", render(view)):
value = int(token)
tick = (value - PRICE if value >= PRICE
else value - VOLUME if value >= VOLUME else None)
if tick is None:
continue
sentinels += 1
assert tick < view.seen, \
f"look {view.turn} renders tick {tick} with {view.seen} observed — the " \
"held-out window is on screen"
return Reference()(view)
play(dataclass_replace(leaky, items=items), watch)
assert sentinels > 0, "the leak check scanned nothing — no graded tick was ever rendered"
# --- a reply nothing can be read out of is a hold, and never raises ------------------
# Parsing runs inside the env's `run()`, so a raise fails the whole episode and the
# trace lands with `rewards: {}` — a formatting slip would be recorded as an
# infrastructure error. `json.loads` is not strict JSON: it takes bare NaN and Infinity
# and overflows 1e309 to inf, and int(inf) raises OverflowError.
victim = LADDER[0].items[0].name
for hostile in (
"", "no trades today", "```json\n{not json}\n```", "```json\n[]\n```",
'```json\n{"orders": {"item": "x"}}\n```',
'```json\n{"expected_profit": 0, "orders": []}\n```',
'```json\n{"orders": [{"item": "%s", "quantity": Infinity, "buy": 1, "sell": 2}]}\n```' % victim,
'```json\n{"orders": [{"item": "%s", "quantity": 1e309, "buy": 1, "sell": 2}]}\n```' % victim,
'```json\n{"orders": [{"item": "%s", "quantity": NaN, "buy": -Infinity, "sell": 2}]}\n```' % victim,
"```json\n" + "[" * 20_000 + "]" * 20_000 + "\n```",
):
legs = parse_sheet(hostile).legs
assert legs == [], f"hostile reply {hostile[:40]!r} invented {len(legs)} legs"
played = play(LADDER[0], lambda view, l=legs: l)
assert score(played, live_reference(LADDER[0]))["total"] <= 1e-9, \
f"hostile reply {hostile[:40]!r} scores above zero"
# --- and the rungs stay in the order the environment claims they are in --------------
for lower, upper in (("crude (market orders, one look)",
"plausible (mean anchor, no filter)"),
("plausible (mean anchor, no filter)",
"impatient (one-shot reference, one look)"),
("restate (re-quote the same book every look)",
"impatient (one-shot reference, one look)")):
assert rows[lower]["total"] < rows[upper]["total"], \
f"'{lower}' scores {rows[lower]['total']:.3f} against '{upper}' at " \
f"{rows[upper]['total']:.3f} — the ladder is upside down"
return {
"inaction": rows["inaction"]["total"],
"crude": rows["crude (market orders, one look)"]["total"],
"plausible": rows["plausible (mean anchor, no filter)"]["total"],
"oracle": rows["oracle (the live reference)"]["total"],
}
def main() -> int: def main() -> int:
discovered = discover() discovered = discover()
if not discovered: if not discovered:
@@ -827,9 +993,13 @@ def main() -> int:
# A WARNING and not exit 1, deliberately. A new environment's `pyproject.toml` lands before # A WARNING and not exit 1, deliberately. A new environment's `pyproject.toml` lands before
# its probe does, several sessions share this repository, and a hard gate here turns CI red # its probe does, several sessions share this repository, and a hard gate here turns CI red
# for all of them the moment someone scaffolds a directory. This becomes `return 1` in the # for all of them the moment someone scaffolds a directory. `exchange_live()` has landed;
# same commit that lands `exchange_live()` — the last environment that has a manifest and no # what is left unregistered is the spatial side, whose taskset layer is being written in
# probe. Until then an ungated environment is loud but not fatal. # another session. This becomes `return 1` in the commit that lands the last of them. Until then an ungated environment is loud but not fatal.
#
# `tests/test_probe.py::PENDING_PROBES` is the floor that keeps this warning honest: an
# environment may only be ungated if it is named there, so degating one by deleting its
# `@probes(...)` decorator fails the suite rather than moving quietly into this line.
missing = sorted(set(discovered) - set(_PROBES)) missing = sorted(set(discovered) - set(_PROBES))
if missing: if missing:
print(f"warning: no probe is registered for {', '.join(missing)} — NOT GATED", print(f"warning: no probe is registered for {', '.join(missing)} — NOT GATED",
-1
View File
@@ -21,7 +21,6 @@ ROW = re.compile(r"^(?P<name>\S(?:.*\S)?)\s{2,}(?:-?\d+\.\d+\s*){4}\s+(?P<verdic
# #
# It only ever shrinks. Delete a name here in the same commit that lands its probe. # It only ever shrinks. Delete a name here in the same commit that lands its probe.
PENDING_PROBES = frozenset({ PENDING_PROBES = frozenset({
"grand-exchange-live",
"tera-crow-nav", "tera-crow-nav",
"tera-drive-101", "tera-drive-101",
"tera-office-nav", "tera-office-nav",