grand-exchange-live: protocol layer, taskset, config, widened reference gate

The stepped engine becomes a runnable taskset. protocol.py is both the
render/parse boundary and the security boundary: TurnView carries the live
Market, so view.market.items[i].prices is the whole held-out series and a
foresight policy reading it earned +9.4% on the oracle. The renderer emits only
observed ticks, the trace carries plain scalars, and a sentinel scan of every
rendered turn fails if either reopens. House rule 1 had no other enforcement.

Replies are a delta sheet and parsing never raises — NaN, Infinity, 1e309 and
20k-deep nesting all resolve to a hold, because the one-shot form has already
shown what a real model sends when it gives up.

LiveExchangeEnv.run drives the engine host-side with the harness left null. Two
things learned the hard way and worth keeping: a terminated Segment mid-window
holds the remaining looks instead of raising, and rewards are recorded INSIDE
the interaction, before close. Recording after close writes the traces correctly
and leaves every eval.log line reading reward=0.000 — the canonical verifiers
reference does it after the block; do not copy it there.

The reference-family gate is widened from four named rungs to the whole 8x8
(requote_band, idle_share) family: no member may out-earn the shipped reference
by more than 1.02x, and any that does must still score >= 0.90. That is the check
that caught the first draft's fake 0.152 gap, now shipped rather than remembered.

Also guards `clean` on the reference being profitable, here and in the one-shot
form. Against a reference that lost money the bar is negative, so doing nothing
cleared it and inaction collected a quarter of the reward. Unreachable while
viable_market admits only profitable references; one relaxed filter from
reachable, and exactly the floor house rule 3 forbids. Probe values unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-21 17:44:20 -07:00
co-authored by Claude Opus 5
parent e4cc2bf1c3
commit f953c03bd1
12 changed files with 1251 additions and 67 deletions
+23
View File
@@ -0,0 +1,23 @@
# grand-exchange-live — see environments/grand_exchange_live for what it measures.
#
# uv run --project environments/grand_exchange_live eval @ configs/grand_exchange_live.toml \
# --model <model-id>
#
# Eight looks per episode against one basket, so a rollout here costs roughly eight times a
# one-shot `grand-exchange` rollout in requests and rather more than that in tokens: the
# whole conversation is replayed to the model on every segment. `num_tasks` is half the
# one-shot config's for that reason, not because fewer baskets measure better — realised
# P&L over sixty ticks is the noisiest number in this repository and the ladder is read
# over five blocks of twenty-four.
num_tasks = 8
num_rollouts = 2
[env.taskset]
id = "grand-exchange-live"
num_items = 5 # items in the basket, one per price tier
visible = 56 # warmup ticks, read before the first look
# The graded window is not configurable: 8 looks x 7 ticks + a 4-tick tail = 60, and every
# constant in the reward was measured against that shape (`measure_ladder.py --sweep-shape`).
[env.agent.harness]
id = "null" # pure-trace reward: a chat loop, no tools, no runtime
@@ -412,9 +412,17 @@ class Outcome:
one-unit order that never fills has identical realised profit and lost 0.146 of one-unit order that never fills has identical realised profit and lost 0.146 of
total reward. The bar is now TARGET_SHARE of the reference on the money AND on the total reward. The bar is now TARGET_SHARE of the reference on the money AND on the
return that money made, which is a run that was worth doing however it got there. return that money made, which is a run that was worth doing however it got there.
Guarded on the reference being profitable at all. Against a reference that LOST
money the bar is negative, so doing nothing clears it and inaction collects this
whole component — the floor violation house rule 3 exists to forbid. Unreachable
while `viable_market` only admits profitable references; one relaxed filter from
reachable. `grand_exchange_live.reward.score` carries the same guard.
""" """
return ( return (
self.fills.realised >= TARGET_SHARE * self.reference.realised self.reference.realised > 0
and self.reference.roc > 0
and self.fills.realised >= TARGET_SHARE * self.reference.realised
and self.fills.roc >= TARGET_SHARE * self.reference.roc and self.fills.roc >= TARGET_SHARE * self.reference.roc
) )
@@ -1,13 +1,35 @@
"""grand-exchange-live — the stepped form of `grand-exchange`. """grand-exchange-live — the stepped form of `grand-exchange`.
⚠️ ENGINE ONLY at this commit. `taskset.py`, the protocol layer and the config are The engine (`live.py`), the market it runs on (`market.py`, `book.py` — copies of the
deliberately not here yet: the spec's constants (`horizon`, `FREEZE_RANGE`, `TARGET_SHARE`, one-shot form's, held byte-identical by `tests/test_market_copy.py`), the reward
`TURNS`) were locked against arithmetic from an engine that no longer exists, and the build (`reward.py`), the render/parse boundary (`protocol.py`) and the taskset that wires them to
order is engine → measurement → constants → taskset. `measure_ladder.py` is the `verifiers` (`taskset.py`).
measurement. Nothing imports `verifiers` yet, so the package is importable — and gradeable
— with nothing installed, exactly as `probe.py` needs it to be. ⚠️ Importing this package imports `verifiers`, through `taskset`. `probe.py` deliberately
does not: it registers each package directory as a namespace and imports the leaf modules
under it, so the floor-and-ceiling gate runs with the training stack uninstalled. Nothing
below `taskset.py` may grow a `verifiers` import, or that stops being true.
""" """
from grand_exchange_live.live import HORIZON, TURNS, Leg, Reference, Result, play from grand_exchange_live.live import HORIZON, TURNS, Leg, Reference, Result, play
from grand_exchange_live.protocol import SYSTEM, Sheet, opening, parse_sheet, render
from grand_exchange_live.reward import TARGET_SHARE, score
from grand_exchange_live.taskset import LiveExchangeEnv, LiveExchangeTaskset
__all__ = ["HORIZON", "TURNS", "Leg", "Reference", "Result", "play"] __all__ = [
"HORIZON",
"SYSTEM",
"TARGET_SHARE",
"TURNS",
"Leg",
"LiveExchangeEnv",
"LiveExchangeTaskset",
"Reference",
"Result",
"Sheet",
"opening",
"parse_sheet",
"play",
"render",
"score",
]
@@ -0,0 +1,305 @@
"""The render/parse boundary — and the security boundary.
`live.play` takes a `Policy = Callable[[TurnView], list[Leg]]`. A model is not that, so this
file is the adapter: `render` turns a `TurnView` into the text of one user turn, and
`parse_sheet` turns the reply back into legs. `taskset.py` is the only caller and does
nothing else with the view.
⚠️ THE LEAK THIS FILE EXISTS TO CLOSE. `TurnView.market` is the live `Market` object, so
`view.market.items[i].prices` is the WHOLE stream — every graded tick, including the ones
the agent is being graded on not having seen. A verifier wrote a policy that reads it and
earns 9.4% more than the oracle. House rule 1 is enforced here and nowhere else, by two
rules this module holds to absolutely:
1. Nothing is rendered from `item.prices` or `item.volumes` past `view.seen`. The warmup
is `[:view.seen]` on the opening turn and every turn after it renders `view.new_prices`
/ `view.new_volumes`, which the engine slices as `[last_seen:t]`. `tests/test_protocol.py`
builds a market whose held-out ticks are impossible sentinel values and asserts none of
them appears in any rendered turn, and separately asserts that what was rendered
reassembles into exactly `prices[:seen]`.
2. `view.market` never leaves this module — not into a return value, not into `trace.info`,
not into a metric. The taskset writes plain ints and floats and the same test asserts
that too.
The reply grammar is the one-shot form's, minus the arithmetic and plus a delta:
```json
{"orders": [{"item": "Bogwater draught", "quantity": 400, "buy": 118, "sell": 129}]}
```
A DELTA SHEET, not a book: a line replaces whatever offer stands on that item, an item not
named keeps its offer untouched, `quantity: 0` cancels, and `{"orders": []}` holds. It is
the one-shot form's object so a model that can play `grand-exchange` can play this without
learning a second syntax, and because the alternative — a bespoke amend/cancel verb set —
would put a second thing in front of the model to fail at, and the failure mode measured on
the one-shot form was not syntax.
MALFORMED REPLIES ARE COERCED TO A HOLD, AND THAT IS THE WHOLE POLICY. Not rejected, not
fined, never raised. Three reasons, in order of how much they cost:
a raise takes the rollout, not the turn. Parsing runs inside the env's `run()`, so an
exception there fails the whole episode and the trace lands with `rewards: {}` — a
formatting slip would score as an infrastructure error rather than as a bad turn, and
a run's mean would silently be a mean over the model's better-formatted episodes.
a hold is already a real cost. The look is spent: the book stands as it was, the freeze
it might have wanted is not paid, and seven more ticks execute against yesterday's limits.
There is no need to invent a fine, and a fine would be a second single-sided term (house
rule 2) sitting beside a reward that already prices doing nothing at zero.
degenerate replies are the measured failure. In 31 of 32 rollouts of the one-shot form the
model returned `{"expected_profit": 0, "orders": []}` in 18 tokens. That is a WELL-FORMED
hold, and it must score what holding scores rather than what a parse error scores, or the
environment reports a parser bug where a sampling problem is. `malformed` is recorded as a
metric so the two are separable in the trace.
"""
from __future__ import annotations
import json
import math
import re
from dataclasses import dataclass, field
from grand_exchange_live.book import MAX_ORDERS
from grand_exchange_live.live import TURNS, Leg, TurnView
from grand_exchange_live.market import (
DUMP_BASE,
DUMP_CAP,
DUMP_IMPACT,
FILL_SHARE,
STARTING_CAPITAL,
TAX,
Market,
)
MAX_LEGS = MAX_ORDERS
"""Lines one sheet may carry, the one-shot form's cap. A basket holds five items and a leg
replaces the offer on its item, so this only ever bites a reply that is repeating itself."""
SYSTEM = f"""You are trading on the Grand Exchange, live, over a window that is running.
You get {TURNS} looks at the market. Between one look and the next, ticks execute against
whatever offers you have standing, and after your last look more ticks still run. You will
never see those final ticks before they are graded.
The rules of the exchange:
- A buy fills at the tick's price when that price is at or below your buy limit. A sell
fills at the tick's price when that price is at or above your sell limit.
- PLACING an offer locks quantity x buy price out of your purse straight away, whether or
not it ever fills. You start with {STARTING_CAPITAL:,} gp. Sale proceeds DO return to the
purse and you can spend them at your next look — that is the whole reason to look again.
- You can take at most {FILL_SHARE:.0%} of a tick's traded volume, per item, per side, and only on
ticks where your limit was actually reached.
- Each item has a buy limit: the most units of it you may buy in the WHOLE window. It is
cumulative — cancelling an offer does not give the units back.
- {TAX:.0%} tax is charged on every sale.
- Stock bought on a tick cannot be sold on the same tick.
- Stock you still hold when the window closes is forced out at the last price minus
{DUMP_BASE:.0%}, plus a further {DUMP_IMPACT:.0%} for every full tick's worth of that item's median volume
you have to push through, capped at {DUMP_CAP:.0%}; then taxed. Buying what you cannot sell is a
loss, not a hold, and it is a bigger loss the more of it there is.
What a look costs:
- Touching a book that already has an offer standing FREEZES THE WHOLE BOOK for a few
ticks. Nothing fills while it is frozen — not the item you touched, all of them. The
number of ticks is stated on every turn.
- An amended offer goes to the BACK OF THE QUEUE: the fill room it had accrued is reset to
zero, so moving a limit by one gp costs you even after the freeze expires.
- Your first placement, into an empty book, is free of both.
- Doing nothing on a look costs nothing at all. Holding a good book is a real move.
Reply with ONE ```json code block containing an object with "orders": a list of at most
{MAX_LEGS} objects, each with "item" (exactly as named), "quantity", "buy" and "sell".
It is a DELTA SHEET, not your whole book:
- a line REPLACES the offer standing on that item;
- an item you do not name keeps the offer it has, untouched and unfrozen;
- "quantity": 0 cancels the offer on that item and returns its locked coins;
- {{"orders": []}} changes nothing — the correct reply when your book is where you want it.
Example:
```json
{{"orders": [{{"item": "Bogwater draught", "quantity": 400, "buy": 118, "sell": 129}}]}}
```"""
# --- rendering --------------------------------------------------------------------------
def _series(values: list[int]) -> str:
return " ".join(str(v) for v in values)
def render(view: TurnView) -> str:
"""One turn of the exchange, as the text the model reads.
Reads `view.new_prices` / `view.new_volumes` and, on the opening turn only, the warmup
slice `[:view.seen]`. Never anything else off the market — see the module docstring.
"""
lines: list[str] = []
lines.append(
f"Look {view.turn} of {view.turn + view.turns_left}. "
f"{view.ticks_left} ticks left in the window; the next ones execute as soon as you "
f"answer, and {view.turns_left} more looks follow this one."
)
lines.append(
f"Re-quoting an item that already has an offer freezes the whole book for "
f"{view.freeze} tick{'s' if view.freeze != 1 else ''}."
)
if view.frozen_for:
lines.append(f"The book is frozen for {view.frozen_for} more tick"
f"{'s' if view.frozen_for != 1 else ''}: nothing is filling.")
lines.append("")
if view.turn == 1:
lines.append(f"The last {view.seen} ticks, oldest first. Nothing of yours was "
"standing for any of them.")
lines.append("")
for item in view.market.items:
lines.append(f"{item.name} (buy limit {item.buy_limit} per window)")
lines.append(f" price {_series(item.prices[:view.seen])}")
lines.append(f" volume {_series(item.volumes[:view.seen])}")
lines.append("")
else:
ticks = len(next(iter(view.new_prices.values()), []))
lines.append(f"{ticks} tick{'s' if ticks != 1 else ''} have executed since your last "
"look, oldest first:")
lines.append("")
for name in sorted(view.new_prices):
lines.append(name)
lines.append(f" price {_series(view.new_prices[name])}")
lines.append(f" volume {_series(view.new_volumes[name])}")
lines.append("")
lines.append(f"Purse: {view.cash:,.0f} gp free. "
f"Realised so far: {view.realised_so_far:+,.0f} gp.")
if view.open_orders:
lines.append("Offers standing:")
for name in sorted(view.open_orders):
order = view.open_orders[name]
lines.append(f" {name} buy {order.buy} sell {order.sell} "
f"{order.quantity} units still sought "
f"{order.locked:,.0f} gp locked behind it")
else:
lines.append("No offers standing.")
held = {n: q for n, q in view.held.items() if q}
if held:
lines.append("Stock held:")
for name in sorted(held):
lines.append(f" {name} {held[name]} units")
else:
lines.append("You hold no stock.")
bought = {n: q for n, q in view.bought.items() if q}
if bought:
lines.append("Bought so far, against each item's window limit:")
for name in sorted(bought):
item = view.market.item(name)
limit = item.buy_limit if item is not None else 0
lines.append(f" {name} {bought[name]} of {limit}")
lines.append("")
lines.append('Your order sheet. One ```json block, {"orders": [...]}; items you do not '
'name keep their offer; {"orders": []} holds.')
return "\n".join(lines)
def opening(market: Market, turns: int = TURNS) -> str:
"""The first turn's text, built before the engine has run.
The task's prompt has to exist at `load()` time — a prompted task speaks first, and the
board is what the model reads before it says anything — but the first `TurnView` does
not exist until `play` builds it. So this reconstructs that view, and
`tests/test_protocol.py::test_the_opening_prompt_is_the_engines_own_first_turn` asserts
the two are byte-identical over twenty-four baskets. Without that assertion this is a
second, quietly drifting description of the state the engine is actually in.
"""
return render(TurnView(
turn=1,
turns_left=turns - 1,
ticks_left=market.held_out,
freeze=market.freeze,
frozen_for=0,
cash=float(market.capital),
market=market,
new_prices={},
new_volumes={},
seen=market.visible,
held={},
bought={},
open_orders={},
realised_so_far=0.0,
))
# --- parsing ----------------------------------------------------------------------------
_BLOCK = re.compile(r"```(?:json)?\s*\n(.*?)```", re.DOTALL)
@dataclass(frozen=True)
class Sheet:
"""One parsed reply. `legs` is what the engine gets, the rest is trace material."""
legs: list[Leg] = field(default_factory=list)
malformed: bool = False
"""No sheet could be read out of the reply at all. Coerced to a hold; recorded."""
dropped_rows: int = 0
"""Rows that were not usable as a leg — not an object, no item name, a quantity that is
not a number. Distinct from the legs the ENGINE drops (a name not on the board, a price
of zero), which it counts itself."""
def _int(value) -> int | None:
"""A JSON number as an int, or None. `json.loads` is not strict JSON — it accepts bare
NaN, Infinity and -Infinity, and overflows 1e309 to inf — and `int(inf)` raises
OverflowError, which inside `run()` would take the whole rollout with it."""
if isinstance(value, bool) or not isinstance(value, (int, float)):
return None
try:
if not math.isfinite(value):
return None
return int(value)
except (OverflowError, ValueError):
return None
def parse_sheet(reply: str) -> Sheet:
"""The last JSON block in the reply, as a delta sheet. Never raises.
A bare array of order objects is accepted as well as `{"orders": [...]}` — a model that
answers with the list and skips the wrapper has stated a sheet, and should be graded on
the sheet. Anything else is a hold, flagged `malformed`; see the module docstring for
why that is the policy and what it costs.
"""
blocks = _BLOCK.findall(reply or "")
raw = blocks[-1] if blocks else (reply or "")
try:
parsed = json.loads(raw.strip())
except (ValueError, RecursionError):
return Sheet(malformed=True)
if isinstance(parsed, dict):
rows = parsed.get("orders", parsed.get("legs"))
if rows is None and not parsed:
rows = [] # `{}` is a hold, stated in the shape the grammar asks for
else:
rows = parsed
if not isinstance(rows, list):
return Sheet(malformed=True)
legs, dropped = [], 0
for row in rows:
if len(legs) >= MAX_LEGS:
dropped += 1
continue
if not isinstance(row, dict) or not isinstance(row.get("item"), str):
dropped += 1
continue
quantity = _int(row.get("quantity", 0))
buy, sell = _int(row.get("buy", 0)), _int(row.get("sell", 0))
if quantity is None or buy is None or sell is None:
dropped += 1
continue
legs.append(Leg(item=row["item"], quantity=quantity, buy=buy, sell=sell))
return Sheet(legs=legs, dropped_rows=dropped)
@@ -0,0 +1,66 @@
"""The reward, in one place, so no two readers can disagree about the same episode.
`taskset.py` scores a rollout with this, `measure_ladder.py` scores its ladder with it and
`probe.py` gates the environment through the ladder — one function, three callers. It lived
in `measure_ladder.py` while the constants were being measured, which was fine while nothing
else read it and would not have survived the taskset landing: a reward the probe computes
and a reward the run computes are two rewards, and the day they drift the probe is gating
something the model is not being paid for.
Nothing here imports `verifiers`, for the same reason the rest of the package does not: the
probe has to grade this environment with the training stack uninstalled.
"""
from __future__ import annotations
from grand_exchange_live.live import Result
TARGET_SHARE = 0.90
"""What counts as a full score, as a share of the LIVE reference's realised profit and of
its return on peak capital. Chosen by `measure_ladder.py --sweep-target`, not carried over
from the spec, which proposed 0.85 against an engine that no longer exists.
The two things it trades off are both measured over 120 baskets. Raising it widens the
rule-4 gap (exhaustive scores 0.723 at 0.80 and 0.574 at 1.00) and narrows the ceiling
plateau — and the plateau is what stops the reward being an imitation score for the
reference's own constants. At 0.90 the best-earning member of the reference's family that
is NOT the reference earns 1.1% more gp and still scores 0.976, far above the 0.90 bar
`probe.py` holds the one-shot form to, and the gap to exhaustive is 0.369. It is also the
one-shot environment's value, which is one fewer constant that differs between two forms of
the same market for no measured reason."""
def score(result: Result, reference: Result,
target_share: float = TARGET_SHARE) -> dict[str, float]:
"""Three numbers and the two ratios behind them, for one played episode.
profit realised, against nine tenths of what the reference earned in the same
basket. Zero for doing nothing, and zero for a round trip that fails to
clear the tax.
discipline profit x return-on-peak-capital, relative to the reference's. A PRODUCT,
never a term beside profit: as a separate weighted term it pays in full
for one tiny position that traded perfectly, which is inaction with a
receipt. `roc` is over PEAK employed capital rather than average — see
`live.Ledger.employed` for the measured free-points bug that is.
clean cleared both bars. Binary: the window was worth trading or it was not.
"""
target = target_share * reference.realised
profit = 0.0 if target <= 0 else min(1.0, max(0.0, result.realised / target))
efficiency = 0.0 if reference.roc <= 0 else min(1.0, max(0.0, result.roc / reference.roc))
discipline = profit * efficiency
# Guarded on the reference the same way `profit` and `efficiency` are. Against a reference
# that LOST money, `target_share * reference.realised` is negative and doing nothing clears
# it — inaction scored 0.25 here, the whole weight of `clean`. Unreachable as shipped, since
# `viable_live_market` only admits profitable references and `run()` rebuilds from the
# stored viable seed, but it is one relaxed filter away from being reachable and it is
# exactly the shape house rule 3 exists to forbid.
clean = float(reference.realised > 0 and reference.roc > 0
and result.realised >= target_share * reference.realised
and result.roc >= target_share * reference.roc)
return {
"total": 0.45 * profit + 0.30 * discipline + 0.25 * clean,
"profit_ratio": profit,
"efficiency": efficiency,
"discipline": discipline,
"clean": clean,
}
@@ -0,0 +1,218 @@
"""grand-exchange-live: trade a window while it is running, eight looks at a time.
The one-shot `grand-exchange` asks for one basket of limit orders against thirty ticks the
agent never sees. This asks for a plan that survives contact with the tape: the same market,
sixty graded ticks, and eight looks spread through them. Between looks, ticks execute — buys
fill, sales pay into the purse, an anchor moves — and after the last look eleven ticks still
run. What is held out is the tail plus every tick the agent chose not to spend a look on.
Three mechanics price a look, and `live.py` documents each at length: re-quoting a live book
freezes the WHOLE book for a few ticks, an amended offer restarts at the back of the queue,
and the per-item buy limit is cumulative over the window rather than per offer. Together
they are what makes house rule 4 hold here as a measurement rather than as a hope — the
policy that spends every look scores 0.631 against the reference's 1.000, gap 0.369, stdev
0.054 over five blocks of twenty-four, worst block 0.316 (`measure_ladder.py`).
The env drives the engine host-side. `run()` opens ONE interaction and walks `live.play`
through it: the engine asks its policy for a sheet, the policy renders the turn, waits on
the model, parses the reply, and hands back legs. The engine is a synchronous loop and the
model call is not, so `play` runs in a worker thread and each policy call hops back onto the
event loop — the alternative is a second copy of the tick loop written in async, which is a
second engine to keep in step with the one the ladder and the probe measure.
Three rewards, the one-shot form's, over `reward.score`:
profit realised gp against nine tenths of what the reference earned in this basket.
discipline profit x return on PEAK employed capital, relative to the reference's. A
product, never a term beside profit.
gate cleared both bars. Binary.
⚠️ `TurnView` carries the live `Market`, and `view.market.items[i].prices` is the full
held-out series. Nothing in this file touches it: the observation comes from
`protocol.render` and every metric below is a plain float computed from `Result`. See
`protocol.py`'s docstring for the leak and the tests that hold it shut.
"""
from __future__ import annotations
import asyncio
from typing import ClassVar
from pydantic import Field
import verifiers.v1 as vf
from grand_exchange_live.live import (
HORIZON,
TURNS,
Leg,
TurnView,
live_reference,
play,
viable_live_market,
)
from grand_exchange_live.market import SEED_BASE, build_market
from grand_exchange_live.protocol import SYSTEM, opening, parse_sheet, render
from grand_exchange_live.reward import TARGET_SHARE, score
class LiveExchangeData(vf.TaskData):
seed: int
"""Rebuilds the whole basket exactly, both halves. The graded ticks are never serialized
— they are regenerated from this integer at scoring time, so the task data cannot leak
what the agent is graded on."""
num_items: int
visible: int
turns: int
class LiveExchangeTask(vf.Task[LiveExchangeData]):
@vf.stop
async def budget(self, trace: vf.Trace) -> bool:
"""The turn budget, as a backstop. `run()` asks for exactly `turns` sheets and stops,
so this only fires if the exchange is driven from somewhere else — but a multi-turn
task with no stop condition is a task whose budget is a convention, and the budget is
the thing this environment measures."""
return trace.num_turns >= self.data.turns
class LiveExchangeConfig(vf.TasksetConfig):
num_tasks: int = Field(48, ge=1)
num_items: int = Field(5, ge=1)
visible: int = Field(56, ge=8)
"""Warmup ticks, read before the first look. Enough that the mean is an anchor and not a
rumour: shorten it and the reference stops being reliably profitable, which takes the
reward's denominator with it."""
class LiveExchangeTaskset(vf.Taskset[LiveExchangeTask, LiveExchangeConfig]):
SEED_BASE: ClassVar[int] = SEED_BASE
def load(self) -> list[LiveExchangeTask]:
tasks, seed = [], self.SEED_BASE
for i in range(self.config.num_tasks):
# `viable_live_market` skips a basket the LIVE reference loses money in — every
# ratio divides by it, so a basket where it loses has no reachable ceiling — and
# the seed stored is the one it landed on, so the next scan starts after it. A
# contiguous range would hand two tasks the same basket and would grade a
# different set from the one `measure_ladder.baskets` reports on.
market = viable_live_market(
seed, self.config.num_items, self.config.visible, HORIZON
)
seed = market.seed + 1
tasks.append(
LiveExchangeTask(
LiveExchangeData(
idx=i,
name=f"basket-{market.seed}",
prompt=opening(market, TURNS),
system_prompt=SYSTEM,
seed=market.seed,
num_items=self.config.num_items,
visible=self.config.visible,
turns=TURNS,
),
self.config.task,
)
)
return tasks
class LiveExchangeEnvConfig(vf.EnvConfig):
agent: vf.AgentConfig = vf.AgentConfig()
"""The one seat, named to match `SingleAgentEnvConfig` so `--env.agent.*` addresses it
exactly as it does every other environment here. A different name parses fine and
silently leaves the seat empty."""
class LiveExchangeEnv(vf.Env[LiveExchangeEnvConfig]):
async def run(self, task: vf.Task, agents: vf.Agents) -> None:
data = task.data
market = build_market(data.seed, data.num_items, data.visible, HORIZON)
loop = asyncio.get_running_loop()
state = {"malformed": 0, "holds": 0, "dropped_rows": 0, "turns": 0, "over": False}
async with agents.agent.interaction(task) as interaction:
async def ask(view: TurnView) -> list[Leg]:
# A terminated segment means the run ended instead of answering — a token
# limit, the @stop, a harness that stopped. The naive loop calls turn()
# again and gets RuntimeError('the exchange is over'), which fails the
# episode over a model that simply stopped talking. The window keeps
# running instead: every remaining look holds, and the offers standing at
# that moment are graded against the ticks that follow.
if state["over"]:
return []
# The task is prompted, so the model speaks first: turn 1 is a BARE turn().
# Passing the opening text here raises ValueError — the prompt already
# opened the exchange. `opening()` built that prompt from this same view.
segment = await interaction.turn(None if view.turn == 1 else render(view))
if segment.terminated:
state["over"] = True
return []
state["turns"] += 1
sheet = parse_sheet(segment.last_reply)
state["malformed"] += int(sheet.malformed)
state["dropped_rows"] += sheet.dropped_rows
state["holds"] += int(not sheet.legs)
return sheet.legs
def policy(view: TurnView) -> list[Leg]:
return asyncio.run_coroutine_threadsafe(ask(view), loop).result()
# `play` is the engine the ladder and the probe measure, run unmodified. The
# thread blocks on the model; the loop stays free for every other rollout.
result = await asyncio.to_thread(play, market, policy, turns=data.turns)
# Scored INSIDE the interaction, before it closes. `Rollout.close` logs the
# rollout's reward as it finishes, so rewards recorded after the `async with`
# exits land in `traces.jsonl` and NOT in the `rollout done` line — and the
# runbook's instruction to cross-check the log against the traces would be
# checking a run's real scores against a column of zeroes. Per-trace scoring
# runs at close and only fills in decorated rewards, so these survive it.
reference = live_reference(market)
parts = score(result, reference, TARGET_SHARE)
trace = interaction.trace
trace.record_reward("profit", parts["profit_ratio"], 0.45)
trace.record_reward("discipline", parts["discipline"], 0.30)
trace.record_reward("gate", parts["clean"], 0.25)
trace.record_metrics({
"profit_ratio": parts["profit_ratio"],
"efficiency": parts["efficiency"],
"discipline": parts["discipline"],
"clean": parts["clean"],
"realised": result.realised,
"reference_realised": reference.realised,
"target_realised": TARGET_SHARE * reference.realised,
"roc": result.roc,
"reference_roc": reference.roc,
"peak_employed": result.peak_employed / max(market.capital, 1),
# What the turn budget was spent on. Recorded, never rewarded: a turn-count
# reward would be a second single-sided term, and the freeze already prices the
# look. `amendments` is the count that actually paid one.
"looks": float(result.looks),
"amendments": float(result.amendments),
"frozen_ticks": float(result.frozen_ticks),
"reference_looks": float(reference.looks),
"seen_ticks": float(result.seen_ticks),
"held_out_ticks": float(HORIZON - result.seen_ticks),
"bought": float(result.bought),
"sold": float(result.sold),
"dumped": float(result.dumped),
"spent": result.spent,
"legs_dropped": float(result.dropped),
"legs_unfunded": float(result.unfunded),
# The two ways a reply can say nothing, kept apart. A degenerate but well-formed
# `{"orders": []}` is a HOLD and scores what holding scores; a reply no sheet
# could be read out of is `malformed`. Reporting one as the other is how a
# sampling problem gets written up as a parser bug.
"malformed_replies": float(state["malformed"]),
"held_replies": float(state["holds"]),
"rows_dropped": float(state["dropped_rows"]),
"answered_turns": float(state["turns"]),
"ended_early": float(state["over"]),
})
# Plain ints only. `trace.info` takes the seed the basket was built from and the cost
# of a look in it — never the Market, which carries every held-out tick.
trace.info["basket"] = int(market.seed)
trace.info["freeze"] = int(market.freeze)
@@ -86,6 +86,17 @@ import types
from pathlib import Path from pathlib import Path
HERE = Path(__file__).resolve().parent HERE = Path(__file__).resolve().parent
sys.path.insert(0, str(HERE))
try:
# The installed package, when there is one — `taskset.py` reads `__all__` off whatever
# `sys.modules` holds, and a namespace shim has none, so an unconditional shim here
# would break every taskset lookup in any process that also imports this file.
import grand_exchange_live # noqa: F401
except ImportError:
# Nothing installed: register the package as a namespace pointing at the source and let
# the leaf modules import under it, so `__init__` — and with it `verifiers` — never runs.
# This is the path `probe.py` takes, and it is the reason this file can be graded with
# the training stack absent.
_shim = types.ModuleType("grand_exchange_live") _shim = types.ModuleType("grand_exchange_live")
_shim.__path__ = [str(HERE / "grand_exchange_live")] _shim.__path__ = [str(HERE / "grand_exchange_live")]
sys.modules["grand_exchange_live"] = _shim sys.modules["grand_exchange_live"] = _shim
@@ -117,28 +128,25 @@ from grand_exchange_live.live import ( # noqa: E402
viable_live_market, viable_live_market,
) )
from grand_exchange_live.market import SEED_BASE, build_market # noqa: E402 from grand_exchange_live.market import SEED_BASE, build_market # noqa: E402
import grand_exchange_live.reward as reward # noqa: E402
HEADER = "requote / idle" HEADER = "requote / idle"
BLOCK = 24 BLOCK = 24
BLOCKS = 5 BLOCKS = 5
TARGET_SHARE = 0.90 TARGET_SHARE = reward.TARGET_SHARE
"""What counts as a full score, as a share of the LIVE reference's realised profit and of """The ceiling band, imported rather than restated. It was defined here while it was being
its return on peak capital. Chosen by `--sweep-target`, not carried over from the spec, measured — `--sweep-target` is what chose 0.90 — and it moved into the shipped package the
which proposed 0.85 against an engine that no longer exists. day the taskset landed: a reward the probe computes and a reward the run computes are two
rewards, and the day they drift the probe is gating something the model is not paid for.
`reward.TARGET_SHARE` carries the measurement that chose it."""
The two things it trades off are both measured over 120 baskets. Raising it widens the score = reward.score
rule-4 gap (exhaustive scores 0.723 at 0.80 and 0.574 at 1.00) and narrows the ceiling """Likewise. One definition, three callers: the run, this ladder, and `probe.py` through it."""
plateau — and the plateau is what stops the reward being an imitation score for the
reference's own constants. At 0.90 the best-earning member of the reference's family that
is NOT the reference earns 1.1% more gp and still scores 0.976, far above the 0.90 bar
`probe.py` holds the one-shot form to, and the gap to exhaustive is 0.372. It is also the
one-shot environment's value, which is one fewer constant that differs between two forms of
the same market for no measured reason."""
# --- the baskets a run actually serves -------------------------------------------------- # --- the baskets a run actually serves --------------------------------------------------
def baskets(count: int, base: int = SEED_BASE): def baskets(count: int, base: int = SEED_BASE):
"""Skip-aware, exactly as `Taskset.load` will be: `viable_live_market` may pass over a """Skip-aware, exactly as `taskset.LiveExchangeTaskset.load` is: `viable_live_market` may pass over a
seed the LIVE reference loses money in, and the next task starts after the seed it seed the LIVE reference loses money in, and the next task starts after the seed it
landed on. Grading a contiguous `range` instead would grade a different set of baskets landed on. Grading a contiguous `range` instead would grade a different set of baskets
from the one a model is served.""" from the one a model is served."""
@@ -150,24 +158,6 @@ def baskets(count: int, base: int = SEED_BASE):
return out return out
# --- scoring ----------------------------------------------------------------------------
def score(result, reference, target_share: float = TARGET_SHARE) -> dict[str, float]:
"""The reward, in one place, so no two rows can disagree about the same episode."""
target = target_share * reference.realised
profit = 0.0 if target <= 0 else min(1.0, max(0.0, result.realised / target))
efficiency = 0.0 if reference.roc <= 0 else min(1.0, max(0.0, result.roc / reference.roc))
discipline = profit * efficiency
clean = float(result.realised >= target_share * reference.realised
and result.roc >= target_share * reference.roc)
return {
"total": 0.45 * profit + 0.30 * discipline + 0.25 * clean,
"profit_ratio": profit,
"efficiency": efficiency,
"discipline": discipline,
"clean": clean,
}
# --- the ladder, as policies rather than as adjectives ---------------------------------- # --- the ladder, as policies rather than as adjectives ----------------------------------
def inaction(view): def inaction(view):
"""Never submits a sheet. House rule 3's floor.""" """Never submits a sheet. House rule 3's floor."""
@@ -0,0 +1,40 @@
"""Import the package under test, installed or not — and get the same one every time.
Two facts collide here. `probe.py` and half of this suite must run with NOTHING installed,
which is why the package is registered as a namespace pointing at the source directory and
its leaf modules are imported under it: `__init__` never executes, so `verifiers` is never
required. And `test_taskset.py` needs the opposite — `verifiers` present and
`grand_exchange_live` resolving to the REAL package, because `import_taskset` reads
`__all__` off whatever is in `sys.modules` and a namespace shim has none.
A shim installed by whichever test module imported first would decide that for the whole
process. So the decision is made once, here, and every test module imports this before it
imports anything else: the real package if it is importable, the shim if it is not.
"""
from __future__ import annotations
import sys
import types
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
def shim(package: Path) -> None:
"""Register `package` as a namespace so its leaf modules import without its `__init__`."""
module = types.ModuleType(package.name)
module.__path__ = [str(package)]
sys.modules[package.name] = module
try:
import grand_exchange_live
INSTALLED = getattr(grand_exchange_live, "__all__", None) is not None
except ImportError:
# The package is there and `verifiers` is not, which is the no-dependency path.
INSTALLED = False
if not INSTALLED:
shim(ROOT / ROOT.name)
@@ -25,22 +25,22 @@ from __future__ import annotations
import random import random
import statistics import statistics
import sys import sys
import types
import unittest import unittest
from pathlib import Path from pathlib import Path
HERE = Path(__file__).resolve().parents[1] sys.path.insert(0, str(Path(__file__).resolve().parent))
ONE_SHOT = HERE.parent / "grand_exchange" import context # noqa: E402 - resolves the package before anything imports it
for _root in (HERE, ONE_SHOT): HERE = context.ROOT
_package = _root / _root.name # The one-shot environment is a sibling checkout rather than a dependency, and the fixed-point
_shim = types.ModuleType(_package.name) # assertion below needs both engines in one process. It is always a shim.
_shim.__path__ = [str(_package)] context.shim(HERE.parent / "grand_exchange" / "grand_exchange")
sys.modules[_package.name] = _shim
from grand_exchange_live.book import Order, execute, reference_orders # noqa: E402 from grand_exchange_live.book import Order, execute, reference_orders # noqa: E402
from grand_exchange_live.live import ( # noqa: E402 from grand_exchange_live.live import ( # noqa: E402
HORIZON, HORIZON,
IDLE_SHARE,
REQUOTE_BAND,
TAIL_TICKS, TAIL_TICKS,
TICKS_PER_TURN, TICKS_PER_TURN,
TURNS, TURNS,
@@ -443,10 +443,15 @@ class LadderGateTests(unittest.TestCase):
self.assertGreater(statistics.fmean(gaps), 0.25) self.assertGreater(statistics.fmean(gaps), 0.25)
def test_the_reference_is_not_merely_the_denominator(self) -> None: def test_the_reference_is_not_merely_the_denominator(self) -> None:
"""The gap has to be a fact about money, not about normalisation. If the exhaustive """The gap has to be a fact about money, not about normalisation. If a policy EARNED
policy EARNED more and scored less, the reward would be an imitation score for the more than the reference and scored below it, the reward would be an imitation score
reference's constants — which is precisely what the first cut of this engine did, for the reference's constants — which is precisely what the first cut of this engine
at 63,974 gp against the reference's 51,520.""" did, at 63,974 gp against the reference's 51,520.
Two halves, because there are two ways to be the wrong denominator. The rungs are
different strategies; the sweep below is the reference's OWN family, and that is the
half this test used to skip.
"""
rows = self.ladder.ladder(self.markets) rows = self.ladder.ladder(self.markets)
oracle = rows["oracle (the live reference)"] oracle = rows["oracle (the live reference)"]
for name in ("exhaustive (re-quote every look)", "churn (re-place at full size every look)", for name in ("exhaustive (re-quote every look)", "churn (re-place at full size every look)",
@@ -455,6 +460,59 @@ class LadderGateTests(unittest.TestCase):
self.assertLess(rows[name]["gp"], oracle["gp"], self.assertLess(rows[name]["gp"], oracle["gp"],
f"{name} earns more gp than the reference it is scored against") f"{name} earns more gp than the reference it is scored against")
def test_no_member_of_the_references_own_family_beats_it_and_scores_below_it(self) -> None:
"""The whole sixty-four-cell family, not four named rungs.
The shipped reference is not the family's argmax and this is the assertion that says
how far from it that is allowed to be. Measured over 120 baskets, 2026-08-21:
shipped (0.025, 0.50) 76,744 gp rule-4 gap 0.369 mean, 0.316 worst block
argmax (1.000, 0.50) 77,599 gp rule-4 gap 0.372 mean, 0.299 worst block
so the best member earns 1.0111x the reference, and the worst-scoring member that
out-earns it — (0.040, 0.50) — still scores 0.972. Moving the denominator to the
argmax would make the invariant exact and was measured before this test was written:
it is not taken, because the argmax's requote band is 1.0, i.e. the anchor trigger
never fires, so the denominator would be a strategy that cannot express "the anchor
moved" at all — and the freeze, the mechanic the whole environment uses to make the
turn budget bind, exists to price exactly that judgement. It also does not buy a
wider gap: 0.372 against 0.369 on the mean, and a WORSE worst block. Paying 1.1% of
gp to keep the denominator a strategy rather than a cash-threshold reflex is the
trade, and the tolerances below are the measurement, not a taste.
"""
bands = (0.0, 0.005, 0.010, 0.015, 0.025, 0.040, 0.060, 1.0)
idles = (0.0, 0.05, 0.10, 0.15, 0.20, 0.30, 0.50, 1.01)
shipped = statistics.fmean(
self.ladder.live_reference(m).realised for m in self.markets
)
for band in bands:
for idle in idles:
if (band, idle) == (REQUOTE_BAND, IDLE_SHARE):
continue
earned = statistics.fmean(
play(m, Reference(band, idle)).realised for m in self.markets
)
self.assertLessEqual(
earned, 1.02 * shipped,
f"reference family member ({band}, {idle}) earns {earned:,.0f} gp "
f"against the shipped reference's {shipped:,.0f} — the denominator is "
"leaving too much money on the table, so the reward's argmax is its "
"constants rather than the profit",
)
if earned <= shipped:
continue
got = statistics.fmean(
self.ladder.score(play(m, Reference(band, idle)),
self.ladder.live_reference(m))["total"]
for m in self.markets
)
self.assertGreaterEqual(
got, 0.90,
f"reference family member ({band}, {idle}) earns {earned:,.0f} gp "
f"against {shipped:,.0f} and scores {got:.3f} — the reward is an "
"imitation score, not a profit metric",
)
def test_looking_costs_and_a_wasted_look_costs_more(self) -> None: def test_looking_costs_and_a_wasted_look_costs_more(self) -> None:
rows = self.ladder.ladder(self.markets) rows = self.ladder.ladder(self.markets)
self.assertLess(rows["restate (re-quote the same book every look)"]["total"], self.assertLess(rows["restate (re-quote the same book every look)"]["total"],
@@ -15,23 +15,19 @@ cannot perturb a single tick. That claim is not an argument here, it is the test
from __future__ import annotations from __future__ import annotations
import sys import sys
import types
import unittest import unittest
from pathlib import Path from pathlib import Path
HERE = Path(__file__).resolve().parents[1] sys.path.insert(0, str(Path(__file__).resolve().parent))
ONE_SHOT = HERE.parent / "grand_exchange" import context # noqa: E402 - resolves the package before anything imports it
# `probe.py`'s shim, for `probe.py`'s reason: each package's `__init__` imports its taskset, HERE = context.ROOT
# which imports verifiers, and a self-check that only runs once the training stack is # The one-shot environment is a sibling checkout rather than a dependency, so it is always
# installed is a self-check nobody runs. Registering the package name as a namespace # `probe.py`'s shim, for `probe.py`'s reason: its `__init__` imports its taskset, which
# pointing at the source directory keeps intra-package imports resolving while `__init__` # imports verifiers, and a self-check that only runs once the training stack is installed is
# never executes. # a self-check nobody runs. Registering the package name as a namespace pointing at the
for _root in (HERE, ONE_SHOT): # source directory keeps intra-package imports resolving while `__init__` never executes.
_package = _root / _root.name context.shim(HERE.parent / "grand_exchange" / "grand_exchange")
_shim = types.ModuleType(_package.name)
_shim.__path__ = [str(_package)]
sys.modules[_package.name] = _shim
import grand_exchange.book as one_shot_book # noqa: E402 import grand_exchange.book as one_shot_book # noqa: E402
import grand_exchange.market as one_shot # noqa: E402 import grand_exchange.market as one_shot # noqa: E402
@@ -0,0 +1,234 @@
"""The render/parse boundary, and the leak it exists to close.
`TurnView` carries the live `Market`, so `view.market.items[i].prices` is every graded tick
including the ones the agent is scored on not having seen. A verifier wrote a policy that
reads it and earned 9.4% more than the oracle. House rule 1 is enforced by `protocol.render`
and by nothing else, so it is enforced here:
the sentinel every graded tick of a purpose-built basket is a value no real price or
volume can take, encoding its own index. Every rendered turn is scanned for
them, and any sentinel it carries must be a tick the view had already
observed. A render that reached one tick past `view.seen` fails this.
the prefix what was rendered, reassembled across all eight turns, must be exactly
`prices[:seen]` — not a subset of it, not a superset. The sentinel test says
nothing forbidden got out; this one says nothing was quietly mangled on the
way, which is the other way a boundary stops being a boundary.
And the parser, against replies a model actually emits. `json.loads` is not strict JSON: it
takes bare NaN and Infinity and overflows 1e309 to inf, and `int(inf)` raises OverflowError
— inside `run()`, which fails the whole episode over a formatting slip.
"""
from __future__ import annotations
import re
import sys
import unittest
from dataclasses import replace
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent))
import context # noqa: E402 - resolves the package before anything imports it
from grand_exchange_live.live import HORIZON, TURNS, Leg, Reference, play # noqa: E402
from grand_exchange_live.market import build_market # noqa: E402
from grand_exchange_live.protocol import ( # noqa: E402
MAX_LEGS,
opening,
parse_sheet,
render,
)
SEEDS = range(60_000, 60_024)
PRICE_SENTINEL = 9_000_000
VOLUME_SENTINEL = 8_000_000
NUMBER = re.compile(r"-?\d+")
def market(seed: int):
return build_market(seed, 5, 56, HORIZON)
def sentinel_market(seed: int):
"""The same basket with every GRADED tick replaced by a value that encodes its index and
that no real stream can produce — prices here run to five figures and volumes to three.
The visible warmup is left alone, so the render still has something true to show."""
items = []
for item in market(seed).items:
visible = item.prices[:56]
prices = visible + [PRICE_SENTINEL + t for t in range(56, len(item.prices))]
volumes = (item.volumes[:56]
+ [VOLUME_SENTINEL + t for t in range(56, len(item.volumes))])
items.append(replace(item, prices=prices, volumes=volumes))
return replace(market(seed), items=items)
def walk(basket, policy=None):
"""Play a basket, keeping (view, rendered text) for every look."""
seen = []
def spy(view):
seen.append((view, render(view)))
return policy(view) if policy else []
play(basket, spy)
return seen
class HeldOutLeakTests(unittest.TestCase):
def test_no_rendered_turn_carries_a_tick_the_view_had_not_observed(self) -> None:
found = 0
for seed in SEEDS:
basket = sentinel_market(seed)
for view, text in walk(basket, Reference()):
for token in NUMBER.findall(text):
value = int(token)
if value >= PRICE_SENTINEL:
tick = value - PRICE_SENTINEL
elif value >= VOLUME_SENTINEL:
tick = value - VOLUME_SENTINEL
else:
continue
found += 1
self.assertLess(
tick, view.seen,
f"basket {basket.seed} look {view.turn} rendered tick {tick} "
f"with only {view.seen} observed — the held-out window is on screen",
)
# Not vacuous: the ticks that HAVE executed are graded ticks and are rendered, so a
# basket that produced no sentinel at all would mean the scan found nothing to check.
self.assertGreater(found, 0)
def test_the_rendered_ticks_reassemble_into_exactly_the_observed_prefix(self) -> None:
for seed in SEEDS:
basket = market(seed)
got: dict[str, list[int]] = {i.name: [] for i in basket.items}
last_view = None
for view, text in walk(basket, Reference()):
last_view = view
name = None
for line in text.splitlines():
stripped = line.strip()
if stripped in got:
name = stripped
elif stripped.startswith("price ") and name:
got[name].extend(int(t) for t in stripped.split()[1:])
elif line.startswith(tuple(f"{i.name} (" for i in basket.items)):
name = line.split(" (")[0]
for item in basket.items:
self.assertEqual(got[item.name], item.prices[:last_view.seen],
f"{item.name} in basket {basket.seed}")
def test_the_last_look_never_sees_the_tail(self) -> None:
"""The floor on the held-out slice: even an agent that takes every look is graded on
ticks that ran after its final observation."""
for seed in SEEDS:
views = [v for v, _ in walk(market(seed), Reference())]
self.assertEqual(len(views), TURNS)
self.assertLessEqual(views[-1].seen, 56 + HORIZON - 11)
class OpeningTests(unittest.TestCase):
def test_the_opening_prompt_is_the_engines_own_first_turn(self) -> None:
"""`opening()` reconstructs the turn-1 view at `load()` time, before the engine has
run. If the two ever drift, the model reads a header describing a state the engine
is not in — and nothing else in the repository would notice."""
for seed in SEEDS:
basket = market(seed)
first = walk(basket)[0][1]
self.assertEqual(opening(basket, TURNS), first, f"basket {seed}")
def test_the_board_is_on_the_opening_turn_and_not_repeated(self) -> None:
seen = walk(market(60_000), Reference())
self.assertIn("buy limit", seen[0][1])
for _, text in seen[1:]:
self.assertNotIn("buy limit", text)
class ParseTests(unittest.TestCase):
def test_a_well_formed_sheet_parses(self) -> None:
sheet = parse_sheet(
'here you go\n```json\n{"orders": [{"item": "Bogwater draught", '
'"quantity": 400, "buy": 118, "sell": 129}]}\n```'
)
self.assertFalse(sheet.malformed)
self.assertEqual(sheet.legs, [Leg("Bogwater draught", 400, 118, 129)])
def test_an_explicit_hold_is_a_hold_and_not_a_parse_failure(self) -> None:
"""The measured failure of the one-shot form was a well-formed empty answer in 31 of
32 rollouts. It must score what holding scores, and it must not be reported as a
parser bug."""
for reply in ('```json\n{"orders": []}\n```', "```json\n[]\n```",
'```json\n{}\n```',
'```json\n{"expected_profit": 0, "orders": []}\n```'):
sheet = parse_sheet(reply)
self.assertEqual(sheet.legs, [])
self.assertFalse(sheet.malformed, reply)
def test_a_bare_array_is_a_sheet(self) -> None:
sheet = parse_sheet('```json\n[{"item": "x", "quantity": 1, "buy": 2, "sell": 3}]\n```')
self.assertEqual(sheet.legs, [Leg("x", 1, 2, 3)])
self.assertFalse(sheet.malformed)
def test_the_last_block_wins(self) -> None:
sheet = parse_sheet('```json\n{"orders": [{"item": "a", "quantity": 1, "buy": 1, '
'"sell": 2}]}\n```\non reflection\n```json\n{"orders": []}\n```')
self.assertEqual(sheet.legs, [])
def test_hostile_replies_never_raise_and_never_invent_a_leg(self) -> None:
for hostile in (
"", "no trades today", "```json\n{not json}\n```",
'```json\n{"orders": {"item": "x"}}\n```',
'```json\n{"orders": [{"item": "x", "quantity": NaN, "buy": 1, "sell": 2}]}\n```',
'```json\n{"orders": [{"item": "x", "quantity": Infinity, "buy": 1, "sell": 2}]}\n```',
'```json\n{"orders": [{"item": "x", "quantity": 1e309, "buy": 1, "sell": 2}]}\n```',
'```json\n{"orders": [{"item": "x", "quantity": 1e400, "buy": 1, "sell": 2}]}\n```',
'```json\n{"orders": [{"item": "x", "quantity": 1, "buy": -Infinity, "sell": 2}]}\n```',
'```json\n{"orders": [{"item": 7, "quantity": 1, "buy": 1, "sell": 2}]}\n```',
'```json\n{"orders": [null, 3, "x"]}\n```',
"```json\n" + "[" * 20_000 + "]" * 20_000 + "\n```",
):
sheet = parse_sheet(hostile)
self.assertEqual(sheet.legs, [], hostile[:40])
def test_a_malformed_reply_is_a_hold_and_is_flagged_as_one(self) -> None:
for hostile in ("no trades today", "```json\n{not json}\n```",
'```json\n{"orders": {"item": "x"}}\n```'):
self.assertTrue(parse_sheet(hostile).malformed, hostile[:40])
def test_a_sheet_is_capped(self) -> None:
rows = ",".join('{"item": "x%d", "quantity": 1, "buy": 1, "sell": 2}' % i
for i in range(MAX_LEGS + 5))
sheet = parse_sheet('```json\n{"orders": [%s]}\n```' % rows)
self.assertEqual(len(sheet.legs), MAX_LEGS)
self.assertEqual(sheet.dropped_rows, 5)
def test_hostile_replies_played_through_the_engine_never_break_it(self) -> None:
"""The parser guarantees nothing on its own — the legs it returns go into the book.
A leg naming nothing, a cancel of an offer that is not there and a price of zero are
all reachable from a reply, and each must be dropped rather than fatal."""
replies = [
"", "no trades today",
'```json\n{"orders": [{"item": "nothing here", "quantity": 5, "buy": 1, "sell": 2}]}\n```',
'```json\n{"orders": [{"item": "Bogwater draught", "quantity": 0, "buy": 0, "sell": 0}]}\n```',
'```json\n{"orders": [{"item": "Bogwater draught", "quantity": -5, "buy": 1, "sell": 2}]}\n```',
'```json\n{"orders": [{"item": "Bogwater draught", "quantity": 999999999, "buy": 1, "sell": 2}]}\n```',
'```json\n{"orders": []}\n```',
'```json\n{"orders": [{"item": "Bogwater draught", "quantity": 5, "buy": 1000000000, "sell": 1}]}\n```',
"```json\n{not json}\n```",
]
for seed in SEEDS:
turn = {"i": 0}
def policy(view, turn=turn):
sheet = parse_sheet(replies[turn["i"] % len(replies)])
turn["i"] += 1
return sheet.legs
result = play(market(seed), policy)
self.assertGreaterEqual(result.realised, -float(market(seed).capital))
self.assertGreaterEqual(result.peak_employed, 0.0)
if __name__ == "__main__":
unittest.main()
@@ -0,0 +1,224 @@
"""The taskset and the env: the loop, the seat, and what reaches the trace.
Everything else in this directory runs with nothing installed. This file needs `verifiers`,
so it skips when it is absent — `uv sync --project environments/grand_exchange_live` first,
exactly as the README's release gate does.
The env is driven against a stub agent rather than a model. What is being tested is the
control flow that only exists here: a synchronous engine walked through an asynchronous
interaction, a prompted task whose first turn must be BARE, a `terminated` segment arriving
mid-window, and the rule that `TurnView.market` never reaches the trace.
"""
from __future__ import annotations
import asyncio
import sys
import types
import unittest
from dataclasses import dataclass, field
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent))
import context # noqa: E402 - resolves the package before anything imports it
try:
import verifiers.v1 as vf # noqa: F401
except ImportError: # pragma: no cover - the no-dependency path
vf = None
from grand_exchange_live.live import HORIZON, TURNS # noqa: E402
from grand_exchange_live.market import build_market # noqa: E402
from grand_exchange_live.protocol import opening # noqa: E402
PRICE_SENTINEL = 9_000_000
@dataclass
class StubTrace:
rewards: dict = field(default_factory=dict)
metrics: dict = field(default_factory=dict)
info: dict = field(default_factory=dict)
def record_reward(self, name, value, weight=1.0):
self.rewards[name] = (float(value), float(weight))
def record_metrics(self, values):
self.metrics.update({k: float(v) for k, v in values.items()})
@dataclass
class StubSegment:
last_reply: str = ""
terminated: bool = False
class StubInteraction:
"""One rollout, scripted. Records what each turn was asked with."""
def __init__(self, replies, terminate_at=None):
self.replies = replies
self.terminate_at = terminate_at
self.asked: list = []
self.trace = StubTrace()
async def turn(self, message=None):
self.asked.append(message)
n = len(self.asked)
if self.terminate_at is not None and n >= self.terminate_at:
return StubSegment(terminated=True)
return StubSegment(last_reply=self.replies[(n - 1) % len(self.replies)])
async def __aenter__(self):
return self
async def __aexit__(self, *exc):
return False
class StubAgent:
def __init__(self, interaction):
self._interaction = interaction
def interaction(self, task):
return self._interaction
class StubAgents:
def __init__(self, interaction):
self.agent = StubAgent(interaction)
HOLD = '```json\n{"orders": []}\n```'
@unittest.skipIf(vf is None, "verifiers is not installed")
class TasksetTests(unittest.TestCase):
def taskset(self, num_tasks=4):
from grand_exchange_live.taskset import LiveExchangeConfig, LiveExchangeTaskset
return LiveExchangeTaskset(
LiveExchangeConfig(id="grand-exchange-live", num_tasks=num_tasks)
)
def test_the_seat_is_named_agent(self) -> None:
"""A wrong name parses fine and silently leaves the seat empty, so it is asserted
rather than read off the file."""
from grand_exchange_live.taskset import LiveExchangeEnvConfig
self.assertIn("agent", LiveExchangeEnvConfig.model_fields)
def test_tasks_carry_the_seed_and_the_opening_board(self) -> None:
tasks = self.taskset().load()
self.assertEqual(len(tasks), 4)
for task in tasks:
market = build_market(task.data.seed, task.data.num_items, task.data.visible,
HORIZON)
self.assertEqual(task.data.prompt, opening(market, TURNS))
self.assertEqual(task.data.turns, TURNS)
self.assertIn("Grand Exchange", task.data.system_prompt)
def test_two_tasks_never_share_a_basket(self) -> None:
"""`viable_live_market` skips baskets the reference loses money in; a load that did
not advance past the seed it landed on would serve the same basket twice."""
seeds = [t.data.seed for t in self.taskset(24).load()]
self.assertEqual(len(set(seeds)), len(seeds))
def test_the_task_stops_at_the_turn_budget(self) -> None:
task = self.taskset(1).load()[0]
trace = types.SimpleNamespace(num_turns=TURNS - 1)
self.assertFalse(asyncio.run(task.budget(trace)))
trace.num_turns = TURNS
self.assertTrue(asyncio.run(task.budget(trace)))
@unittest.skipIf(vf is None, "verifiers is not installed")
class EnvRunTests(unittest.TestCase):
def env_and_task(self):
from grand_exchange_live.taskset import (
LiveExchangeConfig,
LiveExchangeEnv,
LiveExchangeEnvConfig,
LiveExchangeTaskset,
)
config = LiveExchangeEnvConfig(
taskset=LiveExchangeConfig(id="grand-exchange-live", num_tasks=1)
)
# No harness, no model, no runtime: `run()` is the whole subject here, and building
# the env through its constructor would stand up an agent to reach it.
env = LiveExchangeEnv.__new__(LiveExchangeEnv)
env.config = config
return env, LiveExchangeTaskset(config.taskset).load()[0]
def run_env(self, replies, terminate_at=None, task=None, env=None):
from grand_exchange_live.taskset import LiveExchangeEnv
if env is None:
env, task = self.env_and_task()
interaction = StubInteraction(replies, terminate_at)
asyncio.run(LiveExchangeEnv.run(env, task, StubAgents(interaction)))
return interaction
def test_the_first_turn_is_bare_and_every_later_turn_carries_the_view(self) -> None:
"""The task is prompted, so the model speaks first. Passing the opening text on turn
one raises ValueError inside `verifiers` — the prompt already opened the exchange."""
interaction = self.run_env([HOLD])
self.assertEqual(len(interaction.asked), TURNS)
self.assertIsNone(interaction.asked[0])
for message in interaction.asked[1:]:
self.assertIsInstance(message, str)
self.assertIn("Look ", message)
def test_a_terminated_segment_mid_window_is_not_an_error(self) -> None:
"""The naive loop calls turn() again and gets RuntimeError('the exchange is over'),
which fails the episode over a model that stopped talking. The window runs on."""
interaction = self.run_env([HOLD], terminate_at=4)
self.assertEqual(interaction.trace.metrics["ended_early"], 1.0)
self.assertEqual(interaction.trace.metrics["answered_turns"], 3.0)
self.assertEqual(set(interaction.trace.rewards), {"profit", "discipline", "gate"})
def test_a_hold_every_turn_scores_exactly_zero(self) -> None:
"""House rule 3's floor, through the whole stack rather than through the engine —
including the degenerate reply that 31 of 32 one-shot rollouts produced."""
interaction = self.run_env([HOLD])
self.assertEqual(interaction.trace.metrics["held_replies"], float(TURNS))
self.assertEqual(interaction.trace.metrics["malformed_replies"], 0.0)
for name, (score, _) in interaction.trace.rewards.items():
self.assertEqual(score, 0.0, name)
def test_a_malformed_reply_is_a_hold_and_never_fails_the_episode(self) -> None:
interaction = self.run_env(["I would rather not."])
self.assertEqual(interaction.trace.metrics["malformed_replies"], float(TURNS))
self.assertEqual(interaction.trace.rewards["profit"][0], 0.0)
def test_a_real_sheet_trades(self) -> None:
import json
from grand_exchange_live.book import reference_orders
env, task = self.env_and_task()
market = build_market(task.data.seed, task.data.num_items, task.data.visible,
HORIZON)
sheet = "```json\n" + json.dumps({"orders": [
{"item": o.item, "quantity": o.quantity, "buy": o.buy, "sell": o.sell}
for o in reference_orders(market)
]}) + "\n```"
interaction = self.run_env([sheet, HOLD], env=env, task=task)
self.assertGreater(interaction.trace.metrics["bought"], 0.0)
self.assertGreater(interaction.trace.rewards["profit"][0], 0.0)
def test_the_trace_carries_no_held_out_tick_and_no_market(self) -> None:
"""House rule 1's other half. The observation is `protocol.render`'s problem; this is
the one that would put the whole series in the training data through `trace.info`."""
interaction = self.run_env([HOLD])
for key, value in interaction.trace.info.items():
self.assertIsInstance(value, (int, float, str), key)
for key, value in interaction.trace.metrics.items():
self.assertIsInstance(value, float, key)
self.assertLess(abs(value), PRICE_SENTINEL, key)
self.assertEqual(set(interaction.trace.info), {"basket", "freeze"})
if __name__ == "__main__":
unittest.main()