grand-exchange-live: the stepped market engine, and the turn budget binds

Arena's first interactive environment, as its own wheel beside grand-exchange so
the one-shot scores stay comparable rather than conflated. One basket of orders
becomes a loop: quote, see fills and the next tick, re-quote, with inventory,
open orders and cash carrying between turns.

The engine and its measurement only. No taskset, no config, no probe row — the
spec's constants were derived from a throwaway engine and did not reconcile
(impatient computed to 0.612 against a reported 0.623), so every constant here
is chosen from a ladder re-measured in the code that ships.

House rule 4 asks whether the turn budget binds. It does: exhaustive 0.631 vs
oracle 1.000, gap 0.369, stdev 0.054 over five blocks, worst block 0.316, and
0.354-0.405 at four further seed bases. That is 18x the proposed margin and it
clears the 0.25 quantum of the binary gate outright, so the comparison is not a
cliff. The saturation worry does not fire: per-basket clipping is not clipping
the mean, so exhaustive averages 0.849 on profit_ratio rather than 1.000.

The first draft's gap was FAKE — a clean-looking 0.152 produced by an under-tuned
denominator whose every-look sibling earned 22% more gp while scoring below it.
It was caught only by sweeping the reference's own family. Nothing that scores
below the oracle may earn more than it; that check belongs in every interactive
environment that follows this one.

freeze is derived from market.seed inside build_market so viability and scoring
see one value, roc is on peak rather than average capital employed, and the
reference is cached per seed so scoring never re-simulates the played episode.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-21 15:55:12 -07:00
co-authored by Claude Opus 5
parent 8fe348ca82
commit 00d21025ff
8 changed files with 2650 additions and 0 deletions
@@ -0,0 +1,13 @@
"""grand-exchange-live — the stepped form of `grand-exchange`.
⚠️ ENGINE ONLY at this commit. `taskset.py`, the protocol layer and the config are
deliberately not here yet: the spec's constants (`horizon`, `FREEZE_RANGE`, `TARGET_SHARE`,
`TURNS`) were locked against arithmetic from an engine that no longer exists, and the build
order is engine → measurement → constants → taskset. `measure_ladder.py` is the
measurement. Nothing imports `verifiers` yet, so the package is importable — and gradeable
— with nothing installed, exactly as `probe.py` needs it to be.
"""
from grand_exchange_live.live import HORIZON, TURNS, Leg, Reference, Result, play
__all__ = ["HORIZON", "TURNS", "Leg", "Reference", "Result", "play"]
@@ -0,0 +1,532 @@
"""Executing the agent's orders against the held-out window, and scoring the result.
⚠️ A COPY of `grand_exchange/book.py`, held byte-identical in behaviour by
`tests/test_market_copy.py` — see the note at the top of `market.py` for why it is a copy
and not an import. Nothing in this file steps: it is the one-shot engine, kept here because
the live reference bootstraps its first quote out of `reference_orders`, `crossings` and
`MIN_CROSSINGS`, and because `execute` is the fixed point the stepped engine is checked
against (`live.py`'s `test_single_look_matches_one_shot`). The stepped engine is `live.py`.
Realised profit is unbounded above and has no theoretical optimum the data supports, so
scoring it against one would put the ceiling out of reach — which house rule 3 forbids.
Every ratio here is therefore taken against a REFERENCE STRATEGY (`reference_orders`)
that is computed from the visible half and nothing else, and executed through the same
engine as the agent's orders. That is what makes 1.000 reachable by construction, and it
is honest: the model is asked to match a strategy available to anything that can read the
prompt, not to beat hindsight it was never shown.
A reference in the denominator has its own failure mode, and the first cut of this file had
it. If the ceiling is the reference EXACTLY, then above the clip excess profit is worth
nothing while every deviation still costs, so the reward stops being a profit metric and
becomes an imitation score for a strategy whose constants are nowhere in the prompt. It was
measurable: a wider sell band earned sixteen percent more gp and scored 0.812, and one
wasted unit of capital with identical realised profit cost 0.146. Two changes answer it.
The reference's own constants are now the profit-maximising point of its family, swept on
seeds no task is built from. And the ceiling is a BAND — TARGET_SHARE of the reference —
so the whole plateau around it, and everything above it, scores 1.000.
Three numbers come out of a run:
profit_ratio clip(realised / (TARGET_SHARE x reference realised), 0, 1). No orders is
zero, and a round trip that does not clear the tax is negative and also
zero.
efficiency realised profit per coin LOCKED behind an offer, against the reference's
own. This is where the volume limit bites: a fat margin on an item
carrying three thousand gp a tick is an offer that sits there, and a
sitting offer is capital a better item did not get. It also catches the
plan that fills perfectly at prices that were never worth reaching, which
a plain fill rate scores as discipline.
clean cleared the bar on both. Binary, because a trading run either was worth
doing or was not.
`efficiency` is never a reward on its own — one tiny order that traded perfectly would earn
it in full for near-inaction, which is the free-points defect schema-migration shipped with.
It multiplies into profit instead, so it can only ever qualify profit that exists.
"""
from __future__ import annotations
import json
import math
import re
import statistics
from dataclasses import dataclass
from grand_exchange_live.market import (
DUMP_BASE,
DUMP_CAP,
DUMP_IMPACT,
FILL_SHARE,
TAX,
Item,
Market,
build_market,
)
MAX_ORDERS = 12
BUY_BAND = 0.05
SELL_BAND = 0.05
"""The band the reference buys under its anchor and sells over it.
These two numbers are the reward's denominator, so where they sit is not a taste question.
They are the profit-maximising point of the reference's own family, found by sweeping
(buy band, sell band, crossings threshold, purse cap, size multiplier) over 160 baskets
drawn from SEED 400,000 ONWARD — a range no task is ever built from — and they are within
three percent of the sweep's best cell over the whole plateau.
The first cut used a sell band of 0.01, and that was the defect a reviewer found: a strategy
with a wider sell band earned sixteen percent more gp than the reference and scored 0.812,
because every ratio here divides by the reference and clips at one. When the denominator is
a strategy that leaves money on the table, the reward's argmax is the denominator's
hyperparameters rather than the profit. Tuning the reference to its own family's optimum is
half the fix; TARGET_SHARE below is the other half."""
TARGET_SHARE = 0.90
"""What counts as a full score, as a share of the reference's realised profit.
A ratio that divides by the reference EXACTLY makes the ceiling a single point, and a point
ceiling turns the reward into an imitation score: excess profit is worth nothing above the
clip while any deviation is punished, so the gradient near the top points at replicating a
strategy the prompt does not contain rather than at making money. Measured on the first cut,
one wasted unit of capital with byte-identical realised profit cost 0.146 of total reward,
and a one-gp change to every buy limit that EARNED 463 gp a basket more cost 0.195.
Normalising against nine tenths of the reference makes the ceiling a BAND. Everything from
"ten percent short of the reference" upward scores 1.000, so the whole plateau around the
reference — and everything above it — is the argmax, and a rounding difference costs
nothing. The reference still scores exactly 1.000, so house rule 3 is unchanged."""
MIN_CROSSINGS = 0.25
"""Below this share of ticks crossing the mean, the reference will not trade the item at
all. Over 4,000 baskets a reverting series recrosses its own mean 0.362 of the time and a
random walk 0.113, and the threshold sits in the gap: it lets 4.2% of walks through and
turns away 4.1% of reverters. This one line is the difference between a strategy and a
superstition — without it the reference buys the widest-swinging line in the basket, which
is exactly the line that has no anchor to revert to."""
MAX_ITEM_SHARE = 0.40
"""No more than this share of the purse behind one name. The anchor is an ESTIMATE, and one
bad estimate carrying the whole purse is the only way this strategy loses money over a
window. Spread over the 3.4 names it trades on average it profits in 3,976 of 4,000
baskets, and `viable_market` refuses the other twenty-four."""
VIABILITY_TRIES = 64
@dataclass(frozen=True)
class Order:
item: str
quantity: int
buy: int
sell: int
@dataclass
class Fills:
"""What one plan actually did in the window."""
realised: float = 0.0
"""Closing capital minus starting capital: the only number that pays."""
paper: float = 0.0
"""What the orders CLAIM, filled in full at the limit prices. Not part of any reward —
it exists so `arith_ok` can ask whether the agent's stated expected profit matches the
sum its own orders imply."""
bought: int = 0
sold: int = 0
planned: int = 0
spent: float = 0.0
dumped: int = 0
orders: int = 0
committed: float = 0.0
"""Coins actually locked behind offers. Capital left idle earns nothing, which is the
only reason allocation is a decision."""
offered: int = 0
"""Units the purse could fund, against `planned` units asked for."""
@property
def roc(self) -> float:
"""Realised profit per coin locked behind an offer.
This is the counterweight the profit term is qualified by, and it replaced a plain
fill rate. A fill rate only asks whether an offer was reachable; it says nothing
about whether reaching it was worth doing, so a plan that bought badly at prices it
was always going to reach scored as disciplined as one that bought well. Return on
the coins actually committed asks both questions at once, and it cannot be won by
committing nothing: it multiplies into profit, which needs the coins.
"""
if self.committed <= 0.0:
return 0.0
return self.realised / self.committed
@property
def conversion(self) -> float:
"""Of the coins locked behind offers, the share that actually bought stock.
The first version of this was realised profit over CLAIMED profit, and it was
gameable in one line: set the sell limit a hair over the buy limit and the claim
goes to nearly nothing, so the ratio goes to nearly anything. Coins converted
cannot be talked down — the only way to raise it is to offer a quantity the volume
supports at a price the market reaches, which is the lesson the number is for.
"""
if self.committed <= 0.0:
return 0.0
return min(1.0, self.spent / self.committed)
def paper_profit(orders: list[Order]) -> float:
"""The profit the orders assert, at their own limit prices, net of tax on the sale.
Deliberately the naive multiplication — it is the sum the agent is claiming, and
`arith_ok` checks whether the agent can do it."""
return sum(o.quantity * (o.sell * (1.0 - TAX) - o.buy) for o in orders)
def execute(market: Market, orders: list[Order]) -> Fills:
"""Walk the held-out ticks once, filling what the purse and the book support.
Rules, all quoted to the agent:
- COINS ARE LOCKED WHEN THE OFFER IS PLACED, quantity times buy price, in the order
the orders were listed. This is the exchange's own behaviour and it is what makes
the task an allocation: an offer that never fills has still spent the capital a
better offer needed, so ordering the whole basket at its buy limit is not free.
It is also why a buy limit of a billion buys nothing — nobody can fund that offer.
- a buy fills at the tick's price when it is at or under the limit; a sell fills at
the tick's price when it is at or over. You are never filled at your own limit
when the market is better than it, because that is not how a limit works.
- at most FILL_SHARE of a tick's volume per side, per item, and only on the ticks the
offer was eligible on — flow you were not in the market for is not yours to bank.
- stock bought this tick cannot be sold this tick. Without that, a buy limit above a
sell limit is a free round trip on a single price.
- sale proceeds land in the purse but fund nothing: every offer was placed up front.
- stock still held at the close is forced out at DUMP_BASE under the last price plus
DUMP_IMPACT for every tick's worth of the item's median volume being pushed through.
Depth is taken from the window the stock is actually being sold into, which is the
window the agent estimated from the visible median.
"""
fills = Fills(orders=len(orders), planned=sum(o.quantity for o in orders))
fills.paper = paper_profit(orders)
# Placement: fund each offer in submission order out of one purse.
purse = float(market.capital)
live: list[tuple[Order, Item, int]] = []
for order in orders:
item = market.item(order.item)
if item is None or order.buy <= 0:
continue
qty = min(order.quantity, item.buy_limit, int(purse // order.buy))
if qty <= 0:
continue
purse -= qty * order.buy
live.append((order, item, qty))
fills.committed = market.capital - purse
fills.offered = sum(q for _, _, q in live)
remaining = [q for _, _, q in live]
available = [0] * len(live)
pending = [0] * len(live)
# Capacity carries between eligible ticks instead of being truncated at each one. An
# item that trades two units a tick would otherwise be untradeable rather than thin,
# because a quarter of two is zero every time — and thin is what this is about.
buy_room = [0.0] * len(live)
sell_room = [0.0] * len(live)
proceeds = 0.0
for t in range(market.visible, market.visible + market.held_out):
for i in range(len(live)):
available[i] += pending[i]
pending[i] = 0
for i, (order, item, _) in enumerate(live):
price = item.prices[t]
flow = FILL_SHARE * item.volumes[t]
# Room accrues only on the ticks the order was actually eligible on. Accruing it
# every tick banks the flow of ticks the price never reached, which would let an
# offer fill far past the volume that was ever available to it — and the volume
# limit is the whole reason allocation is a decision here.
if price <= order.buy:
buy_room[i] += flow
if price >= order.sell:
sell_room[i] += flow
if price <= order.buy and remaining[i] > 0:
qty = min(remaining[i], int(buy_room[i]))
if qty > 0:
buy_room[i] -= qty
remaining[i] -= qty
pending[i] += qty
fills.spent += qty * price
fills.bought += qty
if price >= order.sell and available[i] > 0:
qty = min(available[i], int(sell_room[i]))
if qty > 0:
sell_room[i] -= qty
proceeds += qty * price * (1.0 - TAX)
available[i] -= qty
fills.sold += qty
for i, (_, item, _) in enumerate(live):
left = available[i] + pending[i]
if left:
depth = statistics.median(item.volumes[market.visible:])
haircut = min(DUMP_CAP, DUMP_BASE + DUMP_IMPACT * (left / max(1.0, depth)))
fills.dumped += left
proceeds += left * item.prices[-1] * (1.0 - haircut) * (1.0 - TAX)
# Coins locked behind an offer that never filled come back when the window closes, so
# the only thing a wasted offer costs is the profit the capital did not make. That is
# the right price for it: an opportunity cost, not a fine.
fills.realised = proceeds - fills.spent
return fills
def crossings(prices: list[int]) -> float:
"""Share of consecutive ticks that straddle the series' own mean.
The whole discrimination, in one countable number, so it is available to anything that
can read the prompt — no variance ratio, no regression, just how often the line cuts
its own average.
"""
if len(prices) < 2:
return 0.0
anchor = statistics.fmean(prices)
above = [p > anchor for p in prices]
return sum(1 for i in range(1, len(above)) if above[i] != above[i - 1]) / (len(above) - 1)
def reference_orders(market: Market) -> list[Order]:
"""The strategy the reward is normalised against, computed from the visible half only.
Per item: anchor on the mean of the visible prices, buy a band under it, sell a band
over it. Then three judgements, and each of them is a way the reward discriminates:
does it revert `crossings` over MIN_CROSSINGS, or the item is skipped. The decoy
swings widest and is worth nothing.
what it pays the average visible price BELOW the buy limit against the average
ABOVE the sell limit — because a limit fills at the market, not at
the limit, so a wide reverting item pays far more than its band.
Ranking by the band alone ranks every item identically.
what it can hold a quarter of a typical tick's volume, over the ticks that touched
the buy limit. Capital committed beyond that is capital locked
behind an offer that will not fill.
Best return on capital first, until the purse is gone. Nothing here reads a held-out
tick; every input is a column the agent was shown.
"""
plans = []
for item in market.items:
prices = item.visible_prices(market.visible)
volumes = item.visible_volumes(market.visible)
if crossings(prices) < MIN_CROSSINGS:
continue
anchor = statistics.fmean(prices)
buy = max(1, round(anchor * (1.0 - BUY_BAND)))
sell = max(buy + 1, round(anchor * (1.0 + SELL_BAND)))
lows = [p for p in prices if p <= buy]
highs = [p for p in prices if p >= sell]
if not lows or not highs:
continue
entry, exit_ = statistics.fmean(lows), statistics.fmean(highs)
expected = (exit_ * (1.0 - TAX) - entry) / entry
if expected <= 0:
continue
reachable = int(
FILL_SHARE * statistics.median(volumes) * (len(lows) / len(prices)) * market.held_out
)
qty = min(item.buy_limit, reachable, int(MAX_ITEM_SHARE * market.capital // buy))
if qty <= 0:
continue
plans.append((expected, item.name, qty, buy, sell))
plans.sort(key=lambda p: -p[0])
orders, purse = [], float(market.capital)
for _, name, qty, buy, sell in plans:
qty = min(qty, int(purse // buy))
if qty <= 0:
continue
purse -= qty * buy
orders.append(Order(item=name, quantity=qty, buy=buy, sell=sell))
return orders
def viable_market(seed: int, num_items: int, visible: int, held_out: int) -> Market:
"""The next basket from `seed` onward in which the reference strategy makes money.
Every ratio in the reward divides by the reference's realised profit, so a basket where
the reference loses has no reachable ceiling and would quietly break house rule 3 for
that task — the oracle would score below 1.000 and nothing would say why. About one
basket in two thousand is like that, from an anchor the visible half happened to
mis-estimate. Skipping it is a guard, not a crutch, and the seed that was used travels
on the Market so scoring rebuilds exactly the basket that was shown.
"""
for offset in range(VIABILITY_TRIES):
market = build_market(seed + offset, num_items, visible, held_out)
if execute(market, reference_orders(market)).realised > 0:
return market
return build_market(seed, num_items, visible, held_out)
@dataclass
class Outcome:
"""One run, beside the reference run it is normalised against."""
fills: Fills
reference: Fills
stated: float | None
dropped: int = 0
"""Orders that named nothing on the board, or asked for a quantity or price of zero."""
@property
def target(self) -> float:
"""The gp a full score is worth in this basket. See TARGET_SHARE."""
return TARGET_SHARE * self.reference.realised
@property
def profit_ratio(self) -> float:
if self.target <= 0.0:
return 0.0
return min(1.0, max(0.0, self.fills.realised / self.target))
@property
def efficiency(self) -> float:
"""Return on committed capital, against the reference's return on its own.
Not discounted by TARGET_SHARE: it is already a rate rather than a total, so a
strategy that trades well on a small book is not penalised for being small, and
giving it a band as well would hand out the term for free."""
if self.reference.roc <= 0.0:
return 0.0
return min(1.0, max(0.0, self.fills.roc / self.reference.roc))
@property
def conversion_ratio(self) -> float:
"""Of the coins locked, the share that bought stock, against the reference's share.
Recorded, never rewarded — `efficiency` is what qualifies profit now. It stays in
the trace because it separates the two ways `efficiency` falls: offers that never
filled, and offers that filled at prices not worth reaching."""
if self.reference.conversion <= 0.0:
return 0.0
return min(1.0, self.fills.conversion / self.reference.conversion)
@property
def discipline(self) -> float:
"""Profit, qualified by what each committed coin earned. A product, not a term
beside it: a term would pay in full for one tiny order that traded perfectly, which
is inaction with a receipt — the free-points defect schema-migration shipped with."""
return self.profit_ratio * self.efficiency
@property
def clean(self) -> bool:
"""A bar, not a knife-edge.
It used to require matching the reference on both ratios to within 1e-9, which is a
demand for replication rather than for a good run: a reference plan with one extra
one-unit order that never fills has identical realised profit and lost 0.146 of
total reward. The bar is now TARGET_SHARE of the reference on the money AND on the
return that money made, which is a run that was worth doing however it got there.
"""
return (
self.fills.realised >= TARGET_SHARE * self.reference.realised
and self.fills.roc >= TARGET_SHARE * self.reference.roc
)
@property
def arith_ok(self) -> bool:
"""Did the agent's stated expected profit match what its own ACCEPTED orders imply.
Not a reward — a probe into the trace. The target model gets simple arithmetic wrong
with thinking off, and without this line a bad multiplication and a bad strategy are
the same low number and nothing in the trace tells them apart. Measured against the
orders that survived parsing, so claiming profit from an order for an item that is
not on the board reads as the arithmetic error it is.
"""
if self.stated is None:
return False
return abs(self.stated - self.fills.paper) <= max(50.0, 0.02 * abs(self.fills.paper))
@property
def arith_error(self) -> float:
if self.stated is None or self.fills.paper == 0.0:
return 0.0
return abs(self.stated - self.fills.paper) / abs(self.fills.paper)
def measure(market: Market, orders: list[Order], stated: float | None) -> Outcome:
"""Run the agent's plan and the reference plan through the same engine."""
clean, seen, dropped = [], set(), 0
for order in orders[:MAX_ORDERS]:
item = market.item(order.item)
# One order per item: holdings pool per item, so two orders on one name would make
# "which sell limit does this unit belong to" a question the engine has to invent an
# answer to. The first one submitted is the one that counts.
if item is None or order.quantity <= 0 or order.buy <= 0 or order.sell <= 0:
dropped += 1
continue
if item.name in seen:
dropped += 1
continue
seen.add(item.name)
clean.append(Order(item.name, order.quantity, order.buy, order.sell))
return Outcome(
fills=execute(market, clean),
reference=execute(market, reference_orders(market)),
stated=stated,
dropped=dropped,
)
_BLOCK = re.compile(r"```(?:json)?\s*\n(.*?)```", re.DOTALL)
def parse_orders(reply: str) -> tuple[list[Order], float | None]:
"""The last JSON block in the reply: an object with "orders" and, optionally,
"expected_profit". A bare array of orders is accepted too — a model that answers the
question and skips the arithmetic has traded, and should be graded on the trade.
A reply that parses to nothing is an empty plan, not an error: it scores what doing
nothing scores. Raising here would turn a formatting slip into a crashed rollout — and
the reward runs inside the metric, so a raise takes the whole rollout with it rather
than scoring zero.
Which is why the except clauses below are wider than they look like they need to be.
Python's `json.loads` is not strict JSON: it accepts the bare literals Infinity,
-Infinity and NaN, and it overflows 1e309 to inf rather than refusing it. RecursionError
is not a ValueError, so twenty thousand nested arrays crashed the decoder; OverflowError
is not a ValueError either, so `int(float("inf"))` crashed the row loop. Both were live
on the first cut and both were reachable from a reply a model can actually emit.
"""
blocks = _BLOCK.findall(reply or "")
raw = blocks[-1] if blocks else (reply or "")
try:
parsed = json.loads(raw.strip())
except (ValueError, RecursionError):
return [], None
stated: float | None = None
if isinstance(parsed, dict):
rows = parsed.get("orders")
value = parsed.get("expected_profit")
# Finite only. `Infinity` and `NaN` parse, and either one propagates through
# `arith_error` into the trace as a non-finite metric, which is a corrupted training
# signal rather than a bad answer. An unusable claim is no claim.
if isinstance(value, (int, float)) and not isinstance(value, bool):
try:
stated = float(value) if math.isfinite(value) else None
except OverflowError:
stated = None
else:
rows = parsed
if not isinstance(rows, list):
return [], stated
orders = []
for row in rows:
if not isinstance(row, dict) or not isinstance(row.get("item"), str):
continue
try:
orders.append(
Order(
item=row["item"],
quantity=int(row.get("quantity", 0)),
buy=int(row.get("buy", 0)),
sell=int(row.get("sell", 0)),
)
)
except (TypeError, ValueError, OverflowError):
continue
return orders, stated
@@ -0,0 +1,667 @@
"""The stepped exchange: quote, watch the tape, re-quote, over a turn budget.
The one-shot environment asks for one basket of limit orders and runs it against a window
the agent never sees. This one hands the agent the SAME market and lets it act inside the
window: eight looks spread over sixty ticks, with inventory, open offers, cumulative buy
limits and one purse carrying between them. What is held out is no longer the whole window
— it is the tail after the last look, plus every tick the agent chose not to spend a look
on. That is a weaker held-out claim than the one-shot form's and it is stated as one: the
agent sees 49 of the 60 graded ticks if it takes every look.
Three mechanics carry the whole design, and each of them exists to price a turn.
the freeze Any look that touches a book which is already live freezes the WHOLE book
for `market.freeze` ticks — one to four, drawn per basket. Nothing fills
while frozen. This is the counterweight in the turn dimension that house
rule 2 demands under multi-turn: without it, re-quoting on every look is
free information and the optimal policy is "use every turn", which is
rule 4's definition of measuring nothing. The FIRST placement is free —
there is nothing to re-quote when the book is empty, and charging for it
would only tax entering the market at all.
the queue An amended offer goes to the back of the queue: its accrued fill room is
reset to zero. In the one-shot engine, room accrues on the ticks an offer
was eligible on and carries between them, which is what makes a thin item
tradeable at all. Re-quoting throws that accrual away, so moving a limit
by one gp is not free even when the freeze has expired.
the cumulative
buy limit The GE's per-item limit is over the WHOLE window, not per offer. Cancel a
half-filled offer and re-place it and the units already bought still
count. Without this, re-quoting resets the limit and the limit stops being
a limit.
`roc` is over PEAK capital employed, never average. Average-employed is a measured
free-points bug in this shape: a policy that commits the purse for two ticks, sells, and
sits in cash for fifty-eight has a tiny average and its realised profit divided by that
average is enormous, so the counterweight pays MORE for doing less. The peak asks the
question the counterweight is for — how much of the purse did this plan need at its
fattest — and it cannot be talked down by exiting early.
Everything is scored against a REFERENCE POLICY (`Reference`) that plays the same stepped
engine through the same interface the agent does, so 1.000 is reachable by construction and
by something the prompt describes rather than by hindsight. The reference re-quotes when
the anchor it can compute from the ticks it has seen has moved materially, and not
otherwise — which is the judgement the freeze exists to make expensive.
"""
from __future__ import annotations
import statistics
from dataclasses import dataclass, field
from typing import Callable
from grand_exchange_live.book import (
BUY_BAND,
MAX_ITEM_SHARE,
MIN_CROSSINGS,
SELL_BAND,
Order,
crossings,
reference_orders,
)
from grand_exchange_live.market import (
DUMP_BASE,
DUMP_CAP,
DUMP_IMPACT,
FILL_SHARE,
TAX,
Item,
Market,
build_market,
)
TURNS = 8
"""Looks the agent gets. The last one is still worth taking — `TAIL_TICKS` run after it —
but it is worth taking only if there is something to change."""
TICKS_PER_TURN = 7
TAIL_TICKS = 4
HORIZON = TURNS * TICKS_PER_TURN + TAIL_TICKS
"""Graded ticks: 8 x 7 + 4 = 60. The tail is what makes the last look a decision rather
than a formality, and it is the floor on the held-out slice — an agent that takes every
look still has TICKS_PER_TURN + TAIL_TICKS = 11 ticks executed after its final observation.
"""
REQUOTE_BAND = 0.025
IDLE_SHARE = 0.50
"""The two triggers that make the reference re-quote.
They are the reward's denominator, so where they sit is not a taste question: they are the
profit-maximising point of the reference's own family, swept over 120 baskets by
`measure_ladder.py --sweep-reference`. REQUOTE_BAND is how far the rolling anchor has to
have moved before chasing it is worth a freeze; IDLE_SHARE is how much of the purse has to
be sitting in cash before putting it back to work is.
The second trigger was missing from the first cut and its absence was measurable in exactly
the way `book.py` warns about. A reference that never redeployed sale proceeds was beaten by
22% on gp and 12% on return-on-capital by a policy that simply re-quoted on every look, so
the reward was an imitation score for an under-tuned strategy rather than a profit metric,
and the rule-4 gap it reported (0.152) was an artifact of the denominator rather than a
measurement of anything. Redeploying is the capability the stepped form EXISTS to reward —
a one-shot plan cannot spend what its own sales earned — and leaving it out meant the
environment measured the freeze instead of the trading. With it, the reference earns 76,744
gp against the every-look policy's 63,802 and dominates it on return as well, and the gap
becomes a real one.
Two measured facts kept here rather than in a commit message:
the anchor trigger is worth about nothing on its own. The sweep's argmax turns it off
entirely (band 1.0, idle 0.50: 77,599 gp) and 0.025 costs 1.1% of gp against it. It is
kept at the top of its own plateau because a reference whose only trigger is a cash
threshold cannot express "the anchor moved" at all, and 1.1% is inside the three-percent
tolerance `probe.py` already holds the one-shot bands to. The cell that earns more scores
0.976, well above the 0.90 that assertion demands.
the idle trigger is flat between 0.0 and 0.5 and falls off a cliff above it — 63,472 gp
at 1.01, where it effectively never fires. There is no knife-edge here to tune."""
# --- what the agent submits ------------------------------------------------------------
@dataclass(frozen=True)
class Leg:
"""One line of a delta sheet: the offer this item should carry from now on.
A leg REPLACES whatever is live on that item — there is one offer per name, exactly as
in the one-shot form, so "which sell limit does this unit belong to" is never a question
the engine has to invent an answer to. `quantity <= 0` cancels.
"""
item: str
quantity: int
buy: int = 0
sell: int = 0
# --- what the agent sees ---------------------------------------------------------------
@dataclass(frozen=True)
class OpenOrder:
item: str
quantity: int
"""Units still sought on the buy side of this offer."""
buy: int
sell: int
locked: float
@dataclass(frozen=True)
class TurnView:
"""Everything true at the moment of a look. Plain types only — this is what a protocol
layer renders and what `trace.info` would carry, and a dataclass or a numpy scalar in
there fails the trace write for the whole rollout."""
turn: int
turns_left: int
ticks_left: int
freeze: int
frozen_for: int
cash: float
market: Market
new_prices: dict[str, list[int]]
"""Ticks that elapsed since the previous look, per item. Empty on turn 1."""
new_volumes: dict[str, list[int]]
seen: int
"""How many ticks of each stream have been observed, warmup included."""
held: dict[str, int]
bought: dict[str, int]
"""Cumulative units bought per item, against the item's window limit."""
open_orders: dict[str, OpenOrder]
realised_so_far: float
# --- the engine's own state ------------------------------------------------------------
@dataclass
class _Live:
"""An offer standing in the book. Not visible to the agent in this shape — `OpenOrder`
is."""
item: Item
buy: int
sell: int
remaining: int
locked: float
buy_room: float = 0.0
sell_room: float = 0.0
@dataclass
class Ledger:
"""The purse, the stock and the cost basis, carried across every turn.
One invariant holds this together and the tests assert it after every tick: cash, plus
coins locked behind offers, plus the cost basis of stock held, plus realised profit,
equals the starting capital. Every mechanic here moves coins between those buckets and
creates none, so a bug that mints money shows up as a broken identity rather than as a
good score.
"""
capital: float
cash: float
locked: float = 0.0
held: dict[str, int] = field(default_factory=dict)
basis: dict[str, float] = field(default_factory=dict)
bought: dict[str, int] = field(default_factory=dict)
incoming: dict[str, int] = field(default_factory=dict)
"""Units bought on the current tick, not yet sellable. Settled at the top of the next
tick — stock bought on a tick cannot be sold on that tick, or a buy limit above a sell
limit is a free round trip on one price."""
incoming_basis: dict[str, float] = field(default_factory=dict)
spent: float = 0.0
proceeds: float = 0.0
sold: int = 0
dumped: int = 0
peak_employed: float = 0.0
@property
def employed(self) -> float:
"""Coins not sitting idle: locked behind an offer, or sunk into stock at what it
cost. The PEAK of this is the denominator of `roc`. Average-employed is a measured
free-points bug — see the module docstring."""
return self.locked + sum(self.basis.values()) + sum(self.incoming_basis.values())
@property
def realised(self) -> float:
return self.cash + self.employed - self.capital
def mark(self) -> None:
self.peak_employed = max(self.peak_employed, self.employed)
def settle(self) -> None:
for name, qty in self.incoming.items():
self.held[name] = self.held.get(name, 0) + qty
self.basis[name] = self.basis.get(name, 0.0) + self.incoming_basis.get(name, 0.0)
self.incoming.clear()
self.incoming_basis.clear()
@dataclass
class Result:
"""One episode, in the shape every reward reads."""
realised: float
peak_employed: float
looks: int
"""Looks that submitted a non-empty sheet. Recorded, never rewarded — a turn-count
reward is a second single-sided term and the freeze already prices the turn."""
amendments: int
"""Looks that touched a live book, i.e. the ones that actually paid a freeze."""
frozen_ticks: int
bought: int
bought_by_item: dict[str, int]
"""Cumulative units per name, for the cumulative-buy-limit invariant and for the trace.
A plain dict of plain ints — `trace.info` fails the whole rollout on anything else."""
sold: int
dumped: int
spent: float
dropped: int
"""Legs that named nothing on the board or asked for a price of zero."""
unfunded: int
"""Legs the purse could not fund at all when they were submitted."""
seen_ticks: int
"""Graded ticks the agent observed. HORIZON minus this is the held-out slice."""
@property
def roc(self) -> float:
if self.peak_employed <= 0.0:
return 0.0
return self.realised / self.peak_employed
Policy = Callable[[TurnView], list[Leg]]
def play(market: Market, policy: Policy, *, turns: int = TURNS,
step: int = TICKS_PER_TURN, tail: int = TAIL_TICKS) -> Result:
"""Run one policy through one basket, look by look.
The tick loop is the one-shot engine's, moved inside a turn loop and given three things
it did not have: a freeze that suppresses fills, a queue reset on amendment, and a purse
that sale proceeds actually return to. With a single look on turn 1 and nothing after
it, the two engines produce byte-identical realised profit and identical committed
capital — `tests/test_live.py::test_single_look_matches_one_shot` is that assertion, and
it is what says the stepped form is the same market rather than a different one.
"""
ledger = Ledger(capital=float(market.capital), cash=float(market.capital))
live: dict[str, _Live] = {}
dropped = unfunded = looks = amendments = frozen_ticks = 0
frozen_until = 0
t = market.visible
end = market.visible + market.held_out
last_seen = market.visible
ledger.mark()
for turn in range(1, turns + 1):
view = TurnView(
turn=turn,
turns_left=turns - turn,
ticks_left=end - t,
freeze=market.freeze,
frozen_for=max(0, frozen_until - t),
cash=ledger.cash,
market=market,
new_prices={i.name: i.prices[last_seen:t] for i in market.items},
new_volumes={i.name: i.volumes[last_seen:t] for i in market.items},
seen=t,
held=dict(ledger.held),
bought=dict(ledger.bought),
open_orders={
name: OpenOrder(name, o.remaining, o.buy, o.sell, o.locked)
for name, o in live.items()
},
realised_so_far=ledger.cash + ledger.employed - ledger.capital,
)
last_seen = t
sheet = policy(view) or []
touched, drop, unfund = _apply(market, ledger, live, sheet)
dropped += drop
unfunded += unfund
if touched:
looks += 1
# The first placement is free: there is nothing to re-quote when the book is
# empty, and charging for entering the market taxes acting at all rather than
# taxing acting AGAIN, which is the thing a turn budget is supposed to price.
if view.open_orders:
amendments += 1
frozen_until = max(frozen_until, t + market.freeze)
ledger.mark()
run = step if turn < turns else step + tail
for _ in range(min(run, end - t)):
frozen_ticks += _tick(market, ledger, live, t, frozen=t < frozen_until)
t += 1
while t < end:
frozen_ticks += _tick(market, ledger, live, t, frozen=t < frozen_until)
t += 1
_close(market, ledger, live)
return Result(
realised=ledger.cash - ledger.capital,
peak_employed=ledger.peak_employed,
looks=looks,
amendments=amendments,
frozen_ticks=frozen_ticks,
bought=sum(ledger.bought.values()),
bought_by_item=dict(ledger.bought),
sold=ledger.sold,
dumped=ledger.dumped,
spent=ledger.spent,
dropped=dropped,
unfunded=unfunded,
seen_ticks=last_seen - market.visible,
)
def _apply(market: Market, ledger: Ledger, live: dict[str, _Live],
sheet: list[Leg]) -> tuple[bool, int, int]:
"""Put a delta sheet into the book. Returns (touched anything, dropped, unfunded)."""
touched = dropped = unfunded = 0
seen: set[str] = set()
for leg in sheet:
item = market.item(leg.item) if isinstance(leg.item, str) else None
if item is None or item.name in seen:
dropped += 1
continue
seen.add(item.name)
existing = live.get(item.name)
if leg.quantity <= 0:
# A cancel. The coins behind the unfilled remainder come back; the stock does
# not, and with no offer on the name it can now only leave through the close-out
# haircut. Cancelling is not free — it is just not charged as a fine.
if existing is None:
dropped += 1
continue
ledger.locked -= existing.locked
ledger.cash += existing.locked
del live[item.name]
touched += 1
continue
if leg.buy <= 0 or leg.sell <= 0:
dropped += 1
continue
# Release the old offer first, so its coins are available to fund the new one. This
# is what makes an amendment a re-quote rather than a second offer.
if existing is not None:
ledger.locked -= existing.locked
ledger.cash += existing.locked
del live[item.name]
room = item.buy_limit - ledger.bought.get(item.name, 0)
qty = max(0, min(leg.quantity, room, int(ledger.cash // leg.buy)))
if qty <= 0 and not ledger.held.get(item.name, 0):
# The purse cannot fund it, or the window's buy limit is already spent, and
# there is no stock for the sell side to work on either. Recorded rather than
# silently ignored: a plan whose offers do not fit is a plan that allocated
# badly, and the trace should say which.
unfunded += 1
if existing is not None:
touched += 1
continue
# qty may be zero while stock is held — that is a SELL-ONLY offer, and it has to be
# expressible. An agent that has bought its whole window limit still needs a way to
# move the price it is asking, and "cancel" is not that way: cancelling strands the
# stock in the close-out haircut.
ledger.cash -= qty * leg.buy
ledger.locked += qty * leg.buy
# Queue room starts at zero. An amended offer is a new offer to the book.
live[item.name] = _Live(item=item, buy=leg.buy, sell=max(leg.buy + 1, leg.sell),
remaining=qty, locked=float(qty * leg.buy))
touched += 1
return bool(touched), dropped, unfunded
def _tick(market: Market, ledger: Ledger, live: dict[str, _Live], t: int,
*, frozen: bool) -> int:
"""One tick of the exchange. Returns 1 if the tick was frozen out.
Identical to `book.execute`'s inner loop, with two additions: a frozen tick fills
nothing at all, and sale proceeds land in the purse where the next look can spend them.
Settlement is NOT frozen — the freeze is on the book, not on stock already paid for, and
stalling settlement too would be a second, unstated penalty on the same action.
"""
ledger.settle()
if frozen:
ledger.mark()
return 1
for name, order in live.items():
item = order.item
price = item.prices[t]
flow = FILL_SHARE * item.volumes[t]
# Room accrues only on the ticks the offer was actually eligible on — banking the
# flow of ticks the price never reached would let an offer fill far past the volume
# that was ever available to it, and the volume limit is the whole reason allocation
# is a decision here.
if price <= order.buy:
order.buy_room += flow
if price >= order.sell:
order.sell_room += flow
if price <= order.buy and order.remaining > 0:
room = item.buy_limit - ledger.bought.get(name, 0)
qty = min(order.remaining, int(order.buy_room), room)
if qty > 0:
order.buy_room -= qty
order.remaining -= qty
order.locked -= qty * order.buy
ledger.locked -= qty * order.buy
# Reserved at the limit, paid at the market: the difference returns to the
# purse, exactly as the exchange does it.
ledger.cash += qty * (order.buy - price)
ledger.spent += qty * price
ledger.bought[name] = ledger.bought.get(name, 0) + qty
ledger.incoming[name] = ledger.incoming.get(name, 0) + qty
ledger.incoming_basis[name] = (
ledger.incoming_basis.get(name, 0.0) + qty * price
)
avail = ledger.held.get(name, 0)
if price >= order.sell and avail > 0:
qty = min(avail, int(order.sell_room))
if qty > 0:
order.sell_room -= qty
ledger.basis[name] -= ledger.basis[name] * (qty / avail)
ledger.held[name] = avail - qty
ledger.cash += qty * price * (1.0 - TAX)
ledger.proceeds += qty * price * (1.0 - TAX)
ledger.sold += qty
ledger.mark()
return 0
def _close(market: Market, ledger: Ledger, live: dict[str, _Live]) -> None:
"""The window shuts. Offers are cancelled, stock is forced out."""
ledger.settle()
for order in live.values():
ledger.locked -= order.locked
ledger.cash += order.locked
live.clear()
for item in market.items:
left = ledger.held.get(item.name, 0)
if not left:
continue
depth = statistics.median(item.volumes[market.visible:])
haircut = min(DUMP_CAP, DUMP_BASE + DUMP_IMPACT * (left / max(1.0, depth)))
ledger.cash += left * item.prices[-1] * (1.0 - haircut) * (1.0 - TAX)
ledger.dumped += left
ledger.held[item.name] = 0
ledger.basis[item.name] = 0.0
# --- the strategy the reward is normalised against --------------------------------------
def desired_orders(items: list[Item], prices: dict[str, list[int]],
volumes: dict[str, list[int]], purse: float, ticks_left: int,
limits: dict[str, int]) -> list[Order]:
"""`book.reference_orders`, generalised to a moment inside the window.
Same three judgements — does it revert, what does it pay, what can it hold — over
whatever ticks have been observed by now, against whatever purse and whatever remaining
per-item window limits are left. On turn 1, with the whole warmup observed, the whole
purse free and no units bought, this returns exactly what `reference_orders` returns:
`tests/test_live.py::test_turn_one_matches_reference_orders` asserts that, and it is
what keeps the live reference recognisably the same strategy as the one-shot one rather
than a second, unswept family of constants.
"""
plans = []
for item in items:
window, flow = prices[item.name], volumes[item.name]
room = limits.get(item.name, item.buy_limit)
if room <= 0 or crossings(window) < MIN_CROSSINGS:
continue
anchor = statistics.fmean(window)
buy = max(1, round(anchor * (1.0 - BUY_BAND)))
sell = max(buy + 1, round(anchor * (1.0 + SELL_BAND)))
lows = [p for p in window if p <= buy]
highs = [p for p in window if p >= sell]
if not lows or not highs:
continue
entry, exit_ = statistics.fmean(lows), statistics.fmean(highs)
expected = (exit_ * (1.0 - TAX) - entry) / entry
if expected <= 0:
continue
reachable = int(
FILL_SHARE * statistics.median(flow) * (len(lows) / len(window)) * ticks_left
)
qty = min(room, reachable, int(MAX_ITEM_SHARE * purse // buy))
if qty <= 0:
continue
plans.append((expected, item.name, qty, buy, sell))
plans.sort(key=lambda p: -p[0])
orders, left = [], float(purse)
for _, name, qty, buy, sell in plans:
qty = min(qty, int(left // buy))
if qty <= 0:
continue
left -= qty * buy
orders.append(Order(item=name, quantity=qty, buy=buy, sell=sell))
return orders
class Reference:
"""The reference POLICY: quote from the warmup, then chase the anchor when it moves.
It is a policy and not a plan because the reward has to be reachable by playing the
game rather than by pre-computing it, and because a denominator that never re-quotes
would make the whole stepped form decorative — every ratio would be maximised by the
one-shot answer and the turn budget would be measuring nothing.
What it does on a look:
1. Recompute each live item's anchor over every tick observed so far.
2. Mark an item for re-quote only if that anchor has moved more than REQUOTE_BAND.
This is the judgement the freeze prices: chasing a two-gp drift costs `freeze`
ticks of fills across the WHOLE book and throws away the offer's queue position,
and it is worth doing only when the old limits are on the wrong side of where the
price now lives.
3. Re-run the same three judgements over the marked items and the items carrying no
offer at all, against the cash that would be free once the marked offers are
released and the buy-limit room actually left.
Stateful, so one instance plays one episode. `play(market, Reference())`.
"""
def __init__(self, requote_band: float = REQUOTE_BAND,
idle_share: float = IDLE_SHARE) -> None:
self.requote_band = requote_band
self.idle_share = idle_share
self.prices: dict[str, list[int]] = {}
self.volumes: dict[str, list[int]] = {}
def _observe(self, view: TurnView) -> None:
if not self.prices:
for item in view.market.items:
self.prices[item.name] = list(item.visible_prices(view.market.visible))
self.volumes[item.name] = list(item.visible_volumes(view.market.visible))
for name, ticks in view.new_prices.items():
self.prices[name].extend(ticks)
self.volumes[name].extend(view.new_volumes[name])
def __call__(self, view: TurnView) -> list[Leg]:
self._observe(view)
items = view.market.items
limits = {i.name: i.buy_limit - view.bought.get(i.name, 0) for i in items}
idle = view.cash >= self.idle_share * view.market.capital
free = [i for i in items if i.name not in view.open_orders]
if not view.open_orders:
purse = view.cash
candidates = items
else:
marked = set()
for name, open_order in view.open_orders.items():
anchor = statistics.fmean(self.prices[name])
buy = max(1, round(anchor * (1.0 - BUY_BAND)))
if abs(buy - open_order.buy) / max(1, open_order.buy) > self.requote_band:
marked.add(name)
elif idle and open_order.quantity == 0 and limits[name] > 0:
# The buy side of this offer is done and the coins it earned are sitting
# in cash. Re-quoting it is not chasing a drift, it is putting the purse
# back to work, and the window limit says there is room to do so.
marked.add(name)
if not marked and not (idle and free):
# Nothing has moved far enough to be worth a freeze and there is nothing
# idle to deploy. Taking the look and doing nothing is the correct play, and
# it is the play the exhaustive policy in the ladder refuses to make.
return []
purse = view.cash + sum(
o.locked for n, o in view.open_orders.items() if n in marked
)
candidates = [i for i in items if i.name in marked or i.name not in view.open_orders]
orders = desired_orders(
candidates,
{i.name: self.prices[i.name] for i in candidates},
{i.name: self.volumes[i.name] for i in candidates},
purse,
max(1, view.ticks_left),
limits,
)
return [Leg(o.item, o.quantity, o.buy, o.sell) for o in orders]
def live_reference(market: Market) -> Result:
"""The denominator, cached per basket.
`viable_live_market` calls this once per candidate seed and `scan()` calls it again for
every episode on the seed that was chosen, so an uncached reference is `VIABILITY_TRIES`
stepped simulations per task before the first token is emitted — 3,072 of them for a
default 48-task load, on every eval AND every `--dry-run`. The cache is keyed on the
seed and the window shape, which is everything the result depends on.
"""
key = (market.seed, market.visible, market.held_out, market.freeze)
hit = _REFERENCE_CACHE.get(key)
if hit is None:
hit = _REFERENCE_CACHE[key] = play(market, Reference())
return hit
_REFERENCE_CACHE: dict[tuple[int, int, int, int], Result] = {}
LIVE_VIABILITY_TRIES = 16
"""Candidate seeds `viable_live_market` will try before giving up, against the one-shot
form's 64. A stepped reference is eight turns and sixty ticks rather than one pass, so the
worst case is the load cost, not the failure rate: at 48 tasks a 64-try cap is 3,072
simulations of a market nobody has looked at yet. Measured skip rate is low enough that 16
never binds in practice — `measure_ladder.py` reports how often it does."""
def viable_live_market(seed: int, num_items: int, visible: int, held_out: int) -> Market:
"""The next basket from `seed` onward in which the LIVE reference makes money.
Retargeted from `book.viable_market`, which asks whether the ONE-SHOT reference profits.
That is the wrong question here: every ratio divides by the live reference, so a basket
where the live reference loses has no reachable ceiling and quietly breaks house rule 3
for that task. The freeze is drawn inside `build_market` off its own RNG precisely so
this function can see it — it used to be drawn after viability had already answered,
which meant viability accepted a basket under one cost of re-quoting and the run played
it under another.
"""
for offset in range(LIVE_VIABILITY_TRIES):
market = build_market(seed + offset, num_items, visible, held_out)
if live_reference(market).realised > 0:
return market
return build_market(seed, num_items, visible, held_out)
@@ -0,0 +1,277 @@
"""The market generator: price and volume streams with structure a reader can find.
⚠️ This file is a COPY of `grand_exchange/market.py`, not an import of it. The one-shot
environment's scores have to stay comparable across time, and an import would mean a change
made for the live form silently moved every historical `grand-exchange` number. The copy is
held to the original byte-for-byte by `tests/test_market_copy.py`, which builds the same
seeds through both packages and compares every price and volume. The ONLY intentional
divergence is the freeze draw below, and it is drawn off a separate RNG so it cannot
perturb a single tick of either stream.
If prices were a random walk this environment would be worthless and would still look
fine — expected profit is zero for every strategy, so the reward is noise, no oracle
exists, and the numbers coming out of it would be plausible and meaningless. So the
structure is put in deliberately and is the whole design:
price_t = fundamental_t * (1 + x_t)
`fundamental_t` drifts by FUND_DRIFT a tick, worth about one percent over a whole stream.
`x_t` is an AR(1) around zero with PHI decay and a stationary spread of NOISE_SD — nine
percent on the deep items, fifteen or sixteen on the thin ones. Noise dominates drift by an
order of magnitude, which is what makes the mean of the visible prices a usable estimate of
the fundamental, and "buy under the estimate, sell over it" a real strategy rather than a
superstition. `book.reference_orders` is that strategy; `probe.py` measures how far it beats
trading at random, and refuses to pass if the margin is thin.
The visible and held-out windows are two halves of ONE stream from ONE seed: the graded
ticks are the next ticks the generator would have produced, not a differently-seeded
population that could have moved somewhere the visible half gave no warning of.
Two things are here to punish reading the price column alone. One or two items per basket do
not revert at all (see DRIFT_STEP and WALKS) and swing widest of everything on screen. And
liquidity — price times volume, gp a tick — is uncorrelated with price, so the fattest
visible margins sit on the items that can absorb the least of the purse.
Item names are invented. Formulas and market mechanics are facts about a kind of game;
item tables are somebody's copyrighted content, and none of it is here.
"""
from __future__ import annotations
import math
import random
from dataclasses import dataclass
# --- the rules of the exchange, quoted to the agent verbatim in the prompt --------------
TAX = 0.01
"""Charged on every sale, the GE's own. It is the counterweight that kills thin flips:
a round trip has to clear it twice over before it is worth doing, so "trade everything"
is a losing strategy rather than a neutral one."""
FILL_SHARE = 0.25
"""Share of a tick's volume one participant can take, and it accrues only on the ticks an
offer was actually eligible on. Without it a fat margin on an item carrying three thousand
gp a tick is worth as much as one carrying a hundred thousand, and the environment stops
measuring allocation."""
DUMP_BASE = 0.03
DUMP_IMPACT = 0.05
DUMP_CAP = 0.40
"""Stock still held when the window closes is forced out at the last price, minus a haircut
of DUMP_BASE plus DUMP_IMPACT for every full tick's worth of the item's median volume that
has to be pushed through, capped at DUMP_CAP.
A FLAT haircut was the first cut of this and it was the wrong shape. Under a flat five
percent, the cost of holding stock at the close was the same whether the leftovers were
twenty units of an item that trades thirty a tick or four thousand units of one that trades
six hundred — so oversizing was only ever punished through the purse, and the purse punishes
it as a cliff: the first over-large offer eats the whole budget and everything after it is
never placed. A cliff is not a gradient. Scaling the haircut by position-over-depth prices
the thing that is actually true — forcing size out costs you in proportion to how much of
the book you are pushing through — and it makes reading the volume column pay smoothly,
which is what the column is here to teach.
It also prices a bad anchor. A buy limit set too high fills fast and leaves the sell limit
out of reach, so the position that a sloppy anchor builds is exactly the position that has
to be dumped, and now it is dumped at a price that scales with its size."""
STARTING_CAPITAL = 250_000
WALKS = (1, 2)
"""How many of the basket's items do not mean-revert, drawn uniformly from this range.
It used to be exactly one, and a fixed count is a free prior: "drop the single widest line
on the board" scores what computing `book.crossings` scores, without computing anything.
It is the same defect `bot_detection` ships an assertion against — a class balance the model
can count on is a class balance it will use instead of the discriminator. With the count
unknown the shape statistic is the only thing that answers the question, and a basket can
punish both over- and under-rejection."""
SEED_BASE = 60_000
"""Where task seeds start. Lives here rather than on the taskset so `probe.py`, which
cannot import the taskset without `verifiers`, grades the baskets a run would actually
serve rather than a different set that happens to share a generator."""
FREEZE_RANGE = (1, 4)
FREEZE_SALT = 0x5F3E_11CE
"""How long the book is frozen after a re-quote, in ticks, drawn per basket.
Drawn HERE, off `random.Random(seed ^ FREEZE_SALT)`, and carried on the Market. It used to
be drawn in the taskset after the basket was chosen, which is circular: `viable_market`
decides whether the reference makes money in a basket, the reference's profit depends on
what re-quoting costs, and the cost was not drawn yet. Viability would then have accepted a
basket under one freeze and the run would have played it under another, so the denominator
of every reward would be a strategy the model never played against. A separate RNG keyed off
the same seed makes the draw deterministic, available before the first simulation, and
provably free of any effect on the price and volume streams — which is what lets the copy
assertion in `tests/` be an equality rather than an approximation.
The value is a per-look cost paid by the WHOLE book: see `live.freeze`. Flat, not
proportional to the number of items re-quoted — the settled decision. Per-item is the
better mechanism and is the identified strengthening, not a thing to build unmeasured on
the critical path."""
# --- stream parameters -----------------------------------------------------------------
PHI = 0.45
"""AR(1) decay of the mispricing. Half-life under a tick, so a visible window holds many
independent draws around the fundamental — which is what makes the mean of it an estimate
rather than a guess — and a held-out window holds many excursions, so the reward is not one
lucky draw."""
FUND_DRIFT = 0.0008
"""Per-tick drift of the fundamental. Small on purpose: the fundamental has to be
ESTIMABLE from the visible half or there is nothing to learn."""
VOLUME_SD = 0.35
DRIFT_STEP = 0.055
"""Per-tick step of the items that do not mean-revert at all — their price is a pure random
walk, so the fundamental IS wherever it last was.
This is the trap the whole environment is built around, put inside the task instead of
left as a hazard the designer has to avoid. A random walk has no anchor, so buying under
its moving average is not a discount, it is a coin flip that pays the tax and the dump
slippage every time. On screen it is the widest-swinging line in the basket and therefore
the most attractive one, because amplitude is what a careless reader ranks by. The two can
only be told apart by SHAPE: a reverting series crosses its own mean constantly, a walk
wanders on one side of it for a dozen ticks at a time. `book.crossings` is that statistic
and the reference strategy will not trade an item that fails it."""
# (names, base price range, base volume range, buy limit, mispricing spread)
#
# What separates these is LIQUIDITY IN GP PER TICK — base price times base volume — and it
# is deliberately uncorrelated with the price. That is the allocation problem: the purse is
# 250,000 gp and a quarter of the flow over thirty ticks is what any one offer can absorb,
# so a deep tier can take a third of the purse and a thin one can take a twentieth of it no
# matter how good the margin looks. An earlier cut of this file made the expensive items the
# thin ones, which sounds right and is not: eight units a tick of a 46,000 gp item is 368,000
# gp of flow, the deepest thing on the board. Thin means small in coins, not small in units.
#
# The two thin tiers also carry the widest mispricing spread, so they show the fattest margin
# and can absorb the least. That is the trap, and it is the same trap either way an agent
# falls into it: ignore the volume column and either the offers sit unfilled or the purse
# sits idle.
TIERS = [
# deep, ~40k-130k gp a tick: this is where the purse actually goes
(["Thornroot poultice", "Chipped bone charm", "Bogwater draught", "Coarse fletching feather"],
(90, 170), (400, 800), 5000, 0.09),
(["Emberglass shard", "Stormrune tablet", "Marrowsteel nail", "Pale grimoire page"],
(900, 1700), (40, 90), 400, 0.09),
# thin, ~3k-15k gp a tick, and the widest swings on the board
(["Gilded harpoon head", "Cinderweave cloak", "Wyrmbone talisman", "Frostbitten ledger"],
(200, 420), (12, 30), 800, 0.15),
(["Duskforged sigil", "Heart of the sunken cairn", "Voidglass lens", "Tideworn crown"],
(6000, 14000), (1.2, 3.0), 40, 0.16),
]
@dataclass(frozen=True)
class Item:
name: str
reverting: bool
"""Whether this item has an anchor at all. Never shown to the agent — it is here so the
probe can assert that the trap is a trap, and that the reference avoids it for a reason
rather than by luck."""
buy_limit: int
"""Units per item per window, the GE's own limit. With finite capital it is what turns
the task into an allocation problem instead of a single pick."""
prices: list[int]
volumes: list[int]
def visible_prices(self, visible: int) -> list[int]:
return self.prices[:visible]
def visible_volumes(self, visible: int) -> list[int]:
return self.volumes[:visible]
@dataclass(frozen=True)
class Market:
seed: int
"""Carried on the basket so the taskset can store the seed it actually used. Baskets are
skipped when the reference strategy is not profitable in them (see `book.viable_market`),
so the seed a task was built from is not always the one it was asked for."""
items: list[Item]
visible: int
held_out: int
capital: int
freeze: int = 0
"""Ticks the book is frozen for after any re-quote. Zero on a market built by the
one-shot generator; drawn from FREEZE_RANGE here. See FREEZE_SALT."""
def item(self, name: str) -> Item | None:
wanted = name.strip().casefold()
for item in self.items:
if item.name.casefold() == wanted:
return item
return None
def _volumes(rng: random.Random, volume: float, ticks: int) -> list[int]:
return [max(1, round(volume * math.exp(rng.gauss(0.0, VOLUME_SD)))) for _ in range(ticks)]
def _walk(rng: random.Random, base: float, ticks: int) -> list[int]:
"""The decoy: no anchor, no reversion, just a wide random walk. Whatever a moving
average says about where this price belongs is a statement about the past only."""
price = base
out = []
for _ in range(ticks):
price *= 1.0 + rng.gauss(0.0, DRIFT_STEP)
out.append(max(1, round(price)))
return out
def _stream(rng: random.Random, base: float, volume: float, noise_sd: float, ticks: int
) -> tuple[list[int], list[int]]:
"""One item's price and volume history, visible and held-out ticks together."""
# eps is scaled so the AR(1) settles at exactly noise_sd rather than drifting toward it
# over the first few ticks; the visible half would otherwise be quieter than the graded
# half and every anchor estimated from it would be too tight.
eps_sd = noise_sd * math.sqrt(1.0 - PHI * PHI)
x = rng.gauss(0.0, noise_sd)
fundamental = base
prices = []
for _ in range(ticks):
fundamental *= 1.0 + rng.gauss(0.0, FUND_DRIFT)
x = PHI * x + rng.gauss(0.0, eps_sd)
prices.append(max(1, round(fundamental * (1.0 + x))))
return prices, _volumes(rng, volume, ticks)
def build_market(seed: int, num_items: int, visible: int, held_out: int) -> Market:
"""One basket, from one seed.
Every tier is represented before any tier repeats, so a thin item and a deep one are
always both on the table: the allocation choice is the task, and a basket that happened
to be all-deep or all-thin would not pose it.
"""
rng = random.Random(seed)
ticks = visible + held_out
order = list(range(len(TIERS)))
rng.shuffle(order)
picks = [order[i % len(order)] for i in range(num_items)]
# Which slots are walks, and how many, are both drawn here. The tier is drawn
# independently of the walk flag, so a basket where the walk was always the cheap item —
# solvable by reading the price column and never the shape — cannot arise.
decoys = set(rng.sample(range(num_items), rng.randint(*WALKS)))
used: set[str] = set()
items = []
for slot, tier_idx in enumerate(picks):
names, (lo, hi), (vlo, vhi), limit, noise_sd = TIERS[tier_idx]
choices = [n for n in names if n not in used] or names
name = rng.choice(choices)
used.add(name)
base, volume = rng.uniform(lo, hi), rng.uniform(vlo, vhi)
if slot in decoys:
prices, volumes = _walk(rng, base, ticks), _volumes(rng, volume, ticks)
else:
prices, volumes = _stream(rng, base, volume, noise_sd, ticks)
items.append(Item(name=name, reverting=slot not in decoys, buy_limit=limit,
prices=prices, volumes=volumes))
items.sort(key=lambda i: i.name)
# Off its own RNG, so `rng` above is left exactly where the one-shot generator leaves
# it and every stream is byte-identical to the original's.
freeze = random.Random(seed ^ FREEZE_SALT).randint(*FREEZE_RANGE)
return Market(seed=seed, items=items, visible=visible, held_out=held_out,
capital=STARTING_CAPITAL, freeze=freeze)
@@ -0,0 +1,540 @@
"""Measure the policy ladder in the SHIPPED engine, before any constant is locked.
The spec this environment came from locked `horizon`, `FREEZE_RANGE`, `TARGET_SHARE` and
`TURNS` against arithmetic from a throwaway engine that no longer exists, and its ladder did
not reconcile with its own reward — impatient computes to 0.612 where 0.623 was reported.
Nothing here is taken from that document. Every row below is produced by `live.play`, the
engine the taskset will actually run, over the baskets a run will actually serve.
Two questions this file exists to answer, in this order:
1. What does each policy score? Nine of them, from doing nothing to the reference itself,
each differing from its neighbour by one judgement.
2. **Does the turn budget bind?** Every reward in Arena saturates — `profit_ratio` and
`efficiency` clip at 1.0 and `clean` is a boolean band — so an exhaustive policy that
clears the band scores 1.000 and house rule 4 is structurally unreachable. That is
risk #2 in the programme's register, it is answered as a NUMBER, and if the number is
zero that is a finding about the reward shape rather than a failure of this lane.
Read block-wise, never as one mean: realised P&L over sixty ticks is the noisiest thing in
the repository, and `probe.py` already documents this generator swinging 0.48 against 0.18
between blocks of twenty-four. A ladder read off one block is a ladder read off the seed.
uv run --no-project python measure_ladder.py
uv run --no-project python measure_ladder.py --sweep-target
uv run --no-project python measure_ladder.py --sweep-requote
uv run --no-project python measure_ladder.py --sweep-reference
uv run --no-project python measure_ladder.py --sweep-freeze
uv run --no-project python measure_ladder.py --sweep-shape
--- what it measured, 2026-08-21, amd-server, 120 baskets from SEED_BASE ------------------
TURNS=8, TICKS_PER_TURN=7, TAIL_TICKS=4, HORIZON=60, FREEZE_RANGE=(1,4), TARGET_SHARE=0.90.
policy total profit effcy discip clean gp looks
inaction 0.000 0.000 0.000 0.000 0.000 0 0.00
staller 0.000 0.000 0.000 0.000 0.000 0 0.00
crude (market orders, one look) 0.126 0.207 0.186 0.095 0.017 -2,684 1.00
spammer 0.158 0.262 0.219 0.114 0.025 13,938 7.89
restate (re-quote the same book) 0.219 0.375 0.362 0.168 0.000 25,244 7.90
plausible (mean anchor, no filter) 0.406 0.563 0.524 0.391 0.142 37,084 1.00
churn (re-place at full size) 0.447 0.684 0.550 0.430 0.042 49,926 8.00
impatient (one-shot ref, one look) 0.538 0.739 0.722 0.560 0.150 50,172 1.00
hindsight (whole-window anchor) 0.598 0.792 0.763 0.630 0.208 54,599 1.00
exhaustive (re-quote every look) 0.631 0.849 0.749 0.669 0.192 63,802 7.97
oracle (the live reference) 1.000 1.000 1.000 1.000 1.000 76,744 3.99
Every row reconciles with 0.45*profit + 0.30*discipline + 0.25*clean to three decimals,
which the spec's ladder did not.
**The budget binds. Oracle - exhaustive = 0.369, stdev 0.054 over five blocks of 24, worst
block 0.316.** That is fifteen times the 0.02 margin B2 proposed and it clears the 0.25
quantum of the binary `gate` term outright, so the comparison is not a cliff. It is also
not an artifact of the denominator: the reference earns 76,744 gp against exhaustive's
63,802 and out-returns it on peak capital, 0.302 against 0.219.
`hindsight` is the row that says the stepped form measures something the one-shot form
cannot. It quotes off the mean of the WHOLE window, held-out ticks included, and still
scores 0.598 — perfect price foresight in a single plan is worth less than playing the
window with none.
Three things this measurement overturned:
the first engine's rule-4 gap was fake. With a reference that never redeployed sale
proceeds, exhaustive earned 22% MORE gp than the reference and scored 0.848: the gap was
the denominator being under-tuned, exactly the imitation-score defect `book.py` documents
paying for once already. Fixed by giving the reference the idle-cash trigger.
the freeze alone is not what binds. At freeze 0 the gap is still 0.180 — the queue reset
and the capital churn carry it. The freeze sets the slope: 0.238 at 1 tick, 0.505 at 4,
0.902 at 8. FREEZE_RANGE=(1,4) is kept because 8 turns a look into a catastrophe rather
than a cost, and because flat-per-look is the settled decision.
more turns bind harder, monotonically: 4x14+4 gives 0.155, 8x7+4 gives 0.369, 12x4+12
gives 0.517. Eight is chosen for the token cost of a rollout, not because the gap needs
it. If the gap ever needs widening, this is the lever, and it is measured.
"""
from __future__ import annotations
import random
import statistics
from dataclasses import replace
import sys
import time
import types
from pathlib import Path
HERE = Path(__file__).resolve().parent
_shim = types.ModuleType("grand_exchange_live")
_shim.__path__ = [str(HERE / "grand_exchange_live")]
sys.modules["grand_exchange_live"] = _shim
from grand_exchange_live.book import ( # noqa: E402
BUY_BAND,
FILL_SHARE,
MAX_ITEM_SHARE,
MIN_CROSSINGS,
SELL_BAND,
TAX,
Order,
crossings,
reference_orders,
)
from grand_exchange_live.live import ( # noqa: E402
HORIZON,
IDLE_SHARE,
LIVE_VIABILITY_TRIES,
REQUOTE_BAND,
TAIL_TICKS,
TICKS_PER_TURN,
TURNS,
Leg,
Reference,
desired_orders,
live_reference,
play,
viable_live_market,
)
from grand_exchange_live.market import SEED_BASE, build_market # noqa: E402
HEADER = "requote / idle"
BLOCK = 24
BLOCKS = 5
TARGET_SHARE = 0.90
"""What counts as a full score, as a share of the LIVE reference's realised profit and of
its return on peak capital. Chosen by `--sweep-target`, not carried over from the spec,
which proposed 0.85 against an engine that no longer exists.
The two things it trades off are both measured over 120 baskets. Raising it widens the
rule-4 gap (exhaustive scores 0.723 at 0.80 and 0.574 at 1.00) and narrows the ceiling
plateau — and the plateau is what stops the reward being an imitation score for the
reference's own constants. At 0.90 the best-earning member of the reference's family that
is NOT the reference earns 1.1% more gp and still scores 0.976, far above the 0.90 bar
`probe.py` holds the one-shot form to, and the gap to exhaustive is 0.372. It is also the
one-shot environment's value, which is one fewer constant that differs between two forms of
the same market for no measured reason."""
# --- the baskets a run actually serves --------------------------------------------------
def baskets(count: int, base: int = SEED_BASE):
"""Skip-aware, exactly as `Taskset.load` will be: `viable_live_market` may pass over a
seed the LIVE reference loses money in, and the next task starts after the seed it
landed on. Grading a contiguous `range` instead would grade a different set of baskets
from the one a model is served."""
out, seed = [], base
for _ in range(count):
market = viable_live_market(seed, 5, 56, HORIZON)
seed = market.seed + 1
out.append(market)
return out
# --- scoring ----------------------------------------------------------------------------
def score(result, reference, target_share: float = TARGET_SHARE) -> dict[str, float]:
"""The reward, in one place, so no two rows can disagree about the same episode."""
target = target_share * reference.realised
profit = 0.0 if target <= 0 else min(1.0, max(0.0, result.realised / target))
efficiency = 0.0 if reference.roc <= 0 else min(1.0, max(0.0, result.roc / reference.roc))
discipline = profit * efficiency
clean = float(result.realised >= target_share * reference.realised
and result.roc >= target_share * reference.roc)
return {
"total": 0.45 * profit + 0.30 * discipline + 0.25 * clean,
"profit_ratio": profit,
"efficiency": efficiency,
"discipline": discipline,
"clean": clean,
}
# --- the ladder, as policies rather than as adjectives ----------------------------------
def inaction(view):
"""Never submits a sheet. House rule 3's floor."""
return []
def staller(view):
"""Takes every look and does nothing with any of them. Scores what inaction scores —
the point of the row is that BURNING the budget is not itself worth anything, which is
the degenerate strategy a time-averaged reward would pay for."""
return []
def spammer(seed: int):
"""Random legs on every look. Acts constantly, reads nothing."""
rng = random.Random(seed)
def policy(view):
items = view.market.items
return [
Leg(i.name, rng.randint(1, i.buy_limit),
round(i.prices[view.seen - 1] * rng.uniform(0.85, 1.15)),
round(i.prices[view.seen - 1] * rng.uniform(0.85, 1.15)))
for i in items if rng.random() < 0.6
]
return policy
def one_look(plan_of):
"""Submit a plan on turn 1, never look again. The one-shot form, played live."""
def make(market):
state = {"done": False}
orders = plan_of(market)
def policy(view):
if state["done"]:
return []
state["done"] = True
return [Leg(o.item, o.quantity, o.buy, o.sell) for o in orders]
return policy
return make
def band(market, items, anchor, sizer, filtered=False, ranked=False):
"""`probe.py`'s strategy family, unchanged, so the live ladder's crude and plausible
rungs are the same strategies the one-shot ladder grades."""
plans, out, purse = [], [], float(market.capital)
for item in items:
prices = item.visible_prices(market.visible)
volumes = item.visible_volumes(market.visible)
if filtered and crossings(prices) < MIN_CROSSINGS:
continue
a = anchor(prices)
buy = max(1, round(a * (1.0 - BUY_BAND)))
sell = max(buy + 1, round(a * (1.0 + SELL_BAND)))
lows = [p for p in prices if p <= buy]
highs = [p for p in prices if p >= sell]
qty = sizer(item, market, buy, prices, volumes)
if qty <= 0:
continue
s = 0.0
if ranked:
if not lows or not highs:
continue
entry, exit_ = statistics.fmean(lows), statistics.fmean(highs)
s = (exit_ * (1.0 - TAX) - entry) / entry
if s <= 0:
continue
plans.append((s, item.name, qty, buy, sell))
plans.sort(key=lambda p: -p[0])
for _, name, qty, buy, sell in plans:
qty = min(qty, int(purse // buy))
if qty <= 0:
continue
purse -= qty * buy
out.append(Order(name, qty, buy, sell))
return out
def by_volume(item, market, buy, prices, volumes):
lows = [p for p in prices if p <= buy]
reach = int(FILL_SHARE * statistics.median(volumes) * (len(lows) / len(prices))
* market.held_out)
return min(item.buy_limit, reach, int(MAX_ITEM_SHARE * market.capital // buy))
def market_orders(market):
"""Buy ten percent over the last price and sell ten under: every offer fills and every
round trip pays the tax. The crude maximiser."""
return [
Order(i.name, i.buy_limit, round(i.prices[market.visible - 1] * 1.10),
round(i.prices[market.visible - 1] * 0.90))
for i in market.items
]
def churn(market):
"""Re-place the ORIGINAL plan at full size on every look.
Not the same policy as re-stating the book: the units already bought release their
reservation and sale proceeds have returned to the purse, so asking for the full size
again spends them. It earns roughly what a single look earns off materially more peak
capital, which is precisely the shape `efficiency` exists to catch.
"""
orders = reference_orders(market)
return lambda view: [Leg(o.item, o.quantity, o.buy, o.sell) for o in orders]
def restate(market):
"""Place the one-shot reference on turn 1, then re-submit the book exactly as it stands
on every look after it.
The pure cost of looking, isolated: same names, same prices, same units still sought,
nothing learned, one freeze and one queue reset per turn. Its gap below `impatient` — an
identical plan that simply stops looking — is the price of a look with no information in
it, in reward units.
"""
orders = reference_orders(market)
def policy(view):
if view.turn == 1:
return [Leg(o.item, o.quantity, o.buy, o.sell) for o in orders]
return [Leg(o.item, o.quantity, o.buy, o.sell) for o in view.open_orders.values()]
return policy
def exhaustive(market):
"""Spend the entire budget: re-quote on EVERY look, off everything observed so far.
This is the rule-4 policy. It makes every judgement the reference makes and one more
look than the reference thinks is worth paying for, so if it scores at the oracle the
budget is free and the environment measures nothing the one-shot form did not.
"""
inner = Reference(requote_band=0.0)
return inner
def perfect_hindsight(market):
"""Not shipped, not a rung: a sanity ceiling. Quotes off the mean of the WHOLE window,
held-out ticks included. If the oracle were close to this, the visible half would not be
the thing being read."""
prices = {i.name: list(i.prices) for i in market.items}
volumes = {i.name: list(i.volumes) for i in market.items}
orders = desired_orders(market.items, prices, volumes, float(market.capital),
market.held_out, {i.name: i.buy_limit for i in market.items})
return one_look(lambda m: orders)(market)
LADDER = [
("inaction", lambda m: inaction),
("staller", lambda m: staller),
("spammer", lambda m: spammer(m.seed)),
("crude (market orders, one look)", one_look(market_orders)),
("plausible (mean anchor, no filter)",
one_look(lambda m: band(m, m.items, statistics.fmean, by_volume))),
("impatient (one-shot reference, one look)", one_look(reference_orders)),
("restate (re-quote the same book every look)", restate),
("churn (re-place at full size every look)", churn),
("exhaustive (re-quote every look)", exhaustive),
("oracle (the live reference)", lambda m: Reference()),
("hindsight (not a rung: whole-window anchor)", perfect_hindsight),
]
def ladder(markets, target_share: float = TARGET_SHARE) -> dict[str, dict[str, float]]:
rows: dict[str, dict[str, float]] = {}
for name, make in LADDER:
totals: dict[str, float] = {}
gp = looks = frozen = 0.0
peak = 0.0
for market in markets:
result = play(market, make(market))
for key, value in score(result, live_reference(market), target_share).items():
totals[key] = totals.get(key, 0.0) + value
gp += result.realised
looks += result.looks
frozen += result.frozen_ticks
peak += result.peak_employed
n = len(markets)
rows[name] = {k: v / n for k, v in totals.items()}
rows[name].update(gp=gp / n, looks=looks / n, frozen=frozen / n, peak=peak / n)
return rows
def table(markets, target_share: float = TARGET_SHARE) -> None:
rows = ladder(markets, target_share)
print(f"{'policy':<46}{'total':>8}{'profit':>8}{'effcy':>8}{'discip':>8}{'clean':>7}"
f"{'gp':>12}{'peak gp':>10}{'looks':>7}{'frozen':>7}")
for name, row in rows.items():
print(f"{name:<46}{row['total']:>8.3f}{row['profit_ratio']:>8.3f}"
f"{row['efficiency']:>8.3f}{row['discipline']:>8.3f}{row['clean']:>7.3f}"
f"{row['gp']:>12,.0f}{row['peak']:>10,.0f}{row['looks']:>7.2f}"
f"{row['frozen']:>7.2f}")
def main() -> None:
flags = set(sys.argv[1:])
started = time.monotonic()
seeds = baskets(BLOCK * BLOCKS)
build = time.monotonic() - started
blocks = [seeds[i:i + BLOCK] for i in range(0, len(seeds), BLOCK)]
skipped = sum(1 for i, m in enumerate(seeds)
if m.seed != (seeds[i - 1].seed + 1 if i else SEED_BASE))
print(f"engine TURNS={TURNS} TICKS_PER_TURN={TICKS_PER_TURN} TAIL={TAIL_TICKS} "
f"HORIZON={HORIZON} REQUOTE_BAND={REQUOTE_BAND:.4f} TARGET_SHARE={TARGET_SHARE}")
print(f"baskets {len(seeds)} over {BLOCKS} blocks of {BLOCK}, "
f"{skipped} seeds skipped by viability (cap {LIVE_VIABILITY_TRIES}), "
f"built in {build:.2f}s")
print(f"freeze {statistics.fmean([m.freeze for m in seeds]):.2f} ticks mean, "
f"distribution {sorted({f: [m.freeze for m in seeds].count(f) for f in {m.freeze for m in seeds}}.items())}")
print()
print("=== the ladder, all 120 baskets ===")
table(seeds)
print()
print("=== does the budget bind? exhaustive against oracle, block by block ===")
print(f"{'block':<22}{'exhaustive':>12}{'oracle':>10}{'gap':>10}"
f"{'exh gp':>12}{'ora gp':>12}{'exh looks':>11}{'ora looks':>11}")
gaps = []
for i, block in enumerate(blocks):
rows = ladder(block)
e, o = rows["exhaustive (re-quote every look)"], rows["oracle (the live reference)"]
gaps.append(o["total"] - e["total"])
print(f"{f'{block[0].seed}-{block[-1].seed}':<22}{e['total']:>12.3f}{o['total']:>10.3f}"
f"{o['total'] - e['total']:>10.3f}{e['gp']:>12,.0f}{o['gp']:>12,.0f}"
f"{e['looks']:>11.2f}{o['looks']:>11.2f}")
print(f"{'mean':<22}{'':>12}{'':>10}{statistics.fmean(gaps):>10.3f}")
print(f"{'stdev':<22}{'':>12}{'':>10}{statistics.stdev(gaps):>10.3f}")
print(f"{'worst block':<22}{'':>12}{'':>10}{min(gaps):>10.3f}")
print()
print("=== every rung, block by block (total only) ===")
per_block = [ladder(b) for b in blocks]
print(f"{'policy':<46}" + "".join(f"{f'b{i}':>9}" for i in range(len(blocks)))
+ f"{'spread':>9}")
for name, _ in LADDER:
values = [rows[name]["total"] for rows in per_block]
print(f"{name:<46}" + "".join(f"{v:>9.3f}" for v in values)
+ f"{max(values) - min(values):>9.3f}")
print()
if "--sweep-target" in flags:
print("=== TARGET_SHARE sweep (the ceiling band) ===")
names = [n for n, _ in LADDER]
print(f"{'share':<8}" + "".join(f"{n.split(' ')[0]:>12}" for n in names))
for share in (0.75, 0.80, 0.85, 0.90, 0.95, 1.00):
rows = ladder(seeds, share)
print(f"{share:<8.2f}" + "".join(f"{rows[n]['total']:>12.3f}" for n in names))
print()
if "--sweep-requote" in flags:
print("=== REQUOTE_BAND sweep (what the reference should pay a freeze for) ===")
print(f"{'band':<8}{'gp':>12}{'roc':>10}{'looks':>8}{'frozen':>8}")
for band_value in (0.0, 0.005, 0.01, 0.015, 0.025, 0.04, 0.06, 0.10, 1.0):
gp = roc = looks = frozen = 0.0
for market in seeds:
result = play(market, Reference(requote_band=band_value))
gp += result.realised
roc += result.roc
looks += result.looks
frozen += result.frozen_ticks
n = len(seeds)
print(f"{band_value:<8.3f}{gp / n:>12,.0f}{roc / n:>10.4f}"
f"{looks / n:>8.2f}{frozen / n:>8.2f}")
print()
if "--sweep-reference" in flags:
# The denominator has to be the profit-maximising point of its OWN family, or some
# other member earns more gp and scores less, and the reward's argmax is a pair of
# constants rather than a strategy. This is the same sweep `book.py` documents for
# the one-shot bands, over the two triggers that make the live reference act.
print("=== reference family sweep: gp per basket ===")
idles = (0.0, 0.05, 0.10, 0.15, 0.20, 0.30, 0.50, 1.01)
print(f"{HEADER:<16}" + "".join(f"{i:>10.2f}" for i in idles))
best = (0.0, None)
for band_value in (0.0, 0.005, 0.010, 0.015, 0.025, 0.040, 0.060, 1.0):
row = []
for idle in idles:
gp = sum(play(m, Reference(band_value, idle)).realised for m in seeds)
row.append(gp / len(seeds))
if row[-1] > best[0]:
best = (row[-1], (band_value, idle))
print(f"{band_value:<16.3f}" + "".join(f"{v:>10,.0f}" for v in row))
print(f"argmax {best[1]} at {best[0]:,.0f} gp per basket; shipped "
f"({REQUOTE_BAND}, {IDLE_SHARE}) is "
f"{sum(play(m, Reference()).realised for m in seeds) / len(seeds):,.0f}")
print()
print("=== the same family, in return on peak capital ===")
print(f"{HEADER:<16}" + "".join(f"{i:>10.2f}" for i in idles))
for band_value in (0.0, 0.005, 0.010, 0.015, 0.025, 0.040, 0.060, 1.0):
row = []
for idle in idles:
roc = sum(play(m, Reference(band_value, idle)).roc for m in seeds)
row.append(roc / len(seeds))
print(f"{band_value:<16.3f}" + "".join(f"{v:>10.4f}" for v in row))
print()
print("=== and what each family member SCORES against the shipped reference ===")
print(f"{HEADER:<16}" + "".join(f"{i:>10.2f}" for i in idles))
for band_value in (0.0, 0.005, 0.010, 0.015, 0.025, 0.040, 0.060, 1.0):
row = []
for idle in idles:
total = sum(score(play(m, Reference(band_value, idle)), live_reference(m))["total"]
for m in seeds)
row.append(total / len(seeds))
print(f"{band_value:<16.3f}" + "".join(f"{v:>10.3f}" for v in row))
print()
if "--sweep-freeze" in flags:
# If no freeze makes looking expensive enough for the optimum to be interior, the
# freeze is decoration and rule 4 is being carried by the denominator alone.
print("=== freeze sweep: what a look has to cost before it is worth declining ===")
print(f"{'freeze':<10}{'exhaust gp':>12}{'oracle gp':>12}{'exh score':>11}"
f"{'gap':>8}{'ora looks':>11}{'ora frozen':>12}")
for forced in (0, 1, 2, 3, 4, 6, 8, 12):
shaped = [replace(m, freeze=forced) for m in seeds]
e = g = 0.0
egp = ogp = looks = frozen = 0.0
for m in shaped:
ref = play(m, Reference())
exh = play(m, Reference(requote_band=0.0, idle_share=0.0))
if ref.realised <= 0:
continue
e += score(exh, ref)["total"]
g += 1.0 - score(exh, ref)["total"]
egp += exh.realised
ogp += ref.realised
looks += ref.looks
frozen += ref.frozen_ticks
n = len(shaped)
print(f"{forced:<10}{egp / n:>12,.0f}{ogp / n:>12,.0f}{e / n:>11.3f}"
f"{g / n:>8.3f}{looks / n:>11.2f}{frozen / n:>12.2f}")
print()
if "--sweep-shape" in flags:
print("=== window shape sweep: does a different budget bind harder? ===")
print(f"{'turns x step + tail':<22}{'horizon':>9}{'exhaust':>9}{'oracle':>9}"
f"{'gap':>9}{'ora looks':>11}{'held out':>10}")
for turns, step, tail in ((4, 14, 4), (6, 9, 6), (8, 7, 4), (8, 5, 20),
(10, 5, 10), (12, 4, 12)):
horizon = turns * step + tail
shaped = [build_market(m.seed, 5, 56, horizon) for m in seeds]
e = o = 0.0
looks = 0.0
for market in shaped:
ref = play(market, Reference(), turns=turns, step=step, tail=tail)
exh = play(market, Reference(requote_band=0.0), turns=turns, step=step,
tail=tail)
if ref.realised <= 0:
continue
e += score(exh, ref)["total"]
o += score(ref, ref)["total"]
looks += ref.looks
n = len(shaped)
print(f"{f'{turns} x {step} + {tail}':<22}{horizon:>9}{e / n:>9.3f}{o / n:>9.3f}"
f"{(o - e) / n:>9.3f}{looks / n:>11.2f}{step + tail:>10}")
print()
print(f"total wall clock {time.monotonic() - started:.1f}s")
if __name__ == "__main__":
main()
@@ -0,0 +1,13 @@
[project]
name = "grand-exchange-live"
version = "0.1.0"
description = "grand-exchange-live — quote, watch the tape, re-quote, over a turn budget."
requires-python = ">=3.11"
dependencies = ["verifiers"]
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
[tool.hatch.build.targets.wheel]
packages = ["grand_exchange_live"]
@@ -0,0 +1,498 @@
"""The stepped engine's invariants, as tests rather than as docstrings.
Five claims are load-bearing and every one of them is a way the environment could be quietly
wrong rather than loudly broken:
the fixed point one look on turn 1 and nothing after it must reproduce the one-shot
engine EXACTLY — same realised profit, same capital committed. If it
does not, the live form is a different market wearing the same
generator, and no number measured here is comparable to a `grand-exchange`
number.
the ledger cash, coins locked, cost basis and realised profit must sum to the
starting capital after every single tick. A bug that mints money reads
as a good score, not as a crash.
the buy limit cumulative across the whole window, so cancel-and-re-place cannot buy
the same item twice over. Without it, re-quoting resets the limit and
the limit stops being a limit.
the freeze a look that touches a live book stops the WHOLE book for `market.freeze`
ticks, and the first placement is free.
the queue an amended offer starts again at zero accrued fill room, so moving a
limit costs even after the freeze has expired.
"""
from __future__ import annotations
import random
import statistics
import sys
import types
import unittest
from pathlib import Path
HERE = Path(__file__).resolve().parents[1]
ONE_SHOT = HERE.parent / "grand_exchange"
for _root in (HERE, ONE_SHOT):
_package = _root / _root.name
_shim = types.ModuleType(_package.name)
_shim.__path__ = [str(_package)]
sys.modules[_package.name] = _shim
from grand_exchange_live.book import Order, execute, reference_orders # noqa: E402
from grand_exchange_live.live import ( # noqa: E402
HORIZON,
TAIL_TICKS,
TICKS_PER_TURN,
TURNS,
Leg,
Ledger,
Reference,
desired_orders,
play,
)
from grand_exchange_live.market import build_market # noqa: E402
SEEDS = range(60_000, 60_024)
def market(seed: int, held_out: int = HORIZON):
return build_market(seed, 5, 56, held_out)
def once(orders):
"""Submit a plan on turn 1 and never look again — the one-shot strategy, played live."""
state = {"done": False}
def policy(view):
if state["done"]:
return []
state["done"] = True
return [Leg(o.item, o.quantity, o.buy, o.sell) for o in orders]
return policy
def restate(view):
"""Re-submit the book exactly as it already stands, on every look.
The true no-op re-quote, and the only honest way to price the freeze and the queue
reset on their own: same names, same prices, same units still sought. Nothing about the
plan changes, so any difference in what it buys or earns IS the cost of looking.
Note that re-submitting the ORIGINAL quantities instead is a different policy and a
bigger one — the units already bought free up their reservation, sale proceeds return
to the purse, and asking for the full size again spends them. That policy is `churn` in
the ladder; it earns about the same gp off a third more peak capital, which is what
`efficiency` is for.
"""
return [Leg(o.item, o.quantity, o.buy, o.sell) for o in view.open_orders.values()]
class FixedPointTests(unittest.TestCase):
def test_single_look_matches_one_shot(self) -> None:
"""The stepped engine, used exactly the way the one-shot engine is used, IS the
one-shot engine. This is the assertion that makes the two forms comparable."""
for seed in SEEDS:
m = market(seed)
plan = reference_orders(m)
stepped = play(m, once(plan))
flat = execute(m, plan)
self.assertAlmostEqual(stepped.realised, flat.realised, places=6, msg=f"seed {seed}")
self.assertAlmostEqual(stepped.peak_employed, flat.committed, places=6,
msg=f"seed {seed}")
self.assertEqual(stepped.bought, flat.bought, f"seed {seed}")
self.assertEqual(stepped.sold, flat.sold, f"seed {seed}")
self.assertEqual(stepped.dumped, flat.dumped, f"seed {seed}")
self.assertEqual(stepped.frozen_ticks, 0, "the first placement is not free")
def test_turn_one_matches_reference_orders(self) -> None:
"""`desired_orders` is `book.reference_orders` generalised to a moment in the
window, and on turn 1 the moment is the one the one-shot form scores."""
for seed in SEEDS:
m = market(seed)
wanted = desired_orders(
m.items,
{i.name: list(i.visible_prices(m.visible)) for i in m.items},
{i.name: list(i.visible_volumes(m.visible)) for i in m.items},
float(m.capital),
m.held_out,
{i.name: i.buy_limit for i in m.items},
)
self.assertEqual(
[(o.item, o.quantity, o.buy, o.sell) for o in wanted],
[(o.item, o.quantity, o.buy, o.sell) for o in reference_orders(m)],
f"seed {seed}",
)
class LedgerTests(unittest.TestCase):
def test_nothing_mints_coins(self) -> None:
"""Cash + locked + cost basis + realised == capital, checked after every tick of
every turn by instrumenting `Ledger.mark`, which the engine calls on each one."""
broken: list[str] = []
original = Ledger.mark
def checked(self) -> None:
total = self.cash + self.locked + sum(self.basis.values()) \
+ sum(self.incoming_basis.values())
if abs((total - self.realised) - self.capital) > 1e-6:
broken.append(f"identity: {total:.6f} against {self.capital:.6f}")
# The purse is the binding constraint on everything: an offer is funded out of
# it or it is not placed. A negative purse is the engine having spent coins that
# were not there, which reads as a good score rather than as a crash.
if self.cash < -1e-9:
broken.append(f"purse: {self.cash:.6f}")
original(self)
Ledger.mark = checked
try:
for seed in SEEDS:
m = market(seed)
for policy in (Reference(), once(reference_orders(m)), restate):
play(m, policy)
finally:
Ledger.mark = original
self.assertEqual(broken, [], "the ledger identity broke")
def test_inaction_is_exactly_zero(self) -> None:
for seed in SEEDS:
result = play(market(seed), lambda view: [])
self.assertEqual(result.realised, 0.0)
self.assertEqual(result.peak_employed, 0.0)
self.assertEqual(result.roc, 0.0)
self.assertEqual(result.looks, 0)
def test_capital_employed_can_exceed_the_starting_purse(self) -> None:
"""A policy that recycles sale proceeds can employ more than the starting capital,
and that is the whole point of a purse the proceeds return to — a one-shot plan
cannot spend what its own sales earned. What it can never do is employ coins that
do not exist, and that is `test_nothing_mints_coins`'s negative-purse check, not a
bound on the final realised profit: gains made mid-window and given back later are
real capital while they are employed."""
peaks = [play(market(s), Reference()).peak_employed for s in SEEDS]
self.assertTrue(all(p > 0.0 for p in peaks))
self.assertTrue(any(p > 250_000 for p in peaks),
"no basket ever employs more than the starting purse — sale "
"proceeds are not being redeployed and the live form is a "
"one-shot form with extra steps")
class BuyLimitTests(unittest.TestCase):
def test_the_window_limit_is_cumulative_across_requotes(self) -> None:
"""Ask for the whole limit on every one of the eight looks. The window limit is the
window's, not the offer's, so the total bought can never exceed it."""
greedy = lambda view: [
Leg(i.name, i.buy_limit, round(i.prices[view.seen - 1] * 1.20),
round(i.prices[view.seen - 1] * 0.80))
for i in view.market.items
]
seen_a_cap = False
for seed in SEEDS:
m = market(seed)
result = play(m, greedy)
for item in m.items:
got = result.bought_by_item.get(item.name, 0)
self.assertLessEqual(got, item.buy_limit,
f"seed {seed}: bought {got} of {item.name} against a "
f"window limit of {item.buy_limit}")
seen_a_cap = seen_a_cap or got == item.buy_limit
self.assertTrue(seen_a_cap, "no basket ever reached a buy limit — the invariant is "
"untested by this policy rather than upheld by the engine")
class FreezeTests(unittest.TestCase):
def test_the_first_placement_is_free_and_the_second_is_not(self) -> None:
for seed in SEEDS:
m = market(seed)
plan = reference_orders(m)
if not plan:
continue
self.assertEqual(play(m, once(plan)).frozen_ticks, 0)
state = {"n": 0}
def twice(view, plan=plan, state=state):
state["n"] += 1
if state["n"] <= 2:
return [Leg(o.item, o.quantity, o.buy, o.sell) for o in plan]
return []
result = play(m, twice)
self.assertEqual(result.frozen_ticks, m.freeze,
f"seed {seed}: one re-quote should freeze {m.freeze} ticks")
self.assertEqual(result.amendments, 1)
self.assertEqual(result.looks, 2)
def test_a_frozen_tick_fills_nothing(self) -> None:
"""The freeze has to bite on the tape, not only on a counter. Re-quoting the
identical plan changes nothing about what is wanted, so any difference in units
bought is the freeze and the queue reset — and it must never be an increase."""
churned = fresh = 0
for seed in SEEDS:
m = market(seed)
plan = reference_orders(m)
fresh += play(m, once(plan)).bought
churned += play(m, restate).bought
self.assertLess(churned, fresh,
f"churning the identical plan bought {churned} units against "
f"{fresh} — re-quoting costs nothing")
def test_churn_earns_less_than_standing_still(self) -> None:
"""The same claim in gp, which is what the reward actually reads."""
churned = fresh = 0.0
for seed in SEEDS:
m = market(seed)
plan = reference_orders(m)
fresh += play(m, once(plan)).realised
churned += play(m, restate).realised
self.assertLess(churned, fresh, f"churn earned {churned:,.0f} gp against {fresh:,.0f}")
class QueueTests(unittest.TestCase):
def test_an_amendment_resets_accrued_fill_room(self) -> None:
"""A thin item accrues a quarter of a tick's volume per eligible tick and carries it
between ticks — that carry is the only reason an item trading two units a tick is
tradeable rather than untradeable. Throwing it away on every look must show up as
fewer units bought even when the freeze is one tick, so the two mechanics are
separable rather than one mechanic counted twice."""
thin = [s for s in SEEDS if build_market(s, 5, 56, HORIZON).freeze == 1]
self.assertTrue(thin, "no one-tick-freeze basket in the block — test proves nothing")
worse = 0
for seed in thin:
m = market(seed)
plan = reference_orders(m)
if play(m, restate).bought < play(m, once(plan)).bought:
worse += 1
self.assertGreater(worse, 0, "with a one-tick freeze, re-quoting cost nothing at all")
class HeldOutTests(unittest.TestCase):
def test_the_tail_after_the_last_look_is_never_observed(self) -> None:
"""House rule 1, in its WEAKER live form, asserted on the size of the slice rather
than on the frontier. An agent that takes every look still has TICKS_PER_TURN +
TAIL_TICKS ticks executed after its final observation."""
seen: list[int] = []
def watcher(view):
seen.append(view.seen - view.market.visible)
return []
for seed in SEEDS:
m = market(seed)
seen.clear()
play(m, watcher)
self.assertEqual(len(seen), TURNS)
self.assertEqual(seen[0], 0, "the first look sees a graded tick")
self.assertEqual(seen[-1], (TURNS - 1) * TICKS_PER_TURN)
self.assertEqual(HORIZON - seen[-1], TICKS_PER_TURN + TAIL_TICKS)
self.assertGreaterEqual(HORIZON - seen[-1], 4)
def test_a_view_never_carries_a_graded_tick_early(self) -> None:
"""The observation handed over on a look must contain only ticks that have already
been executed against — a one-tick lookahead here would hand the agent the answer."""
for seed in (60_000, 60_007, 60_013):
m = market(seed)
def strict(view, m=m):
for item in m.items:
got = view.new_prices[item.name]
start = view.seen - len(got)
self.assertEqual(got, item.prices[start:view.seen])
self.assertLessEqual(view.seen, m.visible + m.held_out)
return []
play(m, strict)
class SheetTests(unittest.TestCase):
def test_a_cancel_returns_the_coins_and_stops_the_buying(self) -> None:
m = market(60_000)
plan = reference_orders(m)
name = plan[0].item
def cancel_after_one(view, plan=plan, name=name):
if view.turn == 1:
return [Leg(o.item, o.quantity, o.buy, o.sell) for o in plan]
if view.turn == 2:
return [Leg(name, 0)]
return []
held = play(m, once(plan)).bought_by_item.get(name, 0)
cut = play(m, cancel_after_one).bought_by_item.get(name, 0)
self.assertLess(cut, held, "cancelling an offer did not stop it buying")
def test_a_leg_naming_nothing_on_the_board_is_dropped_not_fatal(self) -> None:
"""A malformed sheet scores what doing nothing scores. It must never raise: the
reward runs inside the metric, so a raise takes the whole rollout with it rather
than scoring zero."""
junk = lambda view: [
Leg("Nonexistent bauble", 10, 5, 6),
Leg("", 1, 1, 2),
Leg(view.market.items[0].name, 5, 0, 0),
Leg(view.market.items[0].name, -3),
]
result = play(market(60_000), junk)
self.assertEqual(result.realised, 0.0)
self.assertGreaterEqual(result.dropped, 4)
def test_two_legs_on_one_name_keep_the_first(self) -> None:
m = market(60_000)
name = m.items[0].name
def duplicated(view, name=name):
if view.turn > 1:
return []
price = view.market.item(name).prices[view.seen - 1]
return [Leg(name, 10, round(price * 0.95), round(price * 1.05)),
Leg(name, 10_000, round(price * 1.50), round(price * 1.60))]
result = play(m, duplicated)
self.assertEqual(result.dropped, 1)
class ReferencePolicyTests(unittest.TestCase):
def test_the_reference_does_not_requote_every_turn(self) -> None:
"""If it did, the freeze would be a constant cost and the environment would be
measuring the one-shot answer with extra steps."""
amendments = [play(market(s), Reference()).amendments for s in SEEDS]
# Per-basket, the reference will sometimes want every look it is given — a basket
# where the purse keeps freeing up is a basket where redeploying it keeps paying.
# The claim that matters is the average, and it is that a look is declined often
# enough for declining to be a judgement rather than an accident.
self.assertLess(statistics.fmean(amendments), TURNS - 2,
"the reference re-quotes on nearly every look — the budget is a "
"formality")
self.assertGreater(sum(amendments), 0,
"the reference never re-quotes — the stepped form is decorative")
def test_the_reference_makes_money_in_most_baskets(self) -> None:
earned = [play(market(s), Reference()).realised for s in SEEDS]
won = sum(1 for e in earned if e > 0)
self.assertGreaterEqual(won, len(earned) - 2,
f"the reference profits in only {won}/{len(earned)} baskets "
"— the reward's denominator is not reliable")
def test_it_is_deterministic(self) -> None:
for seed in (60_000, 60_011):
m = market(seed)
a, b = play(m, Reference()), play(m, Reference())
self.assertEqual(a, b)
class RandomSheetTests(unittest.TestCase):
def test_random_legs_never_break_the_engine(self) -> None:
"""The hostile battery, cheaply: random garbage on every look, replayed per turn,
must never raise and must never mint a coin."""
rng = random.Random(7)
for seed in SEEDS:
m = market(seed)
def noise(view, rng=rng):
return [
Leg(rng.choice([i.name for i in view.market.items] + ["", "nope"]),
rng.choice([-5, 0, 1, 10 ** 9]),
rng.choice([0, -1, 1, 10 ** 9]),
rng.choice([0, 1, 10 ** 9]))
for _ in range(rng.randint(0, 8))
]
result = play(m, noise)
self.assertLessEqual(result.peak_employed, m.capital + 1e-6)
self.assertGreaterEqual(result.realised, -float(m.capital))
class LadderGateTests(unittest.TestCase):
"""House rules 3 and 4, executable, in this package rather than in `probe.py`.
`probe.py` is B1's file and five workstreams are queued behind it; the gate this
environment lives or dies on should not wait on a merge, and it should fail here first
if the engine drifts. When A's probe row lands it asserts the same three things over the
same helper.
"""
def setUp(self) -> None:
sys.path.insert(0, str(HERE))
import measure_ladder
self.ladder = measure_ladder
self.markets = measure_ladder.baskets(measure_ladder.BLOCK * measure_ladder.BLOCKS)
self.blocks = [self.markets[i:i + measure_ladder.BLOCK]
for i in range(0, len(self.markets), measure_ladder.BLOCK)]
def test_the_floor_is_exactly_zero_and_the_ceiling_is_reachable(self) -> None:
rows = self.ladder.ladder(self.markets)
self.assertEqual(rows["inaction"]["total"], 0.0)
self.assertEqual(rows["staller"]["total"], 0.0)
self.assertEqual(rows["oracle (the live reference)"]["total"], 1.0)
def test_the_budget_binds_in_every_block(self) -> None:
"""Rule 4. Not on the 120-basket mean — `probe.py` documents this generator swinging
0.48 against 0.18 between blocks of twenty-four, and a mean that holds while one
block of the size a run actually grades goes the other way is a mean measured off
the seed."""
gaps = []
for block in self.blocks:
rows = self.ladder.ladder(block)
gap = (rows["oracle (the live reference)"]["total"]
- rows["exhaustive (re-quote every look)"]["total"])
gaps.append(gap)
self.assertGreater(gap, 0.10,
f"baskets {block[0].seed}-{block[-1].seed}: spending every "
f"look scores {gap:.3f} below the reference — the turn "
"budget is a formality")
self.assertGreater(statistics.fmean(gaps), 0.25)
def test_the_reference_is_not_merely_the_denominator(self) -> None:
"""The gap has to be a fact about money, not about normalisation. If the exhaustive
policy EARNED more and scored less, the reward would be an imitation score for the
reference's constants — which is precisely what the first cut of this engine did,
at 63,974 gp against the reference's 51,520."""
rows = self.ladder.ladder(self.markets)
oracle = rows["oracle (the live reference)"]
for name in ("exhaustive (re-quote every look)", "churn (re-place at full size every look)",
"impatient (one-shot reference, one look)",
"restate (re-quote the same book every look)"):
self.assertLess(rows[name]["gp"], oracle["gp"],
f"{name} earns more gp than the reference it is scored against")
def test_looking_costs_and_a_wasted_look_costs_more(self) -> None:
rows = self.ladder.ladder(self.markets)
self.assertLess(rows["restate (re-quote the same book every look)"]["total"],
rows["impatient (one-shot reference, one look)"]["total"],
"re-quoting the identical book every turn is free")
self.assertLess(rows["crude (market orders, one look)"]["total"],
rows["plausible (mean anchor, no filter)"]["total"])
self.assertLess(rows["plausible (mean anchor, no filter)"]["total"],
rows["impatient (one-shot reference, one look)"]["total"])
class LoadCostTests(unittest.TestCase):
def test_a_default_taskset_load_is_not_three_thousand_simulations(self) -> None:
"""`num_tasks` defaults to 48 and viability retries; a live `viable_market` running
a full stepped reference per candidate is thousands of simulations before the first
token, on every eval AND every `--dry-run`. The reference is cached per seed and the
try cap is 16 rather than 64, and this is the assertion that says so."""
import time
sys.path.insert(0, str(HERE))
import measure_ladder
from grand_exchange_live import live
live._REFERENCE_CACHE.clear()
started = time.monotonic()
markets = measure_ladder.baskets(48)
cold = time.monotonic() - started
self.assertEqual(len(markets), 48)
self.assertLess(cold, 2.0, f"a 48-task load took {cold:.2f}s")
hits = len(live._REFERENCE_CACHE)
started = time.monotonic()
for market in markets:
live.live_reference(market)
self.assertEqual(len(live._REFERENCE_CACHE), hits, "scoring re-simulated the "
"reference the model played against")
self.assertLess(time.monotonic() - started, 0.05)
if __name__ == "__main__":
unittest.main()
@@ -0,0 +1,110 @@
"""The vendored generator is a COPY, and this is the assertion that keeps it one.
`grand_exchange_live/market.py` and `book.py` are copies of the one-shot environment's, not
imports of it. The reason is that the one-shot scores have to stay comparable across time:
an import would mean any change made for the live form silently moved every historical
`grand-exchange` number, and nothing would say so. The cost of a copy is drift, and drift is
what this file refuses to allow — every price and every volume of every item, over a block
of seeds, through both packages, compared exactly.
The freeze draw is the one intentional divergence. It is drawn off `random.Random(seed ^
FREEZE_SALT)` rather than off the generator's own `rng`, so it cannot consume a draw and
cannot perturb a single tick. That claim is not an argument here, it is the test below.
"""
from __future__ import annotations
import sys
import types
import unittest
from pathlib import Path
HERE = Path(__file__).resolve().parents[1]
ONE_SHOT = HERE.parent / "grand_exchange"
# `probe.py`'s shim, for `probe.py`'s reason: each package's `__init__` imports its taskset,
# which imports verifiers, and a self-check that only runs once the training stack is
# installed is a self-check nobody runs. Registering the package name as a namespace
# pointing at the source directory keeps intra-package imports resolving while `__init__`
# never executes.
for _root in (HERE, ONE_SHOT):
_package = _root / _root.name
_shim = types.ModuleType(_package.name)
_shim.__path__ = [str(_package)]
sys.modules[_package.name] = _shim
import grand_exchange.book as one_shot_book # noqa: E402
import grand_exchange.market as one_shot # noqa: E402
import grand_exchange_live.book as live_book # noqa: E402
import grand_exchange_live.market as copy # noqa: E402
SEEDS = range(60_000, 60_048)
SHAPES = ((5, 56, 30), (5, 56, 60), (3, 24, 12))
class GeneratorCopyTests(unittest.TestCase):
def test_every_stream_is_byte_identical(self) -> None:
for num_items, visible, held_out in SHAPES:
for seed in SEEDS:
a = one_shot.build_market(seed, num_items, visible, held_out)
b = copy.build_market(seed, num_items, visible, held_out)
self.assertEqual(a.seed, b.seed)
self.assertEqual(a.capital, b.capital)
self.assertEqual(a.visible, b.visible)
self.assertEqual(a.held_out, b.held_out)
self.assertEqual([i.name for i in a.items], [i.name for i in b.items])
for x, y in zip(a.items, b.items):
self.assertEqual(x.reverting, y.reverting)
self.assertEqual(x.buy_limit, y.buy_limit)
self.assertEqual(x.prices, y.prices, f"prices diverge on {seed}/{x.name}")
self.assertEqual(x.volumes, y.volumes, f"volumes diverge on {seed}/{x.name}")
def test_the_constants_did_not_drift(self) -> None:
"""A copy that produces the same streams from a different tax rate is still a
broken copy — the streams would match and every execution would not."""
for name in ("TAX", "FILL_SHARE", "DUMP_BASE", "DUMP_IMPACT", "DUMP_CAP",
"STARTING_CAPITAL", "WALKS", "SEED_BASE", "PHI", "FUND_DRIFT",
"VOLUME_SD", "DRIFT_STEP", "TIERS"):
self.assertEqual(getattr(one_shot, name), getattr(copy, name), name)
for name in ("MAX_ORDERS", "BUY_BAND", "SELL_BAND", "TARGET_SHARE", "MIN_CROSSINGS",
"MAX_ITEM_SHARE", "VIABILITY_TRIES"):
self.assertEqual(getattr(one_shot_book, name), getattr(live_book, name), name)
def test_the_execution_engine_agrees_order_for_order(self) -> None:
"""The one-shot engine is copied too, and `live.play` is checked against it. If the
copy of `execute` drifted, that check would be comparing the stepped engine to the
wrong fixed point and would still pass."""
for seed in SEEDS:
a = one_shot.build_market(seed, 5, 56, 60)
b = copy.build_market(seed, 5, 56, 60)
plan_a = one_shot_book.reference_orders(a)
plan_b = live_book.reference_orders(b)
self.assertEqual(
[(o.item, o.quantity, o.buy, o.sell) for o in plan_a],
[(o.item, o.quantity, o.buy, o.sell) for o in plan_b],
)
fa = one_shot_book.execute(a, plan_a)
fb = live_book.execute(b, plan_b)
self.assertEqual(fa.realised, fb.realised)
self.assertEqual(fa.committed, fb.committed)
self.assertEqual(fa.bought, fb.bought)
self.assertEqual(fa.sold, fb.sold)
self.assertEqual(fa.dumped, fb.dumped)
def test_the_freeze_is_drawn_without_touching_the_streams(self) -> None:
"""The divergence is bounded to one integer, and that integer moves."""
freezes = {copy.build_market(s, 5, 56, 60).freeze for s in SEEDS}
self.assertEqual(freezes - set(range(copy.FREEZE_RANGE[0], copy.FREEZE_RANGE[1] + 1)),
set())
self.assertGreater(len(freezes), 1, "every basket freezes for the same number of "
"ticks — the cost of a look is a free prior")
# Deterministic in the seed alone: viability and scoring must see the same value,
# which is the whole reason it is not drawn in `Taskset.load()`.
self.assertEqual(copy.build_market(60_000, 5, 56, 60).freeze,
copy.build_market(60_000, 5, 56, 60).freeze)
self.assertEqual(one_shot.build_market(60_000, 5, 56, 60).items[0].prices,
copy.build_market(60_000, 5, 56, 60).items[0].prices)
if __name__ == "__main__":
unittest.main()