The first environment shipped through .claude/workflows/new-environment.js:
specification, three adversarial reviews (all 'fixable', none fatal), the
Python environment, the TypeScript port, captured rollouts, and the demo page.
Eleven agents, no errors.
The proof that the platform scales is one line long. Alert Triage has a
completely different shape from Word Five — JSON actions, priced lookups, an
analyst screen instead of a grid — and the only change under
src/components/demo/ is a comment edit, because the isolation lint refused the
word "wordle" there. Zero shell code changed. 415 contract checks now pass
against two demos, up from 206 against one.
The environment is honest by construction. Every alert is synthetic, generated
from the seed, and the banner saying so sits inside the board surface. Two of
the eleven scenario templates are hidden-suspicious: generated by the same code
as their benign twin with the signal overlaid only in lookup data, so the free
screen is identically distributed and a screen-only policy STRUCTURALLY cannot
tell them apart. The probe ladder measures it: `fast` catches 0.0 of hidden
seeds. That is the counterweight made real rather than asserted.
Twelve policies, thirteen ladder assertions, a genuine three-way trade:
fast 0.846 wins hours (0.85), misses every hidden case
targeted 0.894 wins the shipped total
thorough 0.820 wins evidence (1.00), spends 2.9 hours
None dominates. 92 Python tests, 35 TypeScript tests, 65 fixtures replaying at
delta 0, and conformance gated on world + scorer + protocol so the browser shows
the same alert for ?seed= that Python generated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019mt6sHQHEnEYrJZvoMCJSB
32 recorded rollouts, 8 seeds x 4 agents, every one replaying through the Python
engine to a delta of exactly 0.
arm solved economy consistency total
base-off 0/8 0.000 0.122 0.0244
base-on 4/8 0.469 0.479 0.4865
cautious 8/8 0.944 1.000 0.9831
solver 8/8 1.000 0.719 0.9437
Thinking-off solves none of eight. The same model on the same seeds, sampled
with thinking on, solves four. That is the headline, it is ours, and it needed
no training — which is also why the run is labelled an intervention and names
what was done to it, so a sampling change can never read as a training result.
`cautious` is new and it exists to make the reward editor honest. Until now the
page invited you to move a slider and watch the ranking change, and no slider
changed anything, because the solver dominated a model that solved nothing. A
candidate-only player — most informative guess among words that could still win
— takes consistency outright and pays for it in turns. Now:
only speed matters solver 0.9859 cautious 0.9606 -> solver
shipped weights solver 0.9437 cautious 0.9831 -> cautious
punish contradictions solver 0.8312 cautious 0.9972 -> cautious
The ranking really does flip, on recorded data, with one slider. The first
version of this arm picked the alphabetically-first candidate and opened on
'abaci', which made the policy the counterweight exists to reward look like a
straw man.
190 contract checks pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019mt6sHQHEnEYrJZvoMCJSB
The browser engine is a port of the Python one and CI proves it: all 21.2M
(guess, answer) pairs hashed on both sides to the same SHA-256. Six TS tests,
including the duplicate-letter table and the twelve pinned seed vectors that
keep ?seed= permalinks pointing at the same word the recording used.
Word lists are split by how they are used. answers.json is inlined because the
board needs it before first paint to turn a seed into a word, and a fetch there
means a visibly empty board on a cold cache. guesses.json is fetched, because it
is three times larger and only needed the first time somebody presses Enter;
until it lands, validation falls back to the answer list, which accepts strictly
fewer words. The failure mode is 'your real word was briefly rejected', not 'a
non-word was accepted' — the right way round.
The solver runs in a worker constructed from a same-origin module URL, never
Vite's ?worker&inline: that yields a blob:, and production CSP has no
worker-src, so it falls back to default-src 'self' and the worker is blocked
with no console error. It would fail in production only.
deploy.sh smoke-tests the real public hostname from the deploying machine and
fails on a body under 1 kB, because the bind bug's signature is a valid
certificate over an empty 200 and a local --resolve check passes anyway.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019mt6sHQHEnEYrJZvoMCJSB