Commit Graph

2 Commits

Author SHA1 Message Date
karti-ai 331b46b114 Make the page playable: interactive board, live solver, and a usable run picker
Three things the build was quietly missing.

The interactive board was never mounted. The contract has `interactive.init` and
`interactive.Controls`, the demo implemented both, and the shell's `split-play`
beat rendered only the replay — so the beat titled "you and the model get the
same word" showed one board. PlayYourself now renders the visitor's attempt from
the same seed as the run beside it, generically: it knows only the contract, so
any demo shipping an interactive mode gets it and one that does not renders
nothing rather than an empty pane.

solver.worker.ts was dead code — nothing constructed it, which is how CI caught
it: `new Worker(` appeared nowhere in the bundle. It is wired now behind "what
would the best player guess?", and it answers in 92ms from a real worker on
boards no recording covers. That is the difference between a demo and a video.
It also surfaces the moment the solver picks a word that CANNOT win, which is
the counterweight visible in one line instead of explained in a paragraph.

The CI check that found it was itself wrong: it grepped every bundled file for
`blob:`, which React's own code contains in a scheme check, so it failed on a
risk that was not present. It now greps for worker construction from a blob,
which is the thing production CSP actually blocks in silence.

And the run switcher was thirty buttons carrying four distinct labels. Split
into arm and seed, holding the seed across an arm change — comparing two agents
means comparing them on the same hidden word, and silently jumping seeds would
break that while looking fine.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019mt6sHQHEnEYrJZvoMCJSB
2026-08-28 16:40:12 -07:00
karti-ai 408ce4a525 Capture harness, fixture verification, CI, and the public README
The site does no live inference. Rollouts are captured once against spark-1 and
replayed at their recorded wall-clock — a public demo with no auth cannot hold
an API key, and a recorded run can be scrubbed, permalinked, blind-compared and
verified in ways a live one cannot. What stops it being a video is that the
browser re-derives every number from the recorded moves.

verify_fixtures.py is the Python half of that: it replays every committed
fixture through the engine and reproduces its own rewards. All 16 land at
delta 0.0. A fixture that cannot be regenerated is a claim with no receipt.

First real measurement, thinking off, 8 seeds: solved 0/8. The model repeats
guesses it has already played, invents words (trape, slith, postt, boomy),
and contradicts its own feedback — consistency 0.09 to 0.17. That is the
published failure taxonomy showing up in our own data on the first run, and it
is why `consistency` is a reward component rather than a footnote.

A capture failure is recorded as a turn with a null reply, never dropped. A
capture that silently discarded failed turns would be reporting a better model
than the one that ran.

CI gates both halves and four things that fail silently in production: the word
lists must rebuild byte-identically, the prerendered routes must carry their own
baked og tags (crawlers do not run JS, so without them every shared link
previews as the homepage), no blob: URL may reach the bundle (the site's CSP has
no worker-src, so it falls back to default-src 'self' and a blob worker is
blocked with no error), and the conformance digest must match across languages.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019mt6sHQHEnEYrJZvoMCJSB
2026-08-28 15:47:31 -07:00