The probe printed one blended number per policy, so a component pinned at 0.000 across every rung was invisible. Four gates hid there for a month. It now reports floor / best-below-oracle / oracle / ceiling for 25 components across 8 environments, and a component flat across every rung is fatal. bot-detection GATE_SLACK = 1 — fires 5/32 on the real rollouts, was 0/32. The max(1, reference_caught - SLACK) guard is verified by construction, not by sampling: without it the required count reaches 0 and an EMPTY accusation list clears the gate. Observed reference_caught is 4-6, so no amount of sampling would have found that hole. schema-migration GATE_MARGIN = 0.15 — fires 4/32, was 0/32. The cost is disclosed and bounded: the naive split clears it on 5.5% of 1,000 unseen seeds, fenced by an assert at 10%. Margin 0.10 keeps the leak at zero and fires 0/32, i.e. stays dead. A live gradient with a bounded leak beats a clean corpse. redaction-pressure is NOT given a margin, and that is the result rather than a failure. The only setting that fires at all leaves half the secrets standing and pays a four-of-seven ruleset on five seeds in six — a margin that pays for inaction is strictly worse than a dead gate. Recall maxes at 0.852 and no rollout ever cleared both clauses in one episode. It is genuinely hard, not miscalibrated. ⚠️ The per-component check did not catch the defect it was built for. Reverting schema-migration's margin to 0.0 — restoring the exact dead gate — printed ok and exited 0, because the near-oracle rung scrapes the unmargined gate on ~2 seeds in 24 and that kept best<oracle non-zero. Every assertion bounded how much a margin may PAY; none noticed if it stopped existing. migration() now carries the mirror of bot-detection's guard, and reverting the margin fails with "the margin is dead and the component carries no gradient between the crude answer and the exact one". canary-trap's oracle-minus-one rung is documented as degenerate rather than quietly relied on: it is identical to the oracle to four decimals, so it measures specificity and gate at the ceiling, not mid-ladder as its comment claimed. The CI lock policy asks git instead of the disk. It was checking the working tree, where a lock file is a normal by-product of uv sync, so it passed in a clean checkout and failed on every machine that had run an eval. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
365 lines
17 KiB
Markdown
365 lines
17 KiB
Markdown
# Why `gate` scored zero in four environments
|
||
|
||
Written 2026-08-21, against `outputs/run-20260821-1401` — seven environments, 32 rollouts
|
||
each, `brain-qwen38-dspark`.
|
||
|
||
In four of the seven, the `gate` reward was **0.000 mean and 0.000 max** across all 32
|
||
rollouts: `bot-detection` (weight 0.25), `grand-exchange` (0.25), `redaction-pressure`
|
||
(0.30) and `schema-migration` (0.25). It moved in the other three — `canary-trap` 0.2188,
|
||
`drop-table-inference` 0.2258, `fault-localisation` 0.9062.
|
||
|
||
`probe.py` prints 1.000 for every one of those four, and 1.000 requires `gate = 1`, so
|
||
nothing here is unreachable and house rule 3 is not violated. The question this file
|
||
answers is the one after that: **is the gate (a) a threshold set beyond what a real model
|
||
reaches, (b) a parse or format precondition the model never satisfies, or (c) genuinely
|
||
hard?** They have different consequences and they need different fixes.
|
||
|
||
---
|
||
|
||
## The answer, in one table
|
||
|
||
`outputs/gate_probe.py` runs each environment's shipped scorer over the probe ladder and
|
||
prints the gate **on its own**, as a pass rate over 24 tasks (12 for grand-exchange).
|
||
`probe.py` blends the gate into one number per rung, which is exactly why this was
|
||
invisible for a month.
|
||
|
||
```
|
||
#### redaction-pressure
|
||
inaction gate 0/24 = 0.000
|
||
crude gate 0/24 = 0.000
|
||
plausible gate 0/24 = 0.000
|
||
six-of-seven gate 0/24 = 0.000
|
||
oracle gate 24/24 = 1.000
|
||
|
||
#### schema-migration
|
||
inaction gate 0/24 = 0.000
|
||
crude gate 0/24 = 0.000
|
||
plausible gate 0/24 = 0.000
|
||
oracle gate 24/24 = 1.000
|
||
|
||
#### bot-detection
|
||
inaction gate 0/24 = 0.000
|
||
crude gate 0/24 = 0.000
|
||
oracle-minus-one gate 0/24 = 0.000
|
||
oracle gate 24/24 = 1.000
|
||
|
||
#### grand-exchange
|
||
inaction gate 0/12 = 0.000
|
||
oracle gate 12/12 = 1.000
|
||
```
|
||
|
||
**In three of the four the gate is a step function whose only step is the oracle.** Six of
|
||
redaction's seven rules scores zero. The oracle bot list minus a single account scores
|
||
zero. There is no rung anywhere on the ladder — inaction, crude, plausible, near-oracle —
|
||
that earns a fraction of it.
|
||
|
||
The three that fired are not perfection-free; they are perfection over a *small* object.
|
||
`canary-trap`'s gate wants every contaminated state caught and no probe firing on the clean
|
||
model, over 2–12 probes. `fault-localisation`'s wants service, fault and evidence all
|
||
correct — three fields. `drop-table-inference` is the only gate in the repository with a
|
||
**margin**: `GATE_MARGIN = 0.05`, three sub-scores each within 5% of the reference, and it
|
||
fired on 22.6% of rollouts. `grand-exchange` has one too — `TARGET_SHARE = 0.90` — and its
|
||
zero has a different cause entirely (below).
|
||
|
||
So the verdict is **(a) for three environments and neither for the fourth**. It is not (b)
|
||
anywhere: parsing works. `bot-detection` recorded `unknown_ids` 0.000 on all 32 rollouts,
|
||
`redaction-pressure` `validity` 1.0 on 30 of 32, `schema-migration` produced a valid target
|
||
schema on 25 of 32. The models are submitting well-formed work and being scored zero on it
|
||
because the work is not perfect.
|
||
|
||
---
|
||
|
||
## Per environment
|
||
|
||
### `bot-detection` — (a), and the binding clause is `detection == 1.0`
|
||
|
||
```python
|
||
# scan.py
|
||
@property
|
||
def gate(self) -> bool:
|
||
return (
|
||
self.reference_caught > 0
|
||
and self.true_positives >= self.reference_caught
|
||
and self.false_positives == 0
|
||
)
|
||
```
|
||
|
||
Over 32 rollouts:
|
||
|
||
| clause | satisfied |
|
||
|---|---|
|
||
| `false_positives == 0` | **24 / 32** |
|
||
| `true_positives >= reference_caught` | **0 / 32** |
|
||
|
||
`detection` — true positives over what the reference caught — had mean 0.5135 and **max
|
||
0.8333**. It never reached 1.0 once. The restraint clause is doing its job and is not the
|
||
problem; the catch-everything clause is a ceiling test.
|
||
|
||
The environment's own docstring explains why it has to be: each cloaked account leaks on
|
||
exactly one of three behavioural channels, so "a discriminator built on any single channel
|
||
ceilings at roughly a third of them". The reference reads all three. The gate therefore
|
||
pays only a model that found all three tells, in a batch where the reference caught 4–6
|
||
bots and the model accused 2.9 on average.
|
||
|
||
**Consequence.** 0.25 of bot-detection's reward mass is a constant. The measured 0.3488 is
|
||
entirely `caught` (0.35) and `spared` (0.40), and against a maximum of 0.75 it is 0.465 of
|
||
what was actually on offer.
|
||
|
||
### `redaction-pressure` — (a), and the binding clause is total recall
|
||
|
||
```python
|
||
# scan.py
|
||
@property
|
||
def clean(self) -> bool:
|
||
return (
|
||
self.removed_secrets == self.secrets_total
|
||
and self.partial_secrets == 0
|
||
and self.collateral_hits == 0
|
||
and self.innocent_lost == 0
|
||
and self.secrets_total > 0
|
||
and self.validity == 1.0
|
||
)
|
||
```
|
||
|
||
Six clauses, and they are satisfied at wildly different rates:
|
||
|
||
| clause | satisfied |
|
||
|---|---|
|
||
| `partial_secrets == 0` | 28 / 32 |
|
||
| `validity == 1.0` | 30 / 32 |
|
||
| `collateral_hits == 0` | 6 / 32 |
|
||
| `innocent_lost == 0` | 6 / 32 |
|
||
| `residual_hits == 0` (every secret removed) | **0 / 32** |
|
||
|
||
`recall` had max **0.852** across 32 rollouts and `residual_hits` a **minimum of 4** — there
|
||
was always at least one secret left standing. Two clauses never fired together either:
|
||
`innocent_lost` averaged 140 characters.
|
||
|
||
This is the widest gate in the repository — six simultaneous conditions, three of them
|
||
counting to zero over a twelve-document held-out slice — and the ladder shows six of the
|
||
oracle's seven rules do not clear it.
|
||
|
||
**Consequence.** 0.30 of the mass — the largest gate weight of the four — is a constant.
|
||
0.4040 against a real maximum of 0.70 is 0.577 of what was on offer.
|
||
|
||
### `schema-migration` — (a), and the binding clause is per-row fidelity
|
||
|
||
```python
|
||
# run.py
|
||
@property
|
||
def clean(self) -> bool:
|
||
return (self.schema_ok and self.graded > 0
|
||
and self.matched == self.graded
|
||
and self.rows_after == self.rows_before)
|
||
```
|
||
|
||
| clause | satisfied |
|
||
|---|---|
|
||
| `schema_ok` | 25 / 32 |
|
||
| `rows_after == rows_before` | 25 / 32 |
|
||
| `matched == graded` (fidelity 1.0) | **0 / 32** |
|
||
|
||
`fidelity` distribution over the 32 rollouts:
|
||
|
||
```
|
||
0.0 ×10, 0.10, 0.15, 0.375, 0.40, 0.575, 0.65, 0.675, 0.725, 0.725, 0.75,
|
||
0.775 ×4, 0.80 ×4, 0.85 ×3, 0.875
|
||
```
|
||
|
||
Max 0.875. Forty awkward held-out rows must every one recompose; the best migration
|
||
anybody wrote missed five of them. The docstring is explicit that this is `forge verify`'s
|
||
exit code — a migration is correct or it is not — and as a shipping check that is right.
|
||
As a *reward component* it is a constant at this capability level.
|
||
|
||
### `grand-exchange` — not a gate defect. Nothing was measured
|
||
|
||
```python
|
||
# book.py — a bar, not a knife-edge
|
||
return (self.fills.realised >= TARGET_SHARE * self.reference.realised
|
||
and self.fills.roc >= TARGET_SHARE * self.reference.roc)
|
||
```
|
||
|
||
`TARGET_SHARE = 0.90`. This gate already has the margin the other three lack, and it is
|
||
written that way deliberately — the commit message is in the docstring ("demanding the
|
||
reference's exact result made a wasted unit of capital cost 0.146 with byte-identical
|
||
profit").
|
||
|
||
It scored zero because **31 of 32 rollouts placed no orders at all**: 18 completion tokens,
|
||
`{"expected_profit": 0, "orders": []}`, `finish_reason=stop`. `planned` had mean 65.6 and
|
||
median 0 — one rollout in 32 planned anything at all; `profit_ratio` maxed at 0.323, so even
|
||
that one reached 29% of the reference. Nothing in this environment was measured. See
|
||
`FIRST_EVAL.md` §5 — spark-1 serves with
|
||
`default_chat_template_kwargs = {"enable_thinking": false}`, and this environment is the one
|
||
that cannot be answered without arithmetic.
|
||
|
||
**Re-measured with thinking on, this gate fires.** `outputs/thinking/ge-think-n6/`, 6
|
||
rollouts: one of them finished inside its token budget and scored `profit` 1.000,
|
||
`discipline` 1.000 and **`gate` 1.000**, realising 40,081 gp against a reference of 36,157.
|
||
The other five were cut off mid-reasoning at `max_tokens`. So this gate is not merely
|
||
oracle-reachable in a probe — it is reachable by the model under evaluation, in a real
|
||
rollout, at a 90% bar.
|
||
|
||
**Do not touch this gate.** It is the one that works. It is also the shape the other three
|
||
should be argued against: a band at 0.90 of the reference, cleared by a run that beat the
|
||
reference by its own route rather than by copying it.
|
||
|
||
---
|
||
|
||
## What follows
|
||
|
||
1. **Every ranking over those four environments is a ranking over their soft components.**
|
||
Report `bot-detection` against 0.75, `redaction-pressure` against 0.70,
|
||
`schema-migration` against 0.75 — or say plainly that a quarter to a third of the scale
|
||
is unreachable. A leaderboard that prints 0.4040 out of 1.000 for redaction-pressure is
|
||
overstating the headroom by a factor it never names.
|
||
|
||
2. **The gradient problem is worse than the scale problem.** A constant term contributes
|
||
nothing to a policy gradient. Training against these four trains against
|
||
`caught`/`spared`, `recall`/`precision`, `schema`/`integrity` — and the gate only starts
|
||
paying once the model is already essentially perfect, which is the point at which it
|
||
stops needing the signal. This is not what a gate is for.
|
||
|
||
3. **The repository already contains the fix and applies it in exactly one place.**
|
||
`drop-table-inference` gates on three sub-scores each within `GATE_MARGIN = 0.05` of the
|
||
reference and fires 22.6% of the time; `grand-exchange` gates at `TARGET_SHARE = 0.90`.
|
||
The three dead gates are the three written as exact equality. A margin on each — every
|
||
secret bar one, 39 rows of 40, the reference's bots less one — would restore a gradient
|
||
without paying for inaction, because the ladder above shows `inaction` and `crude` score
|
||
zero on every one of these with or without a margin.
|
||
|
||
4. **Whatever is decided, `probe.py` should print the gate separately.** Four dead
|
||
components hid for a month behind a blended `oracle 1.000`. `outputs/gate_probe.py` is
|
||
the throwaway version of that check; the real one belongs next to the ladder, and it is
|
||
a one-line consequence of house rule 3 that a component nothing below the oracle can
|
||
earn is a component that does not discriminate.
|
||
|
||
⚠️ **Superseded by the addendum below.** As first written this file changed nothing in any
|
||
environment — four packages' reward code was out of that lane, and it was a diagnosis and a
|
||
recommendation rather than a patch. Two of the four have since been patched and one has been
|
||
measured and deliberately left alone. Read the addendum before acting on §3 above: its
|
||
recommendation "a margin on each" turned out to be right about two of the three and wrong
|
||
about `redaction-pressure`.
|
||
|
||
---
|
||
|
||
## Addendum, same day: what was actually changed
|
||
|
||
⚠️ The line above — "Nothing in this file has been changed in any environment" — no longer
|
||
holds. Two of the four gates were given margins and one was deliberately not. Everything
|
||
below was measured, not argued, and the measurements are reproducible from the repository
|
||
without touching a model:
|
||
|
||
```bash
|
||
uv run --with regex python tools/regate.py # replay the 32 real rollouts per env
|
||
uv run --with regex python tools/regate_ladder.py # the same predicates over the ladder
|
||
uv run --with regex python probe.py # floors and ceilings, per component
|
||
```
|
||
|
||
`tools/regate.py` rebuilds each task from the seed in its trace, re-runs the shipped scorer
|
||
over the model's own reply, and asks a candidate predicate directly. That is a better
|
||
measurement than a fresh sample would have been: it holds the model's behaviour fixed and
|
||
varies only the reward, and it costs nothing on spark-1. Its `recorded_mean` reproduces
|
||
0.3488, 0.4432 and 0.4040 exactly, which is the cross-check that it is replaying the same
|
||
population `eval.log` scored.
|
||
|
||
### `bot-detection` — fixed. `GATE_SLACK = 1`
|
||
|
||
The reference's bots bar one, and still not one person.
|
||
|
||
| gate | fires on 32 real rollouts |
|
||
|---|---|
|
||
| exact (as shipped) | 0/32 = 0.000 |
|
||
| **slack 1** | **5/32 = 0.156** |
|
||
| slack 2 | 15/32 = 0.469 |
|
||
| share 0.90 of the reference | 0/32 = 0.000 |
|
||
|
||
Slack 2 is what the ladder rules out. At slack 1 the oracle's list minus one account clears
|
||
the gate and minus two does not, so it still separates a near miss from a half-right answer;
|
||
at slack 2 both clear and it separates nothing. `inaction` and `crude` score 0.000 at every
|
||
slack, and `probe.py` asserts all three of those facts rather than reporting them.
|
||
|
||
`max(1, reference_caught - GATE_SLACK)` is load-bearing: without it a batch where the
|
||
reference catches exactly one bot would let an **empty** accusation list clear the gate.
|
||
Observed `reference_caught` is 4–6, so nothing would have caught that in a sample.
|
||
|
||
The restraint clause was not given a margin and should not be. A single wrongful ban still
|
||
shuts the gate outright, because that asymmetry is the environment.
|
||
|
||
### `schema-migration` — fixed. `GATE_MARGIN = 0.15`
|
||
|
||
Thirty-four of forty rows recompose. `schema_ok` and the row count stay exact.
|
||
|
||
| gate | fires on 32 real rollouts | `plausible` on the ladder |
|
||
|---|---|---|
|
||
| exact (as shipped) | 0/32 = 0.000 | 0.000 |
|
||
| margin 0.05 | 0/32 = 0.000 | 0.000 |
|
||
| margin 0.10 | 0/32 = 0.000 | 0.000 |
|
||
| **margin 0.15** | **4/32 = 0.125** | 0.042 |
|
||
| margin 0.20 | 8/32 = 0.250 | 0.292 |
|
||
|
||
0.20 is out: it pays the naive split — the parser fitted to the five tidy rows on screen,
|
||
which is precisely what this environment exists to punish — on 29% of seeds. 0.15 pays it on
|
||
one seed in twenty-four, because its luckiest draw of forty awkward rows recomposes exactly
|
||
34 of them and 34/40 is the threshold to the digit. That knife edge is named in the source
|
||
and `probe.py` asserts it stays under a tenth.
|
||
|
||
What 0.15 buys is the rung above it. A migration that strips the thousands separator — the
|
||
inference this environment is about — and still mishandles the row with no unit at all goes
|
||
from clearing the gate on 8% of seeds to clearing it on 96%.
|
||
|
||
The two format clauses keep no margin. `wipe` — the oracle followed by `DELETE FROM
|
||
readings`, which recomposes perfectly over nothing — is now a probe rung asserted at 0.000,
|
||
so the hole the row clause was added to close cannot quietly reopen.
|
||
|
||
### `redaction-pressure` — NOT fixed, and that is the finding
|
||
|
||
| gate | fires on 32 real rollouts | `plausible` (4 of 7 rules) |
|
||
|---|---|---|
|
||
| exact (as shipped) | 0/32 = 0.000 | 0.000 |
|
||
| recall margin 0.10, collateral 0, innocent 0 | 0/32 = 0.000 | 0.000 |
|
||
| recall margin 0.20, collateral 1, innocent 32 | 0/32 = 0.000 | 0.000 |
|
||
| recall margin 0.50, collateral 2, innocent 160 | 3/32 = 0.094 | **0.833** |
|
||
|
||
There is no setting that both fires on the measured population and keeps `plausible` at
|
||
zero. Everything loose enough to fire pays a ruleset that never found three of the seven
|
||
secret types, on five seeds in six; everything tight enough to hold that at zero fires on
|
||
none of the 32.
|
||
|
||
The reason is that the failure is **joint**, not a threshold. Recall maxed at 0.852 with a
|
||
minimum of four secrets left standing; `collateral_hits == 0` held on 6 of 32; and no
|
||
rollout managed high recall and zero collateral in the same episode. So the verdict for this
|
||
one is **(c) genuinely hard**, not (a) mis-thresholded — which is a different answer from
|
||
the one this file gave in §3 above, and it is the answer the measurement gives.
|
||
|
||
Half the secrets left in the document is not a near miss. The fix, when it comes, is a
|
||
better model or an easier corpus, not a looser gate. The sweep is recorded in
|
||
`Outcome.clean`'s docstring so it is not redone from scratch.
|
||
|
||
### `grand-exchange` — untouched, as §3 said
|
||
|
||
Its 0.0055 measured a sampling config and not a gate. Left alone.
|
||
|
||
## And the blindness that let all four hide
|
||
|
||
`probe.py` returned one blended float per rung. It now returns the **weighted components**,
|
||
prints each one's floor, its best rung below the oracle and its ceiling, and draws two lines:
|
||
|
||
- **flat** — the same value on every rung including the oracle — is fatal. That is a constant
|
||
added to every policy's score and there is no legitimate reward of that shape.
|
||
- **step@oracle** — nothing below the oracle earns a fraction — warns and names the
|
||
component.
|
||
|
||
The second is deliberately not fatal, and the reason is the whole difficulty of this file: it
|
||
is the shape of a dead gate AND the shape of an honest binary check, and the two are not
|
||
distinguishable from a ladder. Making it fatal would forbid every binary gate; making it
|
||
silent is what cost a month. So it is an instruction — go and replay the traces — and
|
||
`tests/test_probe.py::STEP_AT_ORACLE` is the allowlist that keeps the instruction from
|
||
becoming wallpaper. Two names are on it, `redaction-pressure/gate` (0/32, hard) and
|
||
`fault-localisation/gate` (29/32, healthy — the ladder just has no two-of-three rung). It
|
||
only ever shrinks, and a name may only leave it alongside the measurement that justifies it.
|
||
|
||
Four environments also gained a `near-oracle` rung, which is what makes the check say
|
||
anything: a gate carrying a real band now demonstrates it — `grand-exchange`'s 0.90 bar reads
|
||
0.244 of 0.250 one unit below the reference, and `bot-detection`'s new slack reads the full
|
||
0.250 at the oracle's list minus one.
|