gates: give two of them a margin, and make the probe see components
The probe printed one blended number per policy, so a component pinned at 0.000 across every rung was invisible. Four gates hid there for a month. It now reports floor / best-below-oracle / oracle / ceiling for 25 components across 8 environments, and a component flat across every rung is fatal. bot-detection GATE_SLACK = 1 — fires 5/32 on the real rollouts, was 0/32. The max(1, reference_caught - SLACK) guard is verified by construction, not by sampling: without it the required count reaches 0 and an EMPTY accusation list clears the gate. Observed reference_caught is 4-6, so no amount of sampling would have found that hole. schema-migration GATE_MARGIN = 0.15 — fires 4/32, was 0/32. The cost is disclosed and bounded: the naive split clears it on 5.5% of 1,000 unseen seeds, fenced by an assert at 10%. Margin 0.10 keeps the leak at zero and fires 0/32, i.e. stays dead. A live gradient with a bounded leak beats a clean corpse. redaction-pressure is NOT given a margin, and that is the result rather than a failure. The only setting that fires at all leaves half the secrets standing and pays a four-of-seven ruleset on five seeds in six — a margin that pays for inaction is strictly worse than a dead gate. Recall maxes at 0.852 and no rollout ever cleared both clauses in one episode. It is genuinely hard, not miscalibrated. ⚠️ The per-component check did not catch the defect it was built for. Reverting schema-migration's margin to 0.0 — restoring the exact dead gate — printed ok and exited 0, because the near-oracle rung scrapes the unmargined gate on ~2 seeds in 24 and that kept best<oracle non-zero. Every assertion bounded how much a margin may PAY; none noticed if it stopped existing. migration() now carries the mirror of bot-detection's guard, and reverting the margin fails with "the margin is dead and the component carries no gradient between the crude answer and the exact one". canary-trap's oracle-minus-one rung is documented as degenerate rather than quietly relied on: it is identical to the oracle to four decimals, so it measures specificity and gate at the ceiling, not mid-ladder as its comment claimed. The CI lock policy asks git instead of the disk. It was checking the working tree, where a lock file is a normal by-product of uv sync, so it passed in a clean checkout and failed on every machine that had run an eval. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+9
-1
@@ -1,4 +1,12 @@
|
||||
"""What the `gate` component pays the probe ladder, rung by rung.
|
||||
"""What the `gate` component pays the probe ladder, rung by rung. SUPERSEDED — see below.
|
||||
|
||||
⚠️ `probe.py` does this properly now, for every component of every environment rather than
|
||||
the gate of four, and it exits 1 on a component that never moves. Run that instead. This
|
||||
file is kept only because `docs/GATE_DIAGNOSIS.md` quotes its output verbatim as the
|
||||
evidence for the diagnosis, and a quoted table whose generator is gone is not evidence. To
|
||||
sweep a CANDIDATE gate rather than the shipped one, use `tools/regate_ladder.py`, and to
|
||||
measure one against real rollouts rather than against the ladder, `tools/regate.py`.
|
||||
|
||||
|
||||
`probe.py` reports one blended number per rung, so a gate that only ever fires for
|
||||
the oracle is invisible in its table. This prints the gate on its own — the share of
|
||||
|
||||
Reference in New Issue
Block a user