gates: give two of them a margin, and make the probe see components

The probe printed one blended number per policy, so a component pinned at 0.000
across every rung was invisible. Four gates hid there for a month. It now reports
floor / best-below-oracle / oracle / ceiling for 25 components across 8
environments, and a component flat across every rung is fatal.

bot-detection GATE_SLACK = 1 — fires 5/32 on the real rollouts, was 0/32. The
max(1, reference_caught - SLACK) guard is verified by construction, not by
sampling: without it the required count reaches 0 and an EMPTY accusation list
clears the gate. Observed reference_caught is 4-6, so no amount of sampling would
have found that hole.

schema-migration GATE_MARGIN = 0.15 — fires 4/32, was 0/32. The cost is disclosed
and bounded: the naive split clears it on 5.5% of 1,000 unseen seeds, fenced by
an assert at 10%. Margin 0.10 keeps the leak at zero and fires 0/32, i.e. stays
dead. A live gradient with a bounded leak beats a clean corpse.

redaction-pressure is NOT given a margin, and that is the result rather than a
failure. The only setting that fires at all leaves half the secrets standing and
pays a four-of-seven ruleset on five seeds in six — a margin that pays for
inaction is strictly worse than a dead gate. Recall maxes at 0.852 and no rollout
ever cleared both clauses in one episode. It is genuinely hard, not miscalibrated.

⚠️ The per-component check did not catch the defect it was built for. Reverting
schema-migration's margin to 0.0 — restoring the exact dead gate — printed ok and
exited 0, because the near-oracle rung scrapes the unmargined gate on ~2 seeds in
24 and that kept best<oracle non-zero. Every assertion bounded how much a margin
may PAY; none noticed if it stopped existing. migration() now carries the mirror
of bot-detection's guard, and reverting the margin fails with "the margin is dead
and the component carries no gradient between the crude answer and the exact one".

canary-trap's oracle-minus-one rung is documented as degenerate rather than
quietly relied on: it is identical to the oracle to four decimals, so it measures
specificity and gate at the ceiling, not mid-ladder as its comment claimed.

The CI lock policy asks git instead of the disk. It was checking the working
tree, where a lock file is a normal by-product of uv sync, so it passed in a clean
checkout and failed on every machine that had run an eval.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-21 23:14:05 -07:00
co-authored by Claude Opus 5
parent 122a858b61
commit 3305be4ff7
12 changed files with 1107 additions and 95 deletions
@@ -24,8 +24,11 @@ Four quantities come out of one pass:
purity true positives over the accusations made, floored at the reference's count.
The floor is what stops one confident accusation scoring like a full sweep:
without it, naming the single most obvious script earns perfect precision.
gate every bot the reference found, and not one person. Binary; the only thing a
moderation team would actually ship.
gate the reference's bots bar one, and not one person. Binary. It demanded ALL of
them until 2026-08-21, and measured over 32 rollouts it fired zero times —
0.25 of the reward mass with no gradient in it at all. The slack is on the
catching clause only; a single wrongful ban still shuts the gate, because
that is the asymmetry the environment exists to teach. See `GATE_SLACK`.
"""
from __future__ import annotations
@@ -47,6 +50,25 @@ MAX_ACCUSATIONS = 64
CORRECTION_FLOOR = 0.35
LOGIN_SPREAD_FLOOR_MIN = 30
# How many of the reference's bots the gate forgives. It used to forgive none, and over the
# 32 rollouts in `outputs/run-20260821-1401` it fired 0 times with a maximum of 0.000: a
# quarter of this environment's reward mass was a constant, and a constant contributes
# nothing to a policy gradient. `detection` never reached 1.0 in 32 attempts — its maximum
# was 0.833 — so the gate only ever paid a model that was already perfect, which is the
# point at which it has stopped needing the signal.
#
# One is measured, not chosen. Replaying those same 32 rollouts through this scorer
# (`tools/regate.py`): exact 0/32, slack 1 **5/32 = 0.156**, slack 2 15/32 = 0.469. The
# ladder is what rules slack 2 out (`tools/regate_ladder.py`): at slack 1 the oracle's list
# minus one account clears the gate and minus two does not, so the gate still separates a
# near miss from a half-right answer; at slack 2 both clear and it stops separating anything.
# Inaction and accusing everybody score 0.000 at every slack.
#
# The restraint clause is deliberately NOT given a margin. A single wrongful ban still shuts
# the gate outright, because that asymmetry is the whole environment — a missed bot is
# another day of farming, a banned player is a player who does not come back.
GATE_SLACK = 1
def correction_share(account: Account) -> float:
if not account.misclick_next:
@@ -162,11 +184,16 @@ class Outcome:
@property
def gate(self) -> bool:
return (
self.reference_caught > 0
and self.true_positives >= self.reference_caught
and self.false_positives == 0
)
"""The reference's bots bar `GATE_SLACK`, and not one person.
`max(1, ...)` is load-bearing and not defensive dressing. Without it, a batch in
which the reference catches exactly one bot would let an EMPTY accusation list
clear the gate — house rule 3 failing at the one seed nobody thinks to sample.
"""
if self.reference_caught <= 0:
return False
needed = max(1, self.reference_caught - GATE_SLACK)
return self.true_positives >= needed and self.false_positives == 0
def measure(accounts: list[Account], accused: list[str]) -> Outcome: