The probe printed one blended number per policy, so a component pinned at 0.000 across every rung was invisible. Four gates hid there for a month. It now reports floor / best-below-oracle / oracle / ceiling for 25 components across 8 environments, and a component flat across every rung is fatal. bot-detection GATE_SLACK = 1 — fires 5/32 on the real rollouts, was 0/32. The max(1, reference_caught - SLACK) guard is verified by construction, not by sampling: without it the required count reaches 0 and an EMPTY accusation list clears the gate. Observed reference_caught is 4-6, so no amount of sampling would have found that hole. schema-migration GATE_MARGIN = 0.15 — fires 4/32, was 0/32. The cost is disclosed and bounded: the naive split clears it on 5.5% of 1,000 unseen seeds, fenced by an assert at 10%. Margin 0.10 keeps the leak at zero and fires 0/32, i.e. stays dead. A live gradient with a bounded leak beats a clean corpse. redaction-pressure is NOT given a margin, and that is the result rather than a failure. The only setting that fires at all leaves half the secrets standing and pays a four-of-seven ruleset on five seeds in six — a margin that pays for inaction is strictly worse than a dead gate. Recall maxes at 0.852 and no rollout ever cleared both clauses in one episode. It is genuinely hard, not miscalibrated. ⚠️ The per-component check did not catch the defect it was built for. Reverting schema-migration's margin to 0.0 — restoring the exact dead gate — printed ok and exited 0, because the near-oracle rung scrapes the unmargined gate on ~2 seeds in 24 and that kept best<oracle non-zero. Every assertion bounded how much a margin may PAY; none noticed if it stopped existing. migration() now carries the mirror of bot-detection's guard, and reverting the margin fails with "the margin is dead and the component carries no gradient between the crude answer and the exact one". canary-trap's oracle-minus-one rung is documented as degenerate rather than quietly relied on: it is identical to the oracle to four decimals, so it measures specificity and gate at the ceiling, not mid-ladder as its comment claimed. The CI lock policy asks git instead of the disk. It was checking the working tree, where a lock file is a normal by-product of uv sync, so it passed in a clean checkout and failed on every machine that had run an eval. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
redaction-pressure v0.2
A one-turn Verifiers v1 environment. The model returns regex rules; the scanner runs them against later records from the same deterministic corpus stream.
Reward contract
- Recall (0.35): a labelled secret counts only when every original character in its span is gone. Partial edits earn zero span recall.
- Precision (0.35): source characters belonging to fully removed secret spans, divided by all removed source characters. Partial damage stays in the denominator.
- Clean gate (0.30): every secret fully removed, every non-secret character and decoy preserved, and no validation or timeout errors.
- Validity: malformed JSON, bad entries, more than 24 rules, unsafe replacements, invalid regexes, excessive matches, timeouts, or an exhausted episode budget multiply every reward by zero.
Replacement is intentionally limited to "[REDACTED]" or "". That prevents a rule
from matching a complete secret and then leaking it through a capture expansion.
Splits
train and dev have separate, deterministic public seed ranges. private_eval derives
seeds with HMAC from REDACTION_PRESSURE_PRIVATE_SEED, which must contain at least 32
characters. The secret is never task data; traces carry its SHA-256 commitment and reveal
the derived task seed only in the completed trace so the episode can be reproduced by an
authorized evaluator.
Never present public dev scores as private evaluation.
Trace identity
Each task records the environment, corpus, and reward-schema versions, split commitment,
and a SHA-256 digest of the behavior-bearing installed Python sources. Scoring repeats
the identity under trace.info.environment and records validation failures under
trace.info.redaction_errors.
Checks
From the Arena repository root:
uv sync --project environments/redaction_pressure
uv run --project environments/redaction_pressure \
python -m unittest discover -s environments/redaction_pressure/tests -v
uv run --with regex python probe.py
Arena's environments are independently published libraries, so their generated
uv.lock files are intentionally not committed. The runtime contract above is pinned in
this environment's pyproject.toml and exercised on Python 3.11 and 3.12 in CI.
The committed tests demonstrate oracle 1.0 and inaction 0.0, reject the historical first-character exploit across all 64 public development tasks, exercise parser and rule validation, prove deterministic/private split behavior, and interrupt hostile regexes under the shared episode budget.