Files
arena/environments/redaction_pressure
kartiandClaude Opus 5 3305be4ff7 gates: give two of them a margin, and make the probe see components
The probe printed one blended number per policy, so a component pinned at 0.000
across every rung was invisible. Four gates hid there for a month. It now reports
floor / best-below-oracle / oracle / ceiling for 25 components across 8
environments, and a component flat across every rung is fatal.

bot-detection GATE_SLACK = 1 — fires 5/32 on the real rollouts, was 0/32. The
max(1, reference_caught - SLACK) guard is verified by construction, not by
sampling: without it the required count reaches 0 and an EMPTY accusation list
clears the gate. Observed reference_caught is 4-6, so no amount of sampling would
have found that hole.

schema-migration GATE_MARGIN = 0.15 — fires 4/32, was 0/32. The cost is disclosed
and bounded: the naive split clears it on 5.5% of 1,000 unseen seeds, fenced by
an assert at 10%. Margin 0.10 keeps the leak at zero and fires 0/32, i.e. stays
dead. A live gradient with a bounded leak beats a clean corpse.

redaction-pressure is NOT given a margin, and that is the result rather than a
failure. The only setting that fires at all leaves half the secrets standing and
pays a four-of-seven ruleset on five seeds in six — a margin that pays for
inaction is strictly worse than a dead gate. Recall maxes at 0.852 and no rollout
ever cleared both clauses in one episode. It is genuinely hard, not miscalibrated.

⚠️ The per-component check did not catch the defect it was built for. Reverting
schema-migration's margin to 0.0 — restoring the exact dead gate — printed ok and
exited 0, because the near-oracle rung scrapes the unmargined gate on ~2 seeds in
24 and that kept best<oracle non-zero. Every assertion bounded how much a margin
may PAY; none noticed if it stopped existing. migration() now carries the mirror
of bot-detection's guard, and reverting the margin fails with "the margin is dead
and the component carries no gradient between the crude answer and the exact one".

canary-trap's oracle-minus-one rung is documented as degenerate rather than
quietly relied on: it is identical to the oracle to four decimals, so it measures
specificity and gate at the ceiling, not mid-ladder as its comment claimed.

The CI lock policy asks git instead of the disk. It was checking the working
tree, where a lock file is a normal by-product of uv sync, so it passed in a clean
checkout and failed on every machine that had run an eval.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 23:14:05 -07:00
..

redaction-pressure v0.2

A one-turn Verifiers v1 environment. The model returns regex rules; the scanner runs them against later records from the same deterministic corpus stream.

Reward contract

  • Recall (0.35): a labelled secret counts only when every original character in its span is gone. Partial edits earn zero span recall.
  • Precision (0.35): source characters belonging to fully removed secret spans, divided by all removed source characters. Partial damage stays in the denominator.
  • Clean gate (0.30): every secret fully removed, every non-secret character and decoy preserved, and no validation or timeout errors.
  • Validity: malformed JSON, bad entries, more than 24 rules, unsafe replacements, invalid regexes, excessive matches, timeouts, or an exhausted episode budget multiply every reward by zero.

Replacement is intentionally limited to "[REDACTED]" or "". That prevents a rule from matching a complete secret and then leaking it through a capture expansion.

Splits

train and dev have separate, deterministic public seed ranges. private_eval derives seeds with HMAC from REDACTION_PRESSURE_PRIVATE_SEED, which must contain at least 32 characters. The secret is never task data; traces carry its SHA-256 commitment and reveal the derived task seed only in the completed trace so the episode can be reproduced by an authorized evaluator.

Never present public dev scores as private evaluation.

Trace identity

Each task records the environment, corpus, and reward-schema versions, split commitment, and a SHA-256 digest of the behavior-bearing installed Python sources. Scoring repeats the identity under trace.info.environment and records validation failures under trace.info.redaction_errors.

Checks

From the Arena repository root:

uv sync --project environments/redaction_pressure
uv run --project environments/redaction_pressure \
  python -m unittest discover -s environments/redaction_pressure/tests -v
uv run --with regex python probe.py

Arena's environments are independently published libraries, so their generated uv.lock files are intentionally not committed. The runtime contract above is pinned in this environment's pyproject.toml and exercised on Python 3.11 and 3.12 in CI.

The committed tests demonstrate oracle 1.0 and inaction 0.0, reject the historical first-character exploit across all 64 public development tasks, exercise parser and rule validation, prove deterministic/private split behavior, and interrupt hostile regexes under the shared episode budget.