Merge the redaction-pressure hardening with the three OSRS environments

This commit is contained in:
2026-08-19 03:57:10 -07:00
16 changed files with 897 additions and 163 deletions
+42
View File
@@ -0,0 +1,42 @@
name: arena-environments
on:
push:
branches: [main]
pull_request:
workflow_dispatch:
jobs:
validate:
runs-on: ubuntu-latest
strategy:
matrix:
python: ["3.11", "3.12"]
steps:
- uses: actions/checkout@v4
- uses: astral-sh/setup-uv@v6
with:
version: "0.12.5"
enable-cache: false
- uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python }}
- name: Enforce library lock policy
run: test -z "$(find environments -name uv.lock -print -quit)"
- name: Run four-environment floor and ceiling probe
run: uv run --with regex python probe.py
- name: Run root probe integration test
run: uv run --with regex python -m unittest discover -s tests -v
- name: Install Redaction Pressure v0.2
run: uv sync --project environments/redaction_pressure --python ${{ matrix.python }}
- name: Run Redaction Pressure tests
run: >-
uv run --project environments/redaction_pressure
python -m unittest discover -s environments/redaction_pressure/tests -v
- name: Build all environment distributions
run: |
for environment in redaction_pressure canary_trap fault_localisation schema_migration; do
uv build "environments/$environment" --out-dir "dist/$environment"
done
- name: Verify documented Redaction Pressure entrypoint
run: uv run --project environments/redaction_pressure eval --help >/dev/null
+1
View File
@@ -2,3 +2,4 @@ __pycache__/
*.pyc *.pyc
.venv/ .venv/
outputs/ outputs/
environments/*/uv.lock
+17 -3
View File
@@ -33,7 +33,7 @@ what caught it — see the table below, and the git history.
| | What it measures | The trap | | | What it measures | The trap |
|---|---|---| |---|---|---|
| [`redaction-pressure`](environments/redaction_pressure) | Write redaction rules for a ticket corpus, run against the next records the generator would have produced. | Every secret class is shape-matched to a decoy: a live key beside a published test key, a card number beside a build id. A rule keyed on shape earns on one and pays on the other. | | [`redaction-pressure`](environments/redaction_pressure) | Write provenance-safe redaction rules for a ticket corpus, run against the next records the generator would have produced. | Partial edits remain residual secrets, every non-secret source character is protected, and hostile regexes share one episode budget. |
| [`canary-trap`](environments/canary_trap) | Write contamination probes that catch a model trained on your eval set. | Probes are run against a model that memorised the corpus **and** one that learned the same facts from a reworded copy. GUID-recall probes are sound and see only half of what is there. | | [`canary-trap`](environments/canary_trap) | Write contamination probes that catch a model trained on your eval set. | Probes are run against a model that memorised the corpus **and** one that learned the same facts from a reworded copy. GUID-recall probes are sound and see only half of what is there. |
| [`fault-localisation`](environments/fault_localisation) | Read an incident log and name the root cause, its class, and the line proving it. | The service that broke emits one line and goes quiet; its dependents emit a dozen timeouts. Ranking by error volume answers the victim, every time. | | [`fault-localisation`](environments/fault_localisation) | Read an incident log and name the root cause, its class, and the line proving it. | The service that broke emits one line and goes quiet; its dependents emit a dozen timeouts. Ranking by error volume answers the victim, every time. |
| [`schema-migration`](environments/schema_migration) | Split a text column into a number and a unit, executed against held-out rows. | The visible rows are tidy. The graded ones carry thousands separators, negatives, a unit containing a slash, and a value with no unit at all. | | [`schema-migration`](environments/schema_migration) | Split a text column into a number and a unit, executed against held-out rows. | The visible rows are tidy. The graded ones carry thousands separators, negatives, a unit containing a slash, and a value with no unit at all. |
@@ -47,7 +47,7 @@ Reward for the degenerate strategies and for an oracle, from `uv run python prob
| environment | inaction | crude maximiser | plausible attempt | oracle | | environment | inaction | crude maximiser | plausible attempt | oracle |
|---|---|---|---|---| |---|---|---|---|---|
| `redaction-pressure` | 0.000 | 0.380 *(redact everything)* | 0.552 | **1.000** | | `redaction-pressure` | 0.000 | 0.387 *(redact everything)* | 0.549 | **1.000** |
| `canary-trap` | 0.000 | 0.000 *(public-knowledge probes)* | 0.525 *(GUID recall)* | **1.000** | | `canary-trap` | 0.000 | 0.000 *(public-knowledge probes)* | 0.525 *(GUID recall)* | **1.000** |
| `fault-localisation` | 0.000 | 0.250 *(blame the loudest)* | 0.500 | **1.000** | | `fault-localisation` | 0.000 | 0.250 *(blame the loudest)* | 0.500 | **1.000** |
| `schema-migration` | 0.000 | 0.000 *(add columns, touch nothing)* | 0.630 | **1.000** | | `schema-migration` | 0.000 | 0.000 *(add columns, touch nothing)* | 0.630 | **1.000** |
@@ -58,9 +58,23 @@ Reward for the degenerate strategies and for an oracle, from `uv run python prob
## Running one ## Running one
```bash ```bash
uv run eval @ configs/redaction_pressure.toml --model <model-id> uv run --project environments/redaction_pressure eval @ configs/redaction_pressure.toml --model <model-id>
``` ```
Run its release gate from the repository root:
```bash
uv sync --project environments/redaction_pressure
uv run --project environments/redaction_pressure \
python -m unittest discover -s environments/redaction_pressure/tests -v
uv run --with regex python probe.py
```
Arena contains independently published environment libraries, so per-environment
`uv.lock` files are intentionally not committed. Redaction v0.2 pins its runtime contract
in `pyproject.toml`; CI resolves it on Python 3.11 and 3.12, builds every environment, and
runs the shared four-environment probe.
## Publishing ## Publishing
The layout mirrors `verifiers`' own, so an environment goes to the Hub without a fork: The layout mirrors `verifiers`' own, so an environment goes to the Hub without a fork:
+1
View File
@@ -6,6 +6,7 @@ num_rollouts = 2
[env.taskset] [env.taskset]
id = "redaction-pressure" id = "redaction-pressure"
split = "dev" # public deterministic development split; never use it as a private holdout
visible = 4 # records the model reads visible = 4 # records the model reads
held_out = 12 # records it is scored on held_out = 12 # records it is scored on
+56
View File
@@ -0,0 +1,56 @@
# redaction-pressure v0.2
A one-turn Verifiers v1 environment. The model returns regex rules; the scanner runs
them against later records from the same deterministic corpus stream.
## Reward contract
- **Recall (0.35):** a labelled secret counts only when every original character in
its span is gone. Partial edits earn zero span recall.
- **Precision (0.35):** source characters belonging to fully removed secret spans,
divided by all removed source characters. Partial damage stays in the denominator.
- **Clean gate (0.30):** every secret fully removed, every non-secret character and
decoy preserved, and no validation or timeout errors.
- **Validity:** malformed JSON, bad entries, more than 24 rules, unsafe replacements,
invalid regexes, excessive matches, timeouts, or an exhausted episode budget multiply
every reward by zero.
Replacement is intentionally limited to `"[REDACTED]"` or `""`. That prevents a rule
from matching a complete secret and then leaking it through a capture expansion.
## Splits
`train` and `dev` have separate, deterministic public seed ranges. `private_eval` derives
seeds with HMAC from `REDACTION_PRESSURE_PRIVATE_SEED`, which must contain at least 32
characters. The secret is never task data; traces carry its SHA-256 commitment and reveal
the derived task seed only in the completed trace so the episode can be reproduced by an
authorized evaluator.
Never present public `dev` scores as private evaluation.
## Trace identity
Each task records the environment, corpus, and reward-schema versions, split commitment,
and a SHA-256 digest of the behavior-bearing installed Python sources. Scoring repeats
the identity under `trace.info.environment` and records validation failures under
`trace.info.redaction_errors`.
## Checks
From the Arena repository root:
```bash
uv sync --project environments/redaction_pressure
uv run --project environments/redaction_pressure \
python -m unittest discover -s environments/redaction_pressure/tests -v
uv run --with regex python probe.py
```
Arena's environments are independently published libraries, so their generated
`uv.lock` files are intentionally not committed. The runtime contract above is pinned in
this environment's `pyproject.toml` and exercised on Python 3.11 and 3.12 in CI.
The committed tests demonstrate oracle 1.0 and inaction 0.0, reject the historical
first-character exploit across all 64 public development tasks, exercise parser and
rule validation, prove deterministic/private split behavior, and interrupt hostile
regexes under the shared episode budget.
@@ -1,12 +1,16 @@
[project] [project]
name = "redaction-pressure" name = "redaction-pressure"
version = "0.1.0" version = "0.2.0"
description = "redaction-pressure — write redaction rules graded on records the author never saw." description = "Provenance-safe redaction rules graded on records the author never saw."
requires-python = ">=3.11" requires-python = ">=3.11"
dependencies = ["verifiers", "regex"] dependencies = [
"pydantic==2.13.4",
"regex==2026.7.19",
"verifiers==0.3.0",
]
[build-system] [build-system]
requires = ["hatchling"] requires = ["hatchling>=1.27,<2"]
build-backend = "hatchling.build" build-backend = "hatchling.build"
[tool.hatch.build.targets.wheel] [tool.hatch.build.targets.wheel]
@@ -1,3 +1,4 @@
from redaction_pressure.provenance import ENVIRONMENT_VERSION
from redaction_pressure.taskset import RedactionTaskset from redaction_pressure.taskset import RedactionTaskset
__all__ = ["RedactionTaskset"] __all__ = ["ENVIRONMENT_VERSION", "RedactionTaskset"]
@@ -0,0 +1,37 @@
"""Stable environment identity recorded with every task and trace."""
from __future__ import annotations
import hashlib
from functools import lru_cache
from pathlib import Path
ENVIRONMENT_VERSION = "0.2.0"
CORPUS_VERSION = "redaction-pressure-corpus-v2"
REWARD_SCHEMA_VERSION = "redaction-pressure-reward-v2"
SOURCE_FILES = ("corpus.py", "provenance.py", "scan.py", "taskset.py")
@lru_cache(maxsize=1)
def environment_source_sha256() -> str:
"""Hash the behavior-bearing installed sources, independent of checkout paths."""
root = Path(__file__).resolve().parent
digest = hashlib.sha256()
for name in SOURCE_FILES:
payload = (root / name).read_bytes()
digest.update(name.encode("utf-8"))
digest.update(b"\0")
digest.update(len(payload).to_bytes(8, "big"))
digest.update(payload)
return digest.hexdigest()
def trace_provenance(*, split: str, seed_commitment: str) -> dict[str, str]:
return {
"environment_version": ENVIRONMENT_VERSION,
"corpus_version": CORPUS_VERSION,
"reward_schema_version": REWARD_SCHEMA_VERSION,
"source_sha256": environment_source_sha256(),
"split": split,
"seed_commitment": seed_commitment,
}
@@ -1,40 +1,35 @@
"""Applying the agent's rules, and measuring what they did. """Apply hostile redaction rules while preserving source-character provenance.
Everything is counted as *removals*, and that choice is the whole reward. The scorer never infers redaction from an output substring disappearing. Every
character starts with its source offset, replacements carry no source offsets, and a
secret counts as removed only when *all* offsets in its labelled span are gone. This
makes deleting or changing one byte of a secret worth zero recall. The same provenance
requires every ordinary and decoy character to survive for the clean gate.
The obvious framing — secrets caught, decoys survived, prose kept — pays a model that Agent regexes run under one wall-clock budget shared by compilation and every record.
returns no rules at all: it destroys nothing, so it keeps everything. Measured that way Any invalid, excessive, unsafe, over-broad, or timed-out rule invalidates the complete
an empty ruleset and a ruleset that shreds the corpus scored identically, and neither ruleset; valid rules are still measured so the trace explains what happened, but reward
told a trainer anything. Counting removals fixes it, because a model that removes nothing methods apply ``validity`` and therefore pay zero for an invalid submission.
has no true positives and therefore no precision.
removed_secrets true positives — a seeded secret gone from the held-out text.
residual_hits the ones still there. Forge's rule applies to the gate: a residual
hit is exit 1, not a deduction.
collateral_hits decoys destroyed. Each is shape-matched to a secret class, so this is
what a rule keyed on shape pays and one keyed on meaning does not.
innocent_lost ordinary words destroyed. The floor that stops `\\S+ -> [X]`, which
otherwise has perfect recall.
The agent supplies the patterns, so they are hostile input: `regex` is used rather than
`re` for its `timeout=`, since `re` cannot be interrupted and one nested quantifier would
otherwise hang a rollout until the harness killed it.
""" """
from __future__ import annotations from __future__ import annotations
import json import json
import re import re
import time
from dataclasses import dataclass from dataclasses import dataclass, field
from typing import Callable, Iterable
import regex import regex
from redaction_pressure.corpus import Record from redaction_pressure.corpus import Record, Span
MAX_RULES = 24 MAX_RULES = 24
MAX_PATTERN_CHARS = 200 MAX_PATTERN_CHARS = 200
PATTERN_TIMEOUT_S = 0.25 MAX_MATCHES_PER_RULE_RECORD = 128
PER_OPERATION_TIMEOUT_S = 0.025
EPISODE_REGEX_BUDGET_S = 0.75
ALLOWED_REPLACEMENTS = frozenset({"", "[REDACTED]"})
@dataclass(frozen=True) @dataclass(frozen=True)
@@ -43,118 +38,331 @@ class Rule:
replacement: str replacement: str
@dataclass(frozen=True)
class ParsedRules:
rules: tuple[Rule, ...]
errors: tuple[str, ...] = ()
submitted: int = 0
@dataclass(frozen=True)
class LocatedSpan:
kind: str
text: str
start: int
end: int
@property
def offsets(self) -> range:
return range(self.start, self.end)
@dataclass(frozen=True)
class TrackedText:
text: str
origins: tuple[int | None, ...]
@classmethod
def original(cls, text: str) -> "TrackedText":
return cls(text=text, origins=tuple(range(len(text))))
@dataclass(frozen=True)
class CompiledRule:
index: int
pattern: regex.Pattern
replacement: str
@dataclass
class RegexBudget:
seconds: float = EPISODE_REGEX_BUDGET_S
clock: Callable[[], float] = time.monotonic
started: float = field(init=False)
def __post_init__(self) -> None:
self.started = self.clock()
@property
def elapsed(self) -> float:
return max(0.0, self.clock() - self.started)
@property
def remaining(self) -> float:
return max(0.0, self.seconds - self.elapsed)
def operation_timeout(self) -> float:
remaining = self.remaining
if remaining <= 0:
raise RegexBudgetExceeded("episode regex budget exhausted")
return min(PER_OPERATION_TIMEOUT_S, remaining)
class RegexBudgetExceeded(RuntimeError):
pass
@dataclass(frozen=True)
class ApplyResult:
tracked: TrackedText
errors: tuple[str, ...]
@dataclass @dataclass
class Outcome: class Outcome:
"""What one ruleset did to one slice.""" """Measurements for one submitted ruleset over one held-out slice."""
removed_secrets: int
partial_secrets: int
residual_hits: int residual_hits: int
secrets_total: int secrets_total: int
credited_secret_chars: int
removed_secret_chars: int
secret_chars_total: int
collateral_hits: int collateral_hits: int
decoys_total: int decoys_total: int
innocent_lost: int innocent_lost: int
innocent_total: int innocent_total: int
invalid_rules: list[str] invalid_rules: list[str]
budget_seconds: float
@property budget_elapsed_seconds: float
def removed_secrets(self) -> int:
return self.secrets_total - self.residual_hits
@property @property
def recall(self) -> float: def recall(self) -> float:
"""All-or-nothing span recall; partial secret edits earn no recall."""
return self.removed_secrets / max(self.secrets_total, 1) return self.removed_secrets / max(self.secrets_total, 1)
@property @property
def precision(self) -> float: def character_coverage(self) -> float:
"""Of everything this ruleset destroyed, how much of it should have been. """Diagnostic only: never used as reward because partial deletion is unsafe."""
return self.removed_secret_chars / max(self.secret_chars_total, 1)
Zero for an empty ruleset — it destroyed nothing, so none of what it destroyed @property
was a secret. That is the case the survivor-counting version got wrong. def precision(self) -> float:
"""Of removed source characters, the share in fully removed secret spans.
Characters removed from a merely damaged secret stay in the denominator but are
not true positives. This keeps the historical one-byte exploit at precision zero.
""" """
destroyed = self.removed_secrets + self.collateral_hits + self.innocent_lost destroyed = self.removed_secret_chars + self.innocent_lost
return self.removed_secrets / destroyed if destroyed else 0.0 return self.credited_secret_chars / destroyed if destroyed else 0.0
@property
def ordinary_preservation(self) -> float:
return 1.0 - self.innocent_lost / max(self.innocent_total, 1)
@property
def validity(self) -> float:
"""Binary by design: invalid hostile input is never a partially paid policy."""
return float(not self.invalid_rules)
@property @property
def clean(self) -> bool: def clean(self) -> bool:
"""The gate: every secret gone AND no decoy taken with it — the ruleset you could
actually ship. `forge redact`'s exit 1, with the collateral clause it implies."""
return ( return (
self.residual_hits == 0 self.removed_secrets == self.secrets_total
and self.partial_secrets == 0
and self.collateral_hits == 0 and self.collateral_hits == 0
and self.innocent_lost == 0
and self.secrets_total > 0 and self.secrets_total > 0
and self.validity == 1.0
) )
def compile_rules(rules: list[Rule]) -> tuple[list[tuple[regex.Pattern, str]], list[str]]: def compile_rules(
"""Compile what compiles; report the rest rather than failing the rollout. parsed: ParsedRules,
budget: RegexBudget,
A rollout that dies on one bad pattern teaches nothing — the agent's score should ) -> tuple[list[CompiledRule], list[str]]:
reflect that the rule did no work, which is what dropping it does. """Compile bounded rules and retain every validation error."""
""" errors = list(parsed.errors)
compiled, invalid = [], [] rules = list(parsed.rules)
for rule in rules[:MAX_RULES]: if len(rules) > MAX_RULES:
errors.append(f"too many rules: received {len(rules)}, maximum is {MAX_RULES}")
compiled: list[CompiledRule] = []
for index, rule in enumerate(rules[:MAX_RULES]):
label = f"rule[{index}]"
if not rule.pattern:
errors.append(f"{label}: pattern must not be empty")
continue
if len(rule.pattern) > MAX_PATTERN_CHARS: if len(rule.pattern) > MAX_PATTERN_CHARS:
invalid.append(f"{rule.pattern[:40]}: over {MAX_PATTERN_CHARS} characters") errors.append(f"{label}: pattern exceeds {MAX_PATTERN_CHARS} characters")
continue
if rule.replacement not in ALLOWED_REPLACEMENTS:
errors.append(
f"{label}: replacement must be empty or [REDACTED]; captures and arbitrary text are unsafe"
)
continue continue
try: try:
compiled.append((regex.compile(rule.pattern), rule.replacement)) budget.operation_timeout()
pattern = regex.compile(rule.pattern)
budget.operation_timeout()
except RegexBudgetExceeded as exc:
errors.append(f"{label}: {exc}")
break
except regex.error as exc: except regex.error as exc:
invalid.append(f"{rule.pattern[:40]}: {exc}") errors.append(f"{label}: invalid regex: {exc}")
return compiled, invalid
def apply_rules(body: str, compiled: list[tuple[regex.Pattern, str]]) -> str:
"""Rules in the order given, each over the whole text. A pattern that times out is
skipped for that record rather than aborting the measurement."""
text = body
for pattern, replacement in compiled:
try:
text = pattern.sub(replacement, text, timeout=PATTERN_TIMEOUT_S)
except (TimeoutError, regex.error):
continue continue
return text compiled.append(CompiledRule(index=index, pattern=pattern, replacement=rule.replacement))
return compiled, errors
def _innocent_text(record: Record) -> str: def apply_rules(
"""The record with every secret and decoy cut out — the part no rule should touch.""" body: str,
text = record.body compiled: Iterable[CompiledRule],
for span in [*record.secrets, *record.decoys]: budget: RegexBudget,
text = text.replace(span.text, " ") ) -> ApplyResult:
return text """Apply rules atomically per pattern and carry source offsets through unchanged text."""
tracked = TrackedText.original(body)
errors: list[str] = []
for rule in compiled:
label = f"rule[{rule.index}]"
try:
timeout = budget.operation_timeout()
matches = list(rule.pattern.finditer(tracked.text, timeout=timeout))
budget.operation_timeout()
except (TimeoutError, regex.error):
errors.append(f"{label}: regex timed out")
continue
except RegexBudgetExceeded as exc:
errors.append(f"{label}: {exc}")
break
if len(matches) > MAX_MATCHES_PER_RULE_RECORD:
errors.append(
f"{label}: produced {len(matches)} matches in one record; maximum is "
f"{MAX_MATCHES_PER_RULE_RECORD}"
)
continue
text_parts: list[str] = []
origins: list[int | None] = []
cursor = 0
for match in matches:
start, end = match.span()
text_parts.append(tracked.text[cursor:start])
origins.extend(tracked.origins[cursor:start])
text_parts.append(rule.replacement)
origins.extend([None] * len(rule.replacement))
cursor = end
text_parts.append(tracked.text[cursor:])
origins.extend(tracked.origins[cursor:])
tracked = TrackedText(text="".join(text_parts), origins=tuple(origins))
return ApplyResult(tracked=tracked, errors=tuple(errors))
def measure(records: list[Record], rules: list[Rule]) -> Outcome: def _locate(body: str, span: Span) -> LocatedSpan:
compiled, invalid = compile_rules(rules) start = body.find(span.text)
residual = collateral = secrets = decoys = lost = total = 0 if start < 0 or body.find(span.text, start + 1) >= 0:
raise ValueError(f"ground-truth span {span.kind!r} must occur exactly once")
return LocatedSpan(span.kind, span.text, start, start + len(span.text))
def _removed_count(span: LocatedSpan, retained: set[int]) -> int:
return sum(1 for offset in span.offsets if offset not in retained)
def measure(
records: list[Record],
parsed: ParsedRules | list[Rule],
*,
budget_seconds: float = EPISODE_REGEX_BUDGET_S,
clock: Callable[[], float] = time.monotonic,
) -> Outcome:
"""Measure one ruleset under one episode-wide hostile-regex budget."""
if isinstance(parsed, list):
parsed = ParsedRules(tuple(parsed), submitted=len(parsed))
budget = RegexBudget(seconds=budget_seconds, clock=clock)
compiled, errors = compile_rules(parsed, budget)
removed_secrets = partial_secrets = residual = 0
secrets_total = credited_secret_chars = removed_secret_chars = secret_chars_total = 0
collateral = decoys_total = innocent_lost = innocent_total = 0
for record in records: for record in records:
redacted = apply_rules(record.body, compiled) result = apply_rules(record.body, compiled, budget)
secrets += len(record.secrets) errors.extend(result.errors)
decoys += len(record.decoys) retained = {offset for offset in result.tracked.origins if offset is not None}
residual += sum(1 for s in record.secrets if s.text in redacted) secrets = [_locate(record.body, span) for span in record.secrets]
collateral += sum(1 for d in record.decoys if d.text not in redacted) decoys = [_locate(record.body, span) for span in record.decoys]
# Fidelity is counted over whole words rather than characters: a rule that eats secret_offsets = {offset for span in secrets for offset in span.offsets}
# every space would otherwise keep most of its characters and look harmless. decoy_offsets = {offset for span in decoys for offset in span.offsets}
innocent = [w for w in _innocent_text(record).split() if w]
total += len(innocent) for span in secrets:
lost += sum(1 for w in innocent if w not in redacted) removed = _removed_count(span, retained)
return Outcome(residual, secrets, collateral, decoys, lost, total, invalid) removed_secret_chars += removed
secret_chars_total += span.end - span.start
secrets_total += 1
if removed == span.end - span.start:
removed_secrets += 1
credited_secret_chars += removed
elif removed > 0:
partial_secrets += 1
residual += 1
else:
residual += 1
for span in decoys:
decoys_total += 1
if _removed_count(span, retained) > 0:
collateral += 1
# Every non-secret source character is protected. Decoys are included here as well
# as in their span-level gate, so one altered byte is both visible and costly.
protected = set(range(len(record.body))) - secret_offsets
innocent_total += len(protected)
innocent_lost += len(protected - retained)
# Keep this assertion close to the accounting: every decoy must be protected.
if not decoy_offsets <= protected:
raise AssertionError("decoy offsets must be non-secret protected content")
unique_errors = list(dict.fromkeys(errors))
return Outcome(
removed_secrets=removed_secrets,
partial_secrets=partial_secrets,
residual_hits=residual,
secrets_total=secrets_total,
credited_secret_chars=credited_secret_chars,
removed_secret_chars=removed_secret_chars,
secret_chars_total=secret_chars_total,
collateral_hits=collateral,
decoys_total=decoys_total,
innocent_lost=innocent_lost,
innocent_total=innocent_total,
invalid_rules=unique_errors,
budget_seconds=budget.seconds,
budget_elapsed_seconds=budget.elapsed,
)
_BLOCK = re.compile(r"```(?:json)?\s*\n(.*?)```", re.DOTALL) _BLOCK = re.compile(r"```(?:json)?\s*\n(.*?)```", re.DOTALL | re.IGNORECASE)
def parse_rules(reply: str) -> list[Rule]: def parse_rules(reply: str) -> ParsedRules:
"""The last JSON array in the reply. A reply with no parsable ruleset is an empty """Parse the last fenced JSON block and retain every structural error.
ruleset, not an error — it scores what doing nothing scores."""
This parser deliberately lives in the scanner rather than the Verifiers taskset so
the repository-level four-environment probe stays lightweight.
"""
candidates = _BLOCK.findall(reply or "") candidates = _BLOCK.findall(reply or "")
raw = candidates[-1] if candidates else (reply or "") raw = candidates[-1] if candidates else (reply or "")
try: try:
parsed = json.loads(raw.strip()) payload = json.loads(raw.strip())
except json.JSONDecodeError: except json.JSONDecodeError as exc:
return [] return ParsedRules((), (f"response is not valid JSON: {exc.msg}",), 0)
if not isinstance(parsed, list): if not isinstance(payload, list):
return [] return ParsedRules((), ("response JSON must be an array",), 1)
rules = []
for item in parsed: rules: list[Rule] = []
if isinstance(item, dict) and isinstance(item.get("pattern"), str): errors: list[str] = []
rules.append(Rule(item["pattern"], str(item.get("replacement", "")))) for index, item in enumerate(payload):
return rules if not isinstance(item, dict):
errors.append(f"item[{index}] must be an object")
continue
pattern = item.get("pattern")
replacement = item.get("replacement", "")
if not isinstance(pattern, str):
errors.append(f"item[{index}].pattern must be a string")
continue
if not isinstance(replacement, str):
errors.append(f"item[{index}].replacement must be a string")
continue
rules.append(Rule(pattern=pattern, replacement=replacement))
return ParsedRules(tuple(rules), tuple(errors), len(payload))
@@ -1,57 +1,90 @@
"""redaction-pressure: write redaction rules that survive records you never read. """Verifiers v1 taskset for provenance-safe redaction rules."""
The agent reads a handful of support tickets seeded with secrets, and returns a ruleset.
The ruleset is then run against the *next* records the same generator would have produced.
That gap is the environment: rules fitted to the visible examples score well on the
visible examples and are worth nothing, and only the held-out slice can tell the two apart.
Three rewards, all counted over what the ruleset *destroyed*:
recall seeded secrets removed from the held-out text.
precision of everything destroyed, the share that was a secret. An empty ruleset
destroys nothing and therefore scores zero here — which is the point. A
first pass scored survivors instead, and an empty ruleset tied the best
real attempt at 0.500 because it had kept every decoy and every word.
gate binary: every secret gone and no decoy taken with it. `forge redact`'s
exit 1, which is a verdict and not a deduction.
No judge model and no runtime: the reward is a scanner, so a rollout costs one completion
and grades in milliseconds. Deterministic from the task seed.
"""
from __future__ import annotations from __future__ import annotations
import re import hashlib
from typing import ClassVar import hmac
import os
from typing import Literal
from pydantic import Field from pydantic import Field, model_validator
import verifiers.v1 as vf import verifiers.v1 as vf
from redaction_pressure.corpus import build_slices from redaction_pressure.corpus import build_slices
from redaction_pressure.scan import MAX_RULES, Rule, measure, parse_rules from redaction_pressure.provenance import (
CORPUS_VERSION,
ENVIRONMENT_VERSION,
REWARD_SCHEMA_VERSION,
environment_source_sha256,
trace_provenance,
)
from redaction_pressure.scan import MAX_RULES, measure, parse_rules
Split = Literal["train", "dev", "private_eval"]
PUBLIC_SPLIT_BASES = {"train": 100_000, "dev": 200_000}
PRIVATE_SEED_ENV_DEFAULT = "REDACTION_PRESSURE_PRIVATE_SEED"
MAX_TASKS_PER_SPLIT = 10_000
SYSTEM = f"""You are writing redaction rules for a support-ticket corpus. SYSTEM = f"""You are writing redaction rules for a support-ticket corpus.
Return a JSON array of at most {MAX_RULES} rules in one ```json code block. Each rule is Return a JSON array of at most {MAX_RULES} rules in one ```json code block. Each rule is
an object with "pattern" (a Python regular expression) and "replacement" (the text that an object with "pattern" (a Python regex) and "replacement". For safety, replacement
replaces every match). must be exactly "[REDACTED]" or the empty string; capture expansion and arbitrary
replacement text invalidate the complete ruleset.
Your rules will NOT be graded on the tickets below. They will be run against different Your rules are graded on different tickets from the same source. A secret earns credit
tickets from the same source, containing the same kinds of secret and the same kinds of only when every source character in its labelled span is removed. Changing one byte is
lookalike. Removing every secret matters most; destroying text that only resembles a still a residual secret. Every non-secret source character, including public lookalikes,
secret, or shredding ordinary prose, costs you. must survive for the clean gate. Invalid, excessive, over-broad, or timed-out rules make
the rollout invalid and receive zero reward.
Example shape: Example:
```json ```json
[{{"pattern": "\\\\bfoo-\\\\d+\\\\b", "replacement": "[REF]"}}] [{{"pattern": "\\\\bfoo-\\\\d+\\\\b", "replacement": "[REDACTED]"}}]
```""" ```"""
def _private_key(env_name: str) -> bytes:
value = os.environ.get(env_name, "")
if len(value) < 32:
raise ValueError(
f"private_eval requires at least 32 characters in environment variable {env_name}"
)
return value.encode("utf-8")
def seed_for(split: Split, index: int, private_seed_env: str = PRIVATE_SEED_ENV_DEFAULT) -> int:
if split in PUBLIC_SPLIT_BASES:
return PUBLIC_SPLIT_BASES[split] + index
key = _private_key(private_seed_env)
digest = hmac.new(
key,
f"redaction-pressure/v2/private_eval/{index}".encode("utf-8"),
hashlib.sha256,
).digest()
return int.from_bytes(digest[:8], "big") & ((1 << 63) - 1)
def seed_commitment(split: Split, private_seed_env: str = PRIVATE_SEED_ENV_DEFAULT) -> str:
if split in PUBLIC_SPLIT_BASES:
material = f"redaction-pressure/v2/{split}/{PUBLIC_SPLIT_BASES[split]}".encode()
else:
material = _private_key(private_seed_env)
return hashlib.sha256(material).hexdigest()
class RedactionData(vf.TaskData): class RedactionData(vf.TaskData):
seed: int seed: int
"""Rebuilds both slices exactly; the held-out records are never serialized here.""" """Rebuilds both slices exactly; held-out records are generated only for scoring."""
visible: int visible: int
held_out: int held_out: int
split: Split
seed_commitment: str
environment_version: str
corpus_version: str
reward_schema_version: str
source_sha256: str
class RedactionTask(vf.Task[RedactionData]): class RedactionTask(vf.Task[RedactionData]):
@@ -61,59 +94,81 @@ class RedactionTask(vf.Task[RedactionData]):
@vf.metric @vf.metric
async def scan(self, trace: vf.Trace) -> dict[str, float]: async def scan(self, trace: vf.Trace) -> dict[str, float]:
"""Run the ruleset over the held-out slice once; every reward reads this.""" """Rebuild the held-out slice once and expose raw diagnostics in the trace."""
_, held_out = build_slices(self.data.seed, self.data.visible, self.data.held_out) _, held_out = build_slices(self.data.seed, self.data.visible, self.data.held_out)
outcome = measure(held_out, parse_rules(trace.last_reply)) outcome = measure(held_out, parse_rules(trace.last_reply))
trace.info["environment"] = trace_provenance(
split=self.data.split,
seed_commitment=self.data.seed_commitment,
)
trace.info["redaction_errors"] = outcome.invalid_rules
return { return {
"recall": outcome.recall, "recall": outcome.recall,
"precision": outcome.precision, "precision": outcome.precision,
"clean": float(outcome.clean), "clean": float(outcome.clean),
"validity": outcome.validity,
"character_coverage": outcome.character_coverage,
"ordinary_preservation": outcome.ordinary_preservation,
"residual_hits": float(outcome.residual_hits), "residual_hits": float(outcome.residual_hits),
"partial_secrets": float(outcome.partial_secrets),
"collateral_hits": float(outcome.collateral_hits), "collateral_hits": float(outcome.collateral_hits),
"innocent_lost": float(outcome.innocent_lost), "innocent_lost": float(outcome.innocent_lost),
"invalid_rules": float(len(outcome.invalid_rules)), "invalid_rules": float(len(outcome.invalid_rules)),
"regex_budget_elapsed_seconds": outcome.budget_elapsed_seconds,
} }
@vf.reward(weight=0.35) @vf.reward(weight=0.35)
async def recall(self, trace: vf.Trace) -> float: async def recall(self, trace: vf.Trace) -> float:
return trace.metrics.get("recall", 0.0) return trace.metrics.get("recall", 0.0) * trace.metrics.get("validity", 0.0)
@vf.reward(weight=0.35) @vf.reward(weight=0.35)
async def precision(self, trace: vf.Trace) -> float: async def precision(self, trace: vf.Trace) -> float:
return trace.metrics.get("precision", 0.0) return trace.metrics.get("precision", 0.0) * trace.metrics.get("validity", 0.0)
@vf.reward(weight=0.30) @vf.reward(weight=0.30)
async def gate(self, trace: vf.Trace) -> float: async def gate(self, trace: vf.Trace) -> float:
return trace.metrics.get("clean", 0.0) return trace.metrics.get("clean", 0.0) * trace.metrics.get("validity", 0.0)
class RedactionConfig(vf.TasksetConfig): class RedactionConfig(vf.TasksetConfig):
num_tasks: int = Field(64, ge=1) num_tasks: int = Field(64, ge=1, le=MAX_TASKS_PER_SPLIT)
visible: int = Field(4, ge=1) visible: int = Field(4, ge=1, le=64)
"""Records the agent reads.""" held_out: int = Field(12, ge=1, le=256)
held_out: int = Field(12, ge=1) split: Split = "dev"
"""Records it is graded on and never sees.""" private_seed_env: str = PRIVATE_SEED_ENV_DEFAULT
@model_validator(mode="after")
def validate_private_split(self) -> "RedactionConfig":
if self.split == "private_eval":
_private_key(self.private_seed_env)
return self
class RedactionTaskset(vf.Taskset[RedactionTask, RedactionConfig]): class RedactionTaskset(vf.Taskset[RedactionTask, RedactionConfig]):
SEED_BASE: ClassVar[int] = 90_000
def load(self) -> list[RedactionTask]: def load(self) -> list[RedactionTask]:
commitment = seed_commitment(self.config.split, self.config.private_seed_env)
source_hash = environment_source_sha256()
tasks = [] tasks = []
for i in range(self.config.num_tasks): for index in range(self.config.num_tasks):
seed = self.SEED_BASE + i seed = seed_for(self.config.split, index, self.config.private_seed_env)
visible, _ = build_slices(seed, self.config.visible, self.config.held_out) visible, _ = build_slices(seed, self.config.visible, self.config.held_out)
tickets = "\n\n".join(record.body for record in visible) tickets = "\n\n".join(record.body for record in visible)
tasks.append( tasks.append(
RedactionTask( RedactionTask(
RedactionData( RedactionData(
idx=i, idx=index,
name=f"stream-{seed}", name=f"{self.config.split}-{index:05d}",
prompt=f"Tickets from the stream:\n\n{tickets}\n\nWrite the ruleset.", prompt=f"Tickets from the stream:\n\n{tickets}\n\nWrite the ruleset.",
system_prompt=SYSTEM, system_prompt=SYSTEM,
seed=seed, seed=seed,
visible=self.config.visible, visible=self.config.visible,
held_out=self.config.held_out, held_out=self.config.held_out,
split=self.config.split,
seed_commitment=commitment,
environment_version=ENVIRONMENT_VERSION,
corpus_version=CORPUS_VERSION,
reward_schema_version=REWARD_SCHEMA_VERSION,
source_sha256=source_hash,
), ),
self.config.task, self.config.task,
) )
@@ -0,0 +1 @@
"""Tests for redaction-pressure."""
@@ -0,0 +1,36 @@
from __future__ import annotations
import json
from redaction_pressure.scan import Rule
ORACLE_RULES = [
Rule(r"(?<=Customer pasted their key )sk-live-[a-z0-9]{24}", "[REDACTED]"),
Rule(r"(?<=Reply-to was )[a-z]+[.][a-z]+@[a-z]{6}[.]com", "[REDACTED]"),
Rule(r"(?<=Callback number )[+]1-[0-9]{3}-[0-9]{3}-[0-9]{4}", "[REDACTED]"),
Rule(r"(?<=They read the card )[0-9]{4}(?:-[0-9]{4}){3}", "[REDACTED]"),
Rule(r"(?<=Trace pointed at )[a-z]{5}-db-[1-9][.]internal[.]lumbridge", "[REDACTED]"),
Rule(
r"(?<=Escalated by )(?:Dana|Priya|Marcus|Ines|Tomas|Rui|Nadia|Owen) "
r"(?:Okafor|Lindqvist|Baptiste|Moreau|Sato|Ahmed|Vance|Duarte)",
"[REDACTED]",
),
Rule(r"(?<=Session )[0-9a-f]{8}(?:-[0-9a-f]{4}){3}-[0-9a-f]{12}", "[REDACTED]"),
]
FIRST_CHARACTER_HACK = [
Rule(r"(?<=Customer pasted their key )s", ""),
Rule(r"(?<=Reply-to was )[a-z]", ""),
Rule(r"(?<=Callback number )[+]", ""),
Rule(r"(?<=They read the card )[0-9]", ""),
Rule(r"(?<=Trace pointed at )[a-z]", ""),
Rule(r"(?<=Escalated by )[A-Z]", ""),
Rule(r"(?<=Session )[0-9a-f]", ""),
]
def reply_for(rules: list[Rule]) -> str:
return "```json\n" + json.dumps(
[{"pattern": rule.pattern, "replacement": rule.replacement} for rule in rules]
) + "\n```"
@@ -0,0 +1,116 @@
from __future__ import annotations
import unittest
import regex
from redaction_pressure.corpus import build_slices
from redaction_pressure.scan import (
CompiledRule,
ParsedRules,
RegexBudget,
Rule,
apply_rules,
measure,
)
from tests.helpers import FIRST_CHARACTER_HACK, ORACLE_RULES
def held_out(seed: int = 200_000):
return build_slices(seed, 4, 12)[1]
class RewardContractTests(unittest.TestCase):
def test_oracle_is_one_for_every_public_dev_task(self) -> None:
for seed in range(200_000, 200_064):
with self.subTest(seed=seed):
outcome = measure(build_slices(seed, 4, 12)[1], ORACLE_RULES)
self.assertEqual(outcome.recall, 1.0)
self.assertEqual(outcome.precision, 1.0)
self.assertTrue(outcome.clean)
self.assertEqual(outcome.validity, 1.0)
def test_inaction_is_zero(self) -> None:
outcome = measure(held_out(), [])
self.assertEqual(outcome.recall, 0.0)
self.assertEqual(outcome.precision, 0.0)
self.assertFalse(outcome.clean)
self.assertEqual(outcome.validity, 1.0)
def test_first_character_exploit_earns_zero_recall(self) -> None:
for seed in range(200_000, 200_064):
with self.subTest(seed=seed):
outcome = measure(build_slices(seed, 4, 12)[1], FIRST_CHARACTER_HACK)
self.assertEqual(outcome.recall, 0.0)
self.assertEqual(outcome.precision, 0.0)
self.assertEqual(outcome.partial_secrets, outcome.secrets_total)
self.assertEqual(outcome.residual_hits, outcome.secrets_total)
self.assertFalse(outcome.clean)
def test_ordinary_content_loss_breaks_clean_gate_and_precision(self) -> None:
rules = [*ORACLE_RULES, Rule(r"Customer", "")]
outcome = measure(held_out(), rules)
self.assertEqual(outcome.recall, 1.0)
self.assertGreater(outcome.innocent_lost, 0)
self.assertLess(outcome.precision, 1.0)
self.assertFalse(outcome.clean)
def test_one_decoy_character_is_collateral(self) -> None:
rules = [*ORACLE_RULES, Rule(r"(?<=our published sample is )s", "")]
outcome = measure(held_out(), rules)
self.assertGreater(outcome.collateral_hits, 0)
self.assertGreater(outcome.innocent_lost, 0)
self.assertFalse(outcome.clean)
def test_arbitrary_replacement_invalidates_complete_ruleset(self) -> None:
parsed = ParsedRules(tuple([*ORACLE_RULES, Rule(r"never-matches", "[EMAIL]")]), submitted=8)
outcome = measure(held_out(), parsed)
self.assertEqual(outcome.validity, 0.0)
self.assertFalse(outcome.clean)
self.assertTrue(any("replacement" in error for error in outcome.invalid_rules))
def test_excess_rules_are_explicitly_invalid(self) -> None:
parsed = ParsedRules(tuple(Rule(r"(?!)", "") for _ in range(25)), submitted=25)
outcome = measure(held_out(), parsed)
self.assertEqual(outcome.validity, 0.0)
self.assertTrue(any("too many rules" in error for error in outcome.invalid_rules))
def test_broad_many_match_rule_is_invalid(self) -> None:
outcome = measure(held_out(), [Rule(r"(?=.)", "")])
self.assertEqual(outcome.validity, 0.0)
self.assertTrue(any("matches in one record" in error for error in outcome.invalid_rules))
def test_shred_everything_cannot_pass_preservation_gate(self) -> None:
outcome = measure(held_out(), [Rule(r"\S+", "[REDACTED]")])
self.assertGreater(outcome.innocent_lost, 0)
self.assertGreater(outcome.collateral_hits, 0)
self.assertFalse(outcome.clean)
self.assertLess(outcome.precision, 1.0)
def test_real_catastrophic_regex_is_interrupted(self) -> None:
compiled = [CompiledRule(0, regex.compile(r"(?:a|aa)+$"), "")]
result = apply_rules("a" * 20_000 + "!", compiled, RegexBudget(seconds=0.05))
self.assertTrue(any("timed out" in error or "budget" in error for error in result.errors))
self.assertEqual(result.tracked.text, "a" * 20_000 + "!")
def test_episode_budget_is_shared_and_bounded(self) -> None:
class Clock:
def __init__(self) -> None:
self.value = 0.0
def __call__(self) -> float:
self.value += 0.2
return self.value
outcome = measure(held_out(), ORACLE_RULES, budget_seconds=0.5, clock=Clock())
self.equal_zero_reward_invalid(outcome)
self.assertGreaterEqual(outcome.budget_elapsed_seconds, outcome.budget_seconds)
self.assertTrue(any("budget exhausted" in error for error in outcome.invalid_rules))
def equal_zero_reward_invalid(self, outcome) -> None:
self.assertEqual(outcome.validity, 0.0)
self.assertFalse(outcome.clean)
if __name__ == "__main__":
unittest.main()
@@ -0,0 +1,120 @@
from __future__ import annotations
import asyncio
import os
import string
import unittest
from types import SimpleNamespace
from unittest.mock import patch
from pydantic import ValidationError
from redaction_pressure.corpus import build_slices
from redaction_pressure.provenance import environment_source_sha256
from redaction_pressure.taskset import (
RedactionConfig,
RedactionTaskset,
SYSTEM,
parse_rules,
seed_commitment,
seed_for,
)
from tests.helpers import ORACLE_RULES, reply_for
class ParserTests(unittest.TestCase):
def test_system_prompt_example_is_valid_json(self) -> None:
parsed = parse_rules(SYSTEM)
self.assertEqual(parsed.errors, ())
self.assertEqual(len(parsed.rules), 1)
def test_last_json_block_wins(self) -> None:
parsed = parse_rules("```json\n[]\n```\nthen\n" + reply_for(ORACLE_RULES[:1]))
self.assertEqual(parsed.rules, tuple(ORACLE_RULES[:1]))
self.assertEqual(parsed.errors, ())
self.assertEqual(parsed.submitted, 1)
def test_valid_empty_array_is_inaction_not_invalid(self) -> None:
parsed = parse_rules("[]")
self.assertEqual(parsed.rules, ())
self.assertEqual(parsed.errors, ())
self.assertEqual(parsed.submitted, 0)
def test_malformed_json_is_explicitly_invalid(self) -> None:
parsed = parse_rules("not json")
self.assertFalse(parsed.rules)
self.assertTrue(parsed.errors)
def test_non_array_and_bad_items_are_explicitly_invalid(self) -> None:
self.assertTrue(parse_rules('{"pattern": "x"}').errors)
parsed = parse_rules('[1, {"pattern": 2}, {"pattern": "x", "replacement": 4}]')
self.assertEqual(len(parsed.errors), 3)
self.assertEqual(parsed.submitted, 3)
class SeedAndProvenanceTests(unittest.TestCase):
def test_corpus_is_deterministic(self) -> None:
self.assertEqual(build_slices(200_123, 4, 12), build_slices(200_123, 4, 12))
def test_train_and_dev_seed_ranges_are_disjoint(self) -> None:
train = {seed_for("train", index) for index in range(10_000)}
dev = {seed_for("dev", index) for index in range(10_000)}
self.assertTrue(train.isdisjoint(dev))
def test_private_split_requires_external_secret(self) -> None:
with patch.dict(os.environ, {}, clear=True):
with self.assertRaises(ValidationError):
RedactionConfig(split="private_eval")
def test_private_seeds_are_deterministic_keyed_and_committed(self) -> None:
key_a = "a" * 32
key_b = "b" * 32
with patch.dict(os.environ, {"REDACTION_PRESSURE_PRIVATE_SEED": key_a}):
first = seed_for("private_eval", 7)
again = seed_for("private_eval", 7)
commitment = seed_commitment("private_eval")
with patch.dict(os.environ, {"REDACTION_PRESSURE_PRIVATE_SEED": key_b}):
other = seed_for("private_eval", 7)
self.assertEqual(first, again)
self.assertNotEqual(first, other)
self.assertNotIn(key_a, commitment)
self.assertEqual(len(commitment), 64)
def test_source_hash_is_stable_hex(self) -> None:
first = environment_source_sha256()
self.assertEqual(first, environment_source_sha256())
self.assertEqual(len(first), 64)
self.assertTrue(set(first) <= set(string.hexdigits.lower()))
def test_task_data_carries_replay_identity_without_secret_key(self) -> None:
taskset = RedactionTaskset(RedactionConfig(num_tasks=2, split="dev"))
tasks = taskset.load()
self.assertEqual([task.data.name for task in tasks], ["dev-00000", "dev-00001"])
self.assertNotEqual(tasks[0].data.seed, tasks[1].data.seed)
self.assertEqual(tasks[0].data.source_sha256, environment_source_sha256())
self.assertEqual(len(tasks[0].data.seed_commitment), 64)
self.assertEqual(tasks[0].data.environment_version, "0.2.0")
def test_scoring_records_provenance_and_oracle_metrics(self) -> None:
task = RedactionTaskset(RedactionConfig(num_tasks=1, split="dev")).load()[0]
trace = SimpleNamespace(last_reply=reply_for(ORACLE_RULES), info={})
metrics = asyncio.run(task.scan(trace))
self.assertEqual(metrics["recall"], 1.0)
self.assertEqual(metrics["precision"], 1.0)
self.assertEqual(metrics["clean"], 1.0)
self.assertEqual(metrics["validity"], 1.0)
self.assertEqual(trace.info["environment"]["source_sha256"], environment_source_sha256())
self.assertEqual(trace.info["redaction_errors"], [])
def test_invalidity_multiplies_all_rewards_to_zero(self) -> None:
task = RedactionTaskset(RedactionConfig(num_tasks=1, split="dev")).load()[0]
trace = SimpleNamespace(
metrics={"recall": 1.0, "precision": 1.0, "clean": 1.0, "validity": 0.0}
)
self.assertEqual(asyncio.run(task.recall(trace)), 0.0)
self.assertEqual(asyncio.run(task.precision(trace)), 0.0)
self.assertEqual(asyncio.run(task.gate(trace)), 0.0)
if __name__ == "__main__":
unittest.main()
+15 -13
View File
@@ -3,9 +3,9 @@
Run this before believing any reward in this repository. Every environment here was wrong Run this before believing any reward in this repository. Every environment here was wrong
the first time and this is what caught it: the first time and this is what caught it:
redaction-pressure scored SURVIVORS — secrets caught, decoys kept, prose kept. An empty redaction-pressure originally scored SURVIVORS, so an empty ruleset tied a real attempt.
ruleset destroys nothing so it keeps everything, and it tied the best Its v0.2 scanner additionally tracks source offsets: changing one byte
real attempt at 0.500. Counting removals instead put inaction at zero. of a secret earns zero recall, and malformed/unsafe rules earn zero.
canary-trap let the paraphrase restate the corpus verbatim for facts that had no canary-trap let the paraphrase restate the corpus verbatim for facts that had no
distinct rewording, so GUID probes caught it by accident and the distinct rewording, so GUID probes caught it by accident and the
environment taught nothing. Every fact now has a `core` token that is environment taught nothing. Every fact now has a `core` token that is
@@ -102,23 +102,25 @@ def redaction() -> dict[str, float]:
def run(rules: list[dict]) -> float: def run(rules: list[dict]) -> float:
total = 0.0 total = 0.0
for i in range(TASKS): for i in range(TASKS):
_, held = build_slices(90_000 + i, 4, 12) _, held = build_slices(200_000 + i, 4, 12)
o = measure(held, parse_rules("```json\n" + json.dumps(rules) + "\n```")) o = measure(held, parse_rules("```json\n" + json.dumps(rules) + "\n```"))
total += 0.35 * o.recall + 0.35 * o.precision + 0.30 * float(o.clean) total += o.validity * (
0.35 * o.recall + 0.35 * o.precision + 0.30 * float(o.clean)
)
return total / TASKS return total / TASKS
oracle = [ oracle = [
{"pattern": r"sk-live-\w+", "replacement": "[KEY]"}, {"pattern": r"(?<=Customer pasted their key )sk-live-[a-z0-9]{24}", "replacement": "[REDACTED]"},
{"pattern": r"[\w.]+@(?!lumbridgecorp\.com)[\w.]+\.com", "replacement": "[EMAIL]"}, {"pattern": r"(?<=Reply-to was )[a-z]+[.][a-z]+@[a-z]{6}[.]com", "replacement": "[REDACTED]"},
{"pattern": r"\+1-\d{3}-\d{3}-\d{4}", "replacement": "[PHONE]"}, {"pattern": r"(?<=Callback number )[+]1-[0-9]{3}-[0-9]{3}-[0-9]{4}", "replacement": "[REDACTED]"},
{"pattern": r"(?<!ORD-)\b\d{4}-\d{4}-\d{4}-\d{4}\b", "replacement": "[CARD]"}, {"pattern": r"(?<=They read the card )[0-9]{4}(?:-[0-9]{4}){3}", "replacement": "[REDACTED]"},
{"pattern": r"\b[\w-]+\.internal\.lumbridge\b", "replacement": "[HOST]"}, {"pattern": r"(?<=Trace pointed at )[a-z]{5}-db-[1-9][.]internal[.]lumbridge", "replacement": "[REDACTED]"},
{"pattern": r"\b[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}\b", "replacement": "[GUID]"}, {"pattern": r"(?<=Session )[0-9a-f]{8}(?:-[0-9a-f]{4}){3}-[0-9a-f]{12}", "replacement": "[REDACTED]"},
{"pattern": r"\b(Dana|Priya|Marcus|Ines|Tomas|Rui|Nadia|Owen) (Okafor|Lindqvist|Baptiste|Moreau|Sato|Ahmed|Vance|Duarte)\b", "replacement": "[PERSON]"}, {"pattern": r"(?<=Escalated by )(?:Dana|Priya|Marcus|Ines|Tomas|Rui|Nadia|Owen) (?:Okafor|Lindqvist|Baptiste|Moreau|Sato|Ahmed|Vance|Duarte)", "replacement": "[REDACTED]"},
] ]
return { return {
"inaction": run([]), "inaction": run([]),
"crude": run([{"pattern": r"\S+", "replacement": "[X]"}]), "crude": run([{"pattern": r"\S+", "replacement": "[REDACTED]"}]),
"plausible": run(oracle[:4]), "plausible": run(oracle[:4]),
"oracle": run(oracle), "oracle": run(oracle),
} }
+40
View File
@@ -0,0 +1,40 @@
from __future__ import annotations
import subprocess
import sys
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
ENVIRONMENTS = (
"redaction-pressure",
"canary-trap",
"fault-localisation",
"schema-migration",
)
class RootProbeIntegrationTests(unittest.TestCase):
def test_public_root_probe_gates_all_four_environments(self) -> None:
completed = subprocess.run(
[sys.executable, "probe.py"],
cwd=ROOT,
text=True,
capture_output=True,
timeout=30,
check=False,
)
self.assertEqual(
completed.returncode,
0,
msg=f"root probe failed:\nstdout:\n{completed.stdout}\nstderr:\n{completed.stderr}",
)
for environment in ENVIRONMENTS:
with self.subTest(environment=environment):
self.assertIn(environment, completed.stdout)
self.assertEqual(completed.stdout.count(" ok"), len(ENVIRONMENTS))
if __name__ == "__main__":
unittest.main()