Files
PIG-Demo/envs/alert_triage/SPEC.md
T
karti-ai 1239dc7034
ci / web (push) Failing after 1m53s
ci / python (push) Failing after 7s
new-environment pipeline: scripted arms required, model arms captured when spark-1 is up
An environment ships on its scripted agents alone — they are what the reward
flip is built on and they need no model. A recorded model attempt is captured
once from spark-1 for the Watch tab when it is serving, and the environment does
not block on it when it is not.
2026-08-28 18:32:29 -07:00

376 lines
24 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# alert-triage — specification
Vertical: `financial-crime-alert-triage` (`src/content/verticals.ts`, rank 4).
Package: `envs/alert_triage/`. Demo: `src/demos/alert-triage/`. Traces:
`public/traces/alert-triage/`.
This document is written before any code and is meant to be attacked. The two
things it has to survive: a sceptic calling the grader subjective, and a
sceptic finding a policy that maxes the reward while doing something a BSA
officer would fire an analyst for.
**In one paragraph.** One episode is one transaction-monitoring alert on one
synthetic customer. The agent sees what the analyst's screen shows — the alert,
the triggering transactions, the KYC file, a twelve-month summary — and may pay
analyst-hours for lookups (itemised history, a counterparty, prior alerts,
documents on file). It ends the episode by *closing* the alert with a cited
rationale or *escalating* it for a suspicious-activity filing with a typology
and the transactions that constitute it. The generator planted the answer: the
alert is either a benign scenario with a specific explaining fact, or a
laundering typology with a specific set of transactions. The grader compares
the disposition and the citations to what was planted, and charges the hours.
Nothing in it reads prose.
---
## 1. The episode
**One alert.** Deterministic in the seed. The generator draws a scenario
template from the seed (FNV-1a over the decimal seed selects the template and
seeds a hand-written xorshift32 for its parameters — never a built-in RNG,
see AGENTS.md), and produces:
- the **customer** (`C-####`): individual or business, occupation/business
category, stated income or revenue, expected monthly cash, country, account
opened, KYC last refreshed, PEP flag, prior alert and prior SAR counts;
- twelve months of **account history**: a monthly summary table (cash in/out,
wires in/out, ACH in/out, transaction count), and the itemised transactions
behind it, each with a stable id `T-###`, a date, a channel, an amount, a
branch or counterparty id `CP-##`;
- the **alert** (`A-####`): the rule that fired (`R-STR-01` structuring,
`R-VEL-02` rapid movement, `R-CASH-03` cash above KYC expectation,
`R-WIRE-04` wires to a high-risk jurisdiction, `R-DORM-05` dormant-account
reactivation), the fire date, and the triggering transactions itemised;
- the **hidden truth**, never shown: `label ∈ {benign, suspicious}`, the
`typology` if suspicious, and the `planted` citation set (§4).
**The free screen.** Turn 0 shows the alert, its triggering transactions, the
KYC block and the monthly summary. This costs a fixed `SCREEN_HOURS` — the
analyst has to read it whatever they do next.
**Turns.** Up to `MAX_TURNS = 8` replies. Each reply is one lookup or one
disposition (§3). Each turn spent — accepted *or* rejected — costs
`TURN_HOURS`; a lookup adds its own price.
**What ends it.** An accepted `close` or `escalate`, or the eighth turn
without one. The latter is `outcome: aborted` and scores 0.000 on every
component: an alert nobody disposed of is still in the queue, which is the
worst outcome for throughput and, on a suspicious alert, is a miss.
**Scenario templates.** Six benign, five suspicious. Every template is written
from public typology descriptions (FinCEN SAR narrative guidance, FATF
typology reports); none is a real case.
| id | label | rule that fires | what the screen shows | where the answer lives |
|---|---|---|---|---|
| B1 seasonal cash | benign | R-CASH-03 | landscaping / tax-prep business, cash spikes in season | `kyc.expected_activity` says seasonal, `kyc.expected_monthly_cash` covers peak |
| B2 processor settlement | benign | R-VEL-02 | business receives processor wires, pays contractors within 48h | `doc.processor_agreement` on file |
| B3 property sale | benign | R-VEL-02 / R-DORM-05 | individual receives title-company wire, moves it to brokerage | `doc.closing_statement` |
| B4 cash restaurant | benign | R-STR-01 | daily cash deposits, three happen to land in 8,0009,999 | `kyc.expected_monthly_cash` ≥ observed; 12-month cash profile flat |
| B5 family remittance | benign | R-WIRE-04 | monthly wire to one named individual in a high-risk country | `kyc.remittance_declared` names the destination |
| B6 inheritance | benign | R-DORM-05 | dormant account receives estate wire | `doc.probate_letter` |
| S1 structuring | suspicious, **visible** | R-STR-01 | 45k stated income; 8 cash deposits 9,2009,900 in 12 days across 3 branches | the 8 deposits |
| S3 money mule | suspicious, **visible** | R-VEL-02 | student/retiree; 4 inbound P2P from unrelated individuals, each out within 24h to an exchange | the in/out pairs |
| S6 smurfing over time | suspicious, **visible** | R-CASH-03 | cash business; monthly cash tripled over 12 months and no single deposit ever exceeded 9,999 | the deposits in the last 3 months |
| S2 funnel account | suspicious, **hidden** | R-STR-01 or R-CASH-03 | *generated by B4's template* — screen is identically distributed to a benign restaurant | `counterparty` lookup: its 3 outbound-wire counterparties are shared with 614 other alerted accounts; the wires are the evidence |
| S5 serial closer | suspicious, **hidden** | R-VEL-02 | *generated by B2's template* — processor agreement on file, contractors paid | `prior_alerts` lookup: 3 prior alerts in 9 months, each closed by a different analyst with the same boilerplate; `documents` shows the agreement's counterparty is not the payer |
The **hidden** tier is the environment's whole argument. S2 and S5 are
produced by *the same generator code* as B4 and B2 respectively, with the
suspicious signal overlaid only in the data a lookup returns. The free screen
for a hidden case is therefore not merely similar to a benign one; it is drawn
from the same distribution. A policy that decides from the screen cannot tell
them apart, by construction, not by our failing to try hard enough.
Mix by seed: 55% benign, 30% visible-suspicious, 15% hidden. This is
deliberately **not** the real base rate — real programmes close the great
majority of alerts as false positives, and at that rate a 24-seed probe would
contain about one suspicious alert and the ladder would measure noise. The page
must say this (§8, risk 1).
## 2. The data
**Synthetic, entirely, and labelled synthetic on the page — in the surface,
not a footnote.** No real customer, account, transaction, counterparty or
analyst exists behind any of it. Names come from a synthetic name list
committed under `envs/alert_triage/names/` with its own `PROVENANCE.md`;
business names are composed from committed word lists ("Harbor Lane Bistro
LLC"). Amounts, dates, branch numbers and counterparties are generated.
Licence: the generator and its lists are ours, Apache-2.0 with the repository.
The public documents the templates are *written from* — FinCEN's SAR narrative
guidance and FATF typology reports — are cited as references on the evidence
tab, not shipped, and not represented as data.
Why synthetic is the only honest option here: real alert data is customer
financial data under BSA confidentiality; a SAR's existence may not be
disclosed. There is no licensable real dataset for this task, and any page
implying one would be lying.
## 3. The action set
Every reply is exactly **one JSON object** — the first `{…}` the parser can
balance, or the contents of the first ```` ```json ```` fence if present. The
system prompt says so and shows the four shapes.
| action | fields | cost | what comes back |
|---|---|---|---|
| `{"action":"lookup","what":"history","month":"2026-01"}` | a month inside the 12 | `LOOKUP_HOURS` | that month itemised, with `T-` ids and `CP-` ids |
| `{"action":"lookup","what":"counterparty","id":"CP-7"}` | a `CP-` id that appeared on screen or in a fetched month | `LOOKUP_HOURS` | name, type, jurisdiction, first seen, **number of other alerted accounts sharing this counterparty in 12 months** |
| `{"action":"lookup","what":"prior_alerts"}` | — | `LOOKUP_HOURS` | prior alerts: rule, date, disposition, closing analyst, rationale snippet |
| `{"action":"lookup","what":"documents"}` | — | `DOC_HOURS` | documents on file with `doc.` ids and one-line descriptions |
| `{"action":"close","reason":"CONSISTENT_WITH_PROFILE","cites":["kyc.expected_monthly_cash"],"note":"…"}` | `reason` ∈ {`CONSISTENT_WITH_PROFILE`, `DOCUMENTED_SOURCE_OF_FUNDS`, `PREVIOUSLY_REVIEWED`, `RULE_ARTEFACT`}; `cites` a non-empty list of ids; `note` free text | ends episode | — |
| `{"action":"escalate","typology":"STRUCTURING","cites":["T-101","T-102"],"narrative":"…"}` | `typology` ∈ {`STRUCTURING`, `FUNNEL_ACCOUNT`, `MONEY_MULE`, `RAPID_MOVEMENT`, `UNKNOWN`}; `cites` non-empty; `narrative` free text | ends episode | — |
Citable ids: `kyc.<field>` for any field shown in the KYC block, `doc.<id>`,
`T-<n>`, `CP-<n>`, `A-<n>` (a prior alert). An id the agent has not been
shown is accepted and simply will not match anything planted — the grader does
not need to know what the agent saw, only what it cited.
**Malformed replies.** No parseable object, an unknown `action`, an enum value
outside its set, an empty `cites`, a month outside the window, a `CP-` id
that does not exist, a lookup already made: all are **rejected**. A rejection
costs a turn and `TURN_HOURS`, is counted in `rejected_replies`, returns a
one-line reason in the format the prompt promised, and never raises. The
engine is pure and cannot crash on input. A repeated lookup is rejected rather
than re-charged so the transcript stays a record of distinct things learned.
**Hour constants** (the generator's config, shown on the page as such):
`SCREEN_HOURS 0.50`, `TURN_HOURS 0.05`, `LOOKUP_HOURS 0.25`,
`DOC_HOURS 0.10`. These are **assumptions, not measurements**, and the page
must not present them as industry figures (§8, risk 2). The implement stage
may tune them to make the ladder's ordering hold, and must record the values
it ships in `reward.py` where the page quotes them.
## 4. The grader
Everything below is a pure function of `(world, transcript)`. No model, no
human, no prose is read anywhere in it.
From the world (regenerated from the seed): `label`, `typology`, `planted`.
From the transcript: the accepted actions in order, the count of rejected
replies, and the final disposition.
```
hours_spent = SCREEN_HOURS + TURN_HOURS × turns_spent + Σ lookup prices
correct = (label == benign and disposition == close)
or (label == suspicious and disposition == escalate)
```
**`planted`** is the minimal set of ids that resolves the alert, fixed by the
template at generation time:
- benign: the explaining fact(s) — one or two `kyc.`/`doc.` ids (table in §1);
- suspicious: the transaction ids that constitute the typology (S1: the eight
deposits; S2: the outbound wires to the shared counterparties; S3: the
in/out pairs; and so on).
Where a template admits two reasonable rationales (B4 can be closed on the KYC
expectation or on a prior closed alert with the same rationale), `planted` is
a list of **alternates**, and the score takes the best-matching one. Still
deterministic; less opinionated.
**`evidence`** = F1 between `cites` and the best alternate of `planted`. If
`correct` is false, `evidence` is 0 — a false escalation is checked against an
empty set, a missed one against facts that do not exist for that label.
Cite-everything scores what its precision earns (one planted id among eight
cited is F1 0.22); cite-nothing is rejected at parse.
**`reference_hours`** — the `hours` denominator — is *the hours of the cheapest
shipped policy that reached the correct disposition on this seed*, computed by
running the two shipped reference policies (§6) through the engine at
generation time, exactly as wordle computes `reference_depth` by running the
solver. On benign and visible seeds that is `fast`; on hidden seeds it is
`thorough`. It is stored in the trace as a metric so the browser can verify
without running the policies, and its absence makes a run *unverifiable*, never
zero.
**Conformance.** Two digests, both gated in CI: (a) SHA-256 over the canonical
JSON of the generated world for seeds 04095 — the browser must show the same
alert for `?seed=` as the Python generated; (b) SHA-256 over the scorer applied
to a fixed cross-product of synthetic dispositions (every template × every
disposition × a grid of citation sets and hour totals). The generator has to be
ported, not just the scorer, because the Play tab needs it.
## 5. The reward components
Weights sum to 1.0. One counterweight.
| key | label (exec-facing) | weight | role | definition |
|---|---|---|---|---|
| `caught` | Nothing slipped through | **0.55** | counterweight | 1.0 if a disposition was reached and either the alert was benign or it was escalated. 0.0 if the alert was suspicious and closed, or no disposition was reached. |
| `hours` | Cleared it in the hours budgeted | **0.25** | objective | `min(1, reference_hours / hours_spent)` if `correct`, else 0.0. A fast wrong answer saves nothing. |
| `evidence` | Named the facts that decide it | **0.20** | objective | F1 of `cites` against `planted` (best alternate), 0.0 if the disposition was wrong. The prose is shown, never scored. |
**Why this is the asymmetry the vertical promised.** A missed escalation zeroes
`caught` (0.55), and because the call was wrong it also zeroes `hours` and
`evidence`: the episode scores 0.000. The most any policy can gain by being
fast is `hours`'s 0.25. One miss costs more than every hour saved, and you can
read that off the weights rather than off a policy memo. The reward editor is
where a BSA officer argues with it.
**The tension, named.**
- *The crude policy that maxes the objective:* close on turn one, read
nothing. It takes `hours` (1.0 on every benign seed — it spent less than any
reference) and is stopped by `caught`, which it scores at the benign base
rate and no higher.
- *The crude policy that maxes the counterweight:* escalate on turn one,
everything. `caught` 1.0, stopped by `hours` (0 on every benign seed — wrong
call) and `evidence` (0 on every benign seed).
- *The good policy that pays the counterweight's price:* **`fast`** — decides
from the screen and the KYC file, gets every benign and every visible case
right, cites correctly, spends `SCREEN_HOURS` and one turn. It defines the
`hours` denominator on 85% of seeds and scores 1.0 there. It structurally
cannot see a funnel counterparty or a serial-closer history, so it scores
`caught` ≈ 0.85 and is not 1.0.
- *The good policy that takes the counterweight outright:* **`thorough`** —
pulls documents, prior alerts, every counterparty on screen and the itemised
peak months, then decides with everything. `caught` 1.0, `evidence` 1.0,
`hours` ≈ 0.4 (it spends roughly four times the screen cost on the 85% of
seeds where the screen was enough).
Neither dominates: `thorough` wins `caught` and `evidence`; `fast` wins
`hours`. At the shipped weights `thorough` is intended to edge `fast` by a
small margin (≈0.86 vs ≈0.85 on the rough arithmetic); moving `hours` to 0.35
flips it. That the flip is one slider away is the demonstration.
**Metrics** (weight zero — rendered, never summed): `hours_spent`,
`reference_hours`, `lookups`, `turns_used`, `rejected_replies`,
`false_escalation` (0/1), `missed` (0/1), `typology_match` (1/0, `null` on a
close), `label` (rendered as the reveal after the episode ends).
`typology` and `reason` are required for a well-formed action and are
otherwise unscored. `typology_match` is a diagnostic. Scoring it would
double-count `caught` with an enum the agent can guess.
## 6. The probe ladder
Run over 24 seeds; the probe asserts the seed set contains at least one seed of
each tier (benign, visible, hidden) or aborts before measuring anything.
| policy | behaviour | expected |
|---|---|---|
| `inaction` | reasons in prose, never emits an object | **0.000** on every component, exactly |
| `malformed` | emits an object every turn with an unknown action or an empty `cites` | **0.000** exactly; `rejected_replies` = 8 |
| `close_blind` | turn one: `close`, `reason` CONSISTENT_WITH_PROFILE, cites `kyc.business_type` | `caught` = benign share (≈0.55); `hours` = benign share; `evidence` small; total ≈ 0.45 |
| `close_cites_all` | turn one: `close` citing every KYC field and every triggering transaction | same as above with `evidence` ≈ 0.55 × 0.2; must score strictly below `fast` |
| `escalate_blind` | turn one: `escalate`, typology from the rule's nominal mapping, cites all triggering transactions | `caught` 1.0; `hours` = suspicious share (≈0.45); `evidence` ≈ 0.3; total ≈ 0.7; must score strictly below both good policies |
| `fast` | decides from the screen + KYC; closes when the KYC or a screen-visible fact explains the alert, escalates when the screen shows a typology; cites the explaining facts / constituting transactions | `hours` 1.0 on non-hidden seeds, 0 on hidden; `caught` ≈ 0.85 and **< 1.0**; `evidence` ≈ 0.85 |
| `thorough` | lookups: documents, prior_alerts, every `CP-` on screen, itemised history for the two peak months; then decides with everything | `caught` **= 1.0 exactly**; `evidence` = 1.0; `hours` ≈ 0.4 and **< fast's** |
| `thorough_wasteful` | pulls all seven lookups the turn budget allows — including months that carry nothing — then decides exactly as `thorough` | same `caught` and `evidence` as `thorough`; total **strictly below** `thorough` — proves the hour budget binds |
Assertions, in the style of `envs/probe.py`:
1. `inaction` and `malformed` total exactly 0.000.
2. `thorough` scores exactly 1.0 on `caught` — the counterweight is attainable.
3. `fast` scores strictly below 1.0 on `caught` — the counterweight bites the
policy that defines the objective's denominator.
4. `fast` strictly beats `thorough` on `hours` — paying the counterweight's
price is a real price.
5. Neither `fast` nor `thorough` dominates the other on every component.
6. `escalate_blind`, `close_blind` and `close_cites_all` each total strictly
below both good policies. A cheat the reviewers name gets a rung here.
7. `thorough_wasteful` totals strictly below `thorough`.
8. No weighted component is flat across the ladder.
Note `hours` is scored over turns **spent**, rejected replies included, via
`TURN_HOURS`, so a policy that jams the parser and then decides does not read
as cheaper than one that decided cleanly.
## 7. The held-out slice
Three cuts, all deterministic in the seed:
1. **Seed bucket.** `fnv1a(str(seed)) % 8 == 7` is held out — never captured,
never probed, never in a training sweep. About 12.5% of seeds.
2. **Moved threshold.** Within the held-out bucket, structuring and cash
templates (B4, S1, S2, S6) are generated in a second currency with that
jurisdiction's reporting threshold, so the band is not 8,0009,999 USD, and
the rule ids are renamed. An agent that learned "9,xxx and R-STR-01" rather
than "just under the threshold, repeatedly, across branches" is caught
here. This is the honest form of the wordle page's "nobody is pushing back"
limit: the adversary has *moved*, not responded (§8, risk 4).
3. **Re-ordered screen.** The held-out screen renders the KYC block before the
alert and the summary in reverse chronological order. Format memorisation
is not triage.
The page reports the training-slice and held-out numbers side by side, and
never one as the other.
## 8. Risks — stated, not hidden
1. **The ground truth is what the generator planted, and that is easier than
reality in three ways.** Every alert has a knowable answer; real SAR
decisions are human judgements and a filed SAR is not a proven crime. Eleven
templates stand in for a typology space in the hundreds. And the 55/45 mix
is far from the real false-positive rate, which is widely reported to be
above ninety percent; we do not re-quote a figure, we state that our mix is
chosen so every arm meets both labels in a small seed set. The page must say
that "hours per thousand alerts" here is a ratio inside this environment
and not a production number.
2. **The hour constants are assumptions.** 0.50h to read a screen, 0.25h per
lookup, 0.05h per turn. They are shown as the generator's config. No
industry figure is cited for them because we have not measured one.
3. **The narrative prose is unscored.** Only the disposition, the citations and
(as a metric) the typology are graded. A fluent wrong narrative and a terse
right one score the same, and the page says so beside the narrative. This is
the `legal-playbook-redline` move: score the checkable core, display the
judgement half, never put a judge behind a bar.
4. **The adversary is static.** The launderer shaped the data to pass the
screen, and the held-out slice moves the threshold, but nothing responds to
the agent's policy during training. The wordle page's limit "nobody is
pushing back… no counterparty adapting" names this demo as its answer; the
demo stage should claim only what §1 and §7 deliver — an adversary who
designed the data to pass, and moved once — and the ship stage should
consider softening that wordle limit's wording. (That file belongs to
another stage; not edited here.)
5. **The hidden tier is hidden only if the generator keeps it so.** If the
implement stage lets any tell leak into the free screen — a different name
distribution, a rounder amount, a branch number — a policy tuned to the
leak catches S2/S5 from the screen and the tension collapses. Mitigation is
structural (S2 and S5 are B4 and B2's generator with a lookup-only overlay)
and assertion 3 catches a leak the shipped `fast` policy could exploit, but
not one only an RL agent would find. A future review should try to
classify hidden-vs-benign from the screen alone and expect chance.
6. **`caught` is 1.0 on every benign alert by construction.** On a single
benign seed, a close-everything run looks perfect. The page must argue the
counterweight from the ladder over seeds, never from one run — the same
caveat as a lucky first guess in wordle.
7. **The planted set is an opinion about what "the" evidence is.** Alternates
soften it; they do not remove it. Where a reviewer finds a rationale the
grader refuses that a BSA officer would accept, the fix is another alternate
in the template, not a judge.
8. **Capture may find the weak arm cannot emit JSON.** Qwen3-8B without
thinking may reject its way to 0.000 on most seeds. That is a finding to
show, not a bug to hide, but a ladder with three arms at 0.000 teaches
nothing — the capture stage must include `fast` and `thorough` as
`generated` arms so the reward has something to rank.
9. **Terminology is real, cases are not.** BSA, SAR, FinCEN, FATF and the
typology names are public vocabulary and are used correctly. No scenario
is drawn from a real case, and no institution is named or implied.
## 9. What the next stages need from this
- `envs/alert_triage/pyproject.toml` (`alert-triage`, workspace member),
package `alert_triage/` with `generator.py` (templates, xorshift32, FNV-1a
seed), `engine.py` (state, lookups, dispositions, never raises),
`protocol.py` (JSON parse, renderers, system prompt), `reward.py` (with the
single `# region: pig-demo/reward` pair), `policies.py` (`fast`,
`thorough`), `taskset.py`, `tests/`, and `names/PROVENANCE.md`.
- Probe policies registered for taskset `alert-triage`; eval command
`uv run python envs/probe.py`.
- The TypeScript port carries the generator, not only the scorer.
- The demo's surface renders a **"Synthetic data — no real customer, account or
transaction"** label on the board itself, the hour constants as config, the
narrative with "not scored" beside it, and the label reveal only after a
disposition.
- Interactive mode: yes. A person can read the screen, click the four lookups
(each showing its hour price), and close with a citation checklist or
escalate with a typology and a transaction tick-list.
- Suggested `caveat` for the vertical entry at ship time: "The alerts are
synthetic and every one has a planted answer. The narrative prose is shown
and not scored; the disposition and the cited facts are. The adversary
designed the data and moved the threshold once — it does not adapt to the
agent."