From 1239dc70348e1c0d196058cd64e87b3e475711c2 Mon Sep 17 00:00:00 2001 From: karti-ai <176560021+karti-ai@users.noreply.github.com> Date: Fri, 28 Aug 2026 18:32:29 -0700 Subject: [PATCH] new-environment pipeline: scripted arms required, model arms captured when spark-1 is up MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit An environment ships on its scripted agents alone — they are what the reward flip is built on and they need no model. A recorded model attempt is captured once from spark-1 for the Watch tab when it is serving, and the environment does not block on it when it is not. --- .claude/workflows/new-environment.js | 10 +- envs/alert_triage/SPEC.md | 375 +++++++++++++++++++++++++++ 2 files changed, 383 insertions(+), 2 deletions(-) create mode 100644 envs/alert_triage/SPEC.md diff --git a/.claude/workflows/new-environment.js b/.claude/workflows/new-environment.js index f9c40d8..d823a05 100644 --- a/.claude/workflows/new-environment.js +++ b/.claude/workflows/new-environment.js @@ -154,11 +154,17 @@ const captured = await agent(`${BASE} ## YOUR STAGE: real rollouts. -Extend \`envs/capture.py\` for taskset "${slug}" and capture seeds ${seeds} for at least three arms: a weak one, a strong one, and a generated reference policy. +Extend \`envs/capture.py\` for taskset "${slug}". + +**Two kinds of arm, with different rules.** + +*Scripted arms are REQUIRED* — the reference policy and at least one weaker scripted policy (the probe ladder's rungs are the obvious candidates). They are generated by code, need no model, and are what the Reward tab's ranking flip is built on. An environment ships with these alone. + +*Model arms are an ENHANCEMENT* — a real model's recorded attempt for the Watch tab, captured once from spark-1. Check it first: \`curl http://100.127.247.67:8001/v1/models\`. If it is serving, capture seeds ${seeds} for a thinking-off and a thinking-on arm. If it is not, ship without them, say so plainly in your report, and leave a note in \`envs/${pkg}/README.md\` that model arms are still to be captured — do NOT block the environment on it and do NOT invent fixtures. ⚠️ **spark-1 serves one model and is single-stream — run the arms one after another, never concurrently.** Check it is up first: \`curl http://100.127.247.67:8001/v1/models\`. A thinking-enabled arm can take minutes per seed; run it in the background, redirect output to a FILE (piping to head/tail block-buffers Python's stdout and a working run looks hung), and poll the written fixtures rather than the log. -Include at least one run the agent loses. Then \`uv run python envs/build_manifest.py\`. +Include at least one run an agent loses. Then \`uv run python envs/build_manifest.py\`. **Your gate:** \`uv run python envs/verify_fixtures.py\` — every fixture replays and reproduces its own recorded rewards at delta 0. diff --git a/envs/alert_triage/SPEC.md b/envs/alert_triage/SPEC.md new file mode 100644 index 0000000..0c9e2c9 --- /dev/null +++ b/envs/alert_triage/SPEC.md @@ -0,0 +1,375 @@ +# alert-triage — specification + +Vertical: `financial-crime-alert-triage` (`src/content/verticals.ts`, rank 4). +Package: `envs/alert_triage/`. Demo: `src/demos/alert-triage/`. Traces: +`public/traces/alert-triage/`. + +This document is written before any code and is meant to be attacked. The two +things it has to survive: a sceptic calling the grader subjective, and a +sceptic finding a policy that maxes the reward while doing something a BSA +officer would fire an analyst for. + +**In one paragraph.** One episode is one transaction-monitoring alert on one +synthetic customer. The agent sees what the analyst's screen shows — the alert, +the triggering transactions, the KYC file, a twelve-month summary — and may pay +analyst-hours for lookups (itemised history, a counterparty, prior alerts, +documents on file). It ends the episode by *closing* the alert with a cited +rationale or *escalating* it for a suspicious-activity filing with a typology +and the transactions that constitute it. The generator planted the answer: the +alert is either a benign scenario with a specific explaining fact, or a +laundering typology with a specific set of transactions. The grader compares +the disposition and the citations to what was planted, and charges the hours. +Nothing in it reads prose. + +--- + +## 1. The episode + +**One alert.** Deterministic in the seed. The generator draws a scenario +template from the seed (FNV-1a over the decimal seed selects the template and +seeds a hand-written xorshift32 for its parameters — never a built-in RNG, +see AGENTS.md), and produces: + +- the **customer** (`C-####`): individual or business, occupation/business + category, stated income or revenue, expected monthly cash, country, account + opened, KYC last refreshed, PEP flag, prior alert and prior SAR counts; +- twelve months of **account history**: a monthly summary table (cash in/out, + wires in/out, ACH in/out, transaction count), and the itemised transactions + behind it, each with a stable id `T-###`, a date, a channel, an amount, a + branch or counterparty id `CP-##`; +- the **alert** (`A-####`): the rule that fired (`R-STR-01` structuring, + `R-VEL-02` rapid movement, `R-CASH-03` cash above KYC expectation, + `R-WIRE-04` wires to a high-risk jurisdiction, `R-DORM-05` dormant-account + reactivation), the fire date, and the triggering transactions itemised; +- the **hidden truth**, never shown: `label ∈ {benign, suspicious}`, the + `typology` if suspicious, and the `planted` citation set (§4). + +**The free screen.** Turn 0 shows the alert, its triggering transactions, the +KYC block and the monthly summary. This costs a fixed `SCREEN_HOURS` — the +analyst has to read it whatever they do next. + +**Turns.** Up to `MAX_TURNS = 8` replies. Each reply is one lookup or one +disposition (§3). Each turn spent — accepted *or* rejected — costs +`TURN_HOURS`; a lookup adds its own price. + +**What ends it.** An accepted `close` or `escalate`, or the eighth turn +without one. The latter is `outcome: aborted` and scores 0.000 on every +component: an alert nobody disposed of is still in the queue, which is the +worst outcome for throughput and, on a suspicious alert, is a miss. + +**Scenario templates.** Six benign, five suspicious. Every template is written +from public typology descriptions (FinCEN SAR narrative guidance, FATF +typology reports); none is a real case. + +| id | label | rule that fires | what the screen shows | where the answer lives | +|---|---|---|---|---| +| B1 seasonal cash | benign | R-CASH-03 | landscaping / tax-prep business, cash spikes in season | `kyc.expected_activity` says seasonal, `kyc.expected_monthly_cash` covers peak | +| B2 processor settlement | benign | R-VEL-02 | business receives processor wires, pays contractors within 48h | `doc.processor_agreement` on file | +| B3 property sale | benign | R-VEL-02 / R-DORM-05 | individual receives title-company wire, moves it to brokerage | `doc.closing_statement` | +| B4 cash restaurant | benign | R-STR-01 | daily cash deposits, three happen to land in 8,000–9,999 | `kyc.expected_monthly_cash` ≥ observed; 12-month cash profile flat | +| B5 family remittance | benign | R-WIRE-04 | monthly wire to one named individual in a high-risk country | `kyc.remittance_declared` names the destination | +| B6 inheritance | benign | R-DORM-05 | dormant account receives estate wire | `doc.probate_letter` | +| S1 structuring | suspicious, **visible** | R-STR-01 | 45k stated income; 8 cash deposits 9,200–9,900 in 12 days across 3 branches | the 8 deposits | +| S3 money mule | suspicious, **visible** | R-VEL-02 | student/retiree; 4 inbound P2P from unrelated individuals, each out within 24h to an exchange | the in/out pairs | +| S6 smurfing over time | suspicious, **visible** | R-CASH-03 | cash business; monthly cash tripled over 12 months and no single deposit ever exceeded 9,999 | the deposits in the last 3 months | +| S2 funnel account | suspicious, **hidden** | R-STR-01 or R-CASH-03 | *generated by B4's template* — screen is identically distributed to a benign restaurant | `counterparty` lookup: its 3 outbound-wire counterparties are shared with 6–14 other alerted accounts; the wires are the evidence | +| S5 serial closer | suspicious, **hidden** | R-VEL-02 | *generated by B2's template* — processor agreement on file, contractors paid | `prior_alerts` lookup: 3 prior alerts in 9 months, each closed by a different analyst with the same boilerplate; `documents` shows the agreement's counterparty is not the payer | + +The **hidden** tier is the environment's whole argument. S2 and S5 are +produced by *the same generator code* as B4 and B2 respectively, with the +suspicious signal overlaid only in the data a lookup returns. The free screen +for a hidden case is therefore not merely similar to a benign one; it is drawn +from the same distribution. A policy that decides from the screen cannot tell +them apart, by construction, not by our failing to try hard enough. + +Mix by seed: 55% benign, 30% visible-suspicious, 15% hidden. This is +deliberately **not** the real base rate — real programmes close the great +majority of alerts as false positives, and at that rate a 24-seed probe would +contain about one suspicious alert and the ladder would measure noise. The page +must say this (§8, risk 1). + +## 2. The data + +**Synthetic, entirely, and labelled synthetic on the page — in the surface, +not a footnote.** No real customer, account, transaction, counterparty or +analyst exists behind any of it. Names come from a synthetic name list +committed under `envs/alert_triage/names/` with its own `PROVENANCE.md`; +business names are composed from committed word lists ("Harbor Lane Bistro +LLC"). Amounts, dates, branch numbers and counterparties are generated. + +Licence: the generator and its lists are ours, Apache-2.0 with the repository. +The public documents the templates are *written from* — FinCEN's SAR narrative +guidance and FATF typology reports — are cited as references on the evidence +tab, not shipped, and not represented as data. + +Why synthetic is the only honest option here: real alert data is customer +financial data under BSA confidentiality; a SAR's existence may not be +disclosed. There is no licensable real dataset for this task, and any page +implying one would be lying. + +## 3. The action set + +Every reply is exactly **one JSON object** — the first `{…}` the parser can +balance, or the contents of the first ```` ```json ```` fence if present. The +system prompt says so and shows the four shapes. + +| action | fields | cost | what comes back | +|---|---|---|---| +| `{"action":"lookup","what":"history","month":"2026-01"}` | a month inside the 12 | `LOOKUP_HOURS` | that month itemised, with `T-` ids and `CP-` ids | +| `{"action":"lookup","what":"counterparty","id":"CP-7"}` | a `CP-` id that appeared on screen or in a fetched month | `LOOKUP_HOURS` | name, type, jurisdiction, first seen, **number of other alerted accounts sharing this counterparty in 12 months** | +| `{"action":"lookup","what":"prior_alerts"}` | — | `LOOKUP_HOURS` | prior alerts: rule, date, disposition, closing analyst, rationale snippet | +| `{"action":"lookup","what":"documents"}` | — | `DOC_HOURS` | documents on file with `doc.` ids and one-line descriptions | +| `{"action":"close","reason":"CONSISTENT_WITH_PROFILE","cites":["kyc.expected_monthly_cash"],"note":"…"}` | `reason` ∈ {`CONSISTENT_WITH_PROFILE`, `DOCUMENTED_SOURCE_OF_FUNDS`, `PREVIOUSLY_REVIEWED`, `RULE_ARTEFACT`}; `cites` a non-empty list of ids; `note` free text | ends episode | — | +| `{"action":"escalate","typology":"STRUCTURING","cites":["T-101","T-102"],"narrative":"…"}` | `typology` ∈ {`STRUCTURING`, `FUNNEL_ACCOUNT`, `MONEY_MULE`, `RAPID_MOVEMENT`, `UNKNOWN`}; `cites` non-empty; `narrative` free text | ends episode | — | + +Citable ids: `kyc.` for any field shown in the KYC block, `doc.`, +`T-`, `CP-`, `A-` (a prior alert). An id the agent has not been +shown is accepted and simply will not match anything planted — the grader does +not need to know what the agent saw, only what it cited. + +**Malformed replies.** No parseable object, an unknown `action`, an enum value +outside its set, an empty `cites`, a month outside the window, a `CP-` id +that does not exist, a lookup already made: all are **rejected**. A rejection +costs a turn and `TURN_HOURS`, is counted in `rejected_replies`, returns a +one-line reason in the format the prompt promised, and never raises. The +engine is pure and cannot crash on input. A repeated lookup is rejected rather +than re-charged so the transcript stays a record of distinct things learned. + +**Hour constants** (the generator's config, shown on the page as such): +`SCREEN_HOURS 0.50`, `TURN_HOURS 0.05`, `LOOKUP_HOURS 0.25`, +`DOC_HOURS 0.10`. These are **assumptions, not measurements**, and the page +must not present them as industry figures (§8, risk 2). The implement stage +may tune them to make the ladder's ordering hold, and must record the values +it ships in `reward.py` where the page quotes them. + +## 4. The grader + +Everything below is a pure function of `(world, transcript)`. No model, no +human, no prose is read anywhere in it. + +From the world (regenerated from the seed): `label`, `typology`, `planted`. +From the transcript: the accepted actions in order, the count of rejected +replies, and the final disposition. + +``` +hours_spent = SCREEN_HOURS + TURN_HOURS × turns_spent + Σ lookup prices +correct = (label == benign and disposition == close) + or (label == suspicious and disposition == escalate) +``` + +**`planted`** is the minimal set of ids that resolves the alert, fixed by the +template at generation time: + +- benign: the explaining fact(s) — one or two `kyc.`/`doc.` ids (table in §1); +- suspicious: the transaction ids that constitute the typology (S1: the eight + deposits; S2: the outbound wires to the shared counterparties; S3: the + in/out pairs; and so on). + +Where a template admits two reasonable rationales (B4 can be closed on the KYC +expectation or on a prior closed alert with the same rationale), `planted` is +a list of **alternates**, and the score takes the best-matching one. Still +deterministic; less opinionated. + +**`evidence`** = F1 between `cites` and the best alternate of `planted`. If +`correct` is false, `evidence` is 0 — a false escalation is checked against an +empty set, a missed one against facts that do not exist for that label. +Cite-everything scores what its precision earns (one planted id among eight +cited is F1 0.22); cite-nothing is rejected at parse. + +**`reference_hours`** — the `hours` denominator — is *the hours of the cheapest +shipped policy that reached the correct disposition on this seed*, computed by +running the two shipped reference policies (§6) through the engine at +generation time, exactly as wordle computes `reference_depth` by running the +solver. On benign and visible seeds that is `fast`; on hidden seeds it is +`thorough`. It is stored in the trace as a metric so the browser can verify +without running the policies, and its absence makes a run *unverifiable*, never +zero. + +**Conformance.** Two digests, both gated in CI: (a) SHA-256 over the canonical +JSON of the generated world for seeds 0–4095 — the browser must show the same +alert for `?seed=` as the Python generated; (b) SHA-256 over the scorer applied +to a fixed cross-product of synthetic dispositions (every template × every +disposition × a grid of citation sets and hour totals). The generator has to be +ported, not just the scorer, because the Play tab needs it. + +## 5. The reward components + +Weights sum to 1.0. One counterweight. + +| key | label (exec-facing) | weight | role | definition | +|---|---|---|---|---| +| `caught` | Nothing slipped through | **0.55** | counterweight | 1.0 if a disposition was reached and either the alert was benign or it was escalated. 0.0 if the alert was suspicious and closed, or no disposition was reached. | +| `hours` | Cleared it in the hours budgeted | **0.25** | objective | `min(1, reference_hours / hours_spent)` if `correct`, else 0.0. A fast wrong answer saves nothing. | +| `evidence` | Named the facts that decide it | **0.20** | objective | F1 of `cites` against `planted` (best alternate), 0.0 if the disposition was wrong. The prose is shown, never scored. | + +**Why this is the asymmetry the vertical promised.** A missed escalation zeroes +`caught` (0.55), and because the call was wrong it also zeroes `hours` and +`evidence`: the episode scores 0.000. The most any policy can gain by being +fast is `hours`'s 0.25. One miss costs more than every hour saved, and you can +read that off the weights rather than off a policy memo. The reward editor is +where a BSA officer argues with it. + +**The tension, named.** + +- *The crude policy that maxes the objective:* close on turn one, read + nothing. It takes `hours` (1.0 on every benign seed — it spent less than any + reference) and is stopped by `caught`, which it scores at the benign base + rate and no higher. +- *The crude policy that maxes the counterweight:* escalate on turn one, + everything. `caught` 1.0, stopped by `hours` (0 on every benign seed — wrong + call) and `evidence` (0 on every benign seed). +- *The good policy that pays the counterweight's price:* **`fast`** — decides + from the screen and the KYC file, gets every benign and every visible case + right, cites correctly, spends `SCREEN_HOURS` and one turn. It defines the + `hours` denominator on 85% of seeds and scores 1.0 there. It structurally + cannot see a funnel counterparty or a serial-closer history, so it scores + `caught` ≈ 0.85 and is not 1.0. +- *The good policy that takes the counterweight outright:* **`thorough`** — + pulls documents, prior alerts, every counterparty on screen and the itemised + peak months, then decides with everything. `caught` 1.0, `evidence` 1.0, + `hours` ≈ 0.4 (it spends roughly four times the screen cost on the 85% of + seeds where the screen was enough). + +Neither dominates: `thorough` wins `caught` and `evidence`; `fast` wins +`hours`. At the shipped weights `thorough` is intended to edge `fast` by a +small margin (≈0.86 vs ≈0.85 on the rough arithmetic); moving `hours` to 0.35 +flips it. That the flip is one slider away is the demonstration. + +**Metrics** (weight zero — rendered, never summed): `hours_spent`, +`reference_hours`, `lookups`, `turns_used`, `rejected_replies`, +`false_escalation` (0/1), `missed` (0/1), `typology_match` (1/0, `null` on a +close), `label` (rendered as the reveal after the episode ends). + +`typology` and `reason` are required for a well-formed action and are +otherwise unscored. `typology_match` is a diagnostic. Scoring it would +double-count `caught` with an enum the agent can guess. + +## 6. The probe ladder + +Run over 24 seeds; the probe asserts the seed set contains at least one seed of +each tier (benign, visible, hidden) or aborts before measuring anything. + +| policy | behaviour | expected | +|---|---|---| +| `inaction` | reasons in prose, never emits an object | **0.000** on every component, exactly | +| `malformed` | emits an object every turn with an unknown action or an empty `cites` | **0.000** exactly; `rejected_replies` = 8 | +| `close_blind` | turn one: `close`, `reason` CONSISTENT_WITH_PROFILE, cites `kyc.business_type` | `caught` = benign share (≈0.55); `hours` = benign share; `evidence` small; total ≈ 0.45 | +| `close_cites_all` | turn one: `close` citing every KYC field and every triggering transaction | same as above with `evidence` ≈ 0.55 × 0.2; must score strictly below `fast` | +| `escalate_blind` | turn one: `escalate`, typology from the rule's nominal mapping, cites all triggering transactions | `caught` 1.0; `hours` = suspicious share (≈0.45); `evidence` ≈ 0.3; total ≈ 0.7; must score strictly below both good policies | +| `fast` | decides from the screen + KYC; closes when the KYC or a screen-visible fact explains the alert, escalates when the screen shows a typology; cites the explaining facts / constituting transactions | `hours` 1.0 on non-hidden seeds, 0 on hidden; `caught` ≈ 0.85 and **< 1.0**; `evidence` ≈ 0.85 | +| `thorough` | lookups: documents, prior_alerts, every `CP-` on screen, itemised history for the two peak months; then decides with everything | `caught` **= 1.0 exactly**; `evidence` = 1.0; `hours` ≈ 0.4 and **< fast's** | +| `thorough_wasteful` | pulls all seven lookups the turn budget allows — including months that carry nothing — then decides exactly as `thorough` | same `caught` and `evidence` as `thorough`; total **strictly below** `thorough` — proves the hour budget binds | + +Assertions, in the style of `envs/probe.py`: + +1. `inaction` and `malformed` total exactly 0.000. +2. `thorough` scores exactly 1.0 on `caught` — the counterweight is attainable. +3. `fast` scores strictly below 1.0 on `caught` — the counterweight bites the + policy that defines the objective's denominator. +4. `fast` strictly beats `thorough` on `hours` — paying the counterweight's + price is a real price. +5. Neither `fast` nor `thorough` dominates the other on every component. +6. `escalate_blind`, `close_blind` and `close_cites_all` each total strictly + below both good policies. A cheat the reviewers name gets a rung here. +7. `thorough_wasteful` totals strictly below `thorough`. +8. No weighted component is flat across the ladder. + +Note `hours` is scored over turns **spent**, rejected replies included, via +`TURN_HOURS`, so a policy that jams the parser and then decides does not read +as cheaper than one that decided cleanly. + +## 7. The held-out slice + +Three cuts, all deterministic in the seed: + +1. **Seed bucket.** `fnv1a(str(seed)) % 8 == 7` is held out — never captured, + never probed, never in a training sweep. About 12.5% of seeds. +2. **Moved threshold.** Within the held-out bucket, structuring and cash + templates (B4, S1, S2, S6) are generated in a second currency with that + jurisdiction's reporting threshold, so the band is not 8,000–9,999 USD, and + the rule ids are renamed. An agent that learned "9,xxx and R-STR-01" rather + than "just under the threshold, repeatedly, across branches" is caught + here. This is the honest form of the wordle page's "nobody is pushing back" + limit: the adversary has *moved*, not responded (§8, risk 4). +3. **Re-ordered screen.** The held-out screen renders the KYC block before the + alert and the summary in reverse chronological order. Format memorisation + is not triage. + +The page reports the training-slice and held-out numbers side by side, and +never one as the other. + +## 8. Risks — stated, not hidden + +1. **The ground truth is what the generator planted, and that is easier than + reality in three ways.** Every alert has a knowable answer; real SAR + decisions are human judgements and a filed SAR is not a proven crime. Eleven + templates stand in for a typology space in the hundreds. And the 55/45 mix + is far from the real false-positive rate, which is widely reported to be + above ninety percent; we do not re-quote a figure, we state that our mix is + chosen so every arm meets both labels in a small seed set. The page must say + that "hours per thousand alerts" here is a ratio inside this environment + and not a production number. +2. **The hour constants are assumptions.** 0.50h to read a screen, 0.25h per + lookup, 0.05h per turn. They are shown as the generator's config. No + industry figure is cited for them because we have not measured one. +3. **The narrative prose is unscored.** Only the disposition, the citations and + (as a metric) the typology are graded. A fluent wrong narrative and a terse + right one score the same, and the page says so beside the narrative. This is + the `legal-playbook-redline` move: score the checkable core, display the + judgement half, never put a judge behind a bar. +4. **The adversary is static.** The launderer shaped the data to pass the + screen, and the held-out slice moves the threshold, but nothing responds to + the agent's policy during training. The wordle page's limit "nobody is + pushing back… no counterparty adapting" names this demo as its answer; the + demo stage should claim only what §1 and §7 deliver — an adversary who + designed the data to pass, and moved once — and the ship stage should + consider softening that wordle limit's wording. (That file belongs to + another stage; not edited here.) +5. **The hidden tier is hidden only if the generator keeps it so.** If the + implement stage lets any tell leak into the free screen — a different name + distribution, a rounder amount, a branch number — a policy tuned to the + leak catches S2/S5 from the screen and the tension collapses. Mitigation is + structural (S2 and S5 are B4 and B2's generator with a lookup-only overlay) + and assertion 3 catches a leak the shipped `fast` policy could exploit, but + not one only an RL agent would find. A future review should try to + classify hidden-vs-benign from the screen alone and expect chance. +6. **`caught` is 1.0 on every benign alert by construction.** On a single + benign seed, a close-everything run looks perfect. The page must argue the + counterweight from the ladder over seeds, never from one run — the same + caveat as a lucky first guess in wordle. +7. **The planted set is an opinion about what "the" evidence is.** Alternates + soften it; they do not remove it. Where a reviewer finds a rationale the + grader refuses that a BSA officer would accept, the fix is another alternate + in the template, not a judge. +8. **Capture may find the weak arm cannot emit JSON.** Qwen3-8B without + thinking may reject its way to 0.000 on most seeds. That is a finding to + show, not a bug to hide, but a ladder with three arms at 0.000 teaches + nothing — the capture stage must include `fast` and `thorough` as + `generated` arms so the reward has something to rank. +9. **Terminology is real, cases are not.** BSA, SAR, FinCEN, FATF and the + typology names are public vocabulary and are used correctly. No scenario + is drawn from a real case, and no institution is named or implied. + +## 9. What the next stages need from this + +- `envs/alert_triage/pyproject.toml` (`alert-triage`, workspace member), + package `alert_triage/` with `generator.py` (templates, xorshift32, FNV-1a + seed), `engine.py` (state, lookups, dispositions, never raises), + `protocol.py` (JSON parse, renderers, system prompt), `reward.py` (with the + single `# region: pig-demo/reward` pair), `policies.py` (`fast`, + `thorough`), `taskset.py`, `tests/`, and `names/PROVENANCE.md`. +- Probe policies registered for taskset `alert-triage`; eval command + `uv run python envs/probe.py`. +- The TypeScript port carries the generator, not only the scorer. +- The demo's surface renders a **"Synthetic data — no real customer, account or + transaction"** label on the board itself, the hour constants as config, the + narrative with "not scored" beside it, and the label reveal only after a + disposition. +- Interactive mode: yes. A person can read the screen, click the four lookups + (each showing its hour price), and close with a citation checklist or + escalate with a typology and a transaction tick-list. +- Suggested `caveat` for the vertical entry at ship time: "The alerts are + synthetic and every one has a planted answer. The narrative prose is shown + and not scored; the disposition and the cited facts are. The adversary + designed the data and moved the threshold once — it does not adapt to the + agent."