Fill the demo book's motion: eight engagements, one at every stage
CI / verify (push) Successful in 7m20s
CI / publish (push) Has been skipped

The Motion half of the demo book was two engagements, which was enough to show
that the loop works and not enough to show what the product is for. The page
that matters asks whether the motion is repeating, and a stage rail of zeroes
cannot answer it.

Eight engagements now, one at every open stage, hanging off demand deals the
demand book already creates — the pipeline happened to have exactly one open
deal at each of the eight, so no deal was invented and the quoted pipeline
counts are unchanged. Forty-four artefacts, nineteen scores, one promotion.

Three things the book is laid out to prove that a folder of templates cannot.

Every stage is occupied, and every one of the twelve starter templates is
instantiated at least once, so "stages covered 8/8" is a measurement rather
than a claim about the seed.

Scores move, and sometimes move down. Nineteen scores across eight
trajectories, with nine dimensions regressing somewhere — the Verity
fine-tuning record runs 62.5 -> 60.5 -> 84.0 -> 80.8, because a scorecard that
only ever rises is a ratchet and teaches a reader to distrust it. Every score
is computed with `motionScoreBasisPoints` and banded with `motionBand` rather
than written as a literal, so the seed and the product cannot disagree about
what the same dimensions are worth.

The artefact bodies are the customer's own facts — named people, real volumes,
the specific thing going wrong, and a live unresolved risk in each. An artefact
whose body is the template with the blanks still in it is precisely what this
data exists to disprove. Six of the forty-four have no template at all, which
is the honest shape of an engagement and the reason `kind` is carried on the
artefact rather than derived: `engagement_artifacts.kind` is NOT NULL and a
derived kind would have been null for exactly those six.

The bodies live in JSON beside the loader for the same reason the starter
library's do — forty-four markdown bodies as backtick strings is a module
nobody can review.

The promotion copies `body` from the artefact verbatim, as `promoteArtifact`
does, rather than writing a hand-authored version 2. A demo that produced a row
the real path could not have produced would teach the wrong shape of the table.

Two name collisions the authors could not see are fixed: a Quillon contact
shared a full name with a demo seller, and an Aurelian one shared a surname
with another. `usage_count` is raised once per template rather than once per
artefact, so it stays symmetrical with the decrement `clear()` already does —
verified by tearing the book down and confirming the library returns to twelve
templates with every counter back at zero, since a counter left above zero
makes a starter template permanently un-editable.
This commit is contained in:
2026-08-19 01:05:48 -07:00
parent b7d1ffd2d8
commit 666310b264
10 changed files with 2013 additions and 340 deletions
@@ -0,0 +1,341 @@
{
"key": "verity-eu-fine-tuning",
"dealName": "DEMO — EU-resident fine-tuning",
"accountName": "DEMO — Verity Health AI",
"stage": "procurement",
"status": "open",
"summary": "DEMO — Second demand record at Verity, opened out of the same May discovery that produced the inference endpoint: a clinical-coding agent trained against payer-settled DRGs, $1,252,154, technically resolved and commercially stuck. The bundled requisition was rejected on category grounds; the fix is two requisitions, and half of it is raised while the infrastructure half waits on a silent budget holder and an unsettled compute ceiling.",
"openedDaysAgo": 90,
"artefacts": [
{
"templateSlug": "technical-discovery",
"stage": "qualification",
"title": "DEMO — Second discovery: DRG coding inside the Scribe business",
"status": "final",
"authoredDaysAgo": 86,
"body": "## Who was in the room\n\nNinety minutes on 25 May, Verity's Amsterdam office, nineteen days after the Scribe discovery in Lisbon. Their side: Dr Margriet Ruijs (VP Clinical Coding Operations, owns the coding-accuracy number reported to the board), Bram Kooij (staff engineer, built and maintains the coding pipeline), Sabine Brandt (lead clinical coder, twelve years in DRG coding, the only person present who has done the work). Priya Raghunathan sat in for the first twenty minutes and left. Her platform programme and Margriet's coding programme run alongside each other and share a tenancy, a CTO and nothing else. Our side: Sofia Lindgren and Theo Aguirre. No lawyers, deliberately.\n\nMarisol Okonjo asked for this call on 18 May. We are not an incumbent here: the EU inference endpoint is still in redline, nothing of ours is in production at Verity, and I said so in the room rather than let anyone assume a relationship we have not earned yet.\n\n## The workflow\n\nVerity Code is the older half of the business and the one Scribe was built out of. It codes inpatient discharge summaries for nine hospital groups: five German, three Dutch, one Irish. Rheinklinik Verbund — the fourteen-hospital group whose 12 January go-live is driving the endpoint deal — is the largest of them and has been a coding customer since 2023.\n\nVolume is 61,000 cases a month, growing roughly 4% a quarter; it was 58,600 a year ago. Each case produces a primary diagnosis, ordered secondaries, procedures and a resulting DRG, which determines what the hospital is paid. About 12.4% of coded cases are challenged by the payer's medical service, roughly 7,600 a month. A challenge requires a written justification citing the documentation. Sabine's team of nineteen coders spends most of its week on those justifications rather than on coding.\n\n## Repeatable, traced, checkable\n\nRepeatable: yes, and unusually so. Same input shape, same output schema, same rulebook, revised annually.\n\nTraced: yes. Every case since 2023 carries the pipeline's proposed codes, the coder's accepted or corrected codes, the challenge letter where one exists, the justification the coder wrote, and the payer's settled DRG. That last field is the interesting one and Bram had never thought of it as a label.\n\nCheckable: this is the finding of the call. The settled DRG after appeal is an adjudicated label produced by a process that has nothing to do with us and cannot be gamed by the model. Agreement is computable exactly at DRG level and hierarchically at ICD-10-GM code level. The justification text needs a judge; the code assignment does not.\n\n## The failed attempt\n\nBram has run a prompted pipeline on a frontier API for seven months. Primary-DRG agreement moved from 62.3% to 68.1% and has been flat for eleven weeks despite retrieval over the coding rulebook and a two-pass self-critique. Sabine's coders sit at 88.9% against the same settled labels. Bram volunteered, unprompted, that he has run out of ideas. That is the headroom argument made by the customer rather than by us.\n\n## Constraints found\n\nResidency is inherited, not negotiable: the Rheinklinik DPA that forces the Scribe endpoint into the EU covers coding data too, and Verity's contracts with all nine groups permit processing for coding and are silent on training. Margriet believes an amendment is needed and does not know how many of the nine will sign. Largest open question, and it is legal rather than technical.\n\nThey will not accept a hosted model whose weights they do not hold. Consistent with what Marisol said in Lisbon, so this is settled ground rather than an argument.\n\n## What we do not have\n\nNo baseline on a held-out set — 68.1% is computed over rolling production traffic with no frozen split. No named budget: Margriet says the money \"would come from the platform side\", which is a direction, not a source. Katrin Vogel, the fractional DPO who gates the endpoint deal, was not in the room and will gate this one too.",
"kind": "discovery"
},
{
"templateSlug": "data-and-security-brief",
"stage": "legal",
"title": "DEMO — Data and security brief: training rights across nine hospital groups",
"status": "final",
"authoredDaysAgo": 62,
"body": "## Scope of this brief\n\nPost-training on clinical documentation from nine hospital groups across three jurisdictions. Verity is processor to each group; Prime Intellect becomes subprocessor to Verity. Filled in with Bram Kooij and Owen McCabe on 18 June and handed to Katrin Vogel, Verity's fractional DPO, the same afternoon.\n\nThis is an amendment to the DPA Ana Beltrán is negotiating for the inference endpoint, not a fresh negotiation — which sounds like six weeks saved and is not, because that DPA is not signed. Ana is drafting one document to serve two deals. Efficient while it moves; total blockage when it stops. Katrin's endpoint queue comes first and she works two days a week.\n\n## The gap between policy and practice\n\nThirty minutes with Bram found what the policy document does not say. Discharge summaries have never left Verity's Frankfurt tenancy, not because a policy forbids it but because nobody has ever asked. The prohibition Margriet described in Amsterdam turned out to be a habit with a plausible justification attached, and that distinction is the reason this deal is doable at all.\n\nWhat is a genuine constraint: the hospital-group contracts permit Verity to process documentation for the purpose of coding and say nothing about training. Katrin's position, which we accept without argument, is that silence is not permission in this jurisdiction and an amendment is required from each group before its cases enter a training corpus.\n\nStatus: amendments drafted and issued to all nine on 22 May. Seven executed — Rheinklinik Verbund, Sankt-Vitus Klinikgruppe, Moselhöhe Kliniken, Zuiderpoort Ziekenhuizen, Stichting Maasdelta Zorg, Hoogveld Medisch Centrum, St Brendan's Group. Two outstanding, both German university hospitals with their own ethics committees: Universitätsklinikum Bergwald has a committee date in October, Universitätsklinikum Nordharz has not answered in five weeks. Between them they are about 21% of monthly volume and most of the cardiology case mix.\n\nDecision taken: the POC corpus is built from the seven. This costs case-mix breadth and it is stated in the proposal rather than discovered in week four.\n\n## Permitted use\n\nStated by us before Katrin raised it: customer data trains the customer's model and nothing else. No cross-customer training, no shared base, no de-identified reuse. Adapter weights are Verity's property on creation, not on payment. Katrin accepted the clause and asked for the deletion certificate to name the storage system, which we agreed.\n\n## Residency, retention, deletion\n\nTraining and inference run in the Frankfurt region. No failover — the endpoint DPA forbids cross-region replication and a US read replica, which means a Frankfurt outage stops training and stops coding. That constrains our recovery story and we say so in the questionnaire rather than let a security reviewer find it.\n\nTraces used for training are retained for the campaign plus 90 days, then deleted with certificate. Evaluation sets persist for the life of the agreement, because a held-out set that is deleted cannot be re-run. Written as a named exception rather than left implicit.\n\n## Roles, pseudonymisation, subprocessors\n\nDirect identifiers are stripped in Verity's tenancy before anything reaches a training pipeline — the same hash-based method Katrin has already flagged as possibly insufficient on the endpoint deal. If she rules against it there, it moves here too. Free-text re-identification risk is real and neither side pretends otherwise; the mitigation is that the environment never egresses text. Subprocessor list issued unprompted: three named, all EU-resident, fifteen working days to object.\n\n## Open\n\nThe DPIA amendment cannot close while the two groups are outstanding, because the DPIA enumerates the controllers and Katrin will not sign one with a footnote. This does not block a POC scoped to the seven. It does block any expansion to full volume, and it is why the security and legal path will not score full marks on this deal at any point.",
"kind": "discovery"
},
{
"templateSlug": "reference-architectures",
"stage": "scoping",
"title": "DEMO — Coding environment, corpus derivation and the three-layer verifier stack",
"status": "final",
"authoredDaysAgo": 48,
"body": "## What we are building\n\nOne environment, one verifier stack, one held-out set. Nothing here needs tool use, browser control or a live system, which is why environment construction is cheap on this deal and the risk sits in the judge.\n\n## The environment\n\nSingle-turn with an optional second turn. Input: a de-identified discharge summary from a signed group, the structured admission record, and the ICD-10-GM and OPS catalogue version in force for the case year. Output: JSON with primary diagnosis, ordered secondaries, procedures and the derived DRG, plus a justification paragraph citing document spans where a challenge letter exists in the trace.\n\nCases are keyed to the catalogue in force at coding time, not the current one. Getting this wrong silently poisons about a fifth of the corpus. It is the failure Bram's own pipeline made in 2024 and took five weeks to find.\n\n## The corpus, and why it is 214,000 and not 1.25M\n\nSeven signed groups are about 49,300 cases a month; over 26 months that is roughly 1.25M. We are training on 214,000 of them, and the derivation matters because Piotr Sawicki will ask.\n\nAll 148,000 settled challenged cases — challenged by a payer, resolved, settled DRG recorded — because the challenge is where the signal is. Plus a 66,000-case stratified draw from the settled unchallenged remainder, stratified by group, case year and DRG family. Unchallenged cases are 88% of volume; taking all of them would bury the 148,000 that teach anything. Cases still inside the payer's six-month challenge window are excluded because they are not settled and therefore not labelled.\n\nDev split 12,000, held-out 4,000, both stratified the same way. The held-out set is frozen and sealed before the first training run and read at most twice.\n\n## The verifier stack\n\nThree layers, in descending order of trust.\n\nLayer one, deterministic and free: the DRG grouper. Verity licenses the same certified grouper the payers use; given a code set it returns a DRG with no model in the loop. Exact DRG match against the settled label is the primary reward, and it is not gameable because the label is produced outside this system.\n\nLayer two, deterministic: hierarchical code agreement, partial credit for a correct three-character category with a wrong fourth, because a pure DRG signal is too sparse early in training.\n\nLayer three, judged: the justification paragraph. Span citation is checkable deterministically; clinical adequacy is not.\n\nThe failure mode we designed against is the model learning to propose whatever DRG pays most, which correlates with the settled label often enough to look like learning. Mitigation: an upcoding rate reported alongside accuracy on every run, defined as cases whose proposed DRG exceeds the settled DRG in reimbursement weight. If accuracy rises and upcoding rises with it, the run failed regardless of the headline.\n\n## Two numbers Bram measured for us\n\nOn the 12,000-case dev split, 30 June, so that neither appears first in a proposal as if we had invented it. Current pipeline upcoding rate 3.1%, 372 of 12,000. First-pass acceptance of the pipeline's drafted justification by a senior coder, 34%, on a 300-letter sample adjudicated by Sabine Brandt's team.\n\n## Base model and serving\n\nOpen-weights 70B, the same base `verity-scribe-v4` was SFT'd from, so the trained adapter serves on the endpoint Priya is contracting for rather than needing a second migration. Adapter weights are Verity's, held in their Frankfurt tenancy.\n\n## What Applied Research flagged\n\nThe judge is the weak joint. Justification quality is what Sabine cares about most and the only objective without a deterministic check. Pre-committed rule, written here so it cannot be argued about in week six: if the judge's Cohen's kappa against an adjudicated set falls below 0.75 — the human-to-human ceiling on this material is 0.84, measured across Sabine's own seniors — we drop the justification objective from the POC threshold and report it as a secondary observation.\n\nCalibration cannot use customer traces before a PO exists. The set is being assembled instead from published German DRG appeal case law and non-customer teaching cases, adjudicated by Sabine's team on their own time. Any result off it is provisional.",
"kind": "architecture"
},
{
"templateSlug": "proposal-blocks",
"stage": "proposal",
"title": "DEMO — POC proposal: primary-DRG agreement from 67.4% to 80.0%",
"status": "final",
"authoredDaysAgo": 34,
"body": "## Problem restated\n\nVerity Code handles 62,400 inpatient cases a month across nine hospital groups, up from 61,000 in May; payers challenge about 7,700 of them. Your prompted pipeline reaches 67.4% primary-DRG agreement on the sealed 4,000-case split. Your coders reach 88.9% on the same split.\n\nThe 21.5-point gap is absorbed by 11.8 full-time equivalents out of a nineteen-person coding team writing justification letters instead of coding. It also caps onboarding: at 4% growth a quarter you are adding roughly 2,400 cases a month to a team you have already decided not to grow proportionally, which is why two groups came out of the FY26 plan.\n\nSeven months of prompt engineering moved the number 5.8 points and it has now been flat for five months. That is not a criticism of the work. It is evidence about where the remaining gain lives.\n\n## What we will build\n\nFour artefacts, all yours: the coding environment, the three-layer verifier stack, the frozen 4,000-case held-out set, and adapter weights over an open 70B base — the base your Scribe model is already built on. The weights are your property on creation. You hold them in your Frankfurt tenancy and can serve them without us.\n\n## Success and measurement\n\nThe POC succeeds if, on the sealed held-out set, primary-DRG agreement reaches 80.0% or above with the upcoding rate no higher than the pipeline's current 3.1%. Both numbers were measured by Bram Kooij on your data, not by us.\n\nSecondary, reported but not gating: first-pass acceptance of the drafted justification by a senior coder, 34% today, target 60%. It is secondary because its verifier is a judge. If the judge's Cohen's kappa against an adjudicated set comes in below 0.75, against a human-to-human ceiling of 0.84 on the same material, we drop it and say so rather than defend a number we do not trust.\n\n## Measuring it honestly\n\nThe held-out set is frozen before the first training run and read at most twice. Every eval reports accuracy and upcoding together, because a model that has learned to bias upward shows a rising headline and a rising upcoding rate at the same time.\n\n## Data handling\n\nTraining corpus drawn from the seven groups whose amendments are executed: 214,000 settled cases over 26 months. Universitätsklinikum Bergwald and Universitätsklinikum Nordharz are excluded, which removes about 21% of volume and most of your cardiology mix. We would rather state that here than discover it in week four. Everything stays in the Frankfurt region with no failover, per the endpoint DPA.\n\n## What we need from you\n\nBram Kooij for four days across weeks one and two on trace extraction. Sabine Brandt or a delegate for six half-days of adjudication. A named contact for the grouper licence. This is the block most often ignored and the commonest cause of a POC finishing late.\n\n## Commercial structure\n\nTotal $1,252,154 in three deliberately separable lines.\n\nEngagement fee, $520,000. Applied Research, environment construction, verifier and eval build, run triage.\n\nTraining compute, $432,154. A committed floor of 8,000 H100-hours, a blended $54.02 an hour, with on-demand burst above it at $57.75 rather than a block sized to the ceiling.\n\nProduction inference for the coding workload, $300,000 a year: 49,300 cases a month across the seven signed groups, at the per-case rate on the endpoint order form now in redline. The two outstanding groups add roughly $80,000 when they sign. This line attaches to that order form and cannot be ordered before it is signed.\n\n## Timeline and decision\n\nTwelve weeks from PO. New requisitions freeze on 12 September, your financial year closes on 30 September, and the FY27 planning round runs to mid-October — so a compute line not committed against existing spend this month is a January conversation. A PO on 5 September reports on 27 November. Margriet's board paper is due 1 December. Four days is not margin, and we would rather write that than shorten the plan to fit it.",
"kind": "proposal"
},
{
"templateSlug": "pricing-and-packaging",
"stage": "proposal",
"title": "DEMO — Pricing and packaging: three lines, two funding sources, one ceiling",
"status": "final",
"authoredDaysAgo": 30,
"body": "## What this prices\n\nTotal $1,252,154 across three lines that must be separable, because they are not going to be funded from the same place. That is a procurement decision disguised as a pricing decision, and getting it wrong here is what a rejected requisition looks like six weeks later.\n\n| Line | Amount | Category | Recurs |\n| --- | --- | --- | --- |\n| Engagement fee | $520,000 | Professional services | No |\n| Training compute | $432,154 | Infrastructure | Per campaign |\n| Production inference, coding workload | $300,000 | Infrastructure | Yes, annually |\n\n## The compute forecast\n\nBuilt so Piotr Sawicki can rebuild it. He is a finance business partner, not an ML engineer, and a unit he cannot forecast is a unit he escalates.\n\n| Component | Low | Expected | High |\n| --- | --- | --- | --- |\n| Training campaigns (4 x 1,800 h) | 5,400 | 7,200 | 9,000 |\n| Restart factor | x1.20 | x1.35 | x1.55 |\n| Campaigns after restarts | 6,480 | 9,720 | 13,950 |\n| Eval sweeps | 1,900 | 2,400 | 3,400 |\n| Environment iteration | 1,420 | 1,600 | 2,050 |\n| Total H100-hours | 9,800 | 13,720 | 19,400 |\n\nThe restart factor applies to training campaigns only. Eval sweeps and environment iteration do not restart; they are already ranged.\n\nCommercial shape: a committed floor of 8,000 hours at $432,154, a blended $54.02 an hour, with on-demand burst above it at $57.75. The not-to-exceed ceiling on the order form is the high column, 19,400 hours — 11,400 on-demand hours above the floor, $658,350. So the ceiling exposure is $1,090,504 of compute, not $432,154, and any requisition raised at the floor is raised at the wrong number.\n\nVerity has never run a training campaign and cannot defend a commitment, which is why the floor is 8,000 and not 13,720. Our own margin is modelled against the full 8,000 whether consumed or not, because cost is charged on the commitment.\n\n## The anchoring question\n\nAsked of Margriet Ruijs on 25 May: what does the gap cost you today? She had it already, built for a March board paper, and she walked through the derivation rather than quoting the total.\n\nNineteen coders, timesheeted, so this part is measured: 11.8 full-time equivalents spent on challenge justifications rather than coding in the twelve months to April. At €96,000 loaded, €1.13m.\n\nThen the part she computed herself. Universitätsklinikum Bergwald and one group she would not name were deferred out of the FY26 onboarding plan because she could not staff the coding. Contribution roughly €340,000 each, so €680,000. Total €1.81m.\n\nHer board believes the €1.13m — it is in the timesheet system and Greta Halvorsen's team can pull it. It does not believe the €680,000, because deferred contribution is a number a VP can construct and Greta has said so in writing. The defensible anchor is therefore €1.13m a year. $1,252,154 sits above one year of that and below two, which is a survivable place to be but not a comfortable one.\n\n## The marketplace route and its consequence\n\nJan Kowalczyk holds a €4.1m commitment with their hyperscaler, nineteen months remaining. A private offer against it is a drawdown rather than a purchase, which changes three things: the contracting counterparty becomes the marketplace operator, the offer must be listed and accepted before any drawdown counts, on a five working day listing lead time, and the marketplace fee is a margin input rather than a discount. The fee is priced into the two infrastructure lines above and not into the engagement fee.\n\nThat is the load-bearing point. Professional services cannot draw down that commitment. No infrastructure category in Verity's chart of accounts touches them, so the $520,000 has to come from Margriet's operating line.\n\n**Recommendation: two requisitions, not one.** A bundled requisition is rejected by whichever category owner does not expect the other half, and the rejection arrives without a human attached to it.",
"kind": "pricing"
},
{
"templateSlug": null,
"stage": "procurement",
"title": "DEMO — Pre-PO work log: sole-source, questionnaire, and the judge calibration result",
"status": "final",
"authoredDaysAgo": 5,
"body": "## Why this is a separate note\n\nEverything Nadia Osei's list permits before a PO, in one place, so nobody reconstructs it from a requisition thread in October. Two of these changed the deal: the judge calibration returned a result that removes an objective from the POC threshold, and the sole-source page is the only document a procurement analyst will read end to end.\n\n## Sole-source justification\n\nFour paragraphs, sent to Margriet on 11 August to paste under her own name. Paragraph two carries it, verbatim: \"The adapter weights sit over the same open-weights base the EU inference endpoint order form already specifies, so the trained model serves on that endpoint without a second migration. The held-out set was built against our own certified grouper licence and the payers' settled DRGs, and no other supplier has had access to either. Re-deriving both for a competitive process is nine weeks of Bram Kooij's time, which we do not have before the January Rheinklinik go-live.\"\n\n## Vendor security review\n\nRasmus Holm (Vendor Risk Lead, reporting to Owen McCabe) opened the review on 6 August. Twenty working days is typical at Verity, so 3 September, which is inside the 12 September freeze with a week of slack. Not blocking, and we hold Nadia Osei's written list of what may proceed alongside it: supplier onboarding, supplier record creation, the questionnaire, the sole-source page, and the marketplace listing. Everything except work touching a real customer trace.\n\n## The questionnaire\n\nThe questionnaire lives in a portal that opens only once a supplier record exists, so Nadia released the previous vendor's redacted copy instead and we answered against that. Eight of the twelve topics came straight from the standing library. Four went to Applied Research as one batch rather than four emails: base model licence lineage for the open 70B; whether any Verity artefact enters a shared training mix, answered no with the contractual clause cited rather than a policy sentence; judge model provenance and where it runs; and which artefacts cross the customer boundary in each direction. Returned in six days.\n\nOne answer we volunteered because a reviewer will find it anyway: the Frankfurt region has no failover, because the endpoint DPA forbids cross-region replication. A regional outage stops training and stops coding. Rasmus asked a follow-up about recovery time and we gave him a number rather than a paragraph.\n\n## Judge calibration, and the objective we dropped\n\nNadia's rule means no customer traces before a PO, so the judge could not be calibrated against Sabine Brandt's production adjudications. The set was built instead from 900 justifications synthesised from published German DRG appeal decisions and 300 written by Sabine's seniors against non-customer teaching cases, adjudicated twice.\n\nResult on 12 August: Cohen's kappa 0.68 against a human-to-human ceiling of 0.84 measured across the same seniors on the same material. The pre-committed floor in the architecture note was 0.75. It is below it.\n\nSo the justification objective comes out of the POC success threshold and becomes a secondary observation, per the rule written in July precisely so this would not be an argument. Told to Margriet Ruijs and Sabine Brandt on 12 August, before either asked. Sabine's reaction is the part worth recording: justification quality is the number she cares about most, and she now has a POC that does not promise it. She has not withdrawn her six half-days, but she asked what a second calibration on real traces would cost after the PO, which is the right question.\n\nTwo caveats on the 0.68. It is provisional: case law and teaching cases are cleaner than production discharge summaries, so the number could move either way once real traces are in scope. And a single kappa on a three-point clinical-adequacy scale is a coarse instrument. We are not arguing either point into a pass.",
"kind": "pricing"
},
{
"templateSlug": "procurement-and-budget-path",
"stage": "procurement",
"title": "DEMO — Procurement path: splitting REQ-2026-04471 before the 12 September freeze",
"status": "draft",
"authoredDaysAgo": 2,
"body": "## Where this stands\n\nProposal accepted verbally by Margriet Ruijs on 21 July. No PO. REQ-2026-04471 was raised on 30 July for the full $1,252,154 and rejected on 5 August, eight working days ago, by the category system with no human attached to the notice. Margriet read it as \"legal are looking at it\". They are not. The 20 July pricing note recommended two requisitions in bold; it was bundled anyway, which is on us for leaving that in a last paragraph rather than saying it on the phone.\n\n## Budget source\n\nTwo sources, and the failure was assuming one. Jan Kowalczyk (Director of Infrastructure) owns the €4.1m hyperscaler commitment, nineteen months to run. It funds GPU-hours and inference without new approval, because a private offer against it is a drawdown, and it cannot fund professional services. Margriet's business-unit operating line funds the $520,000 engagement fee, and it is the durable source: expansion next year is a line increase, not a fresh fight.\n\n## The unsticking move\n\n**REQ-B**, category 6420, professional services external: $520,000 against Margriet's line, sole-source attached at raise. Raised 13 August. Below the $750,000 CFO threshold, so it needs Margriet and Piotr Sawicki and nothing else — no CFO, no second legal pass, no board. That is the point of the split: half the money moves this week.\n\n**REQ-A**, category 7310, infrastructure committed drawdown, raised at the ceiling rather than the floor: 8,000 committed hours at $432,154, plus 11,400 on-demand at $57.75, $658,350, plus $300,000 of coding inference. $1,390,504. Not raised. It crosses $750,000, so it goes to the Investment Review Board, and the private offer must be listed today, Wednesday 19 August, to be accepted in five working days and counted on 26 August.\n\nNadia Osei (Senior Procurement Manager) was told the split is by funding category, not to duck a threshold. She agreed, then said she would have asked why we were bundling in July had anyone shown her the requisition first. Fair.\n\n## Approval chain\n\nBelow $25,000, budget owner alone. Above, Piotr Sawicki joins as finance business partner. Above $750,000, Greta Halvorsen (CFO), a second legal pass despite the MSA, and the fortnightly Investment Review Board: 26 August, then 9 September. The cadence, not the decision, costs the days.\n\nPiotr rebuilt the ceiling and came out at 22,600 hours, having applied the restart factor to eval sweeps and environment iteration as well as to campaigns. He wants REQ-A at $1,575,304 — $184,800 of ceiling nobody will use and a worse board paper. He thinks a ceiling that binds is not a ceiling. Unresolved, and it is what is stopping REQ-A today. The sole-source page, the security review, the questionnaire and the judge calibration are in the pre-PO work log; none of them is blocking.\n\n## Champion's homework\n\nFour things; she has done two. The sole-source page went in under her name on 12 August, and REQ-B was raised on 13 August with 6420 attached. Not done: raise REQ-A at 7310 once Piotr's ceiling settles, and send Jan the sentence we drafted on 6 August — \"Jan, the coding programme needs $1,390,504 of your hyperscaler commitment this year across compute and inference, as one infrastructure requisition; it is a drawdown, not new spend, and I need a yes or no by Thursday to make the 26 August board.\"\n\n## Where it is stuck\n\nJan has not replied in seven working days. He was told in July this was one requisition against his line and it is now a different, larger one. Margriet has not chased him because she still believes the rejection was legal. The move is three lines from Marisol Okonjo to Jan confirming the drawdown is expected — drafted, sent for signature on Friday, unsigned; she has missed two reviews since the Series C process started. If the offer is not listed today, REQ-A lands on 9 September, three days before the 12 September freeze, with no slack.\n\nWhat REQ-A cannot fix: the $300,000 inference line attaches to an endpoint order form that is not signed. Nadia says a requisition may be approved against an unexecuted contract; Ana Beltrán says it may not. Unresolved, and the second-largest risk here.",
"kind": "pricing"
}
],
"scores": [
{
"daysAgo": 84,
"note": "DEMO — Scored straight off the 25 May Amsterdam call. 62.5, Qualified conditional. The verifier is unusually good for healthcare because the payer's settled DRG is an adjudicated label produced outside the system, and Bram Kooij made the headroom argument himself. Against that: no frozen split, so the 68.1% is rolling production traffic and baseline scores a 1; no named budget beyond \"the platform side\"; and nine hospital-group contracts that are silent on training.",
"dimensions": [
{
"id": "task_definability",
"weight": 10,
"score": 3
},
{
"id": "verifier_quality",
"weight": 13,
"score": 3
},
{
"id": "trace_data_rights",
"weight": 8,
"score": 2
},
{
"id": "baseline_measured",
"weight": 8,
"score": 1
},
{
"id": "headroom",
"weight": 10,
"score": 3
},
{
"id": "env_constructibility",
"weight": 8,
"score": 3
},
{
"id": "metric_owner",
"weight": 9,
"score": 3
},
{
"id": "budget_source",
"weight": 9,
"score": 1
},
{
"id": "exec_sponsor",
"weight": 4,
"score": 3
},
{
"id": "security_legal_path",
"weight": 8,
"score": 2
},
{
"id": "forcing_function",
"weight": 8,
"score": 3
},
{
"id": "expansion_surface",
"weight": 5,
"score": 3
}
]
},
{
"daysAgo": 58,
"note": "DEMO — One move, downward. Amendments went to all nine groups on 22 May; seven executed, and two German university hospitals have not, so the corpus is seven groups and about 21% of volume is out. Katrin Vogel will not sign a DPIA that enumerates controllers who have not signed, so trace and data rights fall 2 to 1. Nothing else changed: still no frozen split, still no budget owner, and the DPA we are amending is itself still in redline on the endpoint deal.",
"dimensions": [
{
"id": "task_definability",
"weight": 10,
"score": 3
},
{
"id": "verifier_quality",
"weight": 13,
"score": 3
},
{
"id": "trace_data_rights",
"weight": 8,
"score": 1
},
{
"id": "baseline_measured",
"weight": 8,
"score": 1
},
{
"id": "headroom",
"weight": 10,
"score": 3
},
{
"id": "env_constructibility",
"weight": 8,
"score": 3
},
{
"id": "metric_owner",
"weight": 9,
"score": 3
},
{
"id": "budget_source",
"weight": 9,
"score": 1
},
{
"id": "exec_sponsor",
"weight": 4,
"score": 3
},
{
"id": "security_legal_path",
"weight": 8,
"score": 2
},
{
"id": "forcing_function",
"weight": 8,
"score": 3
},
{
"id": "expansion_surface",
"weight": 5,
"score": 3
}
]
},
{
"daysAgo": 28,
"note": "DEMO — The technical side moved at once on the architecture and proposal work: environment designed and catalogue keying resolved (task definability 4), Bram measured 67.4% on the 4,000-case split before sealing it (baseline 4), the prompted pipeline is flat five months rather than eleven weeks (headroom 4), no tool use or live system needed (environment 4), and Margriet signed the 80.0% target in writing (metric owner 4). Commercially, pricing named two funding sources for the first time — budget source 1 to 2, and no higher until a PO exists — Ana Beltrán agreed the training-amendment language for the seven signed groups (security and legal 2 to 3, which is the ceiling on this deal), and the 12 September freeze became a hard date (forcing function 4). Data rights recover one point because the seven-group corpus is contractually clean and the POC is scoped to it; verifier holds at 3 until the judge is calibrated.",
"dimensions": [
{
"id": "task_definability",
"weight": 10,
"score": 4
},
{
"id": "verifier_quality",
"weight": 13,
"score": 3
},
{
"id": "trace_data_rights",
"weight": 8,
"score": 2
},
{
"id": "baseline_measured",
"weight": 8,
"score": 4
},
{
"id": "headroom",
"weight": 10,
"score": 4
},
{
"id": "env_constructibility",
"weight": 8,
"score": 4
},
{
"id": "metric_owner",
"weight": 9,
"score": 4
},
{
"id": "budget_source",
"weight": 9,
"score": 2
},
{
"id": "exec_sponsor",
"weight": 4,
"score": 3
},
{
"id": "security_legal_path",
"weight": 8,
"score": 3
},
{
"id": "forcing_function",
"weight": 8,
"score": 4
},
{
"id": "expansion_surface",
"weight": 5,
"score": 3
}
]
},
{
"daysAgo": 1,
"note": "DEMO — 80.75, still Build, and two dimensions fall. Budget source 2 to 1: REQ-2026-04471 was rejected on category grounds on 5 August and Jan Kowalczyk has not replied to the 6 August note in eight working days, so two sources are named and neither has produced a PO. Exec sponsor 3 to 2: Marisol Okonjo has missed two reviews since the Series C process started and has not signed the note to Jan. Holds that are not ratchets — verifier stays at 3 because the judge calibrated at kappa 0.68 against a 0.84 human ceiling and the justification objective has been dropped from the threshold per the pre-committed rule; expansion surface stays at 3 because the DPIA blocks any move to full volume; data rights stay at 2 with Bergwald's ethics committee in October.",
"dimensions": [
{
"id": "task_definability",
"weight": 10,
"score": 4
},
{
"id": "verifier_quality",
"weight": 13,
"score": 3
},
{
"id": "trace_data_rights",
"weight": 8,
"score": 2
},
{
"id": "baseline_measured",
"weight": 8,
"score": 4
},
{
"id": "headroom",
"weight": 10,
"score": 4
},
{
"id": "env_constructibility",
"weight": 8,
"score": 4
},
{
"id": "metric_owner",
"weight": 9,
"score": 4
},
{
"id": "budget_source",
"weight": 9,
"score": 1
},
{
"id": "exec_sponsor",
"weight": 4,
"score": 2
},
{
"id": "security_legal_path",
"weight": 8,
"score": 3
},
{
"id": "forcing_function",
"weight": 8,
"score": 4
},
{
"id": "expansion_surface",
"weight": 5,
"score": 3
}
]
}
],
"promote": null
}