{ "kind": "qualification", "slug": "trainability-qualification", "title": "Trainability and Deal Qualification Scorecard", "summary": "A weighted, anchored scorecard for deciding whether a customer's problem is a post-training problem Prime Intellect can win and whether the deal behind it is real — run it at the end of technical discovery, before anything goes to legal.", "stage": "qualification", "body": "*Playbook stage: qualification, at its close.* Score this from the notes of **First Technical Discovery**, before the deal enters legal. The verifier floor used below is set once, in the **Strategic Deployment Playbook**.\n\n## What this scores\n\nThis is not BANT. BANT tells you whether someone can buy something; it says nothing about whether the thing they want to buy can be built. In post-training, most lost deals are not lost on price or to a competitor. They are lost because the customer's problem was never a trainable problem, and nobody said so in week two.\n\nThis framework scores two independent questions and refuses to average them into one comfortable number:\n\n1. **Technical fit.** Can this workflow be turned into an environment with a verifier, trained against a measured baseline, and shown to improve? (57 points)\n2. **Commercial reality.** Is there a person who owns the metric, money that already exists, and a legal path that closes before compute is committed? (43 points)\n\nRun it at the end of technical discovery, before you write a scoping document. Re-run it after environment design, because the verifier score in particular almost always moves, and usually down.\n\n## How to score\n\nEach dimension is scored 0-4 against its anchors. Anchors describe **observable evidence**, not your confidence. If you cannot point to the artefact the anchor names — the file, the dashboard, the name, the contract clause — you score the level below. \"They said they have it\" is a 1, not a 3.\n\nWeighted score = Σ(weight × score) ÷ 4, giving a 0-100 result. Weights sum to 100 and scores top out at 4, so straight 4s are exactly 100.\n\nTwo derived numbers matter as much as the total:\n\n- **Group subtotals**, reported separately and never only as a blend. Technical is out of 57, commercial out of 43.\n- **Weighted point loss** per dimension: weight × (4 − score). This is what the conditional band acts on, and it is not the same thing as a low weighted score. A perfect 4 on executive sponsor contributes 16 weighted points and loses none; a 3 on verifier quality contributes 39 and loses 13. The dimension to fix is the one losing the most points, not the one contributing the fewest. Point loss across all twelve dimensions always sums to 400 minus the weighted total, which is a useful arithmetic check.\n\n## Score bands and what to do\n\n| Band | Label | Action |\n|---|---|---|\n| 78-100 | Build | Move to scoping. Commit named research support, write the environment and verifier design into the proposal, and put a dated POC gate with the metric owner's agreed threshold on it. |\n| 62-77 | Qualified, conditional | Advance, but name the two dimensions with the largest weighted point loss — max weight × (4 − score), breaking ties toward the technical group — as written POC entry conditions with owners and dates. Do not start compute until both are cleared. |\n| 46-61 | Scope down | Do not sell the workflow as described. Find the narrowest sub-task that would itself score above 62, and sell that as a paid evaluation or a single environment build with a fixed price and a fixed end. |\n| 30-45 | Defer | No POC. Offer evaluations or on-demand compute if genuinely useful, name the artefact that must exist before re-scoring, set a re-qualification date, and stop talking about post-training until then. |\n| 0-29 | Decline | Decline in the next meeting, verbally, naming the missing artefact rather than the customer. Recommend the cheaper thing that would actually work now, and leave an explicit trigger that brings them back. Log the reason in PIG as closed_lost with the disqualifying dimension recorded. |\n\n### On declining\n\nDeclining well is the highest-leverage thing in this document, because a bad POC consumes delivery capacity that cannot be recovered — assume {{FDE-weeks}} per POC, a figure to fill in from your own delivery data rather than one this document can supply — and it produces a reference customer who tells people it did not work. Decline in the room, not in an email chain that dies.\n\nSay roughly this: *\"I don't think this is a post-training problem yet, and I'd rather tell you that than sell you six weeks that end in a shrug. Right now nobody can say what a correct output looks like without {{named senior person}} reading it. Until that judgement is written down as a check something else can run, training has nothing to optimise against. Here's what I'd do instead: {{prompting / retrieval / workflow change}}. When you have {{the specific missing artefact}}, come back and we will move fast.\"*\n\nThree things that does: it names the missing artefact rather than blaming them, it gives them a cheaper thing that actually works, and it leaves a trigger condition that brings them back. Never decline on \"we're not a fit\". That is a door closing with no handle on it.\n\n## The artefact that unblocks each dimension\n\nEvery band above tells you to name an artefact. Naming it is the whole job, so here is the mapping rather than the instruction.\n\n| Dimension | Artefact the customer must produce |\n|---|---|\n| Task definability | A written spec plus ten worked examples, with a second reviewer's independent verdict recorded on each |\n| Verifier availability and quality | At least 200 human-labelled items on a held-out set, stratified so the minority outcome is at least 30 of them, and a measured judge-versus-human agreement number on them |\n| Traces, volume, and rights to use them | Counsel's written confirmation that the traces may be used for model training, plus a counted volume produced by a query rather than from memory |\n| Measured baseline | A held-out set drawn from a time window after the training traces, a scoring script the customer can rerun, and a recorded frontier-model or human number |\n| Headroom over prompting | A categorised failure set from a serious prompted baseline |\n| Environment constructibility | One containerised run of the workflow against seeded state with a scripted reset |\n| Named metric owner | That person's numeric success threshold, stated by them, in writing |\n| Budget line and its source | The budget code, and the named spend it displaces |\n| Executive sponsor | One logged instance of the executive removing a specific blocker |\n| Security and legal path | A comparable AI vendor's completed review, plus deployment topology and weights ownership agreed in writing |\n| External forcing function | The external date, and the quantified cost of missing it stated by the sponsor |\n| Expansion surface | Two named adjacent workflows with named owners |\n\n## Hard disqualifiers\n\nA weighted average is a machine for laundering one fatal problem into an acceptable number. Take a deal scoring 3 on everything, then move budget source, executive sponsor and forcing function to 4 and verifier quality to 0. It scores 70.5: a comfortable Qualified, conditional. It will still fail, because you cannot train against a reward you cannot compute. Averages assume dimensions substitute for each other. Several of these do not; they are gates.\n\nThe four disqualifiers below override the score entirely, and each caps the deal at a named band rather than at one generic outcome:\n\n- **No computable verifier and no path to one**, and **incompatible weights-ownership or deployment expectation**, cap at **Decline (0-29)**. Neither softens with time or effort inside this deal, and routing them to Defer would leave the customer waiting on something that is not coming. Use the decline script.\n- **No rights to train on the data**, and **nobody states the number**, cap at **Defer (30-45)**, with the unblocking artefact from the table above named explicitly and a re-qualification date set. A motivated customer can produce both.\n\nRecord which one fired on the demand record in PIG. The pattern across a quarter of disqualified deals is the most useful thing the field sends back to the people designing environments.\n\n## Worked example\n\nAn anonymised deal, scored across all twelve dimensions at the end of technical discovery.\n\n| Dimension | Weight | Score | Weighted | Point loss |\n|---|---|---|---|---|\n| Task definability | 10 | 3 | 30 | 10 |\n| Verifier availability and quality | 13 | 2 | 26 | 26 |\n| Traces, volume, and rights to use them | 8 | 3 | 24 | 8 |\n| Measured baseline | 8 | 2 | 16 | 16 |\n| Headroom over prompting and workflow design | 10 | 3 | 30 | 10 |\n| Environment constructibility | 8 | 3 | 24 | 8 |\n| Named metric owner | 9 | 3 | 27 | 9 |\n| Budget line and its source | 9 | 4 | 36 | 0 |\n| Executive sponsor | 4 | 4 | 16 | 0 |\n| Security and legal path | 8 | 3 | 24 | 8 |\n| External forcing function | 8 | 4 | 32 | 0 |\n| Expansion surface | 5 | 3 | 15 | 5 |\n| **Total** | **100** | | **300** | **100** |\n\nWeighted total 300, so the score is 300 ÷ 4 = **75**. The subtotals happen to be equal in raw weighted points, 150 each, which is 37.5 of a possible 57 on technical and 37.5 of 43 on commercial — a deal that is commercially stronger than it is trainable, which the blended 75 does not tell you. Point loss sums to 100, which is 400 − 300, so the arithmetic checks.\n\n75 lands in **Qualified, conditional (62-77)**. The two POC entry conditions are the two largest point losses: **verifier quality** (26) and **measured baseline** (16). Note what the discarded rule would have done: the two *lowest weighted scores* are expansion surface (15) and executive sponsor (16), both of which are fine and neither of which can sink the POC. That is the whole reason the band acts on loss.\n\nNow the override. At the environment design workshop the verifier score is re-taken and drops to 1: the rubric exists but the judge has never been measured, and the sponsor is asked directly, on the record, to fund a 200-item labelling pass and declines. The arithmetic barely moves — 287 ÷ 4 = 71.75, still Qualified, conditional — but disqualifier 1 has fired, and it caps at Decline. The deal does not advance, and it gets the decline conversation rather than a re-qualification date, because there is no artefact anyone has agreed to produce.\n\n## Reading the two groups separately\n\nAlways report the two subtotals, never only the blend.\n\n- **High technical, low commercial.** A real training problem inside an organisation that cannot buy it. It burns delivery months on beautiful environments nobody funds. Correct move: an inexpensive paid evaluation that puts a number in front of an executive, so the buyer gets created rather than assumed.\n- **Low technical, high commercial.** Money, sponsor, urgency, and a task that is not trainable. This is the dangerous one, because everything social in the deal pushes you forward. Sell inference or reserved compute if that is genuinely useful, and keep post-training out of the contract. Selling a training programme into an unverifiable task is how you convert a well-funded champion into a detractor.\n- **Both mid.** Usually one workflow doing the work of three. Split it. The narrow version of a sprawling request is often a 3 or 4 across the technical group.\n\n## What the scoring conversation actually sounds like\n\nThe anchors are written so that you can score them from a discovery call if you ask forcing questions. Some that work:\n\n- \"Show me the last ten outputs of this workflow and tell me which ones were wrong.\" (Scores task definability and verifier quality in one move. If they argue about three of the ten, that argument *is* the verifier problem.)\n- \"Who gets asked about this number in their performance review?\" (Metric owner. If the answer is a committee, it is a 1.)\n- \"What is the budget code, and what was it going to be spent on if we don't do this?\" (Budget source. Displacement budget is real; net-new budget in an unapproved plan is a 1.)\n- \"How many of these did you run last month, and where is that logged?\" (Trace availability. A number from memory is a 1; a query against a table is a 3.)\n- \"If we do nothing, what happens on {{date}}?\" (Forcing function. If the honest answer is \"nothing\", it is a 0, whatever the enthusiasm level.)\n- \"What is the latency budget and the per-item cost ceiling for this in production?\" (Not scored — see the weight rationale — but a task that must answer in 300ms for a fraction of a cent changes the whole programme, and you want that on the table before scoping.)\n- \"Can you send me the DPA you signed with your last inference vendor?\" (Security path. An existing executed agreement with a comparable vendor is worth more than any assurance about process.)\n\n## Where the score belongs in the pipeline\n\nScore at **qualification**, before legal. This is deliberate and it matters in this pipeline: legal sits second, so an unqualified deal does not merely waste your time, it consumes counsel's time on an MSA and DPA for an engagement that should never have started. The score is the gate that protects the second stage.\n\nAttach the completed scorecard to the demand record in PIG. On any deal that later reaches closed_lost, the delta between the qualification score and the outcome is the only honest feedback loop this motion has. Deals that scored above 78 and still lost are telling you a dimension is missing or mis-weighted. Deals that scored below 46, were pushed through anyway, and won are telling you the same thing in the other direction, and are rarer than the optimists in the room believe.\n\n## Weight rationale\n\nVerifier quality carries the single largest weight (13) because it is the closest thing in this business to a binary. Everything downstream — environment, reward, eval, the improvement claim in the case study — is built on it, and no amount of budget compensates for its absence. Task definability (10) and headroom over prompting (10) come next: the first determines whether an environment can exist, the second whether anyone should pay for one rather than spending an afternoon on a better prompt. On the commercial side, metric owner (9) and budget source (9) outweigh executive sponsor (4) on purpose. Sponsors change jobs; a named person whose review depends on the number does not stop caring about it, and they are the one who will defend the result internally when it lands short of the deck.\n\nTwo things a reader may expect to find scored here are deliberately not dimensions. **Latency and cost envelope** is a constraint, not a measure of fit: it does not make a task more or less trainable, it decides what model size and serving topology the trained result must fit into, and a task can be perfectly trainable and still uneconomic at 300ms. Constraints belong in scoping, where they change the design, rather than in a score, where a hard limit would be averaged away — which is the same failure the disqualifiers exist to prevent. Ask the question in discovery, record the answer in the scoping document, and if the envelope is genuinely unmeetable, that is a decline on feasibility rather than a low score. **Long-horizon versus single-shot** is not scored separately because it does not vary independently of the dimensions that are: a long-horizon task shows up as a hard verifier (credit assignment over a trajectory), a harder environment to construct (state, resets, episode boundaries) and larger headroom over prompting. Scoring it again would triple-count the same fact. It appears where it is observable, in the top anchors of headroom and environment constructibility.", "fields": { "dimensions": [ { "id": "task_definability", "group": "technical", "name": "Task definability", "weight": 10, "why": "An environment is a task specification with a scoring function attached. If two competent people at the customer disagree about what a correct output is, there is nothing to specify. This is the first thing that breaks and the last thing customers admit is broken.", "anchors": { "0": "The workflow is described only as an outcome (\"handle the ticket queue better\"). No unit of work, no start and end state.", "1": "A unit of work exists but correctness is \"whatever the senior person would have done\". No two examples of correct output have been shown to you.", "2": "Clear input and output types and an informal quality bar, but on the sample you were shown, experienced staff disagree on at least one item in five.", "3": "A written spec or SOP exists with at least ten worked examples, and two reviewers independently agree on correctness for at least eight of the ten.", "4": "Written spec plus a labelled set the customer already uses to judge humans doing this job, with a measured inter-annotator agreement number they can quote." } }, { "id": "verifier_quality", "group": "technical", "name": "Verifier availability and quality", "weight": 13, "why": "The reward signal is the product. A task with a cheap, fast, hard-to-game check trains; a task whose only check is a human reading the output does not, at any budget. This dimension is the strongest single predictor of whether a POC produces a number worth putting in a case study.", "anchors": { "0": "Correctness is a matter of taste. No programmatic check is conceivable and human review takes more than a few minutes per item.", "1": "Human review only, fast and consistent enough (under a minute, one reviewer) that a judge could plausibly be calibrated against it later, or a written rubric exists but the judge has never been run against human labels. An unmeasured rubric is a 1.", "2": "A model-judge is the plausible verifier and it has been measured: at least 200 human-labelled items on a held-out set, with judge-versus-human agreement at or above the library floor set in the Strategic Deployment Playbook - Cohen's kappa 0.7 against adjudicated labels, or within 0.10 of a measured human ceiling that is itself below 0.8.", "3": "A partial programmatic verifier exists or is a few days of work — unit tests, schema validation, a diff against a known-good record, a database assertion — and it decides at least 70% of items unaided, with a measured judge handling the rest. A trivial policy (empty output, retry loop, copy of the input) scores near zero against it.", "4": "Ground truth is computable and already computed in production: tests pass or fail, the trade reconciles, the query returns the right rows, the simulator says solved. The check is cheap, deterministic, difficult to satisfy without doing the task, and a trivial policy scores near zero. Where the task runs over many steps, the check applies at the end state rather than requiring a human to read the trajectory." } }, { "id": "trace_data_rights", "group": "technical", "name": "Traces, volume, and rights to use them", "weight": 8, "why": "Post-training needs both examples and the legal right to train on them. Customers routinely have neither and discover it at the DPA stage, six weeks in, when the data belongs to their own customers under contracts that forbid model training.", "anchors": { "0": "No logged traces, or traces exist but a third-party or end-customer contract clearly prohibits their use for training and nobody is willing to test it.", "1": "Traces exist somewhere (screen recordings, ticket text, email) with no structure, no volume estimate, and unresolved rights.", "2": "Structured logs of the workflow exist at plausible volume, but rights need a review that has not started and residency questions are open.", "3": "Queryable traces with inputs, actions and outcomes, a volume counted by a query rather than estimated, and counsel has confirmed in writing that internal-use training is permitted.", "4": "As level 3, plus outcome labels already attached to the traces, and either the data is first-party or the customer agreements explicitly permit training. Volume is comfortably above what the task needs." } }, { "id": "baseline_measured", "group": "technical", "name": "Measured baseline", "weight": 8, "why": "You cannot prove improvement against a number that did not exist before you arrived. Establishing a baseline mid-POC means the customer can always argue it was set low, which quietly destroys the case study you are building. The held-out set also has to be held out from the traces you will train on, or the number you report at the end measures memorisation.", "anchors": { "0": "No measurement of current performance, human or model, and no agreement on what would be measured.", "1": "Anecdotal claims about current quality. Numbers vary depending on who is in the room.", "2": "A one-off measurement was taken at some point, on a set that no longer exists or cannot be rerun.", "3": "A held-out set and a scoring script the customer can rerun, with a current frontier-model or human number recorded, and the held-out set is disjoint from the traces that would be trained on — drawn from a time window after them — and is not used to calibrate the verifier.", "4": "As level 3, plus the baseline is measured continuously in production, visible on a dashboard someone watches, the disjointness is enforced by the query that builds the set rather than by anyone's discipline, and the customer accepts it as the number the engagement will be judged on." } }, { "id": "headroom", "group": "technical", "name": "Headroom over prompting and workflow design", "weight": 10, "why": "If a better prompt, retrieval, or a fixed tool schema closes the gap, training is an expensive way to lose a customer's trust. Managed RL earns its keep where the failure mode is judgement under distribution the base model has not seen, or credit assignment across a long trajectory, not where it is missing context.", "anchors": { "0": "Nobody has tried a serious prompted baseline. The gap is almost certainly context or tooling, not capability.", "1": "A prompted baseline exists and lands within 5 points of the target on the agreed metric; remaining failures look like retrieval or tool-schema problems.", "2": "Prompting plateaued short of target, but the residual failures have not been categorised, so the cause is unknown.", "3": "Prompting plateaued with a categorised failure set, and the dominant categories are domain judgement, house conventions, or multi-step recovery that examples do not fix.", "4": "As level 3, plus the customer has already tried SFT or a competitor's fine-tune and can show where it stalled. The remaining gap is credit assignment over a multi-step trajectory, where the early decision that caused the failure is several steps from the observable outcome, which is what RL is for." } }, { "id": "env_constructibility", "group": "technical", "name": "Environment constructibility", "weight": 8, "why": "The task must be re-executable in a sandbox without touching production, without a human in the loop, and without a licence key nobody can get. A workflow that only runs against a live system with real side effects cannot be trained on, however trainable it looks on paper.", "anchors": { "0": "The workflow only exists inside a vendor SaaS UI with no API, or every step has irreversible real-world side effects.", "1": "Partial APIs, and a state reset would require manual work between episodes.", "2": "APIs exist for most steps; a staging environment exists but is shared, stateful, and slow to reset.", "3": "Most of the workflow can be containerised with seeded state and a scripted reset. A few steps need stubs the customer agrees to help write.", "4": "The whole task runs headless against a snapshot, resets in seconds, and parallelises — including across a full multi-step episode with clean boundaries, not only single calls. Someone at the customer has already built something close for their own CI." } }, { "id": "metric_owner", "group": "commercial", "name": "Named metric owner", "weight": 9, "why": "Someone must be personally accountable for the number that will move. That person defends the result internally when it lands short of the deck, and their absence is why technically successful POCs die without expanding.", "anchors": { "0": "No number identified. The interest is in the technology.", "1": "A number exists but ownership is spread across a committee or a function rather than a person.", "2": "A named person owns the number, has not been in a meeting with you, and has not agreed to the target.", "3": "Named owner, in the room, agrees the number is theirs and states, in writing, the numeric threshold that would count as success.", "4": "As level 3, and the number appears in their own goals or compensation for this period. They have volunteered a team to work with you." } }, { "id": "budget_source", "group": "commercial", "name": "Budget line and its source", "weight": 9, "why": "Enthusiasm converts to procurement only when the money already exists somewhere. Ask where it comes from, not whether it exists: displaced spend closes, net-new spend in an unapproved plan slips two quarters or dies at the next planning cycle.", "anchors": { "0": "No budget. The hope is that a successful POC creates one.", "1": "Net-new spend that would need approval in a planning cycle that has not started.", "2": "Budget identified in an approved plan, but the amount is a guess and the owner is not the metric owner.", "3": "A specific approved line with an amount, and the sponsor can name what it displaces (an inference bill, a vendor contract, headcount, a competing GPU commitment).", "4": "As level 3, plus money already spent this year on a failed or partial attempt at this problem. The internal argument for spending has been won." } }, { "id": "exec_sponsor", "group": "commercial", "name": "Executive sponsor", "weight": 4, "why": "The sponsor unblocks legal, security and data access, which are the three things that actually delay these engagements. Weighted lower than metric owner deliberately: sponsors move on, and a sponsor without an owner produces a signed pilot nobody runs.", "anchors": { "0": "No executive is aware of the project.", "1": "An executive is aware and supportive in the abstract; has not spent any political capital.", "2": "An executive has been briefed, asks for updates, and would sign off if asked.", "3": "A named executive has intervened at least once to remove a specific blocker (data access, a security review, a budget conversation), and the intervention is logged.", "4": "The executive raised the problem, is publicly associated with solving it, and has said so to their own board or leadership team." } }, { "id": "security_legal_path", "group": "commercial", "name": "Security and legal path", "weight": 8, "why": "Legal sits second in this pipeline because MSA and DPA execution gates the engagement rather than concluding it. Deployment topology and weights-ownership expectations must be surfaced now: a customer who assumes exclusive perpetual rights to the adapted weights and hears otherwise in month three is a renegotiation, not a renewal.", "anchors": { "0": "Data cannot leave their perimeter under any arrangement, or their weights-ownership expectations are incompatible with what can be offered and the sponsor will not move.", "1": "No security review has started, no comparable vendor precedent exists, and requirements are being invented as you ask.", "2": "Requirements are articulated (VPC, region, retention, no-train clauses), and a review process exists but has not begun.", "3": "A comparable AI vendor has been through their process, the path is known, counsel is engaged, and deployment topology and weights ownership have been discussed explicitly and are compatible.", "4": "MSA and DPA are in redlines or executed, security review is scheduled or passed, and the deployment topology is agreed in writing." } }, { "id": "forcing_function", "group": "commercial", "name": "External forcing function", "weight": 8, "why": "Internal deadlines slip without consequence; external ones do not. A date set by a regulator, a customer contract, a competitor launch or a board commitment is the difference between a POC that finishes and one that is still being scheduled two quarters later.", "anchors": { "0": "Nothing happens if this slips a year. The driver is curiosity or a strategy document.", "1": "An internal roadmap date with no external consequence attached.", "2": "An internal date tied to a planning or budget cycle, which creates real pressure at one point in the year.", "3": "An external date: a customer commitment, a contractual SLA, a regulatory deadline, a competitor already shipping the capability.", "4": "As level 3, and missing it has a quantified cost the sponsor can state — churned revenue, a penalty, a lost renewal, a launch that cannot happen." } }, { "id": "expansion_surface", "group": "commercial", "name": "Expansion surface", "weight": 5, "why": "A first environment is worth building when the second, third and fourth are visible behind it — the same verifier pattern, the same traces, adjacent teams. A single isolated task can be technically perfect and still not repay the delivery cost of building it.", "anchors": { "0": "One task, one team, no adjacent workflow shares its data, verifier pattern or tooling.", "1": "Adjacent workflows exist in theory but belong to a different organisation with a different budget and no relationship.", "2": "Two or three adjacent workflows share the trace source or verifier pattern; nobody has discussed them.", "3": "Named adjacent workflows with named owners, and the sponsor treats this engagement as the first of several.", "4": "As level 3, plus a stated intent to build an internal capability — their own environments, their own evals, continuous retraining — rather than to buy a single result." } } ], "bands": [ { "min": 78, "max": 100, "label": "Build", "action": "Move to scoping. Commit named research support, write the environment and verifier design into the proposal, and put a dated POC gate with the metric owner's agreed threshold on it." }, { "min": 62, "max": 77, "label": "Qualified, conditional", "action": "Advance, but name the two dimensions with the largest weighted point loss — max weight × (4 − score), breaking ties toward the technical group — as written POC entry conditions with owners and dates. Do not start compute until both are cleared." }, { "min": 46, "max": 61, "label": "Scope down", "action": "Do not sell the workflow as described. Find the narrowest sub-task that would itself score above 62, and sell that as a paid evaluation or a single environment build with a fixed price and a fixed end." }, { "min": 30, "max": 45, "label": "Defer", "action": "No POC. Offer evaluations or on-demand compute if genuinely useful, name the artefact that must exist before re-scoring, set a re-qualification date, and stop talking about post-training until then." }, { "min": 0, "max": 29, "label": "Decline", "action": "Decline in the next meeting, verbally, naming the missing artefact rather than the customer. Recommend the cheaper thing that would actually work now, and leave an explicit trigger that brings them back. Log the reason in PIG as closed_lost with the disqualifying dimension recorded." } ], "disqualifiers": [ { "name": "No computable verifier and no path to one", "test": "verifier_quality scores 0; or it scores 1 and the funding-or-staffing question for the labelled set was put directly to the sponsor in a meeting and the refusal is recorded on the demand record in PIG.", "why": "Caps the deal at Decline (0-29). Training optimises a reward. With no reward computable without a human reading each output, there is nothing to optimise and no way to demonstrate improvement afterwards. Every other strength in the deal is irrelevant to this fact, which is exactly why an average hides it and why this routes to the decline conversation rather than to a re-qualification date." }, { "name": "No rights to train on the data", "test": "Counsel or the contract owner will not confirm in writing that the traces may be used for model training, and the blocking agreements are with third parties the customer cannot renegotiate inside the deal timeline.", "why": "Caps the deal at Defer (30-45), with counsel's written confirmation named as the artefact that must exist before re-scoring. This surfaces at DPA execution, which in this pipeline is stage two; discovering it after a scoping document exists wastes counsel's time and the sponsor's credibility. There is no engineering workaround for a contract clause, but there is sometimes a first-party dataset, so this is a defer rather than a decline." }, { "name": "Nobody states the number", "test": "The person named as metric_owner declines to state a numeric success threshold after being asked directly in a meeting, or two scheduled meetings across at least three weeks, both logged in PIG, produced no named owner at all.", "why": "Caps the deal at Defer (30-45), with a written threshold from a named person as the artefact that must exist before re-scoring. This is gated on the stated threshold rather than the org chart on purpose: at qualification an owner is often still being found, and a merely unnamed owner should stay a low score on a 9-weight dimension rather than kill the deal. A named owner who will not commit to a number is different — there is then nobody to accept the result, defend it internally, or convert it into expansion, and these POCs succeed technically and then disappear." }, { "name": "Incompatible weights-ownership or deployment expectation", "test": "The customer requires terms on adapted weights, exclusivity, or air-gapped deployment that cannot be offered, and the executive sponsor will not revisit the requirement.", "why": "Caps the deal at Decline (0-29). This is a fixed constraint, not a negotiation about price, and it does not soften after a successful POC. There is no artefact the customer can produce to unblock it, which is what separates it from the two defer-capping disqualifiers." } ] }, "notes": "Operational metadata, not a summary of the body. Version {{version}}, last revised {{date}} — both assumptions to be filled on adoption. Owner: the forward deployed lead on the deal; the completed scorecard is attached to the demand record in PIG, and the score at the POC gate is frozen so it can be compared against the outcome later. Re-run at three points: end of technical discovery (first score), after environment design (verifier quality and environment constructibility move most often, and usually down), and whenever a disqualifier test changes state. Changed in this version: the 62-77 band now selects on weighted point loss, weight × (4 − score), rather than on lowest weighted score, which selected strong low-weight dimensions; disqualifiers now cap at named bands instead of uniformly at Defer; disqualifier 3 is gated on a stated numeric threshold rather than on an unnamed owner; verifier quality and measured baseline anchors now require a measured judge-agreement number, a training-disjoint held-out set, and a trivial-policy check." }