4.0 KiB
Eval data
Three tiers. The split is load-bearing — see docs/DECISIONS.md#d3.
| dir | contents | visibility |
|---|---|---|
public/ |
a few example samples | shown on the site, safe to leak |
private/ |
the real holdout | never leaves this repo |
canary/ |
contamination probes (carrier + detectors) | never leaves this repo |
Sample format
One JSON object per line (JSONL), one file per family:
private/agent_ops.jsonl, private/repo_edit.jsonl, ...
{
"id": "agent_ops/0001",
"family": "agent_ops",
"split": "test",
"canary": "KBENCH-CANARY-3f9a1c2e",
"input": "The prompt sent to the model.",
"target": "The expected answer, or null for rubric-graded samples.",
"scorer": "includes",
"rubric": null,
"metadata": {
"source": "openclaw",
"authored": "2026-07-27",
"tags": ["tool-use", "multi-turn"]
}
}
Fields
-
id— stable and never reused. Per-sample regression tracking across checkpoints keys on this; renumbering breaks every historical comparison. -
split—train|dev|test. Assigned at authoring time and never changed. Atestsample that leaks into fine-tuning makes every subsequent number a lie you have no way to detect. -
canary— provenance only, stamped bykbench add. It is not the contamination control: it never reaches a model, so nothing can reproduce it. The real control is thecanary/tier — see below. -
scorer— how this sample is graded:scorer grading judge exactoutput matches target exactly no includesoutput contains target no regexoutput matches target as a regex no rubricmodel-graded against rubrictextyes Prefer deterministic scorers. Every rubric sample adds judge cost per run and judge drift between runs, which confounds the checkpoint comparisons this bench exists to make. Reach for
rubriconly when the property is genuinely not checkable mechanically.
Contamination probes (canary/)
The leak vector is named in the authoring rules below: prompts reach third-party judge and baseline models no matter how private this repo is. If a provider trains on inference traffic, our eval set is in their next model and every signal score becomes a measurement of memorisation.
The probe tier catches that, and it only works as a pair:
- One carrier. Its prompt contains the canary GUID in full. It always passes, and passing is not the point — its job is to put the GUID into the traffic so a training pipeline can pick it up.
- Two or more detectors. Their prompts withhold the GUID and ask for it cold. A model that has never seen our traffic cannot answer. One that has, can.
Scoring is inverted here. canary_score of 0 is the healthy result.
Anything above 0 means signal scores for that target are void and the eval set
needs regenerating. Never read it as a capability number.
A detector for a GUID that no carrier ever transmitted is unanswerable by construction: it can never fire, so it proves nothing while looking exactly like a working control.
Authoring rules
Scrub before you commit. No API keys, tokens, tailscale hostnames or IPs, customer data, or personal information. This content gets sent to third-party judge and baseline models regardless of how private this repo is. Repo privacy is defence in depth, not the control.
Derive from real work. The value of this bench is that its samples come from things we actually asked a model to do. An invented task measures an invented capability.
Write the target before you see a model's answer. Grading against an answer you just read is how a bench quietly turns into a rubber stamp.
Adding samples
kbench add agent_ops --input-file prompt.txt --target "expected" --scorer includes
kbench add agent_ops --interactive
This assigns the next id, stamps a canary, defaults the split, and validates the record — do not hand-edit the JSONL.