docs: the scope, the plan, the post-review execution plan, and the eval runbook
MULTI_TURN.md scopes the multi-turn gap and records the finding that made it cheap: multi-turn needs no container and no real harness. vf.Env.run programs the control flow host-side and the simulator runs in-process as the user, with the harness left null. PLAN.md carries the settled decisions. EXECUTION.md is the post-adversarial-review plan — every position in it survived a reviewer, and where a design spec was overturned the corrected position is what is written down. Two corrections are marked in place rather than silently rewritten, because both were believed and acted on: - The provenance hole is src/arena/index.ts, not offices/sites.ts. Mutating offices/sites.ts produces a bit-identical episode; index.ts re-exports the scripted baselines that ARE the reward denominator, and a halved office-nav baseline moves return 4.744 -> 4.098 while every tera hash gate prints ok. - office-jobs-v1 is not phantom. It was absent at 13:00 and present at 15:50, because another session fast-forward-merged it into the shared tera tree mid-session. The counts are twelve, five and twenty. docs/FIRST_EVAL.md is the runbook for the first evaluation Arena ever completed: seven environments, 224 episodes, 223 scored, against brain-qwen38-dspark. It warns that the mean is sum(score x weight) — an unweighted mean over components is a different number, and it is how six of seven means were first misreported. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,243 @@
|
||||
# Arena build plan
|
||||
|
||||
Companion to `MULTI_TURN.md` (the scope). This is the sequencing: what gets built,
|
||||
in what order, what can run in parallel, and what collides.
|
||||
|
||||
Status: plan, 2026-08-21. ✅ = settled. ❓ = blocks the work under it. ⛔ = parked by decision.
|
||||
|
||||
## Where Arena sits
|
||||
|
||||
Lumbridge is an open-source company shipping six self-hostable things:
|
||||
**Compute** (governed compute for unified-memory hardware), **Bench** (model quality
|
||||
and serving measured against one pinned host × quant × config tuple), **Simulate**
|
||||
(the Tera/Spaces 3D engine), **Stream** (a headless broadcast rig), **Forge**
|
||||
(post-training as a content-addressed build), and **Arena**. Four of the six are
|
||||
public. `apps/web/src/data/projects.ts` in `lumbridge-v4` is the canonical list.
|
||||
|
||||
**Arena is the RL-environment project — yes, `/arena` is the section.** It is
|
||||
appended last in the project list for numbering reasons, but semantically it sits
|
||||
beside Bench and feeds Forge: Bench measures a model, Arena produces the gradient,
|
||||
Forge runs the training. That triangle is the reason Arena is worth doing well
|
||||
rather than as a side quest — we own all three corners, which nobody on the Hub does.
|
||||
|
||||
Arena spans **two repositories on one page**, and that is deliberate:
|
||||
|
||||
- `git.karti.ai/lumbridge-public/arena` — the Python data environments (this repo)
|
||||
- `git.karti.ai/lumbridge-public/tera` — the TypeScript spatial environments
|
||||
|
||||
## What we take from Prime Intellect, and what stays ours
|
||||
|
||||
All of PI's stack is open source and we use it rather than rebuild it. Checkouts are
|
||||
at `~/vendor/prime-intellect/` (46 repos; see its README).
|
||||
|
||||
| We use, unmodified | We build |
|
||||
|---|---|
|
||||
| `verifiers` — the v1 spec, the taskset/harness/trace model, `vf.Env` | The environments |
|
||||
| `prime-rl` — the trainer traces feed | The house rules |
|
||||
| `prime-agent` (`pi`) — a harness, already embedded in Piggy | `probe.py` — the degenerate-policy gate |
|
||||
| `prime` CLI — `env push`, `env action` | The Tera bridge |
|
||||
| The Environments Hub — distribution | lumbridgecorp.com/arena — documentation |
|
||||
|
||||
We do **not** build a harness, a trainer, or a hub. Every one of those exists,
|
||||
is permissively licensed, and is better than what we would write.
|
||||
|
||||
## Decisions
|
||||
|
||||
**✅ 1 — Hub org handle: `lumbridge`.** Settled 2026-08-21. Environments will
|
||||
publish as `lumbridge/<env-name>`.
|
||||
|
||||
**✅ 2 — Attribution: Lumbridge as author.** Settled 2026-08-21. `karti.ai` stays
|
||||
the code host (`git.karti.ai`), Lumbridge is the author on every surface. One
|
||||
identity across all six projects; no per-project byline.
|
||||
|
||||
**⛔ Hub push is ON HOLD.** Decided 2026-08-21: keep building, publish nothing to
|
||||
the Prime Intellect Hub yet. Workstream E is parked — not cancelled, not blocked on
|
||||
anything technical. `PRIME_API_KEY` is therefore not urgent.
|
||||
|
||||
**✅ 7 — Production brain: Qwen3.8-27B.** Settled 2026-08-21. spark-1 serves
|
||||
`brain-qwen38-dspark`; that is the default. The Nemotron baseline is ARCHIVED, not
|
||||
deleted, and Arena's seven scores get re-run and republished against Qwen.
|
||||
⚠️ **`--model brain` does NOT 404 — it silently succeeds and records `brain` in the
|
||||
trace while the box serves Qwen.** Verified by direct call 2026-08-21. Pin
|
||||
`brain-qwen38-dspark` in every reproduction command, everywhere.
|
||||
|
||||
**✅ 8 — `tera-drive-101` ships labelled.** Settled 2026-08-21. Published as a
|
||||
**bridge-correctness** environment with NO rule-3 and NO rule-4 claim (its return is
|
||||
monotone in throttle; any positive constant scores ~0.95+). Site says **four**
|
||||
spatial environments. It is not listed as a fifth RL environment.
|
||||
|
||||
**✅ 9 — `app.lumbridgecorp.com` is the ONLY authenticated origin.** Settled
|
||||
2026-08-21. The apex loses `/login`, `/signup`, `/app`, `/go`. All existing sessions
|
||||
sign out once on cutover; accepted cost for three accounts.
|
||||
|
||||
**✅ 10 — Adopted defaults** (technical, recommended by the synthesis, settled
|
||||
2026-08-21): the agent seat is named **`agent`** everywhere (matching
|
||||
`SingleAgentEnvConfig`) — it becomes the TOML key in every `configs/*.toml` AND every
|
||||
prime-rl `rl.toml`, and renaming later parses fine while silently having no seat;
|
||||
**D lives in the `arena` repo**, not `tera` (amend the parallelization map);
|
||||
**A ships the FLAT freeze**, with per-item recorded as the identified strengthening;
|
||||
**Arena transcripts are ADMIN-ONLY** (they carry the held-out slice and oracle
|
||||
answers — a contamination question, not an org-politics one); **the dashboard is the
|
||||
archive and `eval-results.json` a curated summary emitted from it**, never two
|
||||
hand-maintained files; **Forge integrates prime-rl with LoRA RL first**.
|
||||
|
||||
**✅ 3 — `office-jobs-v1` EXISTS; the counts are twelve / five / twenty.** Resolved
|
||||
2026-08-21 15:50 when another session fast-forward-merged it into the shared `tera` tree
|
||||
mid-session. It is committed, in `ARENA_MANIFESTS`, and passes `arena:source-hashes`.
|
||||
The site's original copy was correct. ⚠️ `tera` has live concurrent sessions — re-check
|
||||
its HEAD before acting on any claim about its contents.
|
||||
|
||||
~~**❓ 3 — `office-jobs-v1`.** Advertised on the live page and in `seo.ts`; absent
|
||||
from `ARENA_MANIFESTS`. Build it, or correct the page to eleven. Blocks any copy
|
||||
that cites the count, including a Hub listing.
|
||||
|
||||
**✅ 4 — Interactive envs get their own taskset id.** Settled 2026-08-21. The first
|
||||
is **`grand-exchange-live`**, a separate wheel beside `grand-exchange`. One-shot
|
||||
scores stay comparable across time; the two forms are measured, not conflated.
|
||||
|
||||
**✅ 5 — App is a subdomain: `app.lumbridgecorp.com`.** Settled 2026-08-21. Matches
|
||||
the two-surface model (public site + private platform). Public site stays a site.
|
||||
|
||||
**✅ 6 — Compute policy.** Settled 2026-08-21. **Training runs on rented GPUs**
|
||||
(GH200-class, from Prime Intellect) — rented because it is faster, and because
|
||||
renting keeps the hardware choice open per run (the GB300-vs-alternatives call is
|
||||
one we want to keep making). **Inference in production runs on the DGX Spark.**
|
||||
We do **not** train on the Spark.
|
||||
|
||||
## Workstreams
|
||||
|
||||
Five, with the dependency edges called out. **A and D can start immediately and in
|
||||
parallel. B must land before C or E.**
|
||||
|
||||
### A — `grand-exchange` interactive *(start now)*
|
||||
|
||||
The first tier-1 environment: quote → see fills → re-quote over a turn budget,
|
||||
inventory and unfilled orders carrying between turns. Chosen because the market
|
||||
simulation already steps, the counterweight (liquidity in gp/tick, the random-walk
|
||||
trap) already exists, and it is the smallest change that yields a genuinely new
|
||||
measurement.
|
||||
|
||||
Deliverable: a `vf.Env` subclass exported from the wheel's `__init__.py`, a config,
|
||||
and a measured run against spark-1's `brain`. Pattern to copy:
|
||||
`verifiers/v1/tasksets/textarena/taskset.py`.
|
||||
|
||||
This workstream is where **rule 4 — the budget must bind** gets written for real.
|
||||
Whatever it learns about degenerate trajectories is the input to B.
|
||||
|
||||
Depends on: nothing — ✅4 settled.
|
||||
|
||||
### B — `probe.py` grows degenerate *policies* *(after A's first draft)*
|
||||
|
||||
Today the probe scores four fixed strings per environment. Tier 1 needs policies:
|
||||
the **staller** (never acts), the **spammer** (acts every turn at random), the
|
||||
**impatient** (answers turn 1 with no observation), the **exhaustive** (spends the
|
||||
whole budget, then answers). Same gate shape — inaction 0.000, oracle 1.000, exit 1
|
||||
otherwise — plus a new assertion that **exhaustive scores below oracle**. If it does
|
||||
not, rule 4 is violated and the environment does not ship.
|
||||
|
||||
Do this before building a second interactive environment, not after. It is the
|
||||
differentiator: nothing across `prime-envs` (88) or `community-environments` (109)
|
||||
ships a degenerate-policy gate.
|
||||
|
||||
Depends on: A (needs one real interactive env to gate against).
|
||||
Blocks: C, E.
|
||||
|
||||
### C — More interactive environments *(after B)*
|
||||
|
||||
In rough order of value:
|
||||
|
||||
1. `drop-table-inference` — buy observations under budget; measures when to stop looking
|
||||
2. `bot-detection` — request labels one at a time (active learning)
|
||||
3. `fault-localisation` — the log is not given; query it by service, window and level
|
||||
4. `redaction-pressure` — two agents, redactor vs extractor; the counterweight becomes an opponent
|
||||
|
||||
Each is independently parallelizable **once B exists**. `redaction-pressure` is the
|
||||
most interesting and the least certain — a genuine multi-agent env, not a stepped one.
|
||||
|
||||
Depends on: B.
|
||||
|
||||
### D — The Tera bridge *(start now, independent repo)*
|
||||
|
||||
One Python `vf.Env` driving the existing TypeScript `tera.arena/v1` package over a
|
||||
process boundary, turning `step()` observations into `interaction.turn()` messages
|
||||
and existing component rewards into `trace.record_reward` calls.
|
||||
|
||||
Validated by replay: every bridged episode must `replay()` through the TS package to
|
||||
an identical FNV-1a-64 checksum, or the bridge is wrong. That property is rare and
|
||||
the bridge should be built around it.
|
||||
|
||||
Lands four environments at once and makes our strongest existing asset publishable.
|
||||
Larger and riskier than A — the process boundary and the action encoding are the
|
||||
unknowns. **Touches the `tera` repo, not this one, so it does not collide with A/B/C.**
|
||||
|
||||
Depends on: nothing.
|
||||
|
||||
### E — First Hub push *(PARKED)*
|
||||
|
||||
Push once there are three to four interactive environments alongside the seven
|
||||
one-shot ones. `prime env push <name> -v PUBLIC` to our own namespace is primary —
|
||||
we keep authorship and versioning. A PR into `community-environments` is optional
|
||||
extra distribution that hands the namespace to `primeintellect`; secondary.
|
||||
|
||||
Then update `lumbridgecorp.com/arena` to document the same wheels.
|
||||
|
||||
⛔ On hold by decision, 2026-08-21. Resume when Karti says so; ❓3 still gates any
|
||||
copy that cites the environment count.
|
||||
|
||||
## ⚠️ The collision point
|
||||
|
||||
**`tests/test_probe.py` and the wheel-build loop in `.gitea/workflows/ci.yml`
|
||||
hardcode the environment list.** Every workstream that adds an environment edits
|
||||
both files. Two sessions adding environments concurrently *will* conflict there, and
|
||||
CI goes red if either is missed.
|
||||
|
||||
Mitigation, in order of preference: (1) B refactors both to discover environments
|
||||
from `environments/*/pyproject.toml` instead of a literal list — do this as part of
|
||||
B and the collision disappears; (2) failing that, one session at a time touches
|
||||
those two files.
|
||||
|
||||
Separately: this repo already took divergent remote pushes from concurrent sessions
|
||||
on 2026-08-19. **Merge, never force.** Use a worktree per workstream.
|
||||
|
||||
## Parallelization map
|
||||
|
||||
```
|
||||
now ──┬── A (grand-exchange interactive) ──┬── B (probe → policies, de-hardcode CI) ──┬── C1 drop-table
|
||||
│ │ ├── C2 bot-detection
|
||||
│ │ ├── C3 fault-localisation
|
||||
│ │ └── C4 redaction (2-agent)
|
||||
│ │
|
||||
└── D (Tera bridge, tera repo) ──────┴──────────────────────────────── E (Hub push) ⛔ PARKED
|
||||
```
|
||||
|
||||
Two sessions can start today: **A** in this repo, **D** in `tera`. They share no
|
||||
files. Everything downstream of B fans out four ways.
|
||||
|
||||
## Preconditions, verified on amd-server
|
||||
|
||||
- `uv run eval @ configs/<env>.toml --dry-run` → exit 0 ✓
|
||||
- verifiers 0.3.0, all 13 harnesses and 4 runtimes present ✓
|
||||
- Docker 29.1.3 (tier 2/3 only; tier 1 needs none) ✓
|
||||
- spark-1 `brain` at `http://100.127.247.67:8001/v1`, model id `brain` ✓
|
||||
- `prime` CLI 0.6.25 ✓
|
||||
- `PRIME_API_KEY` ✗ — blocks E, which is parked anyway
|
||||
|
||||
---
|
||||
|
||||
## ⚠️ Superseded by `EXECUTION.md`
|
||||
|
||||
The workstream descriptions and the parallelization map above were written before the
|
||||
2026-08-21 design workflow. **`EXECUTION.md` is now authoritative** for sequencing,
|
||||
file ownership and risk. Three structural corrections it makes:
|
||||
|
||||
1. **B splits.** **B1** (de-hardcode environment discovery) does NOT depend on A and
|
||||
ships immediately; the CI env list was never a matrix, just a bash `for` loop.
|
||||
**B2** (degenerate-policy machinery) was found NOT FEASIBLE as written and must be
|
||||
calibrated against A's *measured* engine, not A's first draft.
|
||||
2. **D is no longer file-independent.** It lives in `arena` and edits `probe.py`,
|
||||
`ci.yml`, `test_probe.py`, `README.md` plus two `lumbridge-v4` files.
|
||||
3. **The top risk is that every Arena reward SATURATES**, so house rule 4
|
||||
("exhaustive < oracle") may be structurally unreachable. If it is, every
|
||||
interactive environment is unfalsifiable and the reward shape must change before
|
||||
C1–C4 are built on it. This is also a compute-policy constraint: a non-binding
|
||||
turn budget burns rented GPU hours measuring nothing.
|
||||
Reference in New Issue
Block a user