Files
arena/environments/tera_spatial/README.md
T
kartiandClaude Opus 5 4147de0282 tera-crow-nav: the first spatial taskset, and the hold sweep that sized its budget
The bridge becomes a runnable taskset. The crux is that these are continuous
control environments and the agent emits text: the model issues a control
decision that is held for N steps, and the hold length is measured rather than
guessed. At one decision every 32 steps the scripted baseline returns -2.807 and
reaches the waypoint 0% of the time, against 5.484 and 100% at every frame.

Open loop — one decision for the whole flight — returns -3.764 against an
inaction floor of -4.161. A bird that sets a course and leaves does no better
than one that does nothing, which is why crow-nav carries the strongest rule-4
claim of the five spatial environments. The budget bites from both ends: 8 turns
reach the goal 44% of the time, 12 reach 88%, 14 reach 100%.

32 episodes against Qwen3.8-27B: mean 0.4603, no malformed replies in 412, every
episode replaying to an identical checksum, and 16 of 32 spending all 16 turns.
The gate fires 43.75% of the time and separates the four scenarios cleanly —
whatever is wrong with the four dead data-environment gates, it is not that
binary gates cannot discriminate.

The oracle comes from the worker's own op. Nothing recomputes in Python a number
that came out of TypeScript, least of all the reward denominator.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 17:44:46 -07:00

198 lines
9.9 KiB
Markdown

# tera-spatial
Arena's bridge to Tera's four renderer-independent spatial environments.
The environments themselves live in the `tera` repository, in TypeScript, as
`tera.arena/v1`. They are not reimplemented here and they never will be. This
package vendors their import closure, runs it in a resident `node` worker, and
transports results. The rule the whole thing is built around:
> **Never recompute in Python a number that came out of TypeScript.**
Rewards, per-step state checksums, scenario materialisation and the four
baseline returns that every reward is normalised against are all computed once,
inside Tera, and only ever carried across the pipe.
## What is here
| Path | What it is |
|---|---|
| `tera_spatial/worker.mjs` | the resident NDJSON worker: `reset`, `step`, `snapshot`, `restore`, `trace`, `replay`, `oracle`, `checksum` |
| `tera_spatial/bridge.py` | `TeraWorker` / `TeraEpisode` — transport, and nothing else |
| `tera_spatial/vendor/tera/` | the 23-file, 283 KB import closure of `src/arena/index.ts` |
| `tera_spatial/hashes.py` | SHA-256 of all 23, generated; the worker refuses to start if the tree has drifted |
| `tera_spatial/closure.py` | the import walk both the sync script and the gate use |
| `scripts/sync_tera.py` | reproduces the vendoring from a `tera` checkout |
| `tests/test_replay.py` | the replay gate |
| `tera_spatial/spatial.py` | the shared half of a spatial taskset: the resident worker, the reply grammar, the episode loop, the reward |
| `tera_crow_nav/` | the first taskset — `tera-crow-nav`, fly a crow to a waypoint on sixteen replies |
| `measure_hold.py` | the hold and turn-budget sweep the two integers came from |
| `tests/test_crow_nav.py` | the taskset's gate: the grammar, the reward, and a flight flown off the rendered panel |
`tera-office-nav` and `tera-california-flight` follow. `tera-drive-101` ships
labelled a bridge-correctness environment or not at all: its return is monotone in
throttle and any positive constant scores about 0.95 of the baseline, which is a
number with no house rule 4 behind it.
## The crux: a language model cannot emit sixty control vectors a second
These are continuous-control environments and the agent is text. So a reply is a
DECISION, not a frame: one control vector plus the number of simulation steps to
hold it for, and then the model is shown where the vehicle ended up.
The two integers that decide everything — how many replies, and how long one may
be held — are measured. `measure_hold.py` flies Tera's own scripted baseline at
every refresh rate over sixteen scenario x seed episodes:
```
hold return worst steps calls goal turns hold return goal
1 5.484 5.059 66.1 66.1 100% 1 12 0.031 0%
2 5.484 5.059 66.1 33.2 100% 2 12 0.317 0%
3 5.484 5.060 66.1 22.4 100% 4 12 1.039 0%
5 5.479 5.056 66.1 13.5 100% 6 11 1.700 0%
8 4.906 0.416 97.6 12.6 94% 8 11 3.420 44%
12 4.401 -0.467 127.6 11.2 88% 10 8 4.595 75%
16 3.643 -0.577 153.4 10.1 75% 12 7 4.974 88%
24 1.392 -0.689 280.1 12.1 44% 14 6 5.409 100%
32 -2.807 -2.923 300.0 10.0 0% 16 5 5.479 100%
64 -3.137 -3.658 238.3 4.0 0% 20 4 5.483 100%
300 -3.764 -4.906 149.7 1.0 0% 32 3 5.484 100%
inaction floor: -4.161 crow-nav-v1, 16 episodes per row
```
(the two tables `measure_hold.py` prints, set side by side; rows at hold 48, 100
and 150 are dropped for width.)
The left table is house rule 4's premise: **open-loop fails outright.** One
decision held for the whole flight returns -3.764 against an inaction floor of
-4.161 and reaches the waypoint none of the time — a bird that sets a course and
leaves does no better than a bird that does nothing. A plan is not a policy in a
world with momentum, which is why the grammar is one vector per reply and not a
program.
The right table is the conclusion, and it is the reason `max_turns = 16`: the
budget BINDS all the way up to it. At eight turns the best possible constant hold
reaches the waypoint 44% of the time, at twelve 88%, and only at fourteen does it
reach every one. Sixteen is the first budget at which the ceiling is reachable,
and it is one turn of slack, not fifteen.
`max_hold = 12` is the longest hold at which the baseline still reaches most
waypoints. `16 x 12 = 192` against a 300-step simulator cap, so what binds is the
turn budget and never the simulator's own.
## The rewards
Two, both products, both from numbers TypeScript computed:
| | |
|---|---|
| `flight` (0.70) | the episode's return against nine tenths of what Tera's scripted controller returned on the same scenario at the same seed |
| `economy` (0.30) | that quality, multiplied by having arrived and by how directly |
`flight` is smooth because goal attainment is already the biggest term inside it —
the simulator pays +3 of a ~5.5 return for reaching the waypoint — so a near-miss
scores about four tenths and a wander scores zero.
⚠️ **Goal attainment multiplies and is never a term of its own.** It is already
inside the return as the success bonus, and a separate additive `gate` beside the
sum double-counts it. That is the exact failure `probe.py` records having paid
for. Efficiency multiplies too: a flight that never arrived cannot be quick about
arriving, and standing beside `flight` it would pay a crash for being fast.
`gate`, `quality`, `step_ratio`, `ts_return`, `hold_mean`, `malformed_turns` and
`replay_ok` are recorded and never rewarded.
**`replay_ok` is the bridge's claim asserted once per episode**, not once per
commit: every graded flight is replayed in a fresh TypeScript environment and the
metric is whether it reached the same FNV-1a-64 checksum, the same step count and
the same return. A divergence is recorded, never raised — an episode that flew is
data, and losing it to an assertion would hide the thing the metric exists to
surface.
## Running it
```bash
uv run --project environments/tera_spatial eval @ configs/tera_crow_nav.toml \
--model brain-qwen38-dspark \
--client.base-url http://100.127.247.67:8001/v1 \
--client.api-key-var SPARK_API_KEY \
--no-push --no-rich -c 8 -o outputs/run-<stamp>/tera-crow-nav
```
⚠️ `eval` exits 0 even when every rollout errors. Count non-empty `rewards`; see
`docs/FIRST_EVAL.md`.
## What it measured
`outputs/run-20260821-1627/tera-crow-nav`, 16 tasks x 2 rollouts, model
`brain-qwen38-dspark` on spark-1, 12 minutes at `-c 8`. **32 episodes attempted,
32 scored, 0 errored, 0 provider errors.**
| | |
|---|---|
| `reward_mean_attempted` | **0.4603** |
| `flight` (0.70) | 0.4894 |
| `economy` (0.30) | 0.3926 |
| waypoint reached | 14 of 32 |
| flight-envelope contact | 2 of 32 |
| out of turns, still flying | 16 of 32 |
| malformed replies | **0** of 412 |
| `replay_ok` | **1.000 on all 32** |
Three things in that table are worth more than the headline.
**The turn budget binds, in the run and not only in the sweep.** Turns used were
7, 8, 9, 10, 12, 14 and 16 — mean 12.9, and **sixteen of the thirty-two rollouts
spent every reply they had** and were still airborne when the budget ran out. The
one that scored zero on `train-east-crosswind` did so by orbiting: 128 steps
flown, yaw -90°, and 52.8 m still between it and the waypoint on its last reply.
**The gate carries a gradient.** Goal attainment fired on 43.75% of rollouts.
Across the first seven Arena environments a `gate` component scored exactly 0.000,
max 0.000, in four of them — a quarter to a third of the reward mass with nothing
in it. Here it discriminates, and per-scenario it discriminates sharply:
`train-north-climb` 0.863, `dev-west-return` 0.728, `dev-south-descent` 0.250,
`train-east-crosswind` 0.000.
**`eval.log` and `traces.jsonl` agree.** The mean over the log's own 32 `reward=`
lines is 0.4603, the same number as `sum(score x weight)` over the file. That is
not free: rewards recorded after `interaction.close()` land in the file and not in
the log, and the first cut of this env printed `reward=0.000` for an episode that
scored 1.000. Grade before closing.
## The gate
```bash
uv run --project environments/tera_spatial \
python -m unittest discover -s environments/tera_spatial/tests -v
```
32 episodes — 4 environments x 4 public scenarios x 2 seeds — each driven from
Python one JSON step at a time, then replayed in a fresh TypeScript environment.
All 32 must reach an identical FNV-1a-64 checksum. Then seven forgeries per
environment, each tried twice: raw, where the envelope checksum catches it, and
**re-sealed with a checksum Tera itself recomputed**, where only `replay()`
re-running the simulator can. All must be rejected.
`test_crow_nav.py` adds fifteen more. The reward's floor and ceiling are flown
rather than asserted from a fixture, and the strongest of them is the rendering's
own claim: a policy that reads **nothing but the rendered panel** — the same
characters the model is shown, parsed back out of them with a regex — reaches all
four waypoints inside the shipped budget. An observation encoding that stops
carrying enough to fly on stops passing that test.
`node` >= 22.18 is required — the vendored sources are raw `.ts` and are
type-stripped, not compiled. Set `TERA_NODE` to point at a specific binary.
## Re-vendoring
```bash
uv run python scripts/sync_tera.py --tera ~/repos/gitea/tera # copy + regenerate hashes
uv run python scripts/sync_tera.py --check # CI: is the tree clean?
```
The closure is walked, not declared, so a new import in `tera` travels with it.
A bare specifier is a hard failure: the wheel has no `node_modules` and a
simulator that needs one is not a simulator that ships.