# tera-spatial Arena's bridge to Tera's four renderer-independent spatial environments. The environments themselves live in the `tera` repository, in TypeScript, as `tera.arena/v1`. They are not reimplemented here and they never will be. This package vendors their import closure, runs it in a resident `node` worker, and transports results. The rule the whole thing is built around: > **Never recompute in Python a number that came out of TypeScript.** Rewards, per-step state checksums, scenario materialisation and the four baseline returns that every reward is normalised against are all computed once, inside Tera, and only ever carried across the pipe. ## What is here | Path | What it is | |---|---| | `tera_spatial/worker.mjs` | the resident NDJSON worker: `reset`, `step`, `snapshot`, `restore`, `trace`, `replay`, `oracle`, `checksum` | | `tera_spatial/bridge.py` | `TeraWorker` / `TeraEpisode` — transport, and nothing else | | `tera_spatial/vendor/tera/` | the 23-file, 283 KB import closure of `src/arena/index.ts` | | `tera_spatial/hashes.py` | SHA-256 of all 23, generated; the worker refuses to start if the tree has drifted | | `tera_spatial/closure.py` | the import walk both the sync script and the gate use | | `scripts/sync_tera.py` | reproduces the vendoring from a `tera` checkout | | `tests/test_replay.py` | the replay gate | | `tera_spatial/spatial.py` | the shared half of a spatial taskset: the resident worker, the reply grammar, the episode loop, the reward | | `tera_crow_nav/` | the first taskset — `tera-crow-nav`, fly a crow to a waypoint on sixteen replies | | `measure_hold.py` | the hold and turn-budget sweep the two integers came from | | `tests/test_crow_nav.py` | the taskset's gate: the grammar, the reward, and a flight flown off the rendered panel | `tera-office-nav` and `tera-california-flight` follow. `tera-drive-101` ships labelled a bridge-correctness environment or not at all: its return is monotone in throttle and any positive constant scores about 0.95 of the baseline, which is a number with no house rule 4 behind it. ## The crux: a language model cannot emit sixty control vectors a second These are continuous-control environments and the agent is text. So a reply is a DECISION, not a frame: one control vector plus the number of simulation steps to hold it for, and then the model is shown where the vehicle ended up. The two integers that decide everything — how many replies, and how long one may be held — are measured. `measure_hold.py` flies Tera's own scripted baseline at every refresh rate over sixteen scenario x seed episodes: ``` hold return worst steps calls goal turns hold return goal 1 5.484 5.059 66.1 66.1 100% 1 12 0.031 0% 2 5.484 5.059 66.1 33.2 100% 2 12 0.317 0% 3 5.484 5.060 66.1 22.4 100% 4 12 1.039 0% 5 5.479 5.056 66.1 13.5 100% 6 11 1.700 0% 8 4.906 0.416 97.6 12.6 94% 8 11 3.420 44% 12 4.401 -0.467 127.6 11.2 88% 10 8 4.595 75% 16 3.643 -0.577 153.4 10.1 75% 12 7 4.974 88% 24 1.392 -0.689 280.1 12.1 44% 14 6 5.409 100% 32 -2.807 -2.923 300.0 10.0 0% 16 5 5.479 100% 64 -3.137 -3.658 238.3 4.0 0% 20 4 5.483 100% 300 -3.764 -4.906 149.7 1.0 0% 32 3 5.484 100% inaction floor: -4.161 crow-nav-v1, 16 episodes per row ``` (the two tables `measure_hold.py` prints, set side by side; rows at hold 48, 100 and 150 are dropped for width.) The left table is house rule 4's premise: **open-loop fails outright.** One decision held for the whole flight returns -3.764 against an inaction floor of -4.161 and reaches the waypoint none of the time — a bird that sets a course and leaves does no better than a bird that does nothing. A plan is not a policy in a world with momentum, which is why the grammar is one vector per reply and not a program. The right table is the conclusion, and it is the reason `max_turns = 16`: the budget BINDS all the way up to it. At eight turns the best possible constant hold reaches the waypoint 44% of the time, at twelve 88%, and only at fourteen does it reach every one. Sixteen is the first budget at which the ceiling is reachable, and it is one turn of slack, not fifteen. `max_hold = 12` is the longest hold at which the baseline still reaches most waypoints. `16 x 12 = 192` against a 300-step simulator cap, so what binds is the turn budget and never the simulator's own. ## The rewards Two, both products, both from numbers TypeScript computed: | | | |---|---| | `flight` (0.70) | the episode's return against nine tenths of what Tera's scripted controller returned on the same scenario at the same seed | | `economy` (0.30) | that quality, multiplied by having arrived and by how directly | `flight` is smooth because goal attainment is already the biggest term inside it — the simulator pays +3 of a ~5.5 return for reaching the waypoint — so a near-miss scores about four tenths and a wander scores zero. ⚠️ **Goal attainment multiplies and is never a term of its own.** It is already inside the return as the success bonus, and a separate additive `gate` beside the sum double-counts it. That is the exact failure `probe.py` records having paid for. Efficiency multiplies too: a flight that never arrived cannot be quick about arriving, and standing beside `flight` it would pay a crash for being fast. `gate`, `quality`, `step_ratio`, `ts_return`, `hold_mean`, `malformed_turns` and `replay_ok` are recorded and never rewarded. **`replay_ok` is the bridge's claim asserted once per episode**, not once per commit: every graded flight is replayed in a fresh TypeScript environment and the metric is whether it reached the same FNV-1a-64 checksum, the same step count and the same return. A divergence is recorded, never raised — an episode that flew is data, and losing it to an assertion would hide the thing the metric exists to surface. ## Running it ```bash uv run --project environments/tera_spatial eval @ configs/tera_crow_nav.toml \ --model brain-qwen38-dspark \ --client.base-url http://100.127.247.67:8001/v1 \ --client.api-key-var SPARK_API_KEY \ --no-push --no-rich -c 8 -o outputs/run-/tera-crow-nav ``` ⚠️ `eval` exits 0 even when every rollout errors. Count non-empty `rewards`; see `docs/FIRST_EVAL.md`. ## What it measured `outputs/run-20260821-1627/tera-crow-nav`, 16 tasks x 2 rollouts, model `brain-qwen38-dspark` on spark-1, 12 minutes at `-c 8`. **32 episodes attempted, 32 scored, 0 errored, 0 provider errors.** | | | |---|---| | `reward_mean_attempted` | **0.4603** | | `flight` (0.70) | 0.4894 | | `economy` (0.30) | 0.3926 | | waypoint reached | 14 of 32 | | flight-envelope contact | 2 of 32 | | out of turns, still flying | 16 of 32 | | malformed replies | **0** of 412 | | `replay_ok` | **1.000 on all 32** | Three things in that table are worth more than the headline. **The turn budget binds, in the run and not only in the sweep.** Turns used were 7, 8, 9, 10, 12, 14 and 16 — mean 12.9, and **sixteen of the thirty-two rollouts spent every reply they had** and were still airborne when the budget ran out. The one that scored zero on `train-east-crosswind` did so by orbiting: 128 steps flown, yaw -90°, and 52.8 m still between it and the waypoint on its last reply. **The gate carries a gradient.** Goal attainment fired on 43.75% of rollouts. Across the first seven Arena environments a `gate` component scored exactly 0.000, max 0.000, in four of them — a quarter to a third of the reward mass with nothing in it. Here it discriminates, and per-scenario it discriminates sharply: `train-north-climb` 0.863, `dev-west-return` 0.728, `dev-south-descent` 0.250, `train-east-crosswind` 0.000. **`eval.log` and `traces.jsonl` agree.** The mean over the log's own 32 `reward=` lines is 0.4603, the same number as `sum(score x weight)` over the file. That is not free: rewards recorded after `interaction.close()` land in the file and not in the log, and the first cut of this env printed `reward=0.000` for an episode that scored 1.000. Grade before closing. ## The gate ```bash uv run --project environments/tera_spatial \ python -m unittest discover -s environments/tera_spatial/tests -v ``` 32 episodes — 4 environments x 4 public scenarios x 2 seeds — each driven from Python one JSON step at a time, then replayed in a fresh TypeScript environment. All 32 must reach an identical FNV-1a-64 checksum. Then seven forgeries per environment, each tried twice: raw, where the envelope checksum catches it, and **re-sealed with a checksum Tera itself recomputed**, where only `replay()` re-running the simulator can. All must be rejected. `test_crow_nav.py` adds fifteen more. The reward's floor and ceiling are flown rather than asserted from a fixture, and the strongest of them is the rendering's own claim: a policy that reads **nothing but the rendered panel** — the same characters the model is shown, parsed back out of them with a regex — reaches all four waypoints inside the shipped budget. An observation encoding that stops carrying enough to fly on stops passing that test. `node` >= 22.18 is required — the vendored sources are raw `.ts` and are type-stripped, not compiled. Set `TERA_NODE` to point at a specific binary. ## Re-vendoring ```bash uv run python scripts/sync_tera.py --tera ~/repos/gitea/tera # copy + regenerate hashes uv run python scripts/sync_tera.py --check # CI: is the tree clean? ``` The closure is walked, not declared, so a new import in `tera` travels with it. A bare specifier is a hard failure: the wheel has no `node_modules` and a simulator that needs one is not a simulator that ships.