Files
arena/environments/tera_spatial

tera-spatial

Arena's bridge to Tera's four renderer-independent spatial environments.

The environments themselves live in the tera repository, in TypeScript, as tera.arena/v1. They are not reimplemented here and they never will be. This package vendors their import closure, runs it in a resident node worker, and transports results. The rule the whole thing is built around:

Never recompute in Python a number that came out of TypeScript.

Rewards, per-step state checksums, scenario materialisation and the four baseline returns that every reward is normalised against are all computed once, inside Tera, and only ever carried across the pipe.

What is here

Path What it is
tera_spatial/worker.mjs the resident NDJSON worker: reset, step, snapshot, restore, trace, replay, oracle, checksum
tera_spatial/bridge.py TeraWorker / TeraEpisode — transport, and nothing else
tera_spatial/vendor/tera/ the 23-file, 283 KB import closure of src/arena/index.ts
tera_spatial/hashes.py SHA-256 of all 23, generated; the worker refuses to start if the tree has drifted
tera_spatial/closure.py the import walk both the sync script and the gate use
scripts/sync_tera.py reproduces the vendoring from a tera checkout
tests/test_replay.py the replay gate
tera_spatial/spatial.py the shared half of a spatial taskset: the resident worker, the reply grammar, the episode loop, the reward
tera_crow_nav/ the first taskset — tera-crow-nav, fly a crow to a waypoint on sixteen replies
measure_hold.py the hold and turn-budget sweep the two integers came from
tests/test_crow_nav.py the taskset's gate: the grammar, the reward, and a flight flown off the rendered panel

tera-office-nav and tera-california-flight follow. tera-drive-101 ships labelled a bridge-correctness environment or not at all: its return is monotone in throttle and any positive constant scores about 0.95 of the baseline, which is a number with no house rule 4 behind it.

The crux: a language model cannot emit sixty control vectors a second

These are continuous-control environments and the agent is text. So a reply is a DECISION, not a frame: one control vector plus the number of simulation steps to hold it for, and then the model is shown where the vehicle ended up.

The two integers that decide everything — how many replies, and how long one may be held — are measured. measure_hold.py flies Tera's own scripted baseline at every refresh rate over sixteen scenario x seed episodes:

 hold    return     worst   steps   calls   goal          turns  hold    return   goal
    1     5.484     5.059    66.1    66.1   100%              1    12     0.031     0%
    2     5.484     5.059    66.1    33.2   100%              2    12     0.317     0%
    3     5.484     5.060    66.1    22.4   100%              4    12     1.039     0%
    5     5.479     5.056    66.1    13.5   100%              6    11     1.700     0%
    8     4.906     0.416    97.6    12.6    94%              8    11     3.420    44%
   12     4.401    -0.467   127.6    11.2    88%             10     8     4.595    75%
   16     3.643    -0.577   153.4    10.1    75%             12     7     4.974    88%
   24     1.392    -0.689   280.1    12.1    44%             14     6     5.409   100%
   32    -2.807    -2.923   300.0    10.0     0%             16     5     5.479   100%
   64    -3.137    -3.658   238.3     4.0     0%             20     4     5.483   100%
  300    -3.764    -4.906   149.7     1.0     0%             32     3     5.484   100%

inaction floor: -4.161      crow-nav-v1, 16 episodes per row

(the two tables measure_hold.py prints, set side by side; rows at hold 48, 100 and 150 are dropped for width.)

The left table is house rule 4's premise: open-loop fails outright. One decision held for the whole flight returns -3.764 against an inaction floor of -4.161 and reaches the waypoint none of the time — a bird that sets a course and leaves does no better than a bird that does nothing. A plan is not a policy in a world with momentum, which is why the grammar is one vector per reply and not a program.

The right table is the conclusion, and it is the reason max_turns = 16: the budget BINDS all the way up to it. At eight turns the best possible constant hold reaches the waypoint 44% of the time, at twelve 88%, and only at fourteen does it reach every one. Sixteen is the first budget at which the ceiling is reachable, and it is one turn of slack, not fifteen.

max_hold = 12 is the longest hold at which the baseline still reaches most waypoints. 16 x 12 = 192 against a 300-step simulator cap, so what binds is the turn budget and never the simulator's own.

The rewards

Two, both products, both from numbers TypeScript computed:

flight (0.70) the episode's return against nine tenths of what Tera's scripted controller returned on the same scenario at the same seed
economy (0.30) that quality, multiplied by having arrived and by how directly

flight is smooth because goal attainment is already the biggest term inside it — the simulator pays +3 of a ~5.5 return for reaching the waypoint — so a near-miss scores about four tenths and a wander scores zero.

⚠️ Goal attainment multiplies and is never a term of its own. It is already inside the return as the success bonus, and a separate additive gate beside the sum double-counts it. That is the exact failure probe.py records having paid for. Efficiency multiplies too: a flight that never arrived cannot be quick about arriving, and standing beside flight it would pay a crash for being fast.

gate, quality, step_ratio, ts_return, hold_mean, malformed_turns and replay_ok are recorded and never rewarded.

replay_ok is the bridge's claim asserted once per episode, not once per commit: every graded flight is replayed in a fresh TypeScript environment and the metric is whether it reached the same FNV-1a-64 checksum, the same step count and the same return. A divergence is recorded, never raised — an episode that flew is data, and losing it to an assertion would hide the thing the metric exists to surface.

Running it

uv run --project environments/tera_spatial eval @ configs/tera_crow_nav.toml \
  --model brain-qwen38-dspark \
  --client.base-url http://100.127.247.67:8001/v1 \
  --client.api-key-var SPARK_API_KEY \
  --no-push --no-rich -c 8 -o outputs/run-<stamp>/tera-crow-nav

⚠️ eval exits 0 even when every rollout errors. Count non-empty rewards; see docs/FIRST_EVAL.md.

What it measured

outputs/run-20260821-1627/tera-crow-nav, 16 tasks x 2 rollouts, model brain-qwen38-dspark on spark-1, 12 minutes at -c 8. 32 episodes attempted, 32 scored, 0 errored, 0 provider errors.

reward_mean_attempted 0.4603
flight (0.70) 0.4894
economy (0.30) 0.3926
waypoint reached 14 of 32
flight-envelope contact 2 of 32
out of turns, still flying 16 of 32
malformed replies 0 of 412
replay_ok 1.000 on all 32

Three things in that table are worth more than the headline.

The turn budget binds, in the run and not only in the sweep. Turns used were 7, 8, 9, 10, 12, 14 and 16 — mean 12.9, and sixteen of the thirty-two rollouts spent every reply they had and were still airborne when the budget ran out. The one that scored zero on train-east-crosswind did so by orbiting: 128 steps flown, yaw -90°, and 52.8 m still between it and the waypoint on its last reply.

The gate carries a gradient. Goal attainment fired on 43.75% of rollouts. Across the first seven Arena environments a gate component scored exactly 0.000, max 0.000, in four of them — a quarter to a third of the reward mass with nothing in it. Here it discriminates, and per-scenario it discriminates sharply: train-north-climb 0.863, dev-west-return 0.728, dev-south-descent 0.250, train-east-crosswind 0.000.

eval.log and traces.jsonl agree. The mean over the log's own 32 reward= lines is 0.4603, the same number as sum(score x weight) over the file. That is not free: rewards recorded after interaction.close() land in the file and not in the log, and the first cut of this env printed reward=0.000 for an episode that scored 1.000. Grade before closing.

The gate

uv run --project environments/tera_spatial \
  python -m unittest discover -s environments/tera_spatial/tests -v

32 episodes — 4 environments x 4 public scenarios x 2 seeds — each driven from Python one JSON step at a time, then replayed in a fresh TypeScript environment. All 32 must reach an identical FNV-1a-64 checksum. Then seven forgeries per environment, each tried twice: raw, where the envelope checksum catches it, and re-sealed with a checksum Tera itself recomputed, where only replay() re-running the simulator can. All must be rejected.

test_crow_nav.py adds fifteen more. The reward's floor and ceiling are flown rather than asserted from a fixture, and the strongest of them is the rendering's own claim: a policy that reads nothing but the rendered panel — the same characters the model is shown, parsed back out of them with a regex — reaches all four waypoints inside the shipped budget. An observation encoding that stops carrying enough to fly on stops passing that test.

node >= 22.18 is required — the vendored sources are raw .ts and are type-stripped, not compiled. Set TERA_NODE to point at a specific binary.

Re-vendoring

uv run python scripts/sync_tera.py --tera ~/repos/gitea/tera   # copy + regenerate hashes
uv run python scripts/sync_tera.py --check                     # CI: is the tree clean?

The closure is walked, not declared, so a new import in tera travels with it. A bare specifier is a hard failure: the wheel has no node_modules and a simulator that needs one is not a simulator that ships.