tera-crow-nav: the first spatial taskset, and the hold sweep that sized its budget

The bridge becomes a runnable taskset. The crux is that these are continuous
control environments and the agent emits text: the model issues a control
decision that is held for N steps, and the hold length is measured rather than
guessed. At one decision every 32 steps the scripted baseline returns -2.807 and
reaches the waypoint 0% of the time, against 5.484 and 100% at every frame.

Open loop — one decision for the whole flight — returns -3.764 against an
inaction floor of -4.161. A bird that sets a course and leaves does no better
than one that does nothing, which is why crow-nav carries the strongest rule-4
claim of the five spatial environments. The budget bites from both ends: 8 turns
reach the goal 44% of the time, 12 reach 88%, 14 reach 100%.

32 episodes against Qwen3.8-27B: mean 0.4603, no malformed replies in 412, every
episode replaying to an identical checksum, and 16 of 32 spending all 16 turns.
The gate fires 43.75% of the time and separates the four scenarios cleanly —
whatever is wrong with the four dead data-environment gates, it is not that
binary gates cannot discriminate.

The oracle comes from the worker's own op. Nothing recomputes in Python a number
that came out of TypeScript, least of all the reward denominator.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-21 17:44:46 -07:00
co-authored by Claude Opus 5
parent 94ab513eda
commit 4147de0282
10 changed files with 1482 additions and 18 deletions
+11 -11
View File
@@ -1,25 +1,25 @@
[project]
name = "tera-spatial"
version = "0.1.0"
description = "tera-spatial — Arena's bridge to Tera's four renderer-independent spatial environments, executed in TypeScript and graded in Python."
description = "tera-spatial — Arena's bridge to Tera's four renderer-independent spatial environments, executed in TypeScript and graded in Python, and the tasksets that sit on it."
requires-python = ">=3.11"
dependencies = ["verifiers"]
# One directory, four tasksets: the environments share a vendored simulator and
# One directory, several tasksets: the environments share a vendored simulator and
# a single worker, so splitting them into four wheels would ship the same 283 KB
# of TypeScript four times. B1's discovery reads this key instead of assuming
# one wheel is one taskset.
# of TypeScript four times. Discovery reads this key instead of assuming one wheel
# is one taskset.
#
# Only the ones that exist are listed. `tera-office-nav` and `tera-california-flight`
# follow; `tera-drive-101` ships labelled a bridge-correctness environment or not at
# all, because its return is monotone in throttle and any positive constant scores
# ~0.95 of the baseline — a number with no house rule 4 behind it.
[tool.arena]
tasksets = [
"tera-drive-101",
"tera-office-nav",
"tera-crow-nav",
"tera-california-flight",
]
tasksets = ["tera-crow-nav"]
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
[tool.hatch.build.targets.wheel]
packages = ["tera_spatial"]
packages = ["tera_spatial", "tera_crow_nav"]