The bridge becomes a runnable taskset. The crux is that these are continuous control environments and the agent emits text: the model issues a control decision that is held for N steps, and the hold length is measured rather than guessed. At one decision every 32 steps the scripted baseline returns -2.807 and reaches the waypoint 0% of the time, against 5.484 and 100% at every frame. Open loop — one decision for the whole flight — returns -3.764 against an inaction floor of -4.161. A bird that sets a course and leaves does no better than one that does nothing, which is why crow-nav carries the strongest rule-4 claim of the five spatial environments. The budget bites from both ends: 8 turns reach the goal 44% of the time, 12 reach 88%, 14 reach 100%. 32 episodes against Qwen3.8-27B: mean 0.4603, no malformed replies in 412, every episode replaying to an identical checksum, and 16 of 32 spending all 16 turns. The gate fires 43.75% of the time and separates the four scenarios cleanly — whatever is wrong with the four dead data-environment gates, it is not that binary gates cannot discriminate. The oracle comes from the worker's own op. Nothing recomputes in Python a number that came out of TypeScript, least of all the reward denominator. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
26 lines
1.1 KiB
TOML
26 lines
1.1 KiB
TOML
[project]
|
|
name = "tera-spatial"
|
|
version = "0.1.0"
|
|
description = "tera-spatial — Arena's bridge to Tera's four renderer-independent spatial environments, executed in TypeScript and graded in Python, and the tasksets that sit on it."
|
|
requires-python = ">=3.11"
|
|
dependencies = ["verifiers"]
|
|
|
|
# One directory, several tasksets: the environments share a vendored simulator and
|
|
# a single worker, so splitting them into four wheels would ship the same 283 KB
|
|
# of TypeScript four times. Discovery reads this key instead of assuming one wheel
|
|
# is one taskset.
|
|
#
|
|
# Only the ones that exist are listed. `tera-office-nav` and `tera-california-flight`
|
|
# follow; `tera-drive-101` ships labelled a bridge-correctness environment or not at
|
|
# all, because its return is monotone in throttle and any positive constant scores
|
|
# ~0.95 of the baseline — a number with no house rule 4 behind it.
|
|
[tool.arena]
|
|
tasksets = ["tera-crow-nav"]
|
|
|
|
[build-system]
|
|
requires = ["hatchling"]
|
|
build-backend = "hatchling.build"
|
|
|
|
[tool.hatch.build.targets.wheel]
|
|
packages = ["tera_spatial", "tera_crow_nav"]
|