tera-crow-nav: the first spatial taskset, and the hold sweep that sized its budget
The bridge becomes a runnable taskset. The crux is that these are continuous control environments and the agent emits text: the model issues a control decision that is held for N steps, and the hold length is measured rather than guessed. At one decision every 32 steps the scripted baseline returns -2.807 and reaches the waypoint 0% of the time, against 5.484 and 100% at every frame. Open loop — one decision for the whole flight — returns -3.764 against an inaction floor of -4.161. A bird that sets a course and leaves does no better than one that does nothing, which is why crow-nav carries the strongest rule-4 claim of the five spatial environments. The budget bites from both ends: 8 turns reach the goal 44% of the time, 12 reach 88%, 14 reach 100%. 32 episodes against Qwen3.8-27B: mean 0.4603, no malformed replies in 412, every episode replaying to an identical checksum, and 16 of 32 spending all 16 turns. The gate fires 43.75% of the time and separates the four scenarios cleanly — whatever is wrong with the four dead data-environment gates, it is not that binary gates cannot discriminate. The oracle comes from the worker's own op. Nothing recomputes in Python a number that came out of TypeScript, least of all the reward denominator. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -1,25 +1,25 @@
|
||||
[project]
|
||||
name = "tera-spatial"
|
||||
version = "0.1.0"
|
||||
description = "tera-spatial — Arena's bridge to Tera's four renderer-independent spatial environments, executed in TypeScript and graded in Python."
|
||||
description = "tera-spatial — Arena's bridge to Tera's four renderer-independent spatial environments, executed in TypeScript and graded in Python, and the tasksets that sit on it."
|
||||
requires-python = ">=3.11"
|
||||
dependencies = ["verifiers"]
|
||||
|
||||
# One directory, four tasksets: the environments share a vendored simulator and
|
||||
# One directory, several tasksets: the environments share a vendored simulator and
|
||||
# a single worker, so splitting them into four wheels would ship the same 283 KB
|
||||
# of TypeScript four times. B1's discovery reads this key instead of assuming
|
||||
# one wheel is one taskset.
|
||||
# of TypeScript four times. Discovery reads this key instead of assuming one wheel
|
||||
# is one taskset.
|
||||
#
|
||||
# Only the ones that exist are listed. `tera-office-nav` and `tera-california-flight`
|
||||
# follow; `tera-drive-101` ships labelled a bridge-correctness environment or not at
|
||||
# all, because its return is monotone in throttle and any positive constant scores
|
||||
# ~0.95 of the baseline — a number with no house rule 4 behind it.
|
||||
[tool.arena]
|
||||
tasksets = [
|
||||
"tera-drive-101",
|
||||
"tera-office-nav",
|
||||
"tera-crow-nav",
|
||||
"tera-california-flight",
|
||||
]
|
||||
tasksets = ["tera-crow-nav"]
|
||||
|
||||
[build-system]
|
||||
requires = ["hatchling"]
|
||||
build-backend = "hatchling.build"
|
||||
|
||||
[tool.hatch.build.targets.wheel]
|
||||
packages = ["tera_spatial"]
|
||||
packages = ["tera_spatial", "tera_crow_nav"]
|
||||
|
||||
Reference in New Issue
Block a user