Commit Graph
21 Commits
Author SHA1 Message Date
Metal AgentandClaude Opus 5 beaffbd0a7 Put the layout store behind a boundary, and assert the words it puts on screen
Persistence here has two rules and neither was tested. Losing the layout must
never lose the session, so every failure falls back to the first-run workspace
and carries on; and the fallback must be visible, so each path returns a
status string the footer shows. Nine different strings, one of which is the
only notice a user gets that the arrangement they spent a morning on has been
dropped, and not one of them was asserted anywhere. The way to find out that
a corrupt snapshot reports "invalid layout ignored" was to corrupt one.

Ten tests now cover the branches a real machine reaches: nowhere to write, a
data directory that cannot be created, a database SQLite refuses to open, a
fresh database that is ready rather than restored, a round trip that restores
a panel created before the save, and both ways a snapshot can be unusable --
malformed JSON and well-formed JSON describing a workspace with no attached
panel, since parsing is not validation and only the second is easy to write by
accident.

workspace_database_path now reads through EnvSource rather than std::env, for
exactly the reason that trait was introduced in lumbridge-settings: the
workspace forbids unsafe, set_var is unsafe in Rust 2024, and a precedence
rule that cannot be exercised without mutating the process running the test is
a precedence rule that stays untested. The one behavioural consequence is that
a non-UTF-8 value in LUMBRIDGE_SPIKE_DB, XDG_DATA_HOME or HOME is now treated
as unset rather than used as a path; that is what every other setting in the
shell already does with such a value.

persist_panels stays a method, reduced to the one thing the shell owns: a
workspace with no store is not a save failure. It is memory-only, the footer
has said so since startup, and replacing that standing message with an error
every time a pane moved would say less, not more.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SPYebLiN2w4TqnHUYGdECq
2026-09-01 13:00:46 -07:00
Metal AgentandClaude Opus 5 ae22f5522b Stop CI from skipping the gates it exists to run
The workflow installed neither cargo-deny nor cargo-nextest, and ci.sh treated
both as optional. A run therefore reported green having never checked the
licence closure that DISTRIBUTION.md depends on, and having run cargo test
where the repository believes it runs nextest. Both tools are installed here
from prebuilt binaries -- compiling cargo-deny on this runner would cost more
than the job it guards -- and both jobs set LUMBRIDGE_CI_STRICT=1, which makes
a missing tool a failure rather than a warning.

The toolchain version was also written here twice while rust-toolchain.toml
declared it a third time. It is now read from that file, because a skew
between the compiler CI uses and the one a developer uses is precisely the
kind of difference that is invisible until it is expensive.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SPYebLiN2w4TqnHUYGdECq
2026-09-01 12:59:57 -07:00
Metal AgentandClaude Opus 5 44c2391199 Give the frame-time measurement its own file, and its first tests
RenderTiming is the shell measuring itself, and until now nothing measured
it. The percentile index, the 256-sample bound and the eviction that keeps it
bounded had no test at all, which is a strange place for a codebase to have a
blind spot: this is the number the footer shows a user when they ask whether
the app is slow.

The eight tests state the decisions the code already made, so that changing
one is a choice rather than an accident. Percentiles are computed over sorted
samples, not arrival order. Ninety-nine good frames and one 40ms stall keep a
p50 of 100µs and the stall shows in the p95, which is the entire reason this
is a distribution and not the mean it would be so much easier to compute. A
full window of new frames retires every stale sample, and eviction is from the
front, so the figure describes the last four seconds rather than the session.
An unpaired build records nothing, because a duration with no start is not a
fast frame -- it is no data, and averaging a zero into the p50 would report
the shell as faster than it is. Two dispatches before one build is one sample
and it is the later one, since the earlier action's tree was never built.

The percentile tests build the buffer directly rather than going through the
clock. Driving them through mark_dispatch would make every assertion depend on
how loaded the machine running CI happens to be, which is how a timing test
becomes the flaky one everybody reruns.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SPYebLiN2w4TqnHUYGdECq
2026-09-01 12:58:16 -07:00
Metal AgentandClaude Opus 5 6daee84ca6 Move the keystroke table out of the renderer that reported the keystroke
Deciding what a key means to a shell needs the key name, the composed
character if the platform produced one, and the modifier flags. It does not
need a window, and keeping it beside one made the least forgiving table in the
shell the hardest to read and the least obvious to test.

It is the least forgiving because it has already failed silently. Requiring a
printable character before anything was sent swallowed every control byte --
ctrl-c, ctrl-d, ctrl-a, ctrl-r and ctrl-k -- with no crash and no log, so the
symptom was a shell that ignored you. Those tests move with the functions and
keep their names, which describe the mistake rather than the function.

key_modifiers stays in main.rs. Copying KeyDownEvent.keystroke.modifiers field
by field is a statement about a gpui type and belongs where gpui types live;
what it produces is a KeyModifiers, and that crosses the boundary fine.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SPYebLiN2w4TqnHUYGdECq
2026-09-01 12:57:00 -07:00
Metal AgentandClaude Opus 5 5564063aa5 Take the window arithmetic out of the window
main.rs is 3,884 lines and 56% of the application, so everything inside it is
as expensive to read as the renderer around it. The first thing to leave is
the part that never needed a renderer at all: how tall the workspace is once
the header, tab bar and footer are subtracted, how many panes fit beside the
rail, which slice of the attached panes is on screen, and how many rows and
columns of a measured cell that leaves.

This arithmetic is the contract with the PTY -- a program lays itself out from
the columns it is told it has, so an error of one column here is a wrapped
line in vim and a broken table in git log. That makes it exactly the code that
should be tested against numbers rather than against a window, and decision
0009's pane capacities were already wrong once because they were quoted from a
guessed cell width.

CellMetrics::measure stays behind. Asking the text system for a glyph advance
needs a live App, so the measurement remains a renderer's job and only the
answer crosses over, as two plain f32s. Size<Pixels> stops crossing at all:
geometry works in a local WindowSize and main.rs converts on the way in
through one From impl, which is the entire boundary. The module imports no
gpui, matching sidebar::model, so its two tests run in the headless job.

One thing removed rather than moved: lossless_f32 carried an
allow(clippy::unreadable_literal) with the reason "six-digit colour hex reads
whole". It had drifted up from the terminal colour tables below it and applied
to a function containing no colour and no literal it could suppress.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SPYebLiN2w4TqnHUYGdECq
2026-09-01 12:55:59 -07:00
Metal AgentandClaude Opus 5 9a29e8e335 Make the gate structural, and correct what it tells an agent
CI / rust-ui (push) Failing after 6m20s
CI / rust-headless (push) Successful in 6m38s
Three gates in this repository were decorative, and each was discovered by
being wrong rather than by failing.

A crate directory in neither members nor exclude is silently not built, which
is how lumbridge-devices shipped 1,127 lines that had never compiled.
scripts/workspace-guard.sh refuses that state, and asserts the gpui source
and version out of Cargo.lock rather than the manifest, because a manifest
states an intent while the lockfile states what would actually be compiled --
and a caret requirement accepts a version nobody reviewed. It needs no
compiler, so it runs first and in the headless job, which unlike the UI job
is not continue-on-error and can therefore actually fail a push.

deny.toml's source policy had never been executed: ci.sh ran `check
licenses` alone, and `check sources` failed immediately on the rev-pinned
buzz-sdk. The permitted Git sources are now named one by one and the check
runs, so a fourth is a decision rather than an accident.

cargo-deny and cargo-nextest being absent was a warning that let a run report
green having skipped the licence gate DISTRIBUTION.md depends on. Under
LUMBRIDGE_CI_STRICT=1 a missing tool now fails; locally it stays a warning so
a contributor is not blocked.

skills/lumbridge-development/SKILL.md told every agent that GPUI and Floem
live in spikes/ and that no framework may be selected until both pass the
hard gates. Decision 0017 settled that a month ago in the opposite direction.
The entry point an agent is meant to read was the least accurate document in
the repository.

Decision 0023 records where the GPUI dependency actually goes. Published gpui
has not been released since 2025-10-22, Zed's main still declares 0.2.2 with
no bump pending, the platform backends moved to crates that inherit
publish = false, gpui's own x11 and wayland features are now empty markers,
and 0.2.2 has no accesskit dependency at all -- so "published now, migrate
later" was never available. The adapter 0017 promised was never written and
the call sites grew from few to 147 against 20 identities, so the adapter is
written first, on 0.2.2, before the dependency moves.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SPYebLiN2w4TqnHUYGdECq
2026-09-01 12:51:25 -07:00
Metal AgentandClaude Opus 5 28e455f968 Layer settings over a file that documents itself
Compiled default, then the file, then the environment -- with the
environment winning, which is the reverse of the usual arrangement and is the
point: LUMBRIDGE_CLAUDE_OAUTH=0 is documented as one switch off, and a switch
a configuration file can silently turn back on is not a switch. Where a
variable has pinned a value the settings pane shows that row disabled and
names the variable rather than accepting an edit that would do nothing.

default.toml carries every key Lumbridge understands, commented out, showing
the compiled default, and a test uncomments them and asserts the key set is
exactly the set the binary knows -- so neither half can drift from the other.
toml reads and toml_edit writes, because serialising the parsed model back
over the file would strip every comment and every key this build does not
recognise, and here the comments are the documentation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SPYebLiN2w4TqnHUYGdECq
2026-09-01 12:51:25 -07:00
Metal AgentandClaude Opus 5 401760d670 Generate the theme catalog from real themes instead of a hand table
The derivation landed in decision 0018 with the catalog and the terminal ANSI
palette still fixed tables written by hand. tools/theme-gen reads the
TextMate themes under assets/themes/ and emits the catalog and the reference
vectors, so the anchors a palette is derived from are the ones the theme
actually ships rather than the ones somebody transcribed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SPYebLiN2w4TqnHUYGdECq
2026-09-01 12:51:25 -07:00
Metal AgentandClaude Opus 5 7bee985279 Finish the devices crate, and stop it being invisible
crates/lumbridge-devices was in neither workspace.members nor
workspace.exclude, which is not a build error: Cargo simply never looked at
it. Its 19 tests had never run, it never inherited unsafe_code = "forbid" or
pedantic Clippy, and `cargo check` inside it refused outright with "current
package believes it's in a workspace when it's not". Under that cover its
lib.rs had been declaring `mod manage;` and re-exporting five items from a
manage.rs that did not exist, so the crate did not compile at all.

manage.rs is written here to the contract lib.rs already specified.
available_actions reads neither DeviceReachability nor Device::presence: an
offline device keeps its workspace action and an online one does not gain
one, because the registry is the axis and reachability is Tailscale's
separate claim. DeviceAction has three variants and no more -- install,
reboot and upgrade are absent from the type rather than rejected at runtime,
since a variant that exists is eventually rendered as a greyed-out button
reading "coming soon" instead of "impossible". A test walks every operation
in the module over every fixture device and asserts none of them ever
produces LumbridgePresence::Confirmed, which stays unproducible until a
lumbridge-remote runtime can answer for itself.

RemoteTransport had been declared twice, here and in lumbridge-core, with
byte-identical storage strings, because this crate had no dependency on that
one. Two enumerations of one choice persisted through the same strings is a
drift waiting to happen, so core keeps the single definition -- gaining the
default and the picker phrase -- and this crate depends on core and
re-exports it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SPYebLiN2w4TqnHUYGdECq
2026-09-01 12:51:25 -07:00
Metal AgentandClaude Opus 5 73df4fa679 Correct the column figure decision 0009 quoted from a guessed cell width
"Approximately 71 columns by 42 rows per panel" was derived from a hardcoded
8.4 px advance that nothing had measured. The measured advance is about 7.3 px,
so the figure was roughly 13% short — a full-width pane on the reference display
measures 140 columns, not 122. The paragraph now says the number is a
consequence of the measurement rather than a property of the design.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-01 00:14:23 -07:00
Metal AgentandClaude Opus 5 f9f4f85402 Measure the terminal cell instead of guessing it, and stop telling the PTY zero
TERMINAL_CELL_WIDTH was 8.4 — a number nobody had measured. Asking the text
system for the advance of `0` in the face actually being painted gives ~7.3, so
the guess was 13% wide and the terminal was losing eighteen columns: the same
window that reported 122 columns now reports 140. Layout and paint now read the
same measurement, so they cannot drift apart again.

The plan claimed ws_xpixel disagreed with the painted width by 0.4 px per
column. It did not: the app only ever called TerminalSize::new, which passes no
pixel dimensions, so ws_xpixel and ws_ypixel were both *zero*. Every program
doing pixel arithmetic — sixel, the kitty graphics protocol, anything sizing an
image to the viewport — was being told the window has no size at all. Both the
spawn and the resize paths now report the real extent.

TerminalDimensions gains cell_width/cell_height accessors: it was already
carrying the values and nothing could read them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-01 00:13:30 -07:00
Metal AgentandClaude Opus 5 e5d7a3efd5 Add layered settings, and fix a migration mechanism that silently lied
Two things, because the second could not be built on the first.

The schema stamp was part of the same execute_batch as the CREATE TABLE IF NOT
EXISTS statements, and it wrote unconditionally. Opening an older file therefore
added no columns but flipped the version forward anyway; opening a *newer* file
stamped it back down and then wrote rows the newer build could not read. Both
produced a database whose recorded version was a lie, and every future schema
change would have inherited it.

Now the version is read before anything is applied, migrations are ordered and
forward-only inside one transaction, a newer file is refused with SchemaTooNew
rather than downgraded, and a supported version raised without a step to reach
it fails at the first open instead of claiming success. Tested by stamping a
file at version 99 and asserting both the refusal and that the stamp is left
untouched.

lumbridge-settings resolves compiled default -> settings.toml -> environment.
The environment sits above the file deliberately: decision 0016 calls
LUMBRIDGE_CLAUDE_OAUTH=0 "one switch off", and a switch a config file can
silently re-enable is not a switch. A pinned value renders disabled and names
the variable, rather than accepting an edit that would do nothing.

Every field carries a WriteAuthority. Routing all writes through Configure is
the obvious design and would hand a layout-only agent the program every future
pane launches — the guarantee decision 0006 exists to make. Anything naming a
program, path or destination is Human-only, asserted by a test that reads the
path rather than trusting the author.

Four paths are permanently not settings, with the reason recorded beside each
and a test asserting their absence: the usage endpoint URL, the credentials
path, the client identity, and the shell program. A configuration file that can
redirect where an access token is sent is a credential exfiltration path with a
friendly name.

Environment access is a trait rather than std::env, because the workspace forbids
unsafe, set_var is unsafe in Rust 2024, and the layering rule has to be testable
without mutating the process running the test.

Verified live with LUMBRIDGE_CLAUDE_OAUTH=0: the account-endpoint row reads off,
greyed, "pinned by LUMBRIDGE_CLAUDE_OAUTH". The Advanced page names every file,
endpoint and child process Lumbridge touches and states that nothing is sent
anywhere else — as a fact, not as a toggle nobody can flip.

File loading, comment-preserving writes and editable controls are not in this
pass; 0022 records why that order is the honest one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-01 00:06:06 -07:00
Metal AgentandClaude Opus 5 72887cb4ab Rebuild the sidebar around the four questions it exists to answer
The rail showed a frozen attention count over a worktree list backed by a crate
that does not exist. What replaces it starts from a question rather than from a
list of things we happened to know: what needs me, what am I running, what did I
set aside, where does it run and what will stop me.

Layout is data. sidebar/model.rs holds no renderer types, so which sections
exist, what collapsing hides, what the filter keeps, and where the keyboard
cursor lands are ordinary tests in CI; sidebar/view.rs renders and decides
nothing. Eleven model tests, none of which need a window.

The cursor is a RowKey rather than an index, because an index is wrong the moment
a row above it disappears and silently pointing at a different row is worse than
losing the cursor. Every header renders even when its section is empty, so
positions never move under the pointer. The filter's empty state does not quote
what was typed — the sidebar is the part of the window people screenshot.

One selection language everywhere: before this, attention cards darkened on hover
while worktree rows lightened, so the same gesture meant two different things a
hundred pixels apart.

Two defects the screenshots caught that review had not. Flexbox shrinks
proportionally, so the longer string wins: the attention row rendered as
"Te… Exited with code 7 · observed", having discarded the one word that says
which pane to look at. And three quota rows all read "CLAUDE CODE" with the scope
truncated away, naming the same thing three times and identifying none of them.
Titles now have a floor and the harness name prints once per group.

WORKSPACE is deliberately flat: a Repository → Worktree → Pane tree would need
lumbridge-git, and every level above Pane would be a second fixture. The depth
field and disclosure column are reserved for when it is real. HOSTS has two
states, live or not — connecting and unreachable are unbuildable until
lumbridge-remote exists, and shipping them would be the Buzz card again in a
Rust enum.

The rail drags between 200 and 480 px, applied live so the workspace reflows
under the pointer; decision 0009 measures pane thresholds after the sidebar, so
widening really can drop three panes to one. PTYs are resized on release only, or
every mouse-move is a SIGWINCH storm through the runtime's bounded queues.

While the sidebar owns the keyboard, on_key_down returns before encoding
anything. Without that guard a bare `j` would be written into whatever pane
happened to be selected while the user believed they were walking a list.

Not persisted yet, not virtualised, and describe() has nothing to attach to until
the accessibility adapter from decision 0017 lands. Recorded in 0021.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 23:56:17 -07:00
Metal AgentandClaude Opus 5 3a8a100ea5 Make the attention count real, and make it carry a source
The sidebar had said ATTENTION · 1 since the first commit. It was reading a
fixture's needs_input flag that no live pane ever set, so the number was frozen
at whatever the demo data said.

Attention is now built from observed signals, and every signal carries an
AttentionSource. The rule, enforced by is_countable on the source rather than by
a filter at the call site: a guess may draw a row and sort it, but may not
increment the count. That is decision 0012's provenance rule applied to a
different claim, for the same reason — a wrong count teaches people to ignore
the number, and the number is the whole point of the section.

RuntimeObserved is the only source that produces signals today, from process
exits and runtime faults. The two ACP kinds are declared and never constructed,
so there is a shape for the ACP client to fill and nobody is tempted to
approximate "asked you a question" by watching output for a question mark.

RuntimeEvent::Exited carries a u32 code that was being formatted into a sentence
and discarded; AttentionKind::Finished keeps it, which is why a row can say
"Exited with code 42" instead of "needs attention".

Verified live: exiting a pane with code 42 produces ATTENTION · 1, a row reading
"Exited with code 42 · observed", a dimmed pane banner offering restart, and a
live-PTY count that drops from 3/3 to 2/3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 23:38:55 -07:00
Metal AgentandClaude Opus 5 d76da3babb Make the terminal usable: control keys, paste, scroll, restart, terminate
The largest defect was not the one the plan named. Every control character was
being dropped before it reached the PTY — ctrl-c, ctrl-d, ctrl-a, ctrl-r, not
just ctrl-k — because terminal_key_from_parts required a key_char and GPUI
reports none for a control chord, since ctrl-k produces no printable character.
Verified with `cat -v`, which now prints ^K^A^R; ctrl-c interrupts a sleep and
ctrl-d ends a heredoc. The engine had always encoded these correctly; nothing
ever handed them to it.

The binding shadowing was real too. OpenPalette was on secondary-k, which is
ctrl-k on Linux, and GPUI stops dispatching once a binding claims an event, so
readline's kill-line was unreachable in every pane. Pane selection sat on
alt-1..6, which readline reads as a digit argument, and focus movement on
alt-arrows, which is word motion in most terminals.

Bindings now live in keymap.rs with the rule written down and tested: no binding
may be a bare control character or a bare Meta sequence, because those are what
a terminal application actually receives. A leader chord was considered and
rejected — GPUI parks a chord prefix for a second and drops it if focus moves.

Also in this pass:

- Paste on secondary-shift-v, through the engine's bracketed-paste path so a
  shell that asked for bracketed paste is told this is a paste. secondary-v
  would have been ctrl-v, which readline reads as quoted-insert. There is no
  matching copy: the engine has no selection yet, and a key that copied the
  whole screen would not be the same feature under the same name.
- A scroll wheel on the terminal surface. Shift+PageUp was the only route to
  scrollback, which is not something anyone guesses.
- Restart and Terminate. RuntimeRegistry::shutdown existed and was called only
  from its own crate's tests, so nothing in the application could ever stop a
  PTY. Terminate is the literal words with a confirmation naming the pid, per
  decision 0010, never a close icon; restart keeps the pane and replaces the
  process, per decision 0011.
- A dead or faulted pane now says so over its stale screen instead of looking
  idle, and typing into a pane with no terminal explains where the keystroke
  went instead of silently discarding it.
- The twelve reachable .expect panics on live-terminal state are gone. A pane
  can outlive its runtime — failed spawn, terminate, restored snapshot — and
  every one of those paths used to be a panic in the middle of a paint.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 23:35:32 -07:00
Metal AgentandClaude Opus 5 1556b87f37 Derive the interface palette instead of hardcoding eleven colours
main.rs held eleven `const … : u32` colours, and spikes/floem-shell held a
byte-identical copy of the same eleven. Every one was a judgement call made once,
and no user could change any of them without recompiling.

lumbridge-theme takes a syntax theme's five anchors — background, foreground,
comment, and the git added/deleted/modified colours where the theme has them —
and derives the whole role set. The frame is the editor background pushed one
logarithmic contrast step away from the content, so the work surface is the
brightest thing on screen; a theme already at black lifts its surface instead of
sinking its frame, which is why a pitch-black theme still shows a seam.

Adapted from Buzz's adaptive-theme.ts (block/buzz, Apache-2.0) as a
specification, not as copied code. The golden vectors were taken by running the
original under Node — a research pass had supplied Python-derived vectors and
claimed they reproduced it byte-exactly, and they did not: Python rounds
half-to-even, JavaScript rounds half-up, they disagree on exactly one channel
value of 22.5, and that decides whether the luminance bisection converges a step
early. github-dark's chrome is #171a1d, not #191c20.

Provenance colours are separate roles from state colours, with a test holding
them pairwise distinct in every theme, because decision 0012 colours a usage
reading by where its number came from and never by how alarming it is.

This changed no pixels, and that was verified rather than asserted: the only
difference between before-and-after screenshots is the digits of a process ID.
The check earned its keep — the mechanical rename had rewritten three user-facing
strings, turning the sidebar's "ATTENTION · 0" into "theme.attention · 0" and
"+ ADD PANEL" into "+ ADD theme.surface". A literal-by-literal diff now confirms
zero strings changed.

The default theme pins its roles to the previous constants to make that true;
the anchors underneath are real, and a test bounds how far the pure derivation
sits from them. The terminal ANSI palette keeps its own table, so 29 colour
literals remain in main.rs, all terminal. The catalog, its attribution, and the
picker are separate work.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 23:23:30 -07:00
Metal AgentandClaude Opus 5 2e349282c8 Split the Gitea job so graduation does not break CI on the runner
The single rust job ran cargo clippy --workspace, which after graduation pulls
GPUI onto a runner that has neither the X11 development packages nor a measured
disk budget. It would have failed on the next push.

rust-headless gates every push and stays fast. rust-ui installs the native
dependencies, reports df before and after so the runner's actual budget gets
measured, and is continue-on-error until it has run green a few times — a job
that OOMs on every push trains people to ignore CI, which is worse than one that
reports without blocking. The flag comes off on evidence, not on assumption.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 23:05:54 -07:00
Metal AgentandClaude Opus 5 316fa32745 Graduate the shell out of spikes/ into apps/lumbridge
The product was spikes/gpui-shell: a cargo workspace of its own, named in the
root manifest's exclude list. It inherited neither unsafe_code = "forbid" nor
clippy pedantic, and ./scripts/ci.sh never compiled it. Every test written into
it silently never ran, and apps/lumbridge was an eleven-line stub printing a
version string.

Four separate research passes over the sidebar, settings, devices, and theme
work independently discovered they were about to write substantial new code into
that directory. Graduating first means writing it once.

- apps/lumbridge is the product; spikes/ui-shell-model becomes
  crates/lumbridge-ui-fixture and joins the workspace.
- scripts/ci.sh takes --headless and --ui. The headless pass excludes the two UI
  crates by name, so a contributor changing lumbridge-core does not wait on a
  window toolkit, and a runner that cannot carry GPUI still gates everything
  else. A new crate is headless by default rather than silently joining the slow
  job.
- scripts/native-libs.sh replaces the ad-hoc symlink in the launcher, and says
  which apt package actually fixes the problem instead of working around it
  silently. The stale libxcb/libxkbcommon symlinks in the old spike target
  directory are gone; only libxkbcommon-x11.so was ever needed.
- deny.toml and cargo deny check licenses. spikes/README.md called GPUI's
  licence closure a hard gate and the scorecard scored it pending; graduation
  makes it the product's closure, so it is enforced rather than described. Two
  rejections were reviewed and allowed with the reasoning recorded in the file:
  webpki-roots under CDLA-Permissive-2.0 (Mozilla's CA store, data not code,
  reached through ureq) and libfuzzer-sys under NCSA (reached only under
  all-features via gpui's image decoder; no shipped build links it).

Clippy pedantic across both crates is clean at -D warnings. render was 353
lines; render_sidebar, render_tabs, and render_root come out of it, which the
sidebar rework needed anyway. The remaining over-length functions are single
declarative element trees and carry per-function allows with reasons, not a
blanket suppression.

Decision 0017 records the two calls this forces: published gpui 0.2.2 behind an
accessibility adapter rather than an unpinned Zed revision and an MSRV bump, and
Floem frozen rather than maintained in parity or deleted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 23:05:12 -07:00
Metal AgentandClaude Opus 5 834b73e831 Stop the shell asserting things that are not true
The footer was rebuilt on a real ledger two commits ago. The rest of the shell
was never audited the same way, and a multi-agent pass over it found the same
class of defect everywhere else:

- A sidebar card reading "Buzz · lumbridgecode / connected · signed identity"
  in the success colour. lumbridge-buzz is not a dependency of this binary.
- A saved host "amd-server", and a WORKTREES section with five entries and a
  working selector, backed by a lumbridge-git crate that does not exist.
- A first-run workspace of six panes announcing "Codex · runtime / metal ·
  Tailscale SSH", "Claude Code · UI / MacBook Air · local" and a Pi pane on a
  saved host. Every one of them was a /bin/sh, and the machine names were this
  developer's.
- A declared "Pi · spark-1 · laguna-s-2.1" usage profile with no probe of any
  kind behind it. Declaring a profile promises the gap is real; that one could
  never be filled.
- FOOTER_CENTER = "Codex · ChatGPT subscription · 62% window remaining",
  rendered by the Floem shell. Decision 0013 names that exact form as the thing
  that must never be shown.
- The header's PTY count painted green unconditionally, so "0/5 LIVE PTYS" read
  as success. runtime_rows already had the right rule three hundred lines away.
- A browser panel describing itself as "An isolated system-web-engine surface"
  on the chooser screen where you pick it. There is no web engine in this build.

First run is now three real local shells, and a pane claims a harness when one
has actually been launched into it. The seed mapping stays for when that is
possible.

Also removes the only unsafe block in the shell: a test set LUMBRIDGE_*_PROBE
through the environment, which needs unsafe under edition 2024 and silently
disabled both probes for every other test in the binary. Replaced with
UsageFeedOptions passed to start_with.

Clippy pedantic on the spike goes 79 -> 15 against root CI's -D warnings, so
graduating it into the workspace is not gated on a warning cleanup. The four
remaining too_many_lines are the render split, which the sidebar work needs to
do anyway.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 22:54:24 -07:00
Metal AgentandClaude Opus 5 219c674aea Read Claude Code's own quota endpoint, not just its status line
Decision 0015 rejected the account usage endpoint because AGENTS.md forbade
reading a harness's credential. The rule was written to stop one program
helping itself to another's secrets, and it was catching a legitimate use with
it: the user asking about their own subscription, through software they
installed to do that. AGENTS.md now states the narrow allowance instead of an
absolute the project does not hold, and 0016 records it.

The status line stays. It is free and it speaks every turn. What it cannot do
is report the per-model weekly limits a Max plan meters separately, or answer
at all before a session has taken a turn. The first live reading found the
account-wide seven-day window at 38% left and a per-model weekly window at 77%
left — a second ceiling the footer previously could not see.

Constraints the credential is read under, all enforced in code: access token
only, never the refresh token; zeroed on drop, along with the file buffer it
was borrowed out of; unprintable by construction, since HarnessError carries no
owned strings and AccessToken's Debug is hand-written; identified as
lumbridge/<version>, because sending claude-code/2.1.0 would make our traffic
indistinguishable from the harness's in Anthropic's logs; and off entirely
under LUMBRIDGE_CLAUDE_OAUTH=0.

The request runs on a detached thread with a slow refresh and a 429 backoff, so
a ten-second round trip cannot stall the transcript follower or make quitting
wait on the network, and one surface failing does not fault the other two.

Footer polish on top: the harness name prints once per group instead of in
front of each of its four windows, each quota carries a short scope pill
(5h, 7d, Fable wk, tokens) where an invisible BORDER-weight label used to be,
quotas sort ahead of spend, and a window under ten percent turns its headline
amber — value colour on the number, provenance colour on the meter, never
mixed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 22:04:34 -07:00
Metal AgentandClaude Opus 5 ef52aa7ce2 Replace the footer's placeholder usage with a real observation ledger
The footer showed invented percentages. It now shows what two harnesses
actually report, or says it does not know.

lumbridge-core gains an append-only per-profile UsageLedger and a projection
that labels every derived value estimated, withholds a burn rate from a single
sample, withholds a window fraction with no reported ceiling, withholds an
exhaustion estimate that lands after the reset, and reports an expired window
as rolled over rather than freezing its last percentage. A missing fact renders
as missing, never as zero. (0012)

lumbridge-harness is the impure side: processes, clocks, and untrusted wire
text in, observations out. Three adapters:

- Codex's account/rateLimits/read over the app-server's JSON-RPC stdio. The
  client cannot express a request outside a two-variant enum and answers every
  server-to-client request with -32601, so a harness asking Lumbridge for a
  credential is refused by construction. (0013)
- Claude Code's session transcripts, as a byte-offset tail follower that
  reports nothing until the backlog is read to EOF — a partially-read backlog
  is indistinguishable from a burst of spend, and the first run against 20 MB
  reported forty-six billion tokens an hour. The parser models four counters,
  so the conversations in those files are not representable. (0014)
- Claude Code's five-hour and seven-day subscription windows, via a bridge
  installed as its statusLine command. 0014 had claimed no such surface
  existed; it does, and the record is corrected in place rather than quietly
  edited. Lumbridge does not read the OAuth credential to call the account
  usage endpoint, which is what comparable tools do — AGENTS.md forbids it,
  and 0015 says so rather than leaving the gap unexplained.

Also in here: a capability-check ordering fix in the workspace reducer, where
the applied-request replay table was consulted before the capability check and
so answered questions the caller had no right to ask; the GPUI spike wired to
the live probes with per-harness gauges and provenance chips; and a launcher
that matches its own window by PID, because GPUI sets WM_NAME but not
_NET_WM_NAME and a title match never succeeded.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 21:47:11 -07:00