Read Claude Code's own quota endpoint, not just its status line

Decision 0015 rejected the account usage endpoint because AGENTS.md forbade
reading a harness's credential. The rule was written to stop one program
helping itself to another's secrets, and it was catching a legitimate use with
it: the user asking about their own subscription, through software they
installed to do that. AGENTS.md now states the narrow allowance instead of an
absolute the project does not hold, and 0016 records it.

The status line stays. It is free and it speaks every turn. What it cannot do
is report the per-model weekly limits a Max plan meters separately, or answer
at all before a session has taken a turn. The first live reading found the
account-wide seven-day window at 38% left and a per-model weekly window at 77%
left — a second ceiling the footer previously could not see.

Constraints the credential is read under, all enforced in code: access token
only, never the refresh token; zeroed on drop, along with the file buffer it
was borrowed out of; unprintable by construction, since HarnessError carries no
owned strings and AccessToken's Debug is hand-written; identified as
lumbridge/<version>, because sending claude-code/2.1.0 would make our traffic
indistinguishable from the harness's in Anthropic's logs; and off entirely
under LUMBRIDGE_CLAUDE_OAUTH=0.

The request runs on a detached thread with a slow refresh and a 429 backoff, so
a ten-second round trip cannot stall the transcript follower or make quitting
wait on the network, and one surface failing does not fault the other two.

Footer polish on top: the harness name prints once per group instead of in
front of each of its four windows, each quota carries a short scope pill
(5h, 7d, Fable wk, tokens) where an invisible BORDER-weight label used to be,
quotas sort ahead of spend, and a window under ten percent turns its headline
amber — value colour on the number, provenance colour on the meter, never
mixed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Metal Agent
2026-08-31 22:04:34 -07:00
co-authored by Claude Opus 5
parent ef52aa7ce2
commit 219c674aea
13 changed files with 1674 additions and 73 deletions
+19 -7
View File
@@ -233,13 +233,25 @@ consumption against none. The parser models four token counters and nothing
else, so the conversations in those files are not representable in a Lumbridge
value. See decision 0014.
Claude Code's subscription windows come from a different surface: the CLI pipes
a `rate_limits` object carrying the five-hour and seven-day windows to whatever
`statusLine` command the user has configured, on every turn. A small installed
bridge writes those fields — and only those — to a local feed the probe tails,
so the windows are `ProviderReported` and no credential is ever read. Reading
Claude Code's OAuth token would be more capable and is what comparable tools
do; Lumbridge does not, because AGENTS.md forbids it. See decision 0015.
Claude Code's subscription windows come from two further surfaces, because
neither answers the whole question. The CLI pipes a `rate_limits` object
carrying the five-hour and seven-day windows to whatever `statusLine` command
the user has configured, on every turn; a small installed bridge writes those
fields — and only those — to a local feed the probe tails, for free and without
a credential (decision 0015). The account usage endpoint supplies what the
status line cannot: the per-model `weekly_scoped` limits a Max plan meters
separately, and an answer on a cold start before any session has taken a turn.
Reaching it means reading Claude Code's stored access token, which `AGENTS.md`
now permits under a narrow named allowance — one documented question about the
user's own account, never persisted, never logged, never in argv, no refresh
token, identified as Lumbridge rather than as the harness, and switchable off
(decision 0016).
The endpoint is the authority whenever it answers; the status line covers the
interval between its deliberately slow refreshes. Both are `ProviderReported`
and share the window arithmetic, so one quota read two ways cannot produce two
numbers. Per-model limits arrive as profiles the probe announces at runtime,
since their names come from the response.
The GPUI shell now runs both probes; a probe contributes its own profiles on
top of the declared ones, so the Claude windows appear once they report. Its