Commit Graph

62 Commits

Author SHA1 Message Date
karti 075cd765c5 Stop seeding illustrative Cap rows onto Learn
CI / verify (push) Successful in 7m14s
CI / publish (push) Has been skipped
The `supply` and `demand` concept tracks each carried one `DEMO — ` card
embedding a public recording from the Cap instance at video.karti.ai. Both are
gone, for two reasons that arrived together.

One of the recordings is no longer there. `sjqqvthbfma27bm` now answers 404 on
both `/s/` and `/embed/`, so "DEMO — What a hold takes off the board" promised
six seconds of teaching and played a dead frame. That is precisely the failure
the comment above these rows was written to record the last time it happened,
and it came back — because the footage lives on an instance this repository does
not control, so no amount of care in this file can keep the claim true. Checked
just now rather than assumed: the surviving id, `1rqq9rk4dpp71fd`, still answers
200.

The other reason is judgement. /learn is the page deliberately shown to people
outside this company, and a `DEMO — ` card sitting beside genuine product
footage makes the whole page read as half-placeholder to exactly the audience it
exists to convince. The previous version of this file argued the opposite — that
an empty track hides the shape of the page — and that was the wrong trade. An
empty concept track renders an empty state saying nothing has been published to
it yet, which is true, and true beats furnished.

`seedLearn` and its call site go with them rather than being left as an empty
array behind live machinery. If those tracks get purpose-shot recordings they
belong in `HOSTED_LEARN_MANIFEST`, rendered by `scripts/learn-film.mjs` like the
platform track, and served by PIG itself.

The two rows were also deleted from production directly, matching what
`clear()` does for prefixed learn rows, so the page is correct now rather than
at the next reseed. The underlying Cap recording is untouched.
2026-08-19 02:22:15 -07:00
karti e0473e7258 Cut two vertical Motion films, and write down the pipeline that makes them
CI / verify (push) Successful in 7m23s
CI / publish (push) Has been skipped
THE PIPELINE. The five existing platform videos were cut by an ad-hoc process
that was never committed, so the first time the UI moved nobody could re-shoot
them — which is the same failure `scripts/screenshots.mjs` was written to stop.
`scripts/learn-film.mjs` is that pipeline, and it reuses both of the screenshot
script's hard-won lessons: the appearance preference is stored server-side and
adopted after hydration, so the theme has to be forced by rewriting the profile
response rather than by seeding localStorage; and per-device layout state has to
be pinned or the framing is whatever a human last left behind.

Order is load-bearing. The narration is rendered first and its MEASURED duration
drives every shot length, because a shot list timed by guess leaves the narrator
talking over a frozen frame. Shot durations are then proportional to the words
spoken over them, so the cut lands on the sentence.

Typography is composed in the browser rather than in ffmpeg. The product's face
is Manrope Variable and drawtext would have fallen back to DejaVu, which reads
as a different company. Playwright renders a transparent chrome layer per shot
and ffmpeg only moves pixels.

Detail shots crop to the content column rather than the whole viewport. The
first attempt cropped a box around the element being talked about, which is
narrower than the column, so it sliced the cards either side of centre and the
frame read as broken rather than as close. Width first, height from the aspect.

CRF is 22, not 18. The frame is a static UI with a slow zoom, and 18 spent
3.7 Mbps — a 15 MB download for a video whose whole point is that somebody opens
it on a phone between meetings.

THE STAGE RAIL BUG, which the films found. Every stage label was losing its last
letters — "QUALIFICATIO", "PROCUREMEN" — with no ellipsis to show for it. Three
attempts to fix that by widening the card did nothing, because the card was
never the thing being measured: the LI holding it had `min-w-0` and no
`shrink-0`, so it took its flex share of 88px while the card inside stayed 160,
and every card was overpainted by the next one. The DOM reported no overflow the
whole time, because there wasn't any — the clipping was one level up. `shrink-0`
moves to the LI, where it belongs, and the rail scrolls as it was always meant
to. The labels also wrap rather than truncate now: they are a fixed vocabulary
of eight words we control, and losing a letter is worse than taking a line.

THE FILMS. Two 9:16 clips for the team, narrated in Karti's cloned voice through
Chatterbox at 1.20x and scored with the platform's own ambient bed, ducked and
loudness-normalised so they do not jump against the five already on the page.
Both open dark and switch to light at the midpoint. Shot lists live beside the
narration in `docs/learn-films/`, because a script and its shot list timed
against each other are one object and splitting them is how they drift.
2026-08-19 01:58:35 -07:00
karti 666310b264 Fill the demo book's motion: eight engagements, one at every stage
CI / verify (push) Successful in 7m20s
CI / publish (push) Has been skipped
The Motion half of the demo book was two engagements, which was enough to show
that the loop works and not enough to show what the product is for. The page
that matters asks whether the motion is repeating, and a stage rail of zeroes
cannot answer it.

Eight engagements now, one at every open stage, hanging off demand deals the
demand book already creates — the pipeline happened to have exactly one open
deal at each of the eight, so no deal was invented and the quoted pipeline
counts are unchanged. Forty-four artefacts, nineteen scores, one promotion.

Three things the book is laid out to prove that a folder of templates cannot.

Every stage is occupied, and every one of the twelve starter templates is
instantiated at least once, so "stages covered 8/8" is a measurement rather
than a claim about the seed.

Scores move, and sometimes move down. Nineteen scores across eight
trajectories, with nine dimensions regressing somewhere — the Verity
fine-tuning record runs 62.5 -> 60.5 -> 84.0 -> 80.8, because a scorecard that
only ever rises is a ratchet and teaches a reader to distrust it. Every score
is computed with `motionScoreBasisPoints` and banded with `motionBand` rather
than written as a literal, so the seed and the product cannot disagree about
what the same dimensions are worth.

The artefact bodies are the customer's own facts — named people, real volumes,
the specific thing going wrong, and a live unresolved risk in each. An artefact
whose body is the template with the blanks still in it is precisely what this
data exists to disprove. Six of the forty-four have no template at all, which
is the honest shape of an engagement and the reason `kind` is carried on the
artefact rather than derived: `engagement_artifacts.kind` is NOT NULL and a
derived kind would have been null for exactly those six.

The bodies live in JSON beside the loader for the same reason the starter
library's do — forty-four markdown bodies as backtick strings is a module
nobody can review.

The promotion copies `body` from the artefact verbatim, as `promoteArtifact`
does, rather than writing a hand-authored version 2. A demo that produced a row
the real path could not have produced would teach the wrong shape of the table.

Two name collisions the authors could not see are fixed: a Quillon contact
shared a full name with a demo seller, and an Aurelian one shared a surname
with another. `usage_count` is raised once per template rather than once per
artefact, so it stays symmetrical with the decrement `clear()` already does —
verified by tearing the book down and confirming the library returns to twelve
templates with every counter back at zero, since a counter left above zero
makes a starter template permanently un-editable.
2026-08-19 01:05:48 -07:00
karti b7d1ffd2d8 Speak the scorecard's own band labels, not a second vocabulary
The Trainability and Deal Qualification Scorecard publishes five bands and an
action for each — Decline, Defer, Scope down, Qualified conditional, Build —
and `MOTION_BANDS` published four different ones with different edges. So a
reader could read the scorecard, score a deal against the exact dimensions it
defines, and be told "Strategic" by a band table that document has never heard
of. Two answers to the same question from the same product.

The scorecard wins, on two grounds. Its edges were chosen alongside the
dimension weights they sit on top of, so 78 means something there and 7500 was
a round number here. And every one of its labels is a verb the reader can act
on: "Qualified" describes a deal, "Scope down" says what to do about it, which
is the only reason to band a score rather than show it.

The labels are now duplicated between the JSON a customer reads and the table
the product renders, because a rendered label cannot reach into a seeded row.
That duplication gets a test asserting the whole table verbatim, so
re-authoring one copy alone fails rather than drifts.

`apps/api/test/motion.test.ts` asserted the literal 'Strategic'. It now derives
the band through the shared function, so a band-table change is caught by the
test that owns the decision instead of by a write-path test that does not.
2026-08-19 00:26:21 -07:00
karti 2f32186d22 Fix five defects found by running Motion rather than reading it
CI / verify (push) Successful in 7m21s
CI / publish (push) Has been skipped
The markdown parser was in the eager entry chunk. `manualChunks` in its
object form does not leave an unlisted vendor package to Vite's async
splitting, so react-markdown was hoisted into the entry even though its only
importers are lazy routes — 327.70 kB gzip against a 314 kB baseline, on the
one download every route pays for. Naming it as its own chunk puts it back
behind the Motion pages and takes the entry to 282.21 kB, below where it was
before Motion existed.

The starter library could never be improved. Seeding was insert-only, so a
deployment seeded in August was frozen on August's wording for ever with no
upgrade path short of editing production rows by hand — for a feature whose
entire premise is that the library gets better. A second run now refreshes a
starter row, but only while it is still ours: `is_system`, `usage_count = 0`
and no owner. That is the same condition §7a already enforces on the API, so
a template an engagement was cut from is left alone and reported by name
rather than silently overwritten.

The refresh was not idempotent, and the seed lied about it. `jsonb` does not
preserve key order — Postgres sorts keys by length then bytewise — so
comparing `JSON.stringify(stored)` against `JSON.stringify(authored)` marked
every template as changed on every run, and the seed rewrote nine rows each
time while reporting itself clean. Comparison is now canonical. Found by
running the seed three times and reading the counts.

`motion-overflow-check.mjs` measured less than it claimed. It seeded
`pig.sidebar` and `pig.piggy.dock`, neither of which anything reads (the keys
are `pig.sidebarOpen` and `pig.piggyDockOpen`), so the layout it pinned was
whatever the last run left. Its dark pass set `colorScheme` only, and the
appearance preference is stored server-side and adopted after hydration, so
the dark pass measured the light palette a beat after first paint. It now
rewrites the profile response as `screenshots.mjs` does, asserts the rendered
`data-theme`, and fails a page that renders almost no text — a page that
throws inside its own body otherwise measures zero overflow and passes.

The stage rail rendered "1 templates", in the visible label and in every
aria-label. Singular and plural are now both passed.

Also normalised `artifact` to `artefact` in the seeded prose, which had
drifted American in the playbook. The `artifacts` field key is untouched:
FieldsView reads it, and already labels it in British.
2026-08-19 00:07:08 -07:00
karti 15c72ade1c Fix twenty findings from the Motion review
CI / verify (push) Successful in 4m47s
CI / publish (push) Failing after 3s
Each was raised by a reviewer and then survived an independent attempt to
refute it. The four that mattered most:

- A third of the starter library was invisible. Three templates authored
  `fields` shapes no renderer read — decisions, blockingSet, checks,
  steps and the rest — so about forty records rendered as no DOM at all,
  in the library and again on the engagement that instantiated them.
  Nothing failed: a renderer returns null for a key set it does not
  recognise, and a header-plus-body page looks like a template written
  that way. FieldsView now reads every key the seeds carry.
- "Add a framework" opened a picker that could never match, because the
  dialog was seeded with both the forced kind and the deal's stage, and
  qualification serves only the qualification stage. The stage is now
  dropped when MOTION_KIND_STAGES says the pair is incoherent.
- Piggy reported the promotion count as an exact figure capped at 8,
  against a tile showing the true count beside it. It is now counted in
  SQL, and all three motion tools carry a ResultScope whose denominator
  is shared lineages — never rows, never private drafts.
- No Motion test went through createApp, so the whole feature could be
  unmounted with a green suite. That is the AGENTS.md §5 trap that
  already cost this project read-guards.ts and learn.ts.

Also: both sides of the instantiate/edit race now lock, so a template
cannot be rewritten under an artefact that has copied it; concurrent
engagement opens queue on the deal row and get the 409 the handler
already promised rather than a 500; latestScore uses DISTINCT ON instead
of losing engagements past a 200-row cap; the migration adds the
scored_by_user_id foreign key the schema declares; and the demo clear
refunds usage_count for engagements it reaches by cascade, which
otherwise left starter templates permanently un-editable.

Verified on a fresh database: 16 migrations apply and re-apply as a
no-op, both seeds idempotent, usage_count back to zero after --clear.
564 unit tests pass. Every Motion route measures zero horizontal
overflow at 393 and 1440 in both themes, and all twelve seeded field
trees are asserted onto the screen by scripts/motion-fields-check.mjs.

One thing left open deliberately: the shipped qualification scorecard's
five bands and MOTION_BANDS' four are calibrated differently. The
framework's table is now titled as its own guidance rather than the
product's verdict, which removes the contradiction on screen. Making the
framework's calibration authoritative over the persisted band column is
a product decision nobody has made.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
release-2026-08-17
2026-08-17 19:08:55 -07:00
karti 376ef3d597 Answer 404 for an id that cannot name a row, rather than 500
Measured against a running server: five of the eleven Motion `:id` routes
answered `500 {"error":"Internal error"}` for an id like `nope`, and the
other six only answered 400 because their body schema happened to be
checked first — a valid body would have reached the same cast.

Nothing was wrong with the not-found handling. That branch was never
reached: every id column is a `uuid`, so Postgres refuses the parameter
with `22P02` several layers below it, and the error is not a
MutationError so it leaves as a 500.

404 rather than 400, because a 400 for a malformed id and a 404 for a
well-formed one tells anyone probing which of their guesses are the right
shape — and this feature already routes "somebody else's private draft"
through the same 404 so that no answer distinguishes the reasons a row is
not yours to see.

Also adds the AGENTS.md §5 393px check for the five Motion routes, which
scripts/screenshots.mjs does not photograph. All five measure zero
horizontal overflow at 393 and 1440, light and dark.

The same 500 is reachable on /api/accounts/:id and /api/contracts/:id,
which predates this branch and is left alone here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 18:41:28 -07:00
karti 7a6852e33a Merge gitea/main into the Motion branch
Motion was written against a base five commits behind main, so the
integration is the interesting part of this commit:

- The migration is renumbered 0014 -> 0015. Main shipped
  0014_piggy_conversations, and two migrations sharing an index is a
  journal that applies one of them.
- The seed-idempotency gate keeps main's all-tables diff rather than the
  motion_templates counter this branch added; the general check subsumes
  the specific one.
- Nav gains a Motion group alongside main's new Workspace group, and
  Piggy keeps the mark main gave it.
- Stat keeps main's container-scaled figure, which already carries the
  min-w-0 this branch added for the same reason.
- Piggy's page labels keep main's refusal wording for the four pages with
  no tool of their own, and gain the three Motion routes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 18:30:45 -07:00
karti 516685526c Add Motion: the go-to-market operating system on top of the ledger
The ledger answers which contracted capacity is sold, to whom, at what
margin. It says nothing about the motion — the repeatable practice that
turns a customer conversation into a scoped deployment, and turns that
deployment into something the next one reuses.

Motion is deliberately not a parallel entity tree. DEMAND_STAGES already
is the motion, so Motion binds reusable artefacts to the stages of a
demand deal that already exists: an engagement hangs off one deal,
cascade deleted, one per deal by unique constraint.

Nine closed kinds, each declaring which stages it serves, and a starter
library of twelve templates covering all eight open stages.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 18:27:03 -07:00
claude 6aadf1423c Count what the seed wrote, instead of claiming it
CI / verify (push) Successful in 7m31s
CI / publish (push) Has been skipped
The demo seed's summary was hardcoded when the book was split into modules,
and it drifted the moment the book grew: it reported 12 demand deals and 5
capacity commitments against a database holding 13 and 6. Caught by reading
the seed's own output beside the table it had just written.

Nobody would have noticed for a while, because the numbers were close enough
to look right — which is the whole problem. That output is the only feedback
`pnpm db:demo` gives an operator, and a command that misreports what it did
teaches you to stop reading it.

Counted in SQL now, including the number of distinct stages, so the line
about the pipeline covering every stage is a measurement rather than a
promise the seed makes about itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 20:22:21 -07:00
claude a3b1298257 Count the accounts, so Piggy stops counting deals instead
CI / verify (push) Successful in 7m9s
CI / publish (push) Has been skipped
Asked how many accounts were on the book, Piggy answered "7 demand deals
(accounts)". Production holds 17 accounts and 7 demand deals. The number was
real and the payload had scoped it correctly as deals; the prose relabelled
it on the way out.

This is the other half of the scope fix. That one stopped a filtered count
being read as a total. This one is a total that was simply absent being
filled from the nearest available noun: /accounts resolves to the workspace
summary, which carried commitments, deals, margin and idle capacity and no
count of accounts anywhere. The route's own label admitted it — "Piggy reads
the book here, not the account rows" — which named the gap without closing
it, and a model given a question about accounts and a payload with no
account figure will always find something else to count.

So the summary now counts accounts and contacts in SQL, and the headline
leads with them, because the defective answer was assembled from the first
countable thing in that sentence. Archived accounts are excluded to match
what /api/accounts returns — Piggy disagreeing with the list on screen is the
failure that costs the tool its credibility — but they are reported
separately so the difference stays reconcilable. The side breakdown ships
with a note saying the tabs do not partition, since supply and demand tabs
each include "both" and therefore do not sum to the total: that is the next
reconciliation bug, pre-empted.

Five routes that genuinely have no data tool now say so in their guide
rather than naming a subject they cannot reach. Proven live: /accounts
answers 23 of 23; a question about geography is refused rather than guessed;
/team refuses without substituting a nearby number.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 19:32:54 -07:00
claude f2ef403ee9 Make every tool result say what it counted
CI / verify (push) Successful in 7m44s
CI / publish (push) Has been skipped
Asked how many capacity commitments were on the book, Piggy answered "3".
Production holds 5. It had called the idle-capacity tool, which filters to
blocks above an idle threshold, and read the length of that list as the size
of the book.

The system prompt already forbade this in terms — "never report a filtered
count as a total; pig_get_idle_capacity returns the blocks with idle hours,
not the book" — and the model did it anyway. That is the second time this
argument has been lost in the prompt, so it is settled in the payload
instead: a result that cannot describe its own scope will be misread
eventually, however firmly the prompt objects.

Every tool that returns a count or a collection now carries one shape:
what it covers, how many matched, out of how many, under which filters, and
whether the list was truncated. The denominators are read from the database
rather than inferred. The pre-formatted headline states the scope too, since
that is the sentence a small model quotes most readily — the idle tool now
opens "3 of 5 live capacity commitments on the book", which is the sentence
that makes the original mistake impossible to phrase.

Two details worth keeping. Record reads enumerate rather than filter, so
their scope states a boundary instead of a ratio: these are that record's own
figures, never book-wide totals. And the workspace summary's idle threshold
is deliberately recorded as 0, distinct from the idle tool's 0.25 — that
mismatch is why three different idle figures appeared across the UI, and
naming it in the data is how it stops being invisible.

Verified against the live model: the failing question now answers 5, demand
deals 13 and contracts 20 — each drawn from a payload whose filtered figure
was smaller — while "which blocks are sitting idle" still names exactly the
blocks that are.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 19:03:40 -07:00
claude 18d5f5bfc0 Make Piggy part of the product rather than a guest in it
CI / verify (push) Successful in 7m10s
CI / publish (push) Has been skipped
Piggy arrived as a chat panel bolted onto a CRM and then grew a workspace
around it. The layout was already right — the audit found the approval card
to be the best-designed object in the repo, and the account page's empty
panels less finished than anything in the workspace. What was wrong was
vocabulary: nobody had written the small things down, so both halves kept
inventing them.

Piggy was drawn with five different marks — a pig in the dock, a sparkle in
the sidebar and again on the model picker, a speech bubble on the Ask
buttons, and a stock robot glyph on every assistant message, which is the
one people look at most. There is now one mark. The composer, which is the
first control in the product since sign-in lands on /piggy, was the only
un-adapted shadcn field left: 6px radius against a 12px Send button it sat
8px from. A stat tile had been reinvented six times at three numeral scales,
and the same uppercase micro-label existed in five variants, two of them one
tab apart in the same rail. There were 63 hand-written font sizes: not a
scale, sixty-three opinions.

Underneath that, the focus ring was invisible. The global rule used
ring-accent, which Tailwind deliberately aliases onto the hover tint, so the
ring measured 1.01:1 against the light canvas — no visible focus indicator
anywhere in the product, for any accent, in either theme. It is ring-brand
now and measures 17:1. The warning, positive and info tones were darkened
until each clears 4.5:1 on a card, on inset and on its own chip, and the
light canvas moved to 98% so a card lifts without leaning on its shadow.

The mobile work is the part worth reading. A landscape phone gave the
transcript 28% of the viewport and a keyboard-up phone 16%, against a 45%
floor — and the fixed tab bar painted over the composer, covering the safety
sentence and half the Send button, because two source comments asserted the
bar stood down on short viewports and it never had. Both fixed and measured
by hit-testing rather than by screenshot. The composer itself was 64px tall
for a blank second line nobody typed, because the auto-resize effect sizes
to scrollHeight and scrollHeight counts rows — a CSS height could not win
against an inline style, so the attribute was the honest lever.

Verified across both themes driven through the app's own control: no
horizontal overflow on 15 routes at four viewports, 672 stat values that fit,
297 labels at exactly 11px/500, Escape returning focus to its opener rather
than the body on every overlay, and a rejected write no longer reporting
"Succeeded" with a green check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 18:22:15 -07:00
claude f0173440e4 Put Piggy on Prime Agent, and let it write to the book
CI / verify (push) Successful in 7m6s
CI / publish (push) Has been skipped
Piggy was a hand-rolled OpenAI tool loop. It is now a Prime Agent session —
Prime Intellect's own harness, embedded as a Node library — answering from
PIG's tools and, for the first time, able to put information into the CRM
rather than only read it out.

The harness is a coding agent, so the first job was taking the coding agent
away from it. `noTools: 'all'` plus an explicit allowlist leaves the model
with PIG's ten `pig_*` tools and no bash, no filesystem, no IPython. That
holds under attack: a hostile extension, a skill and a settings file planted
in the agent's own directory, then `setActiveToolsByName` called with every
built-in, still leaves ten tools, all ours. Both lines are load-bearing —
`noTools` alone registers nothing, and the allowlist is what admits our own.

Writing is gated rather than assumed. A change is proposed, not made: the
tool returns a description, the transcript renders a diff card, and nothing
reaches the database until someone presses Apply. Contracts, commitments,
allocations and compliance always stop for a human whatever the mode. Every
write runs through `executeMutation` as the calling user, so their
capabilities and the audit trail apply exactly as they would to a human's.

Four things about the SDK are wrong in its own documentation and cost a
debugging cycle each: models.json does not resolve an env var name for
`apiKey`, it sends the literal string; there is no built-in prime-inference
provider in 0.84.1; a ResourceLoader you pass in is never reloaded for you;
and the stock system prompt is a coding-assistant prompt that must be
replaced — but replacing it also silently removes the tool list, because the
harness only renders that section when it owns the prompt. AGENTS.md records
all four.

The expensive one was thinking level. The harness defaults to `medium`, and
nemotron spent an entire 4,096-token budget reasoning and returned an empty
answer. `low` was worse; `off` omits the parameter so the endpoint's default
wins. An explicit `reasoning_effort: none` via `thinkingLevelMap` took a turn
from 6,195 output tokens to 149.

And a turn is now bounded. The harness loop is `while (true)` with no
iteration cap; a runaway on a frontier model would have eaten the credit it
is supposed to report on. Ceilings on model calls and tokens, enforced both
through the harness hook and independently from the event stream, plus a
per-user daily spend limit — and the ledger now records spend on turns that
fail, which it previously discarded.

Signing in lands on /piggy, which is a workspace: conversations down one
side, the agent in the middle, what it did and what it cost beside it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 05:26:28 -07:00
claude 99d165b5e5 Rebuild Piggy's interface, and give the demo book a business to describe
CI / verify (push) Successful in 4m57s
CI / publish (push) Has been skipped
Piggy answered in raw markdown, threw away every tool result it streamed,
and fought the reader's scroll on every token. The three surfaces that
made it worth having — what it read, how it reasoned, what it cost — were
all on the wire and none of them reached the screen.

The transcript is now composed of five parts under components/piggy:
answers render through streamdown, the container sticks to the bottom
without pinning the reader there, tool steps say what they read and link
to the record, and each turn carries its model and token count. Three
lifecycle bugs went with them: Stop left a permanent spinner, a truncated
stream was indistinguishable from thinking, and a failed send destroyed
the message it failed to send.

Underneath, the inference path grew timeouts, jittered retries on 429 and
5xx, tolerance of the malformed frames a 30B model emits, and an
agent_runs row per turn so chat spend is observable. The system prompt now
states that a field ending in Cents is cents — without it nemotron renders
costPerGpuHourCents: 189 as "$189 per GPU-hour", which is a 100x error on
the most scrutinised number in the room.

The demo book was arithmetically incoherent: every deal's value
contradicted its own allocation revenue by up to 3.6x, nothing had ever
closed, no customer had any paper, and the marketplace was empty. Deal
value is now derived from the allocation, the book clears 5.3% across five
blocks with one deliberately underwater, and the renewal, compliance and
agent-provenance machinery finally has rows to act on. A --clear that
deleted every obligation, SLA term and capacity request in the database
regardless of origin is scoped to the demo's own ids.

Around that: accounts have a detail page, ⌘K searches the book, Settings
can mint the API keys it always claimed to, and deploy.sh actually ships
the agent instead of silently skipping its compose profile.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 00:34:18 -07:00
karti 76e3caa1cb Drop the Comp AI CRM acknowledgement
CI / verify (push) Successful in 3m33s
CI / publish (push) Has been skipped
Nothing in PIG derives from that repository. The fact model, the leased
agent task queue and the `agentBrief` field are our own designs, and MIT's
attribution condition reaches copied source, not ideas — so the credit was
a courtesy that misstated where this code came from.

The one line worth keeping was never a credit: AGENTS.md's rule against
lifting component files out of somebody else's repo. It is restated
generically, and the shadcn-from-upstream guidance stays.

Buzz keeps its NOTICE entry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 23:41:10 -07:00
karti d7e0cbeccc Shoot the README screenshots, and make them regenerable
CI / verify (push) Successful in 2m42s
CI / publish (push) Has been skipped
The screenshot section had been a placeholder since the shell was rebuilt as
three panes, because there was no cheap way to re-shoot and a stale image is
worse than no image. So this ships the capture, not just the captures:
scripts/screenshots.mjs takes all ten pages at 1440x900 and 393x852, in light
and dark, and screenshots-encode.py halves and re-encodes them to WebP — 16MB
of PNG becomes 1.9MB in the tree.

Two things would silently ruin a run, and the script exists to encode both.
Seeding localStorage['pig.themeMode'] is not enough: the appearance preference
is authoritative server-side and adopted after hydration, so every dark capture
snapped back to light a beat after first paint. The /api/me/profile response is
rewritten instead. And pig.sidebarOpen / pig.piggyDockOpen are per-device, so
whatever the last human left behind would otherwise leak in. The run also fails
on a wrong theme, a horizontal scrollbar or a console error — three things a
screenshot cannot show you.

Piggy is not pictured mid-conversation. It is off by default and no inference
credential exists, so such an image would be a staged transcript rather than a
capture. The README says that rather than implying the feature is missing.

Correcting what the README asserted while shooting against the running code:
read authorisation IS enforced — createReadGuardRoutes is mounted ahead of the
feature routes, and routes/learn.ts and routes/activities.ts are mounted too,
so only the HubSpot pair is still unreachable. The remaining read gap is that a
grant cannot be narrowed, there being no row-level team filter in the query
layer. Counts refreshed against the tree: 275 tests, ~47k lines, 14 migrations.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 21:24:24 -07:00
karti e82d5a90bf Theme the sidebar scrollbar
CI / verify (push) Successful in 3m48s
CI / publish (push) Has been skipped
2026-08-13 21:17:41 -07:00
karti 9db53cb36f Polish the public auth experience
CI / verify (push) Successful in 3m26s
CI / publish (push) Has been skipped
Replace the stacked sign-in card with a responsive public shell, a local monochrome compute field, restrained grain and reduced-motion-safe drift. Share the anonymous header with Learn so public navigation stays consistent, and carry the corrected spacing through registration and profile setup.
2026-08-13 20:51:34 -07:00
karti b63ed0181f Fix the music never starting in Safari
CI / verify (push) Successful in 3m20s
CI / publish (push) Has been skipped
It played in Chrome and was silent on every Apple device. Two false signals,
compounding.

Neither the play() promise nor a synchronous `paused` check tells you whether
audio is playing. Chrome REJECTS the promise when autoplay is blocked, which is
the behaviour the obvious implementation is written against. WebKit RESOLVES
it, reports `paused === false` for an instant, and quietly pauses the element a
moment later. So the code concluded it had started, faded the volume up on a
silent element, and — because it thought it had succeeded — never armed the
gesture listener that was the entire fallback. Music could then never start at
all, no matter how many times the visitor clicked.

Traced by hooking HTMLMediaElement.prototype.play and addEventListener before
the app booted: exactly one play() call, "resolved paused=false", and no
pointerdown listener ever registered on window.

Two changes. Playing state is now driven by the element's own `playing` and
`pause` events, which are the only honest source. And the gesture listener is
armed UNCONDITIONALLY rather than only on a detected failure — play() on an
already-playing element is a no-op, so the redundant case costs nothing and the
broken case is fixed. The listeners are capture-phase so a component calling
stopPropagation cannot swallow them, and cover touchend and keydown as well.

Verified 12/12: Chrome and WebKit x desktop and iPhone x /, /learn and /margin
all reach t>2.5s at volume 0.14, and mute still fades and persists in both.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 20:12:43 -07:00
karti 30171767f7 Refine the music: fades, and a mute control on every public screen
CI / verify (push) Successful in 3m43s
CI / publish (push) Has been skipped
THE BUG. An anonymous visitor to / — the sign-in page, which is what the link
in an email opens — got music with ZERO mute controls. The provider moved above
the router so it covers the signed-out screens, but those render outside Shell
and therefore have no app header to host the toggle. Audio a visitor cannot
switch off is the worst version of this feature. Every unauthenticated screen
now carries the control pinned bottom-right; Learn keeps the one in its own
chrome rather than getting a second.

FADES. Volume ramps 0 -> 0.14 over 1.1s on start and back down over 0.42s on
mute, easeOutQuad so a mute feels prompt while a start feels like the room was
already there. Snapping to full volume on the first click reads as a glitch.
Measured: 0.076 at +0.4s, 0.14 at +2.2s, 0.025 at +0.25s after mute, paused by
+0.85s.

NO CHANGE TO THE TRACKS, and this reverses what I said earlier. I claimed
pig-tech would splice audibly every 28 seconds and offered to crossfade the
loop. That came from comparing 0.4-second mean levels, which measures musical
content rather than continuity. Measured properly — the wrap discontinuity
against each track's own 99.9th-percentile sample delta — all three already
loop cleanly (0.03-0.10x, i.e. quieter than their own ordinary transients), and
a folded crossfade made them WORSE (0.09-1.32x). The originals ship unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 19:59:45 -07:00
karti afce0dda28 Play the platform music on the shared Learn page too
CI / verify (push) Successful in 3m41s
CI / publish (push) Has been skipped
Reversing yesterday's call at Karti's direction. The provider moves from Shell
up above the router in App, so a share-code visitor gets the same character as
a member rather than a silent page.

The reason that is safe is the same reason "autoplay" was never really
autoplay: the browser refuses audio until the page has had a real gesture, so
nothing plays the instant a link opens — it starts once someone is actually
using the page.

The anonymous page renders OUTSIDE Shell and therefore has no app header, so
the mute control is added to its own chrome. Music with no way to stop it is
the worst version of this feature, and a visitor who cannot find the switch
does not conclude the site has taste.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 19:04:31 -07:00
karti 0108a70131 Add platform music: three tracks, looped, mutable
CI / verify (push) Successful in 3m32s
CI / publish (push) Has been skipped
Three ~28s tracks Karti generated, re-encoded from 160kbps to 96kbps and
pulled from -15 LUFS to -26. They were mastered at foreground level; the
failure mode of background music in a tool someone has open for eight hours
is not "too quiet". Playback volume is a further 0.14 on top.

AUTOPLAY DOES NOT MEAN AUTOPLAY. Every current browser refuses audio with
sound until the page has had a real user gesture — Chrome sometimes relents
for a site with a high Media Engagement Index, Safari essentially never does
on a first visit. So `play()` rejects on mount, and the naive version of this
looks like a bug: the control says playing and nothing is audible. This tries
immediately, and on refusal arms a one-shot listener and starts on the first
click or keypress. Verified in Chrome: paused on load, playing 2.2s after the
first click. That also happens to be the kind behaviour for someone who opened
six tabs at once.

Mounted in Shell, NOT in App. The anonymous Learn page renders outside Shell,
and a share-code visitor opening a link someone sent them should not get
unexpected audio — that is the one context where it reads as a fault rather
than as character. Asserted by there being no <audio> element on that page at
all.

Preference is localStorage and deliberately not mirrored to the server, on the
same reasoning as the sidebar: whether you want music depends on whether you
are wearing headphones, not on who you are.

Also pauses on tab hide, because music from a tab nobody is looking at is the
thing people hunt through twenty tabs to kill.

The element is rendered rather than `new Audio()` so it is inspectable in
devtools and in a test; `preload="none"` means a user who mutes it never
downloads a track.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 18:52:47 -07:00
karti 19dd30acbe Give the Learn cards a real frame instead of a gradient
CI / verify (push) Successful in 3m23s
CI / publish (push) Has been skipped
The preview cards led with a generated gradient. It was a deliberate fallback —
nothing renders a frame of a Cap embed without loading the embed, and loading
nine embeds to decorate a grid is how a page becomes unusable on a phone — but
for videos PIG serves itself the frame is right there in the file.

The poster is named after the VIDEO's content hash, not its own:
`overview.4d4581ae.mp4` -> `overview.4d4581ae.jpg`. Re-rendering a clip changes
both names together, so a thumbnail cannot outlive what it claims to show. It
needs no schema column and no manifest entry, because the name is derivable.

`learnPoster.sh` cuts the frame with `thumbnail=90` starting four seconds in
rather than taking frame 0: the first frame of a Playwright capture is often
mid-paint, and a poster of a half-rendered page is worse than no poster.

The resolver ASSERTS the poster rather than verifying it — @pig/core is pure and
has no filesystem. That is safe in both directions: a missing poster 404s, which
`<video poster>` renders exactly as it renders no poster, and which the card
falls back from via onError. Claiming a poster that is absent is free; omitting
one that exists would cost every card its thumbnail.

Cap-hosted rows are unchanged and still get the gradient, verified by there
being exactly five <img> elements on a page with nine resources.

Also widens the media allowlist to jpg/webp. The filename pattern, the
traversal rules and the symlink check are untouched and still cover them,
because extension is the only axis that changed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 17:53:30 -07:00
karti 45b70b17f0 Redesign Learn, and give it five real videos in Karti's voice
CI / verify (push) Successful in 3m32s
CI / publish (push) Has been skipped
THE PAGE. The anonymous route rendered outside Shell, so it sat flush against
the viewport edge and read as a form rather than a product — which is the first
thing anyone at Prime Intellect sees when the link is shared. It now brings its
own chrome and leads with a hero; the platform track is a numbered course, the
concept tracks are a poster grid, and admin add/archive moved behind one Manage
toggle so they stop competing with the content. Verified in Chrome at 1440 and
393, light and dark: horizontal overflow is 0 in all three access states.

THE VIDEOS. Five ~30s walkthroughs, narrated in Karti's cloned voice through
Chatterbox and cut against real screen capture of the seeded demo book. The
audio is rendered FIRST and its measured duration drives the capture, because a
shot list that runs short leaves the narrator talking over a frozen frame and
one that runs long gets cut mid-sentence. Levels are loudness-normalised so
clips do not jump between videos.

Cap cannot take a programmatic upload — video.karti.ai needs an interactive
login — so PIG serves these itself. A native <video> on this origin needs no
iframe and therefore no CSP frame-src at all; Karti's own Cap recordings still
render through the existing iframe path, which is why the resolver is now a
discriminated union.

THREE THINGS THE VERIFIERS CAUGHT, all of which shipped green:

  - createMediaRoutes was never mounted. Every layer landed — migration, seed,
    both feeds, the bind mount, the docs — except the one that serves the bytes,
    so /media/learn/* fell through to the SPA fallback and answered HTTP 200
    text/html. The player showed a black box with working controls and no error.
    The tests certified the route factory in isolation, which proves the handler
    and says nothing about whether it is wired in. There is now an assertion
    against the ASSEMBLED app, and it fails loudly on content-type — the failure
    mode is a 200, not a 404.
  - A symlink in the media directory escaped the root. resolve() is lexical and
    stat() follows links, so the containment check this file's own header
    promised did not hold. realpath before the check closes it.
  - Vite proxied only /api, so self-hosted playback broke for anyone running the
    app the documented way — in the same invisible 200-text/html manner.

Also: a duplicate media slug used to throw from the middle of seedDemo() and
take out every later section; it now reports and skips that one entry. And the
player has an onError state, because content-addressed filenames mean a
re-render deliberately leaves the old row pointing at a file that is gone.

The three DEMO platform rows are dropped — five real recordings supersede them,
and placeholders sitting under real ones made the page read as half-finished to
the audience it is meant to convince. The supply and demand concept rows stay:
there are no real recordings for those tracks yet, and an empty track hides the
shape of the page.

Tests 275, typecheck clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 17:30:08 -07:00
karti a21ecf9e53 Stop deploy.sh from rewriting itself while bash is reading it
CI / verify (push) Successful in 2m42s
CI / publish (push) Has been skipped
`git reset --hard origin/main` replaces this script mid-execution. bash does
not slurp a script — it reads incrementally and remembers a byte OFFSET, so
after the reset it resumes at that offset into different content.

This is not theoretical. The 13dec6b deploy hit it: the new public-origin gate
and the rollback were on disk and never ran, because bash was still executing
the buffered previous version. That deploy exited 0 and the release is healthy,
so it cost nothing this time. The failure mode when it does bite is a spliced
or half-executed line, part way through a deployment.

scripts/autodeploy.sh has always re-exec'd from a mktemp copy for exactly this
reason. deploy.sh needed the same guard.

PIG_REPO_ROOT is resolved before the re-exec and exported across it: after the
re-exec `$0` is the copy in /tmp, so `dirname "$0"` would cd to the wrong tree.
Verified with a harness that rewrites the original mid-run and asserts the
child keeps both its content and its working directory.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 15:08:22 -07:00
karti 13dec6b4b8 Rebuild the shell, add Calendar and Learn, and govern reads
CI / verify (push) Successful in 3m45s
CI / publish (push) Has been skipped
Seven parallel agents and an adversarial verification pass. The three things
worth knowing before reading the diff:

RBAC WAS ALREADY BUILT. docs/build-plan.md marks F2 and F3 outstanding and is
stale — packages/core/src/permissions.ts and lib/mutation.ts shipped long ago.
So this does not rebuild them; it closes the gaps an audit found. The big one
is that reads were entirely ungoverned: every GET was "any authenticated
member", so a junior demand rep and a research contractor could both pull
per-block supplier cost and break-even prices from /api/capacity/margin, and
every contract's negotiated terms. For a company whose margin is the business,
that was the hole that mattered. Adds book:read / economics:read / team:read,
a readGuard middleware, and a `viewer` role below member.

THE BUTTON AND THE 403 DISAGREED — the exact thing F3 said must never happen.
Contracts.tsx never called can() at all, so its save button was always enabled
against a server requiring contract:sign; Capacity.tsx gated commitment
creation on deal:write/demand while the server wanted commitment:write/supply.

POST /api/activities was the one write bypassing executeMutation: no capability
check, and any member could mutate accounts.lastActivityAt as a side effect.
It is now a proper mutation() behind activity:write.

The shell becomes three panes — a collapsible shadcn sidebar with an account
switcher on the Piggy accent, a header with real search, and Piggy docked to
the right, page-aware and persistent across navigation. The phone keeps its
bottom tab bar, which is the thing this product already beat trycompai/crm on,
and gains the sidebar as a sheet.

Calendar is a projection over thirteen dated sources rather than a new table,
because a table would duplicate dates that already live on contracts, deals and
commitments and would drift — and one ledger answering the question is the
whole argument. It surfaces export_authorizations and compliance_artifacts,
which had indexed expires_at columns, schema comments saying they must be
alerted on, and no read endpoint or UI anywhere.

Learn carries two tracks. Concepts are members-only; the platform track can be
opened with a share code by someone with no account. The code mints a scoped
learn-only token and never a Principal — every route here resolves a principal
and then checks capabilities, so a principal-minting code would be one missing
check away from leaking the book. "Only platform-track rows may be code-visible"
is a database CHECK constraint as well as a write-path rule, and a test asserts
a valid learn token still gets 401 on /api/dashboard, /api/accounts and
/api/contracts — the same invariant scripts/deploy.sh refuses to ship without.

CD becomes tag-to-ship. CI publishes an image to the Gitea registry on a
release-* tag and cloud-2 pulls it, so no credential on the shared runner can
execute anything on production — by construction rather than by policy. Both
halves of deploy.sh's original rule survive: nothing on the runner reaches the
host, and a human still decides when it ships. deploy.sh gains a rollback and a
public-origin check, and PIG_IMAGE now reaches compose through `sudo env`,
without which sudo's env_reset silently resolved every release to pig:local.

Tests 141 -> 261.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 15:02:48 -07:00
karti 6cf80747cc Keep PIG out of search results until it is meant to be found
CI / verify (push) Successful in 3m11s
The app is pre-launch and shared by link with a handful of people at Prime
Intellect. It should not be accumulating a search footprint yet.

Three layers, because each covers a gap the others leave:

  - robots.txt asks well-behaved crawlers not to fetch at all.
  - The <meta name="robots"> tag covers the HTML document for anything that
    fetched anyway.
  - X-Robots-Tag covers everything that is NOT the HTML document — og.png,
    the manifest, the built assets — which the meta tag cannot reach.

noarchive and nosnippet are there so a cache or an excerpt cannot outlive
the page once this is reversed.

Deliberately NOT stripped: the og:/twitter: tags. Link unfurlers are not
crawlers — they fetch on behalf of the person pasting the link, and a
rendered card is exactly what we want when this is shared.

The real gate remains authentication: / returns the sign-in screen and every
/api/ route returns 401. This only stops the app being indexed.

To go public: delete robots.txt, drop the meta tag, drop the header.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 12:45:18 -07:00
karti e12d27edd1 Polish every product workflow across desktop and mobile
CI / verify (push) Successful in 3m32s
Reframe each screen around the decisions compute brokers make: sellable capacity, full-cost margin, pipeline movement, contract deadlines, evidence review, staged imports, and controlled agent access. Group the shell by operating domain, strengthen mobile navigation and sheets, add responsive record treatments, and make loading, error, empty, readiness, and retry states explicit.

The visual audit exposed sortable table targets and an unnamed file input only after exercising the rendered app, so this commit also pins those accessibility decisions at their actual interaction boundaries. Manrope is self-hosted as a single Latin variable subset to keep the stronger hierarchy without shipping unused font payloads.
2026-08-13 05:34:23 -07:00
karti 1318c0b841 Ship growth intelligence and demo polish
CI / verify (push) Successful in 3m51s
2026-08-13 04:44:14 -07:00
karti a6167629cc Move from npm to pnpm across the workspace, CI and the image
CI / verify (push) Successful in 3m23s
The monorepo was on npm workspaces. pnpm gives it a content-addressed store
shared between the eight packages, a lockfile that records the whole graph
rather than a flattened view of it, and — the reason this mattered in practice —
`workspace:*`, which makes an internal dependency unambiguous instead of a
version range that npm may satisfy from the registry.

Mechanics:

  - `packageManager: pnpm@11.21.0` pins the version; corepack installs it in CI
    and in the image, so all three environments resolve identically.
  - The npm `workspaces` array is replaced by `pnpm-workspace.yaml`. pnpm
    ignores the former, and keeping both would leave two sources of truth.
  - All six internal dependencies moved to `workspace:*`.
  - Root scripts use `pnpm -r --if-present` and `pnpm -F <pkg>`.

Two findings worth recording, both from running it rather than reading it:

`tsx` was a devDependency, but the server runs TypeScript directly in
production — the container's command is `pnpm exec tsx apps/api/src/server.ts`.
Under npm this was concealed by the runtime stage re-installing tsx by hand
after pruning dev dependencies. Under `pnpm install --prod` that sleight of
hand stops working and the image simply fails to start. tsx is now declared in
`dependencies`, which is what it has always actually been.

The first image build failed with ERR_PNPM_ABORTED_REMOVE_MODULES_DIR_NO_TTY.
That is not a pnpm bug: it had decided the modules directory was stale and
wanted confirmation before deleting it, which a non-interactive build cannot
give. The trigger was the host's `node_modules` reaching the build context —
there was no `.dockerignore` at all. pnpm's tree is symlinks into a
content-addressed store, so copying it into an image produces dangling links
and a directory pnpm rightly considers corrupt. Fixed by adding
`.dockerignore` and setting `CI=true`, which is required in any non-interactive
pnpm build.

`esbuild` is denied install scripts via `allowBuilds`. Its platform binary
arrives through the optional dependency `@esbuild/linux-x64` and the postinstall
only verifies it; confirmed by running the binary directly, which reports
0.25.12.

Verified under pnpm: typecheck clean, 150 tests / 0 failures, e2e passes, web
builds. The image was built and booted against a real Postgres — health ok,
`/api/dashboard` 401 with an issuer configured, `/` and `/capacity` serve the
SPA, `/og.png` serves as image/png, and the migrator runs from the pruned
runtime stage.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 04:15:54 -07:00
karti 2b50797349 Merge remote-tracking branch 'gitea/main' into feat/revenue-intelligence
CI / verify (push) Successful in 2m54s
2026-08-13 03:50:26 -07:00
karti 74e37f3e76 Lay the HubSpot and customer-lifecycle foundation
Work in progress from the Codex session, committed so nothing sits undeployed.
Verified before committing: typecheck clean across all packages, 139 unit tests
and the e2e suite green, migrations apply to an empty Postgres.

Adds the HubSpot integration boundary (OAuth, client, contracts, webhook
signature verification, sync), a growth route, customer-lifecycle service,
Piggy lifecycle tools, a Growth page, and shared lifecycle/hubspot types.

Two things are deliberately incomplete and should not be mistaken for finished:

`packages/db/src/schema/hubspot.ts` is NOT exported from the schema index, so it
is inert — no tables, no migration. That is the correct order (the shape can
settle before it becomes a migration), but it does mean the HubSpot routes have
no persistence behind them yet.

`pnpm-workspace.yaml` and `pnpm-lock.yaml` are left uncommitted on purpose. The
workspace file contains a literal unanswered placeholder — "esbuild: set this
to true or false" — and this repository installs with npm, which is also what
CI runs. Committing a second package manager's lockfile would make the install
ambiguous. If the move to pnpm is intended it should be a deliberate change
that updates CI and the Dockerfile together.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 03:50:18 -07:00
karti c2c7fb9c19 Make mutations confirm themselves, and seed the evidence trail
CI / verify (push) Successful in 2m51s
Two demo gaps, both of which made working features look like they were not
there.

**Toasts fired into nothing.** RecordSheets already called toast.success on
every save, but <Toaster /> was never mounted, so nothing appeared. It could
not be mounted, either: the shadcn original imports next-themes, which PIG does
not use — it has its own provider so a chosen theme is persisted server-side
and follows a user between devices. Rewired to PIG's useTheme, mounted inside
ThemeProvider, and offset clear of the phone tab bar and the home indicator.

Feedback added where the interface otherwise gives none: allocation and hold
report the GPU-hours actually written, because the sheet closes on success and
the only other evidence is a number moving off-screen; releasing a hold says
the capacity is sellable again; fact decisions say what the decision meant, and
that approving evidence is not the same as writing it to a record; the profile
form confirms rather than just clearing itself, which otherwise reads as the
input being discarded.

**The fact table was empty**, so the review queue and every provenance tooltip
had nothing to show — the mechanism that makes an agent-written CRM
trustworthy, invisible. Six agent-derived facts seeded with a deliberate mix:
two applied, showing what a confident agent writes unprompted, and four
proposed, including one weak claim that a reviewer should reject, so the queue
is not a row of obvious approvals. Each carries a score, a band, evidence and
where available a source. Idempotent on subject+field+value; verified over two
runs.

Verified: toast confirmed firing in a real browser on a 393px viewport, 135
unit tests and e2e green, typecheck clean, CSP hash unchanged, 0px horizontal
overflow across 12 routes at both breakpoints.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 03:39:00 -07:00
karti 2763531ce4 Align shadcn's accent token with what shadcn means by it
CI / verify (push) Successful in 2m53s
shadcn uses `bg-accent` for its SUBTLE surfaces — dropdown item hover, command
row selection, ghost and outline button hover, the dialog close affordance. The
brand colour in shadcn is `primary`.

PIG's Tailwind config mapped `accent` to `--accent`, which is the brand. That
inverted the meaning, so every shadcn hover and selection state painted a
full-strength brand block. With the monochrome "pig" palette in dark mode the
brand is near-white, so a selected command row rendered as a white slab against
a near-black sheet. Measured before the change: selected row rgb(250,250,250)
on a rgb(9,9,11) body.

`accent` now aliases `--accent-subtle` and `accent-foreground` aliases
`--accent-fg`, which is what those tokens were created for. The eleven places
where PIG's own components wanted a solid brand fill — filled chips, selected
card borders, progress bars — move to `primary`, which still resolves to
`--accent`. A `brand` alias is added for clarity.

After: selected row rgb(39,39,42) in dark and rgb(244,244,245) in light, both a
subtle tint above the body; the pipeline's active stage chip stays a solid
rgb(250,250,250) fill, unchanged.

Found by opening overlays, which earlier screenshot sweeps never did — every
route had been checked, but a dropdown or a command palette only misbehaves
once it is open. Worth remembering: page-level sweeps do not exercise portals.

Typecheck clean, 135 unit tests and e2e green, CSP hash unchanged, 0px
horizontal overflow across 12 routes at 393px and 1440px.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 03:17:49 -07:00
karti c821b2ca07 Authenticate against any OIDC provider, for on-premises installs
CI / verify (push) Successful in 2m55s
The seam existed with only a Supabase implementation, so an on-prem deployment
had no way to authenticate. A customer running PIG inside their own network
already has Okta, Entra, Keycloak, Auth0 or Google Workspace; asking them to
stand up a second identity system is a serious adoption tax and in a regulated
environment usually refused outright.

Setting PIG_OIDC_ISSUER is normally the whole configuration — the JWKS is
discovered from the issuer's well-known document. PIG_OIDC_JWKS_URI skips
discovery entirely for an air-gapped network. OIDC takes precedence over
Supabase so an on-prem install can leave the hosted values in its environment
file without them quietly taking over.

Three decisions worth stating:

Discovery is resolved lazily and the FAILURE is not cached. Doing it per
request would put the customer's identity provider on the critical path of
every API call; doing it eagerly at boot would mean their IdP rebooting takes
the CRM down with it. So it happens on first use and retries on the next
request.

The audience check is optional but warned about loudly. Without it, a token the
provider issued for ANY other application in the same tenant verifies here — a
token minted for an unrelated internal tool would be accepted as a PIG session.
It cannot be mandatory because some providers legitimately issue
single-audience tokens.

Email falls back through email, preferred_username and upn, because providers
disagree, but a preferred_username without an "@" is ignored — PIG keys
membership on the address, and a bare username must never become an account
identity.

Also fixed a warning that claimed "authentication is DISABLED" on a correctly
configured OIDC deployment. That is worse than silence: an operator who reads
it on a secure install learns to ignore the warnings. The dev bypass itself was
already correct — it keys on the resolved provider rather than on Supabase.

18 new tests, most of them about what the provider must REFUSE: a foreign
signing key, a foreign issuer, a token for a different application, an expired
token, a token with no subject, and a discovery outage that must not become
permanent. Keys are generated per test and the JWKS is served locally, so they
run offline.

Verified: production refuses to start with neither provider, starts with OIDC
alone, enforces 401 on an unauthenticated request, and warns only about the
genuinely missing admin list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 02:32:28 -07:00
karti 54edee30ed Unknown /api paths returned the SPA with HTTP 200
An authenticated GET to any unrecognised API route — a typo, a renamed
endpoint, an older client — fell through to the SPA fallback and returned
200 text/html containing the app shell.

This is close to the worst failure shape for an API consumer. `response.ok` is
true, so nothing treats it as an error; the caller then dies on `JSON.parse`
with "Unexpected token '<'" far from the actual cause. The MCP server, the CLI
and Piggy all consume this API and would all have hit it. It was masked from
casual testing because unauthenticated requests are rejected earlier by the
auth middleware, so it only appears once you hold a valid token.

Found by probing production with Scott's token: GET /api/keys (the real path is
/api/api-keys) returned 200 text/html.

The static-file middleware already carried this guard — added for the same
reason when og.png was being served as HTML — but the SPA fallback beneath it
did not. Same guard, one place missing.

Verified: unknown API paths now return 404 application/json, real API paths
still answer, client-side routes still receive the shell, and static assets
still serve with their own content types. Typecheck clean, 124 unit tests and
the e2e suite green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 02:05:00 -07:00
karti 6bd5526675 Give Card min-w-0 so the page stops scrolling sideways on a phone
The Overview page overflowed 80px at 393px wide. Traced to the "The book" card:
the grid column was a correct 361px, the card inside it was 457px and refused
to shrink. Confirmed by forcing `min-width: 0` on grid children in the live
page, which took the overflow to 0.

Fixed on the Card base class rather than at the call site, because this is the
third time the same trap has been fixed individually — grid and flex children
default to `min-width: auto` and cards routinely hold something unshrinkable, a
tabular-nums figure or a nowrap badge. `min-width: 0` is inert for a
block-level card outside a flex or grid parent, so applying it always costs
nothing and removes the whole class of bug.

Verified by running the stack locally against the demo data: 0px overflow
across all 12 routes at both 393px and 1440px.

AGENTS.md updated to say any NEW container primitive needs the same, with the
one-line browser check to confirm it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 02:01:17 -07:00
karti d4d7095605 Migrate before starting the application
CI / verify (push) Successful in 2m32s
2026-08-13 01:48:05 -07:00
karti 853bde2265 Build the agent-native compute CRM platform
CI / verify (push) Successful in 3m6s
2026-08-13 01:39:01 -07:00
karti bfd2f8d95a Add AGENTS.md — onboarding for whoever picks this up next
CI / verify (push) Successful in 1m32s
Named by convention so a coding agent finds it without being told. Written to
get someone productive from a cold start without having to reconstruct the
reasoning from the diff.

Four sections carry the weight. The architecture rules that must not be broken,
each of which fails silently rather than loudly — intelligence never lives in
the API, authentication is not authorization, cost is charged against the full
commitment, sold and held are different things. The traps that have already
cost time here, with the specific symptom each produces: onConflictDoNothing
being a no-op without a constraint, z.coerce.boolean turning "false" into true,
grid children needing min-w-0, Drizzle emitting a cast Postgres rejects, the
CSP allowing exactly one inline script by hash, and the Prime Intellect API
quoting node totals rather than per-GPU prices. The conventions, including that
comments explain why rather than what. And an explicit start order.

The last section says what not to do, which is the part most easily lost: do
not copy component files from the MIT project we borrowed ideas from, do not
open self-registration on a shared identity provider, do not put a production
key on the CI runner, do not weaken the guard that refuses to serve the CRM
unauthenticated, and do not invent email addresses for real people.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 22:25:57 -07:00
karti cf3117e458 Record verified Prime Intellect API facts, and log a real pricing bug
CI / verify (push) Successful in 1m33s
The token was found on cloud-1 after all, in a Claude memory note. Verifying its
claims against the live API turned up a defect in code already shipped.

**prices.onDemand is the total for the whole node, not per-GPU.** Confirmed:
datacrunch lists 1x A100 at 1.79 and 2x A100 at 3.58, and gpuMemory scales the
same way (640 for 8x 80GB). packages/prime/src/map.ts stores both as if they
were per-GPU, so an 8-GPU node reads eight times too expensive. It would have
silently poisoned inventory search, the max-price filter and every margin
comparison against bought capacity — and nobody would have noticed, because the
numbers still look plausible. Logged rather than fixed, per the instruction to
hold; it needs a regression test built from the real 1x/2x pair.

**Inference is a different host.** api.primeintellect.ai is compute and pods;
inference is api.pinference.ai/api/v1, OpenAI-compatible. PIG's config knows
only the first, so A4 and A13 need both.

**Piggy's default model** is nvidia/nemotron-3-nano-30b-a3b, and the important
detail is that it is a hybrid reasoning model which thinks aloud by default and
truncates under a tight max_tokens. `reasoning_effort: "none"` gives ~1s terse
output for tool use and extraction, which is what Piggy does nearly all of the
time.

One claim did NOT reproduce: the note warns of Cloudflare 403ing non-browser
user-agents, but PIG's own UA and curl's both returned 200. Recorded as history
in case a 403 ever appears.

Also logged: the key is a broad, never-expiring credential sitting in plaintext
in a memory markdown file. PIG's sync should hold a separate narrower key
scoped to availability reads.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 21:52:33 -07:00
karti 3398345109 Plan: add RBAC, capacity tiers, imports, admin settings; expand contracts
CI / verify (push) Successful in 2m36s
Six additions, and one promoted to foundation.

RBAC becomes a Wave 0 task rather than something absorbed into each CRUD
track. Authorization today stops at "is a member", and eight parallel write
tracks plus a bulk-import feature are about to land. Import especially: one bad
column mapping can rewrite thousands of records, so it must not ship before
there is a real answer to who may run it. Suggested default is team leads and
platform admins only — easy to loosen later, unpleasant to tighten.

A government/sovereign capacity tier joins secure and community. It is a schema
change, so it is cheaper before the tables carry real data. Matching must never
satisfy a government requirement with community capacity, and the tier
interacts with the export-control predicate already modelled in compliance.ts.

Imports are split so the work is reusable: a CSV/Excel framework carrying the
upload, mapping, dry-run preview and idempotent commit, with Notion and Google
Sheets as thin OAuth front-ends onto the same mapping step. Notion databases
are tables with typed properties, so treating them as a separate importer would
duplicate the hard part.

Admin settings gains Piggy's model selection, defaulting to a Nemotron model on
Prime Intellect inference, with the endpoint configurable so an on-prem install
can point at the customer's own.

Contracts is expanded from "surface the schema" to the full field set — the
hierarchy with order-form-beats-MSA precedence, negotiated SLA terms including
fee abatement and its trigger, spare-pool scope, maintenance classes and the
reasonable-endeavours carve-out, plus take-or-pay and termination tier. Written
as a first draft to be corrected by someone who negotiates these for a living.

Also records three open questions, including that the Prime Intellect API token
could not be found on cloud-1 — searched ~/.prime, ~/.config/prime, /opt and
/etc.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 21:33:38 -07:00
karti e4698e4d0c Add a build plan, informed by reading Comp AI CRM properly
CI / verify (push) Successful in 1m32s
Cloned trycompai/crm (MIT) and inventoried it rather than assuming. Three
findings shaped the plan.

Their component library is far deeper: 68 primitives to our 9, including
data-table, command, sheet, drawer, combobox, chart, and a set of agent-chat
components that map almost exactly onto what Piggy needs.

Their SourcedValue/Provenance pattern is worth adopting outright — a dotted
underline on any agent-derived value with a tooltip carrying the claim, the
reasons, when it was observed and the source URL. PIG already stores all of
that in `facts` and surfaces none of it.

But we are ahead of them on mobile, not behind. Measured across both repos:
211 responsive utilities across their 329 tsx files (0.6 per file) against 58
across our 16 (3.6 per file); zero safe-area handling to our five; no drawer or
sheet used for navigation, no viewport-fit. Their app is effectively
desktop-only. So the plan takes depth from them, not mobile behaviour.

The plan also says not to copy their component files. Most are shadcn/ui
originals — MIT, and designed to be installed from upstream where they are
canonical and current. Borrow the compositions as ideas; the debt is already
credited in NOTICE.

Structured into waves with real dependency edges so the work can be handed to
several agents without collision. Two foundation tasks must land alone first
(the primitive set, and the API write-path convention) because eight parallel
CRUD tracks would otherwise each invent their own.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 21:09:34 -07:00
karti 14417e34bc Margin table: say "Covered" rather than $0.00 for break-even
CI / verify (push) Successful in 1m51s
Same fix already applied to the capacity cards, missed here. A zero break-even
means the block's cost is fully recovered and any further sale is upside;
printing "$0.00" is technically true and reads like a rendering bug. The demo
book has three such blocks, so it was visible on every screenshot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 20:58:22 -07:00
karti 468979b303 Add a plausible demo dataset
CI / verify (push) Successful in 2m1s
So the product is legible before anyone has entered real data, and so Piggy
has something to reason about while it is being built.

Kept separate from the base seed because that one is publicly-sourced and cited
while this is invented. Two rules, both deliberate:

Every record is prefixed "DEMO — ", so a screenshot can never be mistaken for
real business. And demand-side customers are fictional. Suppliers are real
companies — they are public, and naming the actual market is the point — but
inventing customers with invented contract values against real named businesses
would be fabricating commercial records about them, which is a different thing
and not worth the extra realism.

The numbers are tuned to teach rather than to flatter. The book clears +5.4% at
79% utilisation, which is thin and about right for this industry once capacity
cost is charged honestly. Underneath, the blocks disagree: the large H200 block
carries it, the EU H100 block is underwater at 55% sold because a 46% markup
needs ~69% sold to break even, and the community pool holds a large unconverted
hold — so the difference between "sold" and "held" is visible rather than
theoretical.

An earlier tuning left the whole book at -26%. Honest, but it reads as a broken
product rather than an under-utilised book, so the totals now open healthy and
the problems appear on drill-down.

Also exercises parts of the schema nothing had touched yet: ramped capacity
shapes, negotiated SLAs with fee abatement and spare-pool scope, renewal
obligations with one deliberately near-term, EU data-residency constraints on a
capacity request, and internal research burn.

Two bugs found while testing it, both the same trap as before: the research
allocation duplicated on every run because onConflictDoNothing() is a no-op
without a matching unique constraint, and notes were double-prefixed. Verified
idempotent over three consecutive runs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 20:55:41 -07:00
karti 2a30645c8d CI: Postgres on 127.0.0.1, because the runner uses host networking
CI / verify (push) Successful in 1m59s
The runner is configured with `container.network: host`. That one setting
explains all three earlier failures, and the workflow now records them so
nobody repeats the sequence:

  services:                     not resolvable by name from a host-networked
                                job — "getaddrinfo EAI_AGAIN postgres"
  --network container:$HOSTNAME /etc/hostname is the HOST's name, not a
                                container id, so the join finds nothing
  default-gateway addressing    wrong idea outright: with host networking the
                                default route is the real router, not a bridge

Sharing the host's network namespace means a published port is just on
127.0.0.1. Readiness is now checked over TCP from the job itself rather than
with pg_isready inside the container — the latter proves the server started,
not that this job can reach it, which is the thing that actually failed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 20:32:06 -07:00
karti e6b4c1618e CI: reach Postgres through the gateway instead of a shared namespace
CI / verify (push) Failing after 5s
Second attempt failed differently: /etc/hostname inside the job reports the
HOST's name rather than the container id, so `--network container:$HOSTNAME`
found no such container.

Rather than hunt for our own container id through /proc, publish the port on
the host and connect through the job container's default gateway. That needs
no container identity at all. The port is derived from the run id so two
concurrent runs cannot collide, and DATABASE_URL is exported through GITHUB_ENV
once Postgres is actually accepting connections.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 20:30:24 -07:00
karti 40f6fd993d CI: start Postgres as a step rather than a service container
CI / verify (push) Failing after 5s
The first run got through install, typecheck and all 39 tests, then failed on
`getaddrinfo EAI_AGAIN postgres`. This runner does not attach service
containers to the job's network, so the `services:` hostname never resolves.

Fixed by starting Postgres with `--network container:$HOSTNAME`, sharing the
job container's own network namespace so it appears on 127.0.0.1. That works
regardless of how the runner is configured — which matters here because the
runner is shared with other repositories and should not need reconfiguring to
suit this one.

Also queries row counts through `docker exec` rather than a local psql, since
the runner image is not guaranteed to ship postgresql-client, and removes the
container in an `if: always()` step so a failed run does not leave it behind.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 20:29:25 -07:00
karti 73231a8944 Add CI, a test suite, and a deploy script
CI / verify (push) Failing after 34s
`npm test` did nothing until now. CI that runs no tests is theatre, so the
tests came first — 39 of them, over the two places where an error would be
silent and expensive.

packages/core: the margin arithmetic. Every dashboard figure, idle-capacity
alert and agent answer resolves through it, and wrong numbers still look like
numbers. The cases pin decisions rather than implementation: cost is charged
against the full commitment (a naive version reports the opposite sign on a
loss-making block), aggregation sums cents rather than averaging percentages
(averaging reports +22% on a book that is losing money), break-even prices the
remaining hours and returns null rather than Infinity when there are none, and
internal research burn counts as cost with no revenue.

packages/prime: the upstream mapping. Rounding rather than truncating cents,
because 2.43 is 2.4299999 in binary and a lost cent compounds across millions
of GPU-hours. And interconnect normalisation, where an unrecognised fabric maps
to Unknown rather than Ethernet — guessing low loses a deal, guessing high
sells a training customer a cluster that cannot train.

CI runs on push and pull request: typecheck all six packages, unit tests,
migrations applied twice to a real Postgres, a seed-idempotency assertion that
fails the build if row counts move on a second run, a server boot, the front-end
build, and a Docker build.

It also asserts the inline theme script's hash still matches the CSP the proxy
allows. That script prevents a white flash for dark-mode users; if it changes
without the CSP being updated, the browser silently blocks it and nothing
anywhere reports an error.

Deployment stays a script rather than push-to-deploy. Automating it would put
an SSH key with production write access on the CI runner — a real escalation
for a project this size. The script takes a database dump before migrating and
refuses to finish if an unauthenticated request returns anything but 401.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 20:27:47 -07:00