Put Piggy on Prime Agent, and let it write to the book
CI / verify (push) Successful in 7m6s
CI / publish (push) Has been skipped

Piggy was a hand-rolled OpenAI tool loop. It is now a Prime Agent session —
Prime Intellect's own harness, embedded as a Node library — answering from
PIG's tools and, for the first time, able to put information into the CRM
rather than only read it out.

The harness is a coding agent, so the first job was taking the coding agent
away from it. `noTools: 'all'` plus an explicit allowlist leaves the model
with PIG's ten `pig_*` tools and no bash, no filesystem, no IPython. That
holds under attack: a hostile extension, a skill and a settings file planted
in the agent's own directory, then `setActiveToolsByName` called with every
built-in, still leaves ten tools, all ours. Both lines are load-bearing —
`noTools` alone registers nothing, and the allowlist is what admits our own.

Writing is gated rather than assumed. A change is proposed, not made: the
tool returns a description, the transcript renders a diff card, and nothing
reaches the database until someone presses Apply. Contracts, commitments,
allocations and compliance always stop for a human whatever the mode. Every
write runs through `executeMutation` as the calling user, so their
capabilities and the audit trail apply exactly as they would to a human's.

Four things about the SDK are wrong in its own documentation and cost a
debugging cycle each: models.json does not resolve an env var name for
`apiKey`, it sends the literal string; there is no built-in prime-inference
provider in 0.84.1; a ResourceLoader you pass in is never reloaded for you;
and the stock system prompt is a coding-assistant prompt that must be
replaced — but replacing it also silently removes the tool list, because the
harness only renders that section when it owns the prompt. AGENTS.md records
all four.

The expensive one was thinking level. The harness defaults to `medium`, and
nemotron spent an entire 4,096-token budget reasoning and returned an empty
answer. `low` was worse; `off` omits the parameter so the endpoint's default
wins. An explicit `reasoning_effort: none` via `thinkingLevelMap` took a turn
from 6,195 output tokens to 149.

And a turn is now bounded. The harness loop is `while (true)` with no
iteration cap; a runaway on a frontier model would have eaten the credit it
is supposed to report on. Ceilings on model calls and tokens, enforced both
through the harness hook and independently from the event stream, plus a
per-user daily spend limit — and the ledger now records spend on turns that
fail, which it previously discarded.

Signing in lands on /piggy, which is a workspace: conversations down one
side, the agent in the middle, what it did and what it cost beside it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
claude
2026-08-14 05:26:28 -07:00
parent 99d165b5e5
commit f0173440e4
77 changed files with 28108 additions and 1672 deletions
+288 -11
View File
@@ -35,16 +35,19 @@ packages/db Drizzle schema (47 tables), migrations, seeds
packages/prime Typed client for the Prime Intellect compute API
apps/api Hono HTTP API, auth, capacity/contract/calendar services
apps/web React + Vite + Tailwind + shadcn-idiom components
apps/piggy The agent — lease-based queue worker + private chat server
apps/piggy The agent — a Prime Agent session over the CRM tools behind a
private chat server, plus a lease-based queue worker (§6)
apps/mcp MCP server (stdio) — 9 tools
apps/cli `pig`, the HTTP surface for scripts and agent kernels
docs/ ontology.md, build-plan.md, agents.md, seed-data.md
docs/ ontology.md, build-plan.md, agents.md, seed-data.md,
screenshots.md, learn-scripts.md
deploy/ README.md (deployment), Caddyfile example, autodeploy units
```
~45,000 lines including tests. 261 tests across five packages
(core 62, prime 24, api 157, piggy 13, cli 5), plus a critical-path E2E suite
under `apps/api/e2e`. Node 22+.
~45,000 lines including tests. 403 unit tests across five packages
(core 62, prime 24, api 217, piggy 95, cli 5), plus E2E suites under
`apps/api/e2e` and `apps/piggy/e2e` that need a database — and, for one Piggy
case, a key. Node 22+.
| | |
|---|---|
@@ -224,9 +227,10 @@ anything added after them needs it too.
**Prime Intellect has two API hosts.** `api.primeintellect.ai` is compute and
pods. Inference is `api.pinference.ai/api/v1`, OpenAI-compatible.
**Piggy's default model thinks aloud.** `nvidia/nemotron-3-nano-30b-a3b` is a
hybrid reasoning model; under a tight `max_tokens` it rambles and truncates.
Pass `reasoning_effort: "none"` for tool use, routing and extraction.
**Piggy's default model thinks aloud, and the harness makes it worse.** The
agent SDK defaults `thinkingLevel` to `medium`; on `nvidia/nemotron-3-nano-30b-a3b`
that produced 6,195 output tokens of reasoning and an *empty* answer. The fix is
two halves and both are needed — see [§6](#6-piggy-and-the-harness-it-runs-on).
**A route file with green tests can still be unmounted.** Every route module is
a factory returning a `Hono` app, and `createApp` has to call it. The tests
@@ -243,7 +247,275 @@ block on that host needs `bind 10.0.0.2`, and that the CI runner uses
---
## 6. Conventions
## 6. Piggy, and the harness it runs on
Everything below was learned by running the thing. The product-level account is
in the README under *The agent surface*; this section is the engineering one,
and it exists because much of what follows either contradicts the SDK's own
documentation or is invisible in TypeScript.
### 6.1 The shape
`apps/piggy` embeds **Prime Agent** — Prime Intellect's harness,
`@earendil-works/pi-coding-agent@0.84.1`, MIT — as a Node library. Nothing is
shelled out to, and there is no second process.
```
src/agent/session.ts Builds a turn: runtime, credential, model, prompt,
tools, and the assertions that make the tool set a
fact rather than a hope
src/agent/models.json The provider document the harness reads: five models,
their prices, their context windows, their reasoning
maps. Copied verbatim into PIGGY_AGENT_DIR when the
runtime is first built (once per process)
src/agent/models.ts Validates that file and turns it into the picker's
catalogue. One source for price and size
src/agent/prompt.ts Piggy's system prompt, including the tool list the
harness stops writing (§6.4)
src/agent/tool-bridge.ts PIG's zod `AgentTool`s → harness `ToolDefinition`s
src/chat-tools.ts Read tools ─┐
src/page-tools.ts Page summaries ├─ the product; the harness swap did
src/lifecycle-tools.ts Lifecycle ─┘ not touch a line of them
src/write-tools.ts The five write tools and the approval flow
src/chat-server.ts The NDJSON server, the approval rendezvous, the ledger
src/provider.ts + worker.ts + queue.ts The queue worker, which does NOT use
the harness at all — it still speaks
OpenAI-completions directly
```
The queue worker and the chat agent are different code paths that happen to
share a process. `PIGGY_MODEL` and `PIGGY_INFERENCE_BASE` belong to the worker;
`PIGGY_AGENT_*` and `models.json` belong to the agent. Changing one does not
change the other, which has already confused one person into "fixing" the model
in the wrong place.
### 6.2 No shell, and why the flag is not enough
The session is constructed with `noTools: 'all'` **plus** an explicit `tools`
allowlist (the `createAgentSession` call in `agent/session.ts`). Neither alone
would do:
`noTools: 'all'` removes the built-ins, and the allowlist is the positive
statement of what may exist. But both are *the harness's* configuration, and the
harness composes its tool set from several sources — built-ins, extensions,
skills, custom tools — so a future release that changes the precedence between
them would widen the set without changing a line of PIG. Three gates exist for
that reason:
1. `assertPigToolBoundary` (`src/chat.ts`) — a name must start `pig_` and must
not read like a shell. PIG's own code, PIG's own rule.
2. `assertUniqueToolNames` (`agent/session.ts`) — the harness keeps its tools in
a `Map` keyed by name and *sets* each one in turn
(`dist/core/agent-session.js:1963-1968`), so a duplicate silently overwrites
the other. That is how a read tool ends up answering for a write tool of the
same name, with nothing anywhere saying so.
3. `assertExactToolSet` (`agent/session.ts`) — compares the live
`session.agent.state.tools` against exactly what was handed in and throws at
session construction if they differ. This is the one that would notice a
harness upgrade.
`test/agent-session.test.ts` pins all three, including `pig_bash` and friends.
Extensions, skills, prompt templates, themes and context-file discovery are all
disabled on the `DefaultResourceLoader`, and `PIGGY_AGENT_DIR` is deliberately
not a checkout: the harness reads context files from its cwd, and the cwd is
also appended to the live system prompt verbatim as
`Current working directory: …`.
### 6.3 Four places the SDK's own docs are wrong
Each of these compiles, starts, and fails somewhere else.
**`apiKey` in `models.json` is not an environment variable name.** Writing
`"apiKey": "PRIME_API_KEY"` sends the literal string `PRIME_API_KEY` as the
bearer token, and the endpoint answers 401. The value is a *template*:
`$PRIME_API_KEY` or `${PRIME_API_KEY}` interpolate, a leading `!` executes the
rest as a shell command, and anything else is a literal
(`dist/core/resolve-config-value.js:116-128`). PIG uses none of those forms —
it calls
`modelRuntime.setRuntimeApiKey(PIGGY_PROVIDER_ID, config.PRIME_API_KEY)`
(`agent/session.ts`), which is the only line that authenticates Piggy and keeps
the key out of the file that gets written to disk.
**There is no built-in `prime-inference` provider in 0.84.1.** The published
docs describe a build that is not on npm; `KnownProvider` in
`@earendil-works/pi-ai/dist/types.d.ts:19` lists forty providers and none of
them is Prime Intellect's inference host. PIG registers one itself from
`models.json`, and the id `prime-inference` has to match in three places — the
JSON key, `setRuntimeApiKey`, and `modelRuntime.getModel`. A typo in any of them
surfaces as a 401 or an undefined model, never as "unknown provider".
**A `ResourceLoader` you pass in is never reloaded for you.**
`createAgentSession` constructs and reloads one *only when you do not supply
one* (`dist/core/sdk.js:75-78`). Pass your own and forget `await loader.reload()`
and the session runs on the stock coding-assistant preamble — no error, no
warning, and an agent that offers to read your files.
**The stock prompt is a coding-assistant prompt and must be replaced, not
appended to.** It opens "You are an expert coding assistant operating inside pi"
and cites the SDK's own README paths (`dist/core/system-prompt.js:73`).
Appending does not help: a CRM agent told it edits code reaches for tools it
does not have and apologises for not having them. The replacement goes through
the loader's `systemPromptOverride`, which takes the literal text — the
`systemPrompt` option is a *file source*, and handing it a prompt loads nothing
and says nothing.
### 6.4 Replacing the prompt silently removes the tool list
`buildSystemPrompt` returns early on the `customPrompt` branch
(`dist/core/system-prompt.js:13-33`); the "Available tools" section is only ever
built further down, on the branch where no custom prompt was supplied
(`:40`, `:75`). So the moment the preamble is replaced — which is not optional
here — every tool becomes invisible to the model, `promptSnippet` or not.
`agent/prompt.ts` therefore renders the list itself, in `toolSection`. A 30B
model that cannot see a tool in its prompt answers from the page title instead
of calling it, and that failure is completely silent: the tool is registered,
callable, and never called. If you add a tool, give it a `promptSnippet`, and
check it appears in `session.systemPrompt`.
### 6.5 The thinking-level trap
The one that cost real money.
The harness defaults `thinkingLevel` to `medium`. On the default model that
produced **6,195 output tokens of reasoning and an empty answer**, stopping at
`finish_reason: length` — the budget was gone before a word of the reply was
written, and reasoning bills as output. `low` was worse. After the fix the same
question answered correctly in **149 output tokens**.
The fix is two halves and either alone is silent:
- `PIGGY_AGENT_THINKING` defaults to `off` (`src/config.ts`), and
- the model entry carries a `thinkingLevelMap` mapping `off``"none"`
(`src/agent/models.json`).
Why the second is needed: a thinking level of `off` becomes
`reasoningEffort: undefined` in the provider
(`@earendil-works/pi-ai/dist/api/openai-completions.js:473-474`), and the
request builder then emits `reasoning_effort` **only if the model has a map**:
```js
else if (options?.reasoningEffort && model.reasoning && compat.supportsReasoningEffort) {
params.reasoning_effort = model.thinkingLevelMap?.[options.reasoningEffort] ?? options.reasoningEffort;
}
else if (!options?.reasoningEffort && model.reasoning && compat.supportsReasoningEffort) {
const offValue = model.thinkingLevelMap?.off;
if (typeof offValue === "string") { params.reasoning_effort = offValue; }
}
dist/api/openai-completions.js:657-666
```
Without the map, `off` sends **no reasoning parameter at all** and the
endpoint's own default — thinking on, verbosely — wins. This is per model. The
two nemotron entries have a map; deepseek, opus and gpt-5.6 do not, and were
left to their own defaults deliberately. **If you change `PIGGY_AGENT_MODEL` and
answers start coming back empty or truncated, this is why.**
`test/agent-thinking.test.ts` fails if the default model has no map, and
`e2e/prime-agent.test.ts` counts the tokens against the live endpoint.
### 6.6 Modes, and the one function that decides
`PiggyMode` is `read_only` | `confirm` | `auto`.
- **`read_only`** offers no write tool at all. Not offered-and-refused: absent
(`createPigWriteTools` returns `[]`). A model that can see a capability
narrates using it.
- **`confirm`** — the shipped default — turns every write into a proposal. The
tool emits an `approval_required` card, the turn stays open, the decision
arrives on a separate `POST /internal/approve`, and only then does the
mutation run.
- **`auto`** writes immediately, as the calling user, under their permissions.
Contracts, commitments, allocations and compliance require a human in **every**
mode. That rule is one function — `requiresApproval` in
`packages/core/src/piggy-protocol.ts` — and it is the single source of truth:
the write tools read it, the tests assert against it, and nothing restates it.
If you add a guarded kind, add it to `PIGGY_ALWAYS_CONFIRM_KINDS` and everything
downstream follows.
Two properties of the write path are not negotiable. Every write goes through
`executeMutation` with the caller's own `Principal`, so Piggy holds no privilege
of its own — there is no elevated principal anywhere in `write-tools.ts` and
there must never be one. And a refusal is an *answer*: a missing capability, a
declined card and a rejected input all come back as ordinary tool results whose
first line says `NOT SAVED`. Thrown into the stream they would end the turn on
the user's own permissions, which reads to them as Piggy being broken.
The rendezvous itself (`ApprovalRegistry` in `src/chat-server.ts`) is single-use
— an id is deleted the instant it settles, so a replayed decision cannot apply a
change twice — deadlined at five minutes, and turn-owned: an abandoned turn
rejects every approval it opened, because a pending promise there holds a billed
inference connection open.
### 6.7 The browser never learns which harness this is
Prime Agent emits twenty-three event types. PIG's own protocol
(`PiggyChatEvent` in `packages/core/src/piggy-protocol.ts`) has nine, and
`translateSessionEvent` in `src/chat-server.ts` maps exactly four of the
harness's — `message_update`, `tool_execution_start`, `tool_execution_end`,
`turn_end` — and drops the rest on the server. That is deliberate: a harness
upgrade is then a server change and never a client one.
The risk in a `default: return` is the upgrade that *adds* an event — a
delegated sub-agent, a permission request — which would be dropped in silence
for as long as it took somebody to notice a missing feature.
`test/chat-server.test.ts` therefore writes out both lists and asserts, at
compile time, that they are mutually assignable with `AgentSessionEvent['type']`.
Bump the SDK and `tsc` tells you what is new before anything runs.
### 6.8 Working on Piggy without spending credit
Almost all of it is free, and only one path is not.
- **The unit suite never makes a request.** `createPiggySession` resolves the
model, builds the prompt and registers the tools entirely offline with a fake
key, so the tool set, the prompt, the thinking level and the model's own
ceiling are all inspectable without inference. That is what
`test/agent-session.test.ts` and `test/agent-thinking.test.ts` do.
- **The whole chat protocol is drivable with no model at all.**
`startPiggyChatServer` takes `createSession`, `createReadTools` and
`createWriteTools` as options; the tests hand it a fake harness that emits
real `AgentSessionEvent`s. `e2e/approval-rendezvous.test.ts` does this against
a real database, which is how the approval flow is tested end to end for free.
- **`src/dev/mock-inference.ts`** (`pnpm -F @pig/piggy run dev:mock`, port 8945)
speaks the OpenAI-compatible wire protocol with steering directives —
`/mock error`, `/mock ratelimit`, `/mock cut`, `/mock badtool`. Note what it
serves: the **queue worker**, through `PIGGY_INFERENCE_BASE`. The agent reads
its base URL from `models.json`, so pointing the chat path at the mock means
editing that file.
- **`src/dev/verify-prime-agent.ts`**
(`pnpm -F @pig/piggy exec tsx src/dev/verify-prime-agent.ts [modelId]`) is the
live probe: it asks the real endpoint one question with a seeded tool and
prints the model, the tool set, whether anything shell-shaped survived, the
first 200 characters of the system prompt and the answer. It spends a few
hundred tokens. Nothing in CI runs it.
- **The database.** Anything that writes runs against a scratch database, never
the development book — an activity appearing in somebody's feed because a test
ran is exactly what a CRM must not do. `e2e/write-tools.test.ts` and
`e2e/approval-rendezvous.test.ts` take `PIGGY_WRITE_DATABASE_URL` and refuse
`pig_combined` by name.
- **The one paid test** is `e2e/prime-agent.test.ts`, gated on
`PIGGY_E2E_LIVE=1` *and* a key, because a suite that spends money whenever the
environment happens to be loaded spends money by accident. One turn is about
$0.0003.
```bash
# Unit suite: no database, no key, no network.
pnpm -F @pig/piggy run typecheck && pnpm -F @pig/piggy run test
# E2E: a scratch database of its own. `pig_combined` is refused by name.
docker exec pig-ux-db psql -U pig -d postgres -c "CREATE DATABASE pig_scratch"
DATABASE_URL=postgres://pig:pig@localhost:54330/pig_scratch pnpm -F @pig/db run migrate
DATABASE_URL=postgres://pig:pig@localhost:54330/pig_scratch \
PIGGY_WRITE_DATABASE_URL=postgres://pig:pig@localhost:54330/pig_scratch \
pnpm -F @pig/piggy run test:e2e # the live case skips, and says so
# Add the paid one deliberately, never by default.
PIGGY_E2E_LIVE=1 PRIME_API_KEY=... pnpm -F @pig/piggy run test:e2e
```
---
## 7. Conventions
**Comments explain *why*, never *what*.** The code says what it does. Comments
carry the reasoning that would otherwise be lost — why this treatment and not
@@ -272,7 +544,7 @@ real database. "It should work" has been wrong repeatedly.
---
## 7. Where to start
## 8. Where to start
**Every task in the original three-wave plan has shipped.**
[`docs/build-plan.md`](./docs/build-plan.md) is now an audited record of that
@@ -294,6 +566,11 @@ settled; read them before adding any write.
Piggy does far less than the ontology implies.
4. **Mount the HubSpot routes, or delete them.** Seven tables, OAuth, sync jobs
and webhook verification, all written, tested and unreachable.
5. **Give Piggy's writes their notification.** A stage change made through the
API raises a Slack notification; the same change made in chat does not,
because `write-tools.ts` passes no `NotificationOutbox` — it runs in the
Piggy process and the outbox is wired in the API server. The other open
Piggy items are listed under *Left to do* in the build plan.
`app.ts` is the one shared file. If your change needs a route mounted, a public
path allowlisted or a schema widened there, say so rather than racing another
@@ -301,7 +578,7 @@ agent for it.
---
## 8. What not to do
## 9. What not to do
- Do not copy component files out of other people's repositories. Where a
primitive is a shadcn/ui original, take it from upstream, where it is