Put Piggy on Prime Agent, and let it write to the book
Piggy was a hand-rolled OpenAI tool loop. It is now a Prime Agent session — Prime Intellect's own harness, embedded as a Node library — answering from PIG's tools and, for the first time, able to put information into the CRM rather than only read it out. The harness is a coding agent, so the first job was taking the coding agent away from it. `noTools: 'all'` plus an explicit allowlist leaves the model with PIG's ten `pig_*` tools and no bash, no filesystem, no IPython. That holds under attack: a hostile extension, a skill and a settings file planted in the agent's own directory, then `setActiveToolsByName` called with every built-in, still leaves ten tools, all ours. Both lines are load-bearing — `noTools` alone registers nothing, and the allowlist is what admits our own. Writing is gated rather than assumed. A change is proposed, not made: the tool returns a description, the transcript renders a diff card, and nothing reaches the database until someone presses Apply. Contracts, commitments, allocations and compliance always stop for a human whatever the mode. Every write runs through `executeMutation` as the calling user, so their capabilities and the audit trail apply exactly as they would to a human's. Four things about the SDK are wrong in its own documentation and cost a debugging cycle each: models.json does not resolve an env var name for `apiKey`, it sends the literal string; there is no built-in prime-inference provider in 0.84.1; a ResourceLoader you pass in is never reloaded for you; and the stock system prompt is a coding-assistant prompt that must be replaced — but replacing it also silently removes the tool list, because the harness only renders that section when it owns the prompt. AGENTS.md records all four. The expensive one was thinking level. The harness defaults to `medium`, and nemotron spent an entire 4,096-token budget reasoning and returned an empty answer. `low` was worse; `off` omits the parameter so the endpoint's default wins. An explicit `reasoning_effort: none` via `thinkingLevelMap` took a turn from 6,195 output tokens to 149. And a turn is now bounded. The harness loop is `while (true)` with no iteration cap; a runaway on a frontier model would have eaten the credit it is supposed to report on. Ceilings on model calls and tokens, enforced both through the harness hook and independently from the event stream, plus a per-user daily spend limit — and the ledger now records spend on turns that fail, which it previously discarded. Signing in lands on /piggy, which is a workspace: conversations down one side, the agent in the middle, what it did and what it cost beside it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -35,16 +35,19 @@ packages/db Drizzle schema (47 tables), migrations, seeds
|
||||
packages/prime Typed client for the Prime Intellect compute API
|
||||
apps/api Hono HTTP API, auth, capacity/contract/calendar services
|
||||
apps/web React + Vite + Tailwind + shadcn-idiom components
|
||||
apps/piggy The agent — lease-based queue worker + private chat server
|
||||
apps/piggy The agent — a Prime Agent session over the CRM tools behind a
|
||||
private chat server, plus a lease-based queue worker (§6)
|
||||
apps/mcp MCP server (stdio) — 9 tools
|
||||
apps/cli `pig`, the HTTP surface for scripts and agent kernels
|
||||
docs/ ontology.md, build-plan.md, agents.md, seed-data.md
|
||||
docs/ ontology.md, build-plan.md, agents.md, seed-data.md,
|
||||
screenshots.md, learn-scripts.md
|
||||
deploy/ README.md (deployment), Caddyfile example, autodeploy units
|
||||
```
|
||||
|
||||
~45,000 lines including tests. 261 tests across five packages
|
||||
(core 62, prime 24, api 157, piggy 13, cli 5), plus a critical-path E2E suite
|
||||
under `apps/api/e2e`. Node 22+.
|
||||
~45,000 lines including tests. 403 unit tests across five packages
|
||||
(core 62, prime 24, api 217, piggy 95, cli 5), plus E2E suites under
|
||||
`apps/api/e2e` and `apps/piggy/e2e` that need a database — and, for one Piggy
|
||||
case, a key. Node 22+.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
@@ -224,9 +227,10 @@ anything added after them needs it too.
|
||||
**Prime Intellect has two API hosts.** `api.primeintellect.ai` is compute and
|
||||
pods. Inference is `api.pinference.ai/api/v1`, OpenAI-compatible.
|
||||
|
||||
**Piggy's default model thinks aloud.** `nvidia/nemotron-3-nano-30b-a3b` is a
|
||||
hybrid reasoning model; under a tight `max_tokens` it rambles and truncates.
|
||||
Pass `reasoning_effort: "none"` for tool use, routing and extraction.
|
||||
**Piggy's default model thinks aloud, and the harness makes it worse.** The
|
||||
agent SDK defaults `thinkingLevel` to `medium`; on `nvidia/nemotron-3-nano-30b-a3b`
|
||||
that produced 6,195 output tokens of reasoning and an *empty* answer. The fix is
|
||||
two halves and both are needed — see [§6](#6-piggy-and-the-harness-it-runs-on).
|
||||
|
||||
**A route file with green tests can still be unmounted.** Every route module is
|
||||
a factory returning a `Hono` app, and `createApp` has to call it. The tests
|
||||
@@ -243,7 +247,275 @@ block on that host needs `bind 10.0.0.2`, and that the CI runner uses
|
||||
|
||||
---
|
||||
|
||||
## 6. Conventions
|
||||
## 6. Piggy, and the harness it runs on
|
||||
|
||||
Everything below was learned by running the thing. The product-level account is
|
||||
in the README under *The agent surface*; this section is the engineering one,
|
||||
and it exists because much of what follows either contradicts the SDK's own
|
||||
documentation or is invisible in TypeScript.
|
||||
|
||||
### 6.1 The shape
|
||||
|
||||
`apps/piggy` embeds **Prime Agent** — Prime Intellect's harness,
|
||||
`@earendil-works/pi-coding-agent@0.84.1`, MIT — as a Node library. Nothing is
|
||||
shelled out to, and there is no second process.
|
||||
|
||||
```
|
||||
src/agent/session.ts Builds a turn: runtime, credential, model, prompt,
|
||||
tools, and the assertions that make the tool set a
|
||||
fact rather than a hope
|
||||
src/agent/models.json The provider document the harness reads: five models,
|
||||
their prices, their context windows, their reasoning
|
||||
maps. Copied verbatim into PIGGY_AGENT_DIR when the
|
||||
runtime is first built (once per process)
|
||||
src/agent/models.ts Validates that file and turns it into the picker's
|
||||
catalogue. One source for price and size
|
||||
src/agent/prompt.ts Piggy's system prompt, including the tool list the
|
||||
harness stops writing (§6.4)
|
||||
src/agent/tool-bridge.ts PIG's zod `AgentTool`s → harness `ToolDefinition`s
|
||||
src/chat-tools.ts Read tools ─┐
|
||||
src/page-tools.ts Page summaries ├─ the product; the harness swap did
|
||||
src/lifecycle-tools.ts Lifecycle ─┘ not touch a line of them
|
||||
src/write-tools.ts The five write tools and the approval flow
|
||||
src/chat-server.ts The NDJSON server, the approval rendezvous, the ledger
|
||||
src/provider.ts + worker.ts + queue.ts The queue worker, which does NOT use
|
||||
the harness at all — it still speaks
|
||||
OpenAI-completions directly
|
||||
```
|
||||
|
||||
The queue worker and the chat agent are different code paths that happen to
|
||||
share a process. `PIGGY_MODEL` and `PIGGY_INFERENCE_BASE` belong to the worker;
|
||||
`PIGGY_AGENT_*` and `models.json` belong to the agent. Changing one does not
|
||||
change the other, which has already confused one person into "fixing" the model
|
||||
in the wrong place.
|
||||
|
||||
### 6.2 No shell, and why the flag is not enough
|
||||
|
||||
The session is constructed with `noTools: 'all'` **plus** an explicit `tools`
|
||||
allowlist (the `createAgentSession` call in `agent/session.ts`). Neither alone
|
||||
would do:
|
||||
`noTools: 'all'` removes the built-ins, and the allowlist is the positive
|
||||
statement of what may exist. But both are *the harness's* configuration, and the
|
||||
harness composes its tool set from several sources — built-ins, extensions,
|
||||
skills, custom tools — so a future release that changes the precedence between
|
||||
them would widen the set without changing a line of PIG. Three gates exist for
|
||||
that reason:
|
||||
|
||||
1. `assertPigToolBoundary` (`src/chat.ts`) — a name must start `pig_` and must
|
||||
not read like a shell. PIG's own code, PIG's own rule.
|
||||
2. `assertUniqueToolNames` (`agent/session.ts`) — the harness keeps its tools in
|
||||
a `Map` keyed by name and *sets* each one in turn
|
||||
(`dist/core/agent-session.js:1963-1968`), so a duplicate silently overwrites
|
||||
the other. That is how a read tool ends up answering for a write tool of the
|
||||
same name, with nothing anywhere saying so.
|
||||
3. `assertExactToolSet` (`agent/session.ts`) — compares the live
|
||||
`session.agent.state.tools` against exactly what was handed in and throws at
|
||||
session construction if they differ. This is the one that would notice a
|
||||
harness upgrade.
|
||||
|
||||
`test/agent-session.test.ts` pins all three, including `pig_bash` and friends.
|
||||
Extensions, skills, prompt templates, themes and context-file discovery are all
|
||||
disabled on the `DefaultResourceLoader`, and `PIGGY_AGENT_DIR` is deliberately
|
||||
not a checkout: the harness reads context files from its cwd, and the cwd is
|
||||
also appended to the live system prompt verbatim as
|
||||
`Current working directory: …`.
|
||||
|
||||
### 6.3 Four places the SDK's own docs are wrong
|
||||
|
||||
Each of these compiles, starts, and fails somewhere else.
|
||||
|
||||
**`apiKey` in `models.json` is not an environment variable name.** Writing
|
||||
`"apiKey": "PRIME_API_KEY"` sends the literal string `PRIME_API_KEY` as the
|
||||
bearer token, and the endpoint answers 401. The value is a *template*:
|
||||
`$PRIME_API_KEY` or `${PRIME_API_KEY}` interpolate, a leading `!` executes the
|
||||
rest as a shell command, and anything else is a literal
|
||||
(`dist/core/resolve-config-value.js:116-128`). PIG uses none of those forms —
|
||||
it calls
|
||||
`modelRuntime.setRuntimeApiKey(PIGGY_PROVIDER_ID, config.PRIME_API_KEY)`
|
||||
(`agent/session.ts`), which is the only line that authenticates Piggy and keeps
|
||||
the key out of the file that gets written to disk.
|
||||
|
||||
**There is no built-in `prime-inference` provider in 0.84.1.** The published
|
||||
docs describe a build that is not on npm; `KnownProvider` in
|
||||
`@earendil-works/pi-ai/dist/types.d.ts:19` lists forty providers and none of
|
||||
them is Prime Intellect's inference host. PIG registers one itself from
|
||||
`models.json`, and the id `prime-inference` has to match in three places — the
|
||||
JSON key, `setRuntimeApiKey`, and `modelRuntime.getModel`. A typo in any of them
|
||||
surfaces as a 401 or an undefined model, never as "unknown provider".
|
||||
|
||||
**A `ResourceLoader` you pass in is never reloaded for you.**
|
||||
`createAgentSession` constructs and reloads one *only when you do not supply
|
||||
one* (`dist/core/sdk.js:75-78`). Pass your own and forget `await loader.reload()`
|
||||
and the session runs on the stock coding-assistant preamble — no error, no
|
||||
warning, and an agent that offers to read your files.
|
||||
|
||||
**The stock prompt is a coding-assistant prompt and must be replaced, not
|
||||
appended to.** It opens "You are an expert coding assistant operating inside pi"
|
||||
and cites the SDK's own README paths (`dist/core/system-prompt.js:73`).
|
||||
Appending does not help: a CRM agent told it edits code reaches for tools it
|
||||
does not have and apologises for not having them. The replacement goes through
|
||||
the loader's `systemPromptOverride`, which takes the literal text — the
|
||||
`systemPrompt` option is a *file source*, and handing it a prompt loads nothing
|
||||
and says nothing.
|
||||
|
||||
### 6.4 Replacing the prompt silently removes the tool list
|
||||
|
||||
`buildSystemPrompt` returns early on the `customPrompt` branch
|
||||
(`dist/core/system-prompt.js:13-33`); the "Available tools" section is only ever
|
||||
built further down, on the branch where no custom prompt was supplied
|
||||
(`:40`, `:75`). So the moment the preamble is replaced — which is not optional
|
||||
here — every tool becomes invisible to the model, `promptSnippet` or not.
|
||||
|
||||
`agent/prompt.ts` therefore renders the list itself, in `toolSection`. A 30B
|
||||
model that cannot see a tool in its prompt answers from the page title instead
|
||||
of calling it, and that failure is completely silent: the tool is registered,
|
||||
callable, and never called. If you add a tool, give it a `promptSnippet`, and
|
||||
check it appears in `session.systemPrompt`.
|
||||
|
||||
### 6.5 The thinking-level trap
|
||||
|
||||
The one that cost real money.
|
||||
|
||||
The harness defaults `thinkingLevel` to `medium`. On the default model that
|
||||
produced **6,195 output tokens of reasoning and an empty answer**, stopping at
|
||||
`finish_reason: length` — the budget was gone before a word of the reply was
|
||||
written, and reasoning bills as output. `low` was worse. After the fix the same
|
||||
question answered correctly in **149 output tokens**.
|
||||
|
||||
The fix is two halves and either alone is silent:
|
||||
|
||||
- `PIGGY_AGENT_THINKING` defaults to `off` (`src/config.ts`), and
|
||||
- the model entry carries a `thinkingLevelMap` mapping `off` → `"none"`
|
||||
(`src/agent/models.json`).
|
||||
|
||||
Why the second is needed: a thinking level of `off` becomes
|
||||
`reasoningEffort: undefined` in the provider
|
||||
(`@earendil-works/pi-ai/dist/api/openai-completions.js:473-474`), and the
|
||||
request builder then emits `reasoning_effort` **only if the model has a map**:
|
||||
|
||||
```js
|
||||
else if (options?.reasoningEffort && model.reasoning && compat.supportsReasoningEffort) {
|
||||
params.reasoning_effort = model.thinkingLevelMap?.[options.reasoningEffort] ?? options.reasoningEffort;
|
||||
}
|
||||
else if (!options?.reasoningEffort && model.reasoning && compat.supportsReasoningEffort) {
|
||||
const offValue = model.thinkingLevelMap?.off;
|
||||
if (typeof offValue === "string") { params.reasoning_effort = offValue; }
|
||||
}
|
||||
— dist/api/openai-completions.js:657-666
|
||||
```
|
||||
|
||||
Without the map, `off` sends **no reasoning parameter at all** and the
|
||||
endpoint's own default — thinking on, verbosely — wins. This is per model. The
|
||||
two nemotron entries have a map; deepseek, opus and gpt-5.6 do not, and were
|
||||
left to their own defaults deliberately. **If you change `PIGGY_AGENT_MODEL` and
|
||||
answers start coming back empty or truncated, this is why.**
|
||||
`test/agent-thinking.test.ts` fails if the default model has no map, and
|
||||
`e2e/prime-agent.test.ts` counts the tokens against the live endpoint.
|
||||
|
||||
### 6.6 Modes, and the one function that decides
|
||||
|
||||
`PiggyMode` is `read_only` | `confirm` | `auto`.
|
||||
|
||||
- **`read_only`** offers no write tool at all. Not offered-and-refused: absent
|
||||
(`createPigWriteTools` returns `[]`). A model that can see a capability
|
||||
narrates using it.
|
||||
- **`confirm`** — the shipped default — turns every write into a proposal. The
|
||||
tool emits an `approval_required` card, the turn stays open, the decision
|
||||
arrives on a separate `POST /internal/approve`, and only then does the
|
||||
mutation run.
|
||||
- **`auto`** writes immediately, as the calling user, under their permissions.
|
||||
|
||||
Contracts, commitments, allocations and compliance require a human in **every**
|
||||
mode. That rule is one function — `requiresApproval` in
|
||||
`packages/core/src/piggy-protocol.ts` — and it is the single source of truth:
|
||||
the write tools read it, the tests assert against it, and nothing restates it.
|
||||
If you add a guarded kind, add it to `PIGGY_ALWAYS_CONFIRM_KINDS` and everything
|
||||
downstream follows.
|
||||
|
||||
Two properties of the write path are not negotiable. Every write goes through
|
||||
`executeMutation` with the caller's own `Principal`, so Piggy holds no privilege
|
||||
of its own — there is no elevated principal anywhere in `write-tools.ts` and
|
||||
there must never be one. And a refusal is an *answer*: a missing capability, a
|
||||
declined card and a rejected input all come back as ordinary tool results whose
|
||||
first line says `NOT SAVED`. Thrown into the stream they would end the turn on
|
||||
the user's own permissions, which reads to them as Piggy being broken.
|
||||
|
||||
The rendezvous itself (`ApprovalRegistry` in `src/chat-server.ts`) is single-use
|
||||
— an id is deleted the instant it settles, so a replayed decision cannot apply a
|
||||
change twice — deadlined at five minutes, and turn-owned: an abandoned turn
|
||||
rejects every approval it opened, because a pending promise there holds a billed
|
||||
inference connection open.
|
||||
|
||||
### 6.7 The browser never learns which harness this is
|
||||
|
||||
Prime Agent emits twenty-three event types. PIG's own protocol
|
||||
(`PiggyChatEvent` in `packages/core/src/piggy-protocol.ts`) has nine, and
|
||||
`translateSessionEvent` in `src/chat-server.ts` maps exactly four of the
|
||||
harness's — `message_update`, `tool_execution_start`, `tool_execution_end`,
|
||||
`turn_end` — and drops the rest on the server. That is deliberate: a harness
|
||||
upgrade is then a server change and never a client one.
|
||||
|
||||
The risk in a `default: return` is the upgrade that *adds* an event — a
|
||||
delegated sub-agent, a permission request — which would be dropped in silence
|
||||
for as long as it took somebody to notice a missing feature.
|
||||
`test/chat-server.test.ts` therefore writes out both lists and asserts, at
|
||||
compile time, that they are mutually assignable with `AgentSessionEvent['type']`.
|
||||
Bump the SDK and `tsc` tells you what is new before anything runs.
|
||||
|
||||
### 6.8 Working on Piggy without spending credit
|
||||
|
||||
Almost all of it is free, and only one path is not.
|
||||
|
||||
- **The unit suite never makes a request.** `createPiggySession` resolves the
|
||||
model, builds the prompt and registers the tools entirely offline with a fake
|
||||
key, so the tool set, the prompt, the thinking level and the model's own
|
||||
ceiling are all inspectable without inference. That is what
|
||||
`test/agent-session.test.ts` and `test/agent-thinking.test.ts` do.
|
||||
- **The whole chat protocol is drivable with no model at all.**
|
||||
`startPiggyChatServer` takes `createSession`, `createReadTools` and
|
||||
`createWriteTools` as options; the tests hand it a fake harness that emits
|
||||
real `AgentSessionEvent`s. `e2e/approval-rendezvous.test.ts` does this against
|
||||
a real database, which is how the approval flow is tested end to end for free.
|
||||
- **`src/dev/mock-inference.ts`** (`pnpm -F @pig/piggy run dev:mock`, port 8945)
|
||||
speaks the OpenAI-compatible wire protocol with steering directives —
|
||||
`/mock error`, `/mock ratelimit`, `/mock cut`, `/mock badtool`. Note what it
|
||||
serves: the **queue worker**, through `PIGGY_INFERENCE_BASE`. The agent reads
|
||||
its base URL from `models.json`, so pointing the chat path at the mock means
|
||||
editing that file.
|
||||
- **`src/dev/verify-prime-agent.ts`**
|
||||
(`pnpm -F @pig/piggy exec tsx src/dev/verify-prime-agent.ts [modelId]`) is the
|
||||
live probe: it asks the real endpoint one question with a seeded tool and
|
||||
prints the model, the tool set, whether anything shell-shaped survived, the
|
||||
first 200 characters of the system prompt and the answer. It spends a few
|
||||
hundred tokens. Nothing in CI runs it.
|
||||
- **The database.** Anything that writes runs against a scratch database, never
|
||||
the development book — an activity appearing in somebody's feed because a test
|
||||
ran is exactly what a CRM must not do. `e2e/write-tools.test.ts` and
|
||||
`e2e/approval-rendezvous.test.ts` take `PIGGY_WRITE_DATABASE_URL` and refuse
|
||||
`pig_combined` by name.
|
||||
- **The one paid test** is `e2e/prime-agent.test.ts`, gated on
|
||||
`PIGGY_E2E_LIVE=1` *and* a key, because a suite that spends money whenever the
|
||||
environment happens to be loaded spends money by accident. One turn is about
|
||||
$0.0003.
|
||||
|
||||
```bash
|
||||
# Unit suite: no database, no key, no network.
|
||||
pnpm -F @pig/piggy run typecheck && pnpm -F @pig/piggy run test
|
||||
|
||||
# E2E: a scratch database of its own. `pig_combined` is refused by name.
|
||||
docker exec pig-ux-db psql -U pig -d postgres -c "CREATE DATABASE pig_scratch"
|
||||
DATABASE_URL=postgres://pig:pig@localhost:54330/pig_scratch pnpm -F @pig/db run migrate
|
||||
DATABASE_URL=postgres://pig:pig@localhost:54330/pig_scratch \
|
||||
PIGGY_WRITE_DATABASE_URL=postgres://pig:pig@localhost:54330/pig_scratch \
|
||||
pnpm -F @pig/piggy run test:e2e # the live case skips, and says so
|
||||
|
||||
# Add the paid one deliberately, never by default.
|
||||
PIGGY_E2E_LIVE=1 PRIME_API_KEY=... pnpm -F @pig/piggy run test:e2e
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. Conventions
|
||||
|
||||
**Comments explain *why*, never *what*.** The code says what it does. Comments
|
||||
carry the reasoning that would otherwise be lost — why this treatment and not
|
||||
@@ -272,7 +544,7 @@ real database. "It should work" has been wrong repeatedly.
|
||||
|
||||
---
|
||||
|
||||
## 7. Where to start
|
||||
## 8. Where to start
|
||||
|
||||
**Every task in the original three-wave plan has shipped.**
|
||||
[`docs/build-plan.md`](./docs/build-plan.md) is now an audited record of that
|
||||
@@ -294,6 +566,11 @@ settled; read them before adding any write.
|
||||
Piggy does far less than the ontology implies.
|
||||
4. **Mount the HubSpot routes, or delete them.** Seven tables, OAuth, sync jobs
|
||||
and webhook verification, all written, tested and unreachable.
|
||||
5. **Give Piggy's writes their notification.** A stage change made through the
|
||||
API raises a Slack notification; the same change made in chat does not,
|
||||
because `write-tools.ts` passes no `NotificationOutbox` — it runs in the
|
||||
Piggy process and the outbox is wired in the API server. The other open
|
||||
Piggy items are listed under *Left to do* in the build plan.
|
||||
|
||||
`app.ts` is the one shared file. If your change needs a route mounted, a public
|
||||
path allowlisted or a schema widened there, say so rather than racing another
|
||||
@@ -301,7 +578,7 @@ agent for it.
|
||||
|
||||
---
|
||||
|
||||
## 8. What not to do
|
||||
## 9. What not to do
|
||||
|
||||
- Do not copy component files out of other people's repositories. Where a
|
||||
primitive is a shadcn/ui original, take it from upstream, where it is
|
||||
|
||||
Reference in New Issue
Block a user