Put Piggy on Prime Agent, and let it write to the book
CI / verify (push) Successful in 7m6s
CI / publish (push) Has been skipped

Piggy was a hand-rolled OpenAI tool loop. It is now a Prime Agent session —
Prime Intellect's own harness, embedded as a Node library — answering from
PIG's tools and, for the first time, able to put information into the CRM
rather than only read it out.

The harness is a coding agent, so the first job was taking the coding agent
away from it. `noTools: 'all'` plus an explicit allowlist leaves the model
with PIG's ten `pig_*` tools and no bash, no filesystem, no IPython. That
holds under attack: a hostile extension, a skill and a settings file planted
in the agent's own directory, then `setActiveToolsByName` called with every
built-in, still leaves ten tools, all ours. Both lines are load-bearing —
`noTools` alone registers nothing, and the allowlist is what admits our own.

Writing is gated rather than assumed. A change is proposed, not made: the
tool returns a description, the transcript renders a diff card, and nothing
reaches the database until someone presses Apply. Contracts, commitments,
allocations and compliance always stop for a human whatever the mode. Every
write runs through `executeMutation` as the calling user, so their
capabilities and the audit trail apply exactly as they would to a human's.

Four things about the SDK are wrong in its own documentation and cost a
debugging cycle each: models.json does not resolve an env var name for
`apiKey`, it sends the literal string; there is no built-in prime-inference
provider in 0.84.1; a ResourceLoader you pass in is never reloaded for you;
and the stock system prompt is a coding-assistant prompt that must be
replaced — but replacing it also silently removes the tool list, because the
harness only renders that section when it owns the prompt. AGENTS.md records
all four.

The expensive one was thinking level. The harness defaults to `medium`, and
nemotron spent an entire 4,096-token budget reasoning and returned an empty
answer. `low` was worse; `off` omits the parameter so the endpoint's default
wins. An explicit `reasoning_effort: none` via `thinkingLevelMap` took a turn
from 6,195 output tokens to 149.

And a turn is now bounded. The harness loop is `while (true)` with no
iteration cap; a runaway on a frontier model would have eaten the credit it
is supposed to report on. Ceilings on model calls and tokens, enforced both
through the harness hook and independently from the event stream, plus a
per-user daily spend limit — and the ledger now records spend on turns that
fail, which it previously discarded.

Signing in lands on /piggy, which is a workspace: conversations down one
side, the agent in the middle, what it did and what it cost beside it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
claude
2026-08-14 05:26:28 -07:00
parent 99d165b5e5
commit f0173440e4
77 changed files with 28108 additions and 1672 deletions
+203 -28
View File
@@ -1,8 +1,8 @@
# Deploying PIG
PIG is an API/web container, Postgres, and — when you ask for it — a private
Piggy worker/chat container, behind any reverse proxy that terminates TLS.
Nothing here is specific to a particular host.
Piggy container running a Prime Agent session over the CRM, behind any reverse
proxy that terminates TLS. Nothing here is specific to a particular host.
## 1. DNS
@@ -30,9 +30,15 @@ The values that must be set for a production start:
| `PIG_ADMIN_EMAILS` | Who may administer. **Every address here must already have an account** — an unregistered address listed as an admin is a standing offer of admin rights to whoever claims it first |
| `PIG_SETTINGS_ENCRYPTION_KEY` | Base64-encoded 32 bytes (`openssl rand -base64 32`). Only needed for the Notion and Google OAuth secrets typed into the admin UI, which the API refuses to store without it |
Optional: `PRIME_API_KEY` (scope it to `Availability → Read` only), the Slack
and Buzz credentials, and the whole Piggy block — the agent is off unless you
[turn it on](#turning-piggy-on).
Optional: `PRIME_API_KEY`, the Slack and Buzz credentials, and the whole Piggy
block — the agent is off unless you [turn it on](#turning-piggy-on).
`PRIME_API_KEY` is now one key with two jobs: the API syncs GPU availability
from `api.primeintellect.ai` with it, and Piggy calls models on
`api.pinference.ai` with it. Scope it to `Availability → Read` plus inference,
and nothing that can provision. Piggy still accepts `PIGGY_INFERENCE_API_KEY`
as an alias for the same value, so a host configured before the agent moved onto
Prime Inference keeps starting untouched.
Piggy listens on `piggy:8931` inside the Compose network. The port is exposed to
other containers but never published to the host, and Caddy must not route to
@@ -112,16 +118,42 @@ application's mount point for this reason.
## Turning Piggy on
Piggy is the in-app agent: a queue worker and a private chat server, one image
running a second command. It is **off by default** and nothing about it is
configurable from the admin UI — every value below is read once, when the
container boots.
Piggy is the in-app agent, and it is worth knowing what it is before you run it.
It is a **Prime Agent session** — Prime Intellect's own agent harness,
`@earendil-works/pi-coding-agent`, embedded as a library rather than shelled out
to — holding **PIG's CRM tools and nothing else**. The harness is constructed
with every built-in tool disabled (`noTools: 'all'`) and an explicit allowlist
on top, so the model has **no shell, no filesystem access and no Python**. The
live tool list is compared against that allowlist when a session starts and a
mismatch is a startup error, so a future harness release cannot quietly widen
it.
Three things follow for an operator:
- **It writes.** `PIGGY_AGENT_MODE` decides how: `read_only`, `confirm`
(default — a change is proposed as a card and applied when a person clicks) or
`auto`. Contracts, commitments, allocations and compliance records always
require a click regardless. Every write runs as the calling user's own
principal, so Piggy cannot reach a record its user could not, and the audit
trail names the human.
- **The model is chosen per conversation**, from a five-model picker defined in
`apps/piggy/src/agent/models.json`. `PIGGY_AGENT_MODEL` is only the default
for a user who has not chosen.
- **Conversations are persisted** in `piggy_conversations` and `piggy_messages`
(migration 0014), so an agent turn now depends on the schema being current.
`scripts/deploy.sh` migrates from a one-off container before it starts either
container, which is what keeps that true.
It is **off by default** and nothing about it is configurable from the admin UI
— every value below is read once, when the container boots.
Three keys in `.env`, and all three are needed:
```bash
PIGGY_ENABLED=true
PIGGY_INFERENCE_API_KEY=<an inference key from app.primeintellect.ai>
PRIME_API_KEY=<a Prime Intellect key with inference; PIGGY_INFERENCE_API_KEY
is still accepted as the legacy alias for the same value>
PIGGY_INTERNAL_TOKEN=<openssl rand -hex 32>
```
@@ -155,14 +187,87 @@ last month's code indefinitely.
### A missing API key crash-loops the worker
`PIGGY_INFERENCE_API_KEY` is required by `apps/piggy/src/config.ts`. Without it
the process exits at boot with `Invalid Piggy configuration:
PIGGY_INFERENCE_API_KEY is required.`, and `restart: unless-stopped` starts it
again — so the symptom is a container restarting every few seconds, not an error
anyone sees in the CRM. `PIGGY_INTERNAL_TOKEN` shorter than 32 characters fails
the same way. `deploy.sh` catches both: it waits for the container to report
healthy and exits 3 if it does not, deliberately **without** rolling back,
because the previous image reads the same `.env` and would fail identically.
`PRIME_API_KEY` (or its alias `PIGGY_INFERENCE_API_KEY`) is required by
`apps/piggy/src/config.ts`. Without either the process exits at boot with
`Invalid Piggy configuration: PRIME_API_KEY ... is required.`, and
`restart: unless-stopped` starts it again — so the symptom is a container
restarting every few seconds, not an error anyone sees in the CRM.
`PIGGY_INTERNAL_TOKEN` shorter than 32 characters fails the same way, and so
does a `PIGGY_AGENT_MODEL` that is not one of the five ids in
`apps/piggy/src/agent/models.json` — rejected at boot on purpose, because the
alternative is a model that 404s on a user's first question.
`deploy.sh` catches all of them: it waits for the container to report healthy
and exits 3 if it does not, deliberately **without** rolling back, because the
previous image reads the same `.env` and would fail identically.
**A blank line is not an absent one, and here it is actively misleading.**
`PRIME_API_KEY` and `PIGGY_INFERENCE_API_KEY` are two spellings of one key, and
each is declared `.min(1).optional()`. Compose passes a blank `.env` line
through as the empty string, so a file that sets `PRIME_API_KEY` correctly *and*
carries a leftover empty `PIGGY_INFERENCE_API_KEY=` line crash-loops Piggy with:
```
Invalid Piggy configuration:
PIGGY_INFERENCE_API_KEY: String must contain at least 1 character(s)
```
— a message about the key you did not use. **Comment the unused spelling out.**
Before enabling Piggy on an existing host, check for exactly this:
```bash
grep -nE '^(PRIME_API_KEY|PIGGY_INFERENCE_API_KEY)=$' /opt/pig/.env
```
Any line that prints is one to comment out.
### The thinking-level trap — read this before changing the model
The single setting most likely to make a working deployment look broken.
The harness defaults `thinkingLevel` to `medium`, which is tuned for a coding
agent. On the default model that produced **6,195 output tokens of reasoning and
an empty answer**: the turn hit its token ceiling while still thinking and came
back with `finish_reason: length`. `low` was worse. `off` is Piggy's default, and
for the nemotron models it maps to the endpoint's `reasoning_effort: none` — the
same question then answered correctly in **149 output tokens**.
The mapping is **per model** and lives in `thinkingLevelMap` in
`apps/piggy/src/agent/models.json`. The two nemotron entries have one; deepseek,
opus and gpt-5.6 do not, and for them `off` omits `reasoning_effort` entirely so
the endpoint's own default applies.
So if you change `PIGGY_AGENT_MODEL` and start getting empty answers, truncated
answers or a surprising bill, this is where to look — not at the agent, the
tools or the network. Give the new model a `thinkingLevelMap` before raising
`PIGGY_AGENT_THINKING`.
### Where the harness is allowed to look at the filesystem
`PIGGY_AGENT_DIR` is the harness's own directory. `docker-compose.yml` pins it
to `/var/lib/piggy-agent`, which the image creates owned by the unprivileged
`node` user at mode 0700. Leave it alone.
Two reasons it is not the default `~/.pig/piggy-agent`:
- **Writability.** Under `docker run` with `USER node`, `~` resolves to
`/home/node` and works. That is incidental: a runtime that starts this image
with a numeric user and no matching passwd entry (`runAsUser: 1000` under
Kubernetes) leaves `HOME` unset, `os.homedir()` falls back to `/`, and the
agent dies creating its directory — on the first turn, long after the deploy
reported success.
- **Prompt containment.** The harness discovers extensions, skills and context
files from its cwd, and Piggy hands it this directory as cwd. Point it at the
checkout, or bind-mount a repository over it, and source files become
reachable from a CRM agent's prompt. **Never bind-mount anything here.**
`deploy.sh` reads the value back off the running container and refuses to
report success if it sits inside `/app`.
The directory is **not persisted**, deliberately. Nothing in it is worth keeping
across a restart: Piggy rewrites `models.json` there from the image at every
boot, the credential store is in-memory by design, sessions are in-memory, and
the conversations live in Postgres. A cold start costs nothing measurable, and a
volume would only be a way for a file to outlive the image that wrote it.
### Health
@@ -189,9 +294,41 @@ only from inside the Compose network, the API authenticates the user before
forwarding anything, and the bearer token goes in a header — never a query
string, where a proxy or an access log would keep it.
### What the image carries for the agent
Two things about the production image are worth knowing before you debug a
container that will not start.
**`models.json` is a runtime file, not a compiled-in constant.**
`apps/piggy/src/agent/models.json` is read from disk at boot, validated, and
copied into the agent directory for the harness to register its provider from.
It reaches the image inside `COPY apps/piggy`, and the Dockerfile parses it
during the build so that a narrowed `COPY` or a new `.dockerignore` rule fails
there rather than at 03:00 in a crash loop.
**Production dependencies are installed with `--ignore-scripts`.** The Prime
Agent SDK drags in a large transitive tree, including `@google/genai` and
`protobufjs`; their install scripts — and esbuild's — are denied in
`pnpm-workspace.yaml` on purpose, so nothing a dependency pulls in can execute
code at install time. Piggy talks to exactly one provider over an
OpenAI-compatible API and none of that tree is on a path it executes. The
Dockerfile imports the SDK during the build to prove the scriptless install
still yields a loadable agent, and the container has been run end to end against
a live key: health, a tool call, and a correct answer.
### Tuning
Everything else has a working default and exists to be lowered:
The agent's own settings, all with defaults in `apps/piggy/src/config.ts`:
| Variable | Default | What it does |
|---|---|---|
| `PIGGY_AGENT_MODEL` | `nvidia/nemotron-3-nano-30b-a3b` | The default answer model. Must be one of the ids in `apps/piggy/src/agent/models.json`; anything else is refused at boot |
| `PIGGY_AGENT_MODE` | `confirm` | `read_only`, `confirm` or `auto`. Contracts, commitments, allocations and compliance always confirm regardless |
| `PIGGY_AGENT_THINKING` | `off` | See [the thinking-level trap](#the-thinking-level-trap--read-this-before-changing-the-model) before touching it |
| `PIGGY_AGENT_MAX_TOKENS` | `4096` | Output tokens per agent turn, reasoning included. Clamped down to the chosen model's own ceiling |
| `PIGGY_AGENT_DIR` | `~/.pig/piggy-agent` | Pinned to `/var/lib/piggy-agent` by `docker-compose.yml`. Do not override under Compose |
The pre-agent settings still apply to the queue worker and exist to be lowered:
`PIGGY_MAX_TOKENS` (per queued task), `PIGGY_CHAT_MAX_TOKENS` (per interactive
answer), `PIGGY_MAX_TURNS`, `PIGGY_POLL_INTERVAL_MS`, `PIGGY_LEASE_SECONDS`,
`PIGGY_REASONING_EFFORT` and the two `PIGGY_PRICE_*_CENTS_PER_MTOK` values that
@@ -200,6 +337,26 @@ commented out, and that is not decoration: an empty `PIGGY_MAX_TOKENS=` line is
passed to the container as the empty string, which coerces to 0 and refuses to
start. Leave a key commented to get its default; do not leave it blank.
### The database is the only undo for an agent write
Piggy can modify CRM records, so a release that ships a broken write tool can
corrupt data no migration ever touched. `scripts/deploy.sh` dumps the database
before it migrates and before the new image starts, to `backups/` next to the
checkout:
```
backups/pig-20260814-030201.sql.gz
```
That directory is in `.gitignore`, so `autodeploy.sh`'s forced checkout at a
release tag leaves it alone. The dump is verified rather than assumed: the
script decompresses it and refuses to migrate if a database with tables in it
produced less than 4 KB of SQL, because `pg_dump | gzip` turns silence into a
plausible-looking 20-byte file and a deploy log that claims a backup was taken.
Nothing prunes these. They are the only copy — copy them off the host if the
data matters as much as the uptime does.
Changing any of them means restarting the container — `bash scripts/deploy.sh`,
or `docker compose -p pig --profile piggy up -d piggy` if the release is
otherwise unchanged.
@@ -311,10 +468,16 @@ to seek.
bash scripts/deploy.sh
```
It fetches `origin/main`, dumps the database, builds, migrates from a one-off
container, starts the app — and Piggy, when `PIGGY_ENABLED` is on — and refuses
to call the deploy done until the health endpoint, the unauthenticated-401 gate,
the piggy container's image and health, and the public origin all agree.
It fetches `origin/main`, starts the database, dumps it and checks the dump is
usable, builds, migrates from a one-off container, starts the app — and Piggy,
when `PIGGY_ENABLED` is on — and refuses to call the deploy done until the
health endpoint, the unauthenticated-401 gate, the piggy container's image,
agent directory and health, and the public origin all agree.
The database is started **before** the dump rather than after it, which is a
recent fix: the dump `exec`s into that container, so on a host where the stack
was down — a reboot, a `compose down`, a first-ever deploy — the backup step
failed and took the release with it.
### By tag — the normal path
@@ -383,16 +546,28 @@ sentence for each, so the journal never claims a rollback that did not happen.
Two things it deliberately does **not** do:
- **It does not roll the database back.** Migrations are additive, so the
previous image runs against the new schema. The dump taken before the
migration is the escape hatch for when that is not true.
previous image runs against the new schema. 0014 is the current example: it
creates `piggy_conversations` and `piggy_messages` and alters nothing, so an
image that predates it never names those tables and cannot notice they exist.
That safety is a property of the migrations, not of this script — a migration
that drops a column, renames one or tightens a constraint would break the
restored image, and the dump taken before the migration is the escape hatch
for that.
- **It does not roll back when the public origin answers with an EMPTY body.**
Something terminated TLS and replied, so the fault is the proxy — see `bind`
below — and the previous image would fail the same check. Exit 3.
- **It does not roll back when Piggy is enabled but does not come up.** Same
reasoning: the previous image reads the same `.env`, so restoring it churns a
healthy CRM without fixing the agent. Exit 3, and the log names
`PIGGY_INFERENCE_API_KEY` because that is nearly always the cause. This is the
one exit-3 case where the site itself is fine.
healthy CRM without fixing the agent. Exit 3, and the log lists the three
configuration faults that cause it — a missing or rejected `PRIME_API_KEY`, a
short `PIGGY_INTERNAL_TOKEN`, a `PIGGY_AGENT_MODEL` outside the catalogue.
- **It does not roll back when `PIGGY_AGENT_DIR` points inside `/app`.** The
container is healthy; the problem is that the harness's cwd would be the
application checkout, so repository files could reach a CRM agent's prompt.
Restoring the previous image changes nothing — it reads the same compose file
— so the script refuses to report success and exits 3.
Those last two are the exit-3 cases where the site itself is fine.
It *does* roll back when the origin answers with a **non-empty** body that lacks
the marker. A proxy fault cannot serve a wrong-but-populated page for this