Seven parallel agents and an adversarial verification pass. The three things worth knowing before reading the diff: RBAC WAS ALREADY BUILT. docs/build-plan.md marks F2 and F3 outstanding and is stale — packages/core/src/permissions.ts and lib/mutation.ts shipped long ago. So this does not rebuild them; it closes the gaps an audit found. The big one is that reads were entirely ungoverned: every GET was "any authenticated member", so a junior demand rep and a research contractor could both pull per-block supplier cost and break-even prices from /api/capacity/margin, and every contract's negotiated terms. For a company whose margin is the business, that was the hole that mattered. Adds book:read / economics:read / team:read, a readGuard middleware, and a `viewer` role below member. THE BUTTON AND THE 403 DISAGREED — the exact thing F3 said must never happen. Contracts.tsx never called can() at all, so its save button was always enabled against a server requiring contract:sign; Capacity.tsx gated commitment creation on deal:write/demand while the server wanted commitment:write/supply. POST /api/activities was the one write bypassing executeMutation: no capability check, and any member could mutate accounts.lastActivityAt as a side effect. It is now a proper mutation() behind activity:write. The shell becomes three panes — a collapsible shadcn sidebar with an account switcher on the Piggy accent, a header with real search, and Piggy docked to the right, page-aware and persistent across navigation. The phone keeps its bottom tab bar, which is the thing this product already beat trycompai/crm on, and gains the sidebar as a sheet. Calendar is a projection over thirteen dated sources rather than a new table, because a table would duplicate dates that already live on contracts, deals and commitments and would drift — and one ledger answering the question is the whole argument. It surfaces export_authorizations and compliance_artifacts, which had indexed expires_at columns, schema comments saying they must be alerted on, and no read endpoint or UI anywhere. Learn carries two tracks. Concepts are members-only; the platform track can be opened with a share code by someone with no account. The code mints a scoped learn-only token and never a Principal — every route here resolves a principal and then checks capabilities, so a principal-minting code would be one missing check away from leaking the book. "Only platform-track rows may be code-visible" is a database CHECK constraint as well as a write-path rule, and a test asserts a valid learn token still gets 401 on /api/dashboard, /api/accounts and /api/contracts — the same invariant scripts/deploy.sh refuses to ship without. CD becomes tag-to-ship. CI publishes an image to the Gitea registry on a release-* tag and cloud-2 pulls it, so no credential on the shared runner can execute anything on production — by construction rather than by policy. Both halves of deploy.sh's original rule survive: nothing on the runner reaches the host, and a human still decides when it ships. deploy.sh gains a rollback and a public-origin check, and PIG_IMAGE now reaches compose through `sudo env`, without which sudo's env_reset silently resolved every release to pig:local. Tests 141 -> 261. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
13 KiB
Deploying PIG
PIG is an API/web container, a private Piggy worker/chat container and Postgres, behind any reverse proxy that terminates TLS. Nothing here is specific to a particular host.
1. DNS
Point the apex and www at the machine. Both must resolve before the proxy can
obtain a certificate.
A primeintellectgrowth.com -> <your IP>
A www.primeintellectgrowth.com -> <your IP>
2. Configuration
cp .env.example .env # then edit
The values that must be set for a production start:
| Variable | Why |
|---|---|
POSTGRES_PASSWORD |
Generate a fresh one; never reuse another service's |
PIG_PUBLIC_URL |
The single origin the app is served from |
SUPABASE_URL / SUPABASE_ANON_KEY |
Authentication. The app refuses to start in production without a Supabase URL, because it would otherwise serve the whole CRM unauthenticated |
PIG_ADMIN_EMAILS |
Who may administer. Every address here must already have an account — an unregistered address listed as an admin is a standing offer of admin rights to whoever claims it first |
PIGGY_INFERENCE_API_KEY |
Model credential held only by the Piggy process |
PIGGY_INTERNAL_TOKEN |
A generated 32+ character bearer token shared only by API and Piggy |
Optional: PRIME_API_KEY (scope it to Availability → Read only),
PIGGY_ENABLED, and the Slack and Buzz credentials.
Piggy listens on piggy:8931 inside the Compose network. The port is exposed to
other containers but never published to the host, and Caddy must not route to
it. The CRM API authenticates the user, forwards only bounded chat context, and
uses PIGGY_INTERNAL_TOKEN in an Authorization header. Never put that token in
a query string, where proxies and access logs can retain it.
3. Start
docker compose -p pig up -d db
docker compose -p pig run --rm --no-deps app pnpm exec tsx packages/db/src/migrate.ts
docker compose -p pig up -d --build app
docker compose -p pig run --rm --no-deps app pnpm exec tsx packages/db/src/seed/index.ts # optional
Migrate from a one-off container, before the app starts — not with exec.
exec needs a running app to attach to, and a release that queries a table its
migration has not yet created crash-loops before you can attach to it. You then
have a container restarting every few seconds and no way in. run --rm --no-deps uses the same image without the app, and without starting its
dependencies twice. This is what commit d4d7095 changed and it is what
scripts/deploy.sh does.
When Piggy is enabled, start its private profile as well:
docker compose -p pig --profile piggy up -d --build
Use -p pig. A compose project that shares a name with a neighbouring stack
will adopt its volumes, which is a memorable way to lose a database.
4. Reverse proxy
See Caddyfile.example. Serve the app and API from the same origin.
Two things that will otherwise cost you an hour:
-
If other sites on the host use
bind <address>, yours must too. Caddy groups site blocks into servers by listen address. A block withoutbindlands in a separate server on:443, and the more specific listener wins for traffic arriving on that address — which is all public traffic after NAT. The symptom is a valid certificate, a 200 response, an empty body, and none of your headers. It looks like the app is broken; it is that the request never reached it. -
The CSP must carry the hash of the inline theme script in
index.html. That script sets light or dark before first paint so dark-mode users do not get a white flash. Editing it changes the hash and CSP will silently block it — the browser console prints the hash it expects.That hash exists in three places, and only two of them are checked.
Copy Checked by .gitea/workflows/ci.yml(theexpectedconstant)itself, on every run deploy/Caddyfile.examplenothing — it is an example the live Caddyfileon the hostnothing at all The live one is the only copy that decides whether a browser runs the script. Nothing in this repository can see it, CI cannot fail on it, and the failure is a white flash for dark-mode users with no error anywhere. Editing that script means editing all three by hand and reloading Caddy.
5. Verify
curl -s https://primeintellectgrowth.com/api/health
# {"ok":true,"service":"pig","version":"0.1.0"}
Check the public origin, not just 127.0.0.1:8920. The bind failure
above answers with a valid certificate, HTTP 200 and an empty body, which
satisfies every check that only asks whether something responded.
scripts/deploy.sh now asserts the body is non-empty and contains the
application's mount point for this reason.
Upgrading
By hand
bash scripts/deploy.sh
It fetches origin/main, dumps the database, builds, migrates from a one-off
container, starts the app, and refuses to call the deploy done until the health
endpoint, the unauthenticated-401 gate and the public origin all agree.
By tag — the normal path
Shipping is two steps and the second one is a human being:
git tag release-2026-08-13 && git push origin release-2026-08-13
That is the entire ship decision. What follows:
- CI runs the full
verifyjob against the tagged commit — the same job a push to main runs. A tag does not skip verification. - Only if that passes, the
publishjob builds and pushesgit.karti.ai/pig/pig:<tag>and:<short-sha>to the Gitea registry. - Within five minutes
pig-autodeploy.timeron the host notices that the newestrelease-*tag has a digest different from the running container, checks the tree out at that tag, and runsscripts/deploy.shwithPIG_IMAGEset — so it pulls the published image instead of rebuilding it.
Push to main deploys nothing. Tagging does.
The direction of travel is the point. No credential on the shared CI runner can
execute anything on this host; the host holds a pull-only registry token and
fetches. That preserves both halves of the constraint written at the top of
scripts/deploy.sh — no production key on the runner, and a human still
choosing when it ships.
Trap: sudo throws PIG_IMAGE away. The default sudoers policy sets
env_reset, so PIG_IMAGE=… sudo docker compose … hands compose an environment
without it and compose interpolates the pig:local fallback from
docker-compose.yml. The pull then dies with "pull access denied for pig" — and
if it had not died, the migrate, the up and the rollback would all have run
the stale local image while the log named the release tag. Every compose
invocation in deploy.sh therefore goes through the dc() wrapper, which uses
sudo env PIG_IMAGE=… docker compose …; sudo -E and bare sudo VAR=val are
both refused by that same policy. Anything new that shells out to compose must
use the wrapper.
Rollback
scripts/deploy.sh records the image the app container was running before it
replaces it. If the health check, the unauthenticated-401 gate, or the
public-origin marker check fails, it re-tags that image, restarts the app on it,
reports whether the restored version is healthy, and exits non-zero. Previously
those exits left the broken release live, which was fine when a human was
watching the terminal and an outage when the poller ran at 04:00.
The exit code says what is serving, because that is the one thing the
on-call needs at 04:00 and autodeploy.sh can see nothing else:
| Exit | Meaning |
|---|---|
| 0 | deployed |
| 1 | a gate failed and the previous image was restored |
| 3 | a gate failed and the release under test is still live |
3 covers the cases where a rollback was not attempted (the fault is not
attributable to the release), where there was no previous image to restore, and
where the restore itself did not come up. autodeploy.sh logs a different
sentence for each, so the journal never claims a rollback that did not happen.
Two things it deliberately does not do:
- It does not roll the database back. Migrations are additive, so the previous image runs against the new schema. The dump taken before the migration is the escape hatch for when that is not true.
- It does not roll back when the public origin answers with an EMPTY body.
Something terminated TLS and replied, so the fault is the proxy — see
bindbelow — and the previous image would fail the same check. Exit 3.
It does roll back when the origin answers with a non-empty body that lacks
the marker. A proxy fault cannot serve a wrong-but-populated page for this
hostname; a bad release can — a changed Vite base, a Dockerfile step that
stopped copying apps/web/dist. Every earlier gate passes in that state, so
this is the only one that fires, and refusing to roll back would leave the
broken release facing the public.
If curl cannot reach the public origin at all (no hairpin for the public
name, egress to 443 filtered), that is not a failed deploy: it is a host that
cannot see itself from outside, and blacklisting the digest over it costs a
release. The script warns and exits 0. On a host where the public name really is
reachable from the host, set PIG_DEPLOY_REQUIRE_PUBLIC=1 to make it fatal
(exit 3).
scripts/autodeploy.sh writes the failed digest to
/var/lib/pig/failed-release and will not retry it, so one bad tag does not
become a five-minute restart loop. Delete that file, or publish a new tag, to
try again.
To pin a specific release by hand:
git checkout --detach refs/tags/release-2026-08-12
PIG_IMAGE=git.karti.ai/pig/pig:release-2026-08-12 bash scripts/deploy.sh
Check the tree out at the tag as well as setting PIG_IMAGE. The compose file
and the migrations must come from the same commit as the image; with
PIG_IMAGE set, deploy.sh deliberately does not touch git, precisely so it
cannot drag you back to main behind your back.
Migrations are additive and safe to re-run; Drizzle tracks what has been
applied. deploy.sh takes a dump before every deploy; take one by hand before
a major upgrade anyway:
docker compose -p pig exec db pg_dump -U pig pig | gzip > pig-$(date +%F).sql.gz
Installing the release poller
Only on the host that serves production, and only once.
# 1. A pull-only credential. read:package scope and NOTHING else — a token here
# that can write packages or push to the repository undoes the reason
# deployment is not automated from CI in the first place.
sudo install -d -m 0755 /etc/pig
printf '%s' 'gitea-token-here' | sudo tee /etc/pig/registry-token > /dev/null
sudo chmod 0600 /etc/pig/registry-token
sudo chown root:root /etc/pig/registry-token
# 2. Anything the defaults get wrong. Optional; the script assumes
# /opt/pig, git.karti.ai and pig/pig.
sudo tee /etc/pig/autodeploy.env > /dev/null <<'EOF'
PIG_REGISTRY_USER=pig-deploy
PIG_REPO_DIR=/opt/pig
EOF
sudo chmod 0600 /etc/pig/autodeploy.env
# 3. The units.
sudo cp /opt/pig/deploy/pig-autodeploy.service /etc/systemd/system/
sudo cp /opt/pig/deploy/pig-autodeploy.timer /etc/systemd/system/
sudo systemctl daemon-reload
# 4. Dry-run it once, in the foreground, before trusting a timer with it.
sudo systemctl start pig-autodeploy.service
sudo journalctl -u pig-autodeploy.service -n 50 --no-pager
# 5. Then arm it.
sudo systemctl enable --now pig-autodeploy.timer
systemctl list-timers pig-autodeploy.timer
The service is Type=oneshot with no Restart=, and the timer is not
Persistent=true: a missed poll is caught at the next tick rather than fired
at boot, which is when nobody is watching.
Watch a deploy:
journalctl -u pig-autodeploy.service -f
On-premises: using your own identity provider
PIG authenticates against any standards-compliant OIDC provider, which is how an install inside your own network works. Set:
PIG_OIDC_ISSUER=https://id.yourcompany.internal
PIG_OIDC_AUDIENCE=pig # the client/app id you registered for PIG
That is usually the whole configuration — the JWKS is discovered from the
issuer. On an air-gapped network, set PIG_OIDC_JWKS_URI too and no discovery
request is made.
PIG_OIDC_ISSUER takes precedence over SUPABASE_URL, so the hosted values
can stay in the environment file without quietly taking over.
Set the audience. Without it, any token your provider issued for any application in the same tenant verifies here — a token minted for an unrelated internal tool would be accepted as a PIG session. PIG warns about this at boot but cannot refuse, because some providers legitimately issue single-audience tokens.
Provisioning stays in PIG. Authenticating proves who someone is; it does not make them a member. They still need an invite, and their team and role live in PIG's database. That is deliberate — your directory should not have to model "supply lead versus demand member" for one application.
A note on the auth project
PIG verifies JWTs but authorizes from its own users table. If the Supabase
project is shared with another application, its users get nothing here until
they are explicitly invited. That is deliberate, and it is why a valid token can
still return 403 needs_profile.