Files
pig/scripts/deploy.sh
T
karti 13dec6b4b8
CI / verify (push) Successful in 3m45s
CI / publish (push) Has been skipped
Rebuild the shell, add Calendar and Learn, and govern reads
Seven parallel agents and an adversarial verification pass. The three things
worth knowing before reading the diff:

RBAC WAS ALREADY BUILT. docs/build-plan.md marks F2 and F3 outstanding and is
stale — packages/core/src/permissions.ts and lib/mutation.ts shipped long ago.
So this does not rebuild them; it closes the gaps an audit found. The big one
is that reads were entirely ungoverned: every GET was "any authenticated
member", so a junior demand rep and a research contractor could both pull
per-block supplier cost and break-even prices from /api/capacity/margin, and
every contract's negotiated terms. For a company whose margin is the business,
that was the hole that mattered. Adds book:read / economics:read / team:read,
a readGuard middleware, and a `viewer` role below member.

THE BUTTON AND THE 403 DISAGREED — the exact thing F3 said must never happen.
Contracts.tsx never called can() at all, so its save button was always enabled
against a server requiring contract:sign; Capacity.tsx gated commitment
creation on deal:write/demand while the server wanted commitment:write/supply.

POST /api/activities was the one write bypassing executeMutation: no capability
check, and any member could mutate accounts.lastActivityAt as a side effect.
It is now a proper mutation() behind activity:write.

The shell becomes three panes — a collapsible shadcn sidebar with an account
switcher on the Piggy accent, a header with real search, and Piggy docked to
the right, page-aware and persistent across navigation. The phone keeps its
bottom tab bar, which is the thing this product already beat trycompai/crm on,
and gains the sidebar as a sheet.

Calendar is a projection over thirteen dated sources rather than a new table,
because a table would duplicate dates that already live on contracts, deals and
commitments and would drift — and one ledger answering the question is the
whole argument. It surfaces export_authorizations and compliance_artifacts,
which had indexed expires_at columns, schema comments saying they must be
alerted on, and no read endpoint or UI anywhere.

Learn carries two tracks. Concepts are members-only; the platform track can be
opened with a share code by someone with no account. The code mints a scoped
learn-only token and never a Principal — every route here resolves a principal
and then checks capabilities, so a principal-minting code would be one missing
check away from leaking the book. "Only platform-track rows may be code-visible"
is a database CHECK constraint as well as a write-path rule, and a test asserts
a valid learn token still gets 401 on /api/dashboard, /api/accounts and
/api/contracts — the same invariant scripts/deploy.sh refuses to ship without.

CD becomes tag-to-ship. CI publishes an image to the Gitea registry on a
release-* tag and cloud-2 pulls it, so no credential on the shared runner can
execute anything on production — by construction rather than by policy. Both
halves of deploy.sh's original rule survive: nothing on the runner reaches the
host, and a human still decides when it ships. deploy.sh gains a rollback and a
public-origin check, and PIG_IMAGE now reaches compose through `sudo env`,
without which sudo's env_reset silently resolved every release to pig:local.

Tests 141 -> 261.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 15:02:48 -07:00

251 lines
11 KiB
Bash
Executable File

#!/usr/bin/env bash
#
# Deploy PIG. Run on the host that serves it.
#
# ./scripts/deploy.sh
#
# Deliberately a script rather than automated push-to-deploy. Automating it
# would mean putting an SSH key with write access to the production host onto
# the CI runner, which is a meaningful escalation for a project this size. CI
# proves the commit is sound; a human decides when it ships.
#
# That constraint still holds under the tag-to-ship flow added later. Nothing
# on the CI runner can reach this host: CI publishes an image, and this host
# PULLS it (scripts/autodeploy.sh). The direction of the credential is the
# whole point — a read-only registry token here, no host credential there. And
# a human still decides when it ships, by choosing to create a `release-*` tag.
#
# PIG_IMAGE unset Build from this working tree, as it always has.
# PIG_IMAGE set Pull that published image instead of building, and leave
# the checkout exactly where the caller put it. The caller
# is responsible for having checked out the matching
# commit, because the compose file, the migrations and the
# image have to agree.
#
# Safe to re-run. Migrations are additive and tracked.
set -euo pipefail
cd "$(dirname "$0")/.."
# What compose will run. Must match the `image:` default in docker-compose.yml,
# because that is the tag a rollback re-points at the previous image.
IMAGE_REF="${PIG_IMAGE:-pig:local}"
# Every compose invocation goes through this. sudo's default `env_reset` drops
# PIG_IMAGE, so a plain `sudo docker compose` interpolates the `pig:local`
# fallback in docker-compose.yml instead of the release tag: the pull then fails
# with "pull access denied for pig", and — worse, if the pull is ever made
# non-fatal — the migrate, the `up` and the rollback all silently run whatever
# `pig:local` happens to be while the log reports the release. Pass it on the
# command line via env(1): `sudo -E` and bare `sudo VAR=val` are both refused by
# the default sudoers policy, `sudo env VAR=val …` is not.
dc() {
sudo env PIG_IMAGE="$IMAGE_REF" docker compose -p pig "$@"
}
# Gate failures come in two kinds and the caller must be able to tell them
# apart: scripts/autodeploy.sh reports to the journal what is serving right now.
# 1 a gate failed and the previous image was restored — the old release is up
# 3 a gate failed and the release under test is STILL LIVE
EXIT_STILL_LIVE=3
# Read one key out of .env WITHOUT sourcing it. Sourcing an environment file
# executes it, and this one holds every secret the deployment has.
env_value() {
[ -f .env ] || return 0
sed -n "s/^[[:space:]]*$1=//p" .env | tail -n1 | sed -e 's/^"\(.*\)"$/\1/' -e "s/^'\(.*\)'\$/\1/"
}
if [ -n "${PIG_IMAGE:-}" ]; then
echo "==> Deploying published image $PIG_IMAGE"
# No git sync. The caller has already detached this checkout at the tag the
# image was built from; resetting to origin/main here would silently deploy a
# compose file and a migration set from a different commit than the image.
AFTER=$(git rev-parse --short HEAD)
echo " tree at $AFTER"
else
echo "==> Fetching"
git fetch -q origin
BEFORE=$(git rev-parse --short HEAD)
git reset --hard -q origin/main
AFTER=$(git rev-parse --short HEAD)
if [ "$BEFORE" = "$AFTER" ]; then
echo " Already at $AFTER"
else
echo " $BEFORE -> $AFTER"
git --no-pager log --oneline "$BEFORE..$AFTER" | sed 's/^/ /'
fi
fi
echo "==> Backing up the database first"
# Cheap insurance. A migration that goes wrong on a database holding real deal
# data is not something to discover without a dump in hand.
mkdir -p backups
BACKUP="backups/pig-$(date +%Y%m%d-%H%M%S).sql.gz"
dc exec -T db pg_dump -U pig pig | gzip > "$BACKUP"
echo " $BACKUP ($(du -h "$BACKUP" | cut -f1))"
if [ -n "${PIG_IMAGE:-}" ]; then
echo "==> Pulling"
dc pull app
else
echo "==> Building"
dc build app
fi
echo "==> Starting the database"
dc up -d db
for _ in $(seq 1 60); do
if dc exec -T db pg_isready -U pig -d pig > /dev/null; then break; fi
sleep 1
done
if ! dc exec -T db pg_isready -U pig -d pig > /dev/null; then
echo "ERROR: database did not become ready within 60 seconds" >&2
exit 1
fi
echo "==> Migrating before the schema-dependent app starts"
# A release may query a newly introduced table during startup. Running the
# migration from a one-off container prevents that app from crash-looping
# before an `exec`-based migration can reach it.
dc run --rm --no-deps app pnpm exec tsx packages/db/src/migrate.ts
# The image the current container is running, captured before it is replaced.
# Without this a failed release stays live: both gates below used to exit 1 and
# leave the broken version serving, which is tolerable when a human is reading
# the terminal and an outage when the poller ran this at 04:00.
PREVIOUS_IMAGE=""
PREVIOUS_CID=$(dc ps -q app 2>/dev/null | head -n1 || true)
if [ -n "$PREVIOUS_CID" ]; then
PREVIOUS_IMAGE=$(sudo docker inspect -f '{{.Image}}' "$PREVIOUS_CID" 2>/dev/null || true)
fi
# Restore the previous image and exit non-zero. Deliberately NOT a database
# rollback: migrations are additive, so the previous code runs against the new
# schema, and the dump taken above is the escape hatch for the case where it
# does not. The re-tag makes $IMAGE_REF point back at the old image locally; the
# next successful pull moves it forward again.
roll_back() {
echo " !! $1"
dc logs app --tail 40 || true
if [ -z "$PREVIOUS_IMAGE" ]; then
echo " NO PREVIOUS IMAGE TO ROLL BACK TO — the broken release is live" >&2
exit "$EXIT_STILL_LIVE"
fi
echo "==> Rolling back to $PREVIOUS_IMAGE"
sudo docker tag "$PREVIOUS_IMAGE" "$IMAGE_REF"
dc up -d --no-build app
for _ in $(seq 1 60); do
if curl -sf http://127.0.0.1:8920/api/health > /dev/null; then break; fi
sleep 1
done
if curl -sf http://127.0.0.1:8920/api/health | grep -q '"ok":true'; then
echo " rolled back, previous release is healthy"
exit 1
fi
# The restore ran but did not come up. Reporting "rolled back" here would tell
# the on-call the previous release is serving when nothing is.
echo " ROLLBACK ALSO UNHEALTHY — this host needs a human" >&2
exit "$EXIT_STILL_LIVE"
}
echo "==> Starting the app"
dc up -d app
echo "==> Waiting for health"
for _ in $(seq 1 60); do
if curl -sf http://127.0.0.1:8920/api/health > /dev/null; then break; fi
sleep 1
done
echo "==> Verifying"
if curl -sf http://127.0.0.1:8920/api/health | grep -q '"ok":true'; then
echo " health ok"
else
echo " HEALTH CHECK FAILED"
roll_back "health check failed"
fi
# Authentication must be enforced. A deploy that accidentally serves the CRM
# unauthenticated is the one failure worth blocking on.
CODE=$(curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:8920/api/dashboard)
if [ "$CODE" != "401" ]; then
echo " UNAUTHENTICATED REQUEST RETURNED $CODE, EXPECTED 401"
roll_back "unauthenticated request returned $CODE, expected 401"
fi
echo " auth enforced"
# Everything above proves the container is well. It proves nothing about what
# the public actually gets, and there is a failure on this host that every
# check so far passes: a Caddy site block missing `bind 10.0.0.2` lands in a
# separate server on :443, wins for traffic arriving on that address, and
# answers with a valid certificate, HTTP 200 and an EMPTY body. The app is
# fine; the request never reached it. So ask the origin, from outside the
# compose network, and insist on seeing something the app actually renders.
PUBLIC_URL="${PIG_DEPLOY_PUBLIC_URL:-$(env_value PIG_PUBLIC_URL)}"
PUBLIC_URL="${PUBLIC_URL:-https://primeintellectgrowth.com}"
PUBLIC_MARKER="${PIG_DEPLOY_PUBLIC_MARKER:-<div id=\"root\">}"
echo "==> Verifying the public origin ($PUBLIC_URL)"
# Three outcomes, and conflating any two of them has already cost a deploy.
# `|| true` used to fold "curl never got an answer" into the empty string, so a
# host where the public name does not hairpin, or where egress to 443 is
# filtered, failed a completely healthy release — and autodeploy.sh then
# blacklists that digest for good. Keep the transport result separate from the
# body.
PUBLIC_ERR=$(mktemp)
trap 'rm -f "$PUBLIC_ERR"' EXIT
if PUBLIC_BODY=$(curl -sS --max-time 20 --location "${PUBLIC_URL%/}/" 2>"$PUBLIC_ERR"); then
CURL_STATUS=0
else
CURL_STATUS=$?
fi
PUBLIC_BYTES=${#PUBLIC_BODY}
if [ "$CURL_STATUS" -ne 0 ]; then
# Could not connect at all: DNS, TLS, no hairpin, filtered egress. This says
# nothing about the release, and nothing about the proxy either — from this
# host the two are indistinguishable. Do not fail a deploy that passed every
# check that can actually be trusted; set PIG_DEPLOY_REQUIRE_PUBLIC=1 on a
# host where the public name IS reachable from here and this must block.
echo " could not reach the public origin (curl exit $CURL_STATUS)" >&2
sed 's/^/ /' "$PUBLIC_ERR" >&2
if [ "${PIG_DEPLOY_REQUIRE_PUBLIC:-0}" = '1' ]; then
echo " PIG_DEPLOY_REQUIRE_PUBLIC=1, so this is fatal." >&2
exit "$EXIT_STILL_LIVE"
fi
echo " Not treated as a failure: the container passed every local gate," >&2
echo " and this host may simply not be able to route to its own name." >&2
echo " Verify $PUBLIC_URL from somewhere else." >&2
elif [ "$PUBLIC_BYTES" -eq 0 ]; then
echo " PUBLIC ORIGIN ANSWERED WITH AN EMPTY BODY" >&2
echo " Something terminated TLS and answered, so this is the proxy, not" >&2
echo " the release. Check that every Caddy site block on this host has" >&2
echo " 'bind 10.0.0.2'." >&2
# Deliberately no rollback: the previous image would fail this same check.
# Rolling back would churn the service and not fix a proxy fault. The release
# under test therefore stays live, which is what EXIT_STILL_LIVE tells the
# poller to say.
exit "$EXIT_STILL_LIVE"
elif ! printf '%s' "$PUBLIC_BODY" | grep -qF "$PUBLIC_MARKER"; then
echo " PUBLIC ORIGIN ANSWERED $PUBLIC_BYTES BYTES WITHOUT '$PUBLIC_MARKER'" >&2
# This one DOES roll back, unlike the empty-body case above. A proxy fault
# cannot serve a wrong-but-non-empty page for this hostname; a bad release
# can — a changed Vite `base`, a Dockerfile that stopped copying
# apps/web/dist. Every earlier gate passes in that state (the API is well,
# /api/dashboard still 401s) and only this one fires, so if it is the release
# then refusing to roll back leaves the broken release serving the public.
# Restoring the previous image is right when it is, and harmless when it is
# not: the marker check is the last gate, so the rollback costs one restart.
roll_back "public origin answered without '$PUBLIC_MARKER'"
else
echo " public origin serving PIG ($PUBLIC_BYTES bytes)"
fi
echo "==> Deployed $AFTER"