Files
pig/scripts/deploy.sh
T
karti a21ecf9e53
CI / verify (push) Successful in 2m42s
CI / publish (push) Has been skipped
Stop deploy.sh from rewriting itself while bash is reading it
`git reset --hard origin/main` replaces this script mid-execution. bash does
not slurp a script — it reads incrementally and remembers a byte OFFSET, so
after the reset it resumes at that offset into different content.

This is not theoretical. The 13dec6b deploy hit it: the new public-origin gate
and the rollback were on disk and never ran, because bash was still executing
the buffered previous version. That deploy exited 0 and the release is healthy,
so it cost nothing this time. The failure mode when it does bite is a spliced
or half-executed line, part way through a deployment.

scripts/autodeploy.sh has always re-exec'd from a mktemp copy for exactly this
reason. deploy.sh needed the same guard.

PIG_REPO_ROOT is resolved before the re-exec and exported across it: after the
re-exec `$0` is the copy in /tmp, so `dirname "$0"` would cd to the wrong tree.
Verified with a harness that rewrites the original mid-run and asserts the
child keeps both its content and its working directory.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 15:08:22 -07:00

274 lines
12 KiB
Bash
Executable File

#!/usr/bin/env bash
#
# Deploy PIG. Run on the host that serves it.
#
# ./scripts/deploy.sh
#
# Deliberately a script rather than automated push-to-deploy. Automating it
# would mean putting an SSH key with write access to the production host onto
# the CI runner, which is a meaningful escalation for a project this size. CI
# proves the commit is sound; a human decides when it ships.
#
# That constraint still holds under the tag-to-ship flow added later. Nothing
# on the CI runner can reach this host: CI publishes an image, and this host
# PULLS it (scripts/autodeploy.sh). The direction of the credential is the
# whole point — a read-only registry token here, no host credential there. And
# a human still decides when it ships, by choosing to create a `release-*` tag.
#
# PIG_IMAGE unset Build from this working tree, as it always has.
# PIG_IMAGE set Pull that published image instead of building, and leave
# the checkout exactly where the caller put it. The caller
# is responsible for having checked out the matching
# commit, because the compose file, the migrations and the
# image have to agree.
#
# Safe to re-run. Migrations are additive and tracked.
set -euo pipefail
# Resolved BEFORE the re-exec below and carried across it. After the re-exec
# `$0` is a copy in /tmp, so `dirname "$0"` would point at the wrong tree.
PIG_REPO_ROOT="${PIG_REPO_ROOT:-$(cd "$(dirname "$0")/.." && pwd)}"
export PIG_REPO_ROOT
cd "$PIG_REPO_ROOT"
# Run from a copy, because this script rewrites itself.
#
# `git reset --hard origin/main` below replaces this very file while bash is
# still reading it. bash does not slurp a script: it reads incrementally and
# remembers a byte OFFSET, so after the reset it resumes at that offset into
# different content — which silently skips or splices steps, and at worst
# executes a fragment of a line. The 13dec6b deploy hit exactly this: the new
# public-origin gate was on disk and never ran, because bash was still
# executing the buffered previous version.
#
# It failed harmlessly that time. It is not guaranteed to. scripts/autodeploy.sh
# has always had this guard; deploy.sh needed it for the same reason.
if [ "${PIG_DEPLOY_REEXEC:-}" != '1' ]; then
_copy=$(mktemp -t pig-deploy.XXXXXX)
cat "$0" > "$_copy"
PIG_DEPLOY_REEXEC=1 exec bash "$_copy" "$@"
fi
trap 'rm -f "$0"' EXIT
# What compose will run. Must match the `image:` default in docker-compose.yml,
# because that is the tag a rollback re-points at the previous image.
IMAGE_REF="${PIG_IMAGE:-pig:local}"
# Every compose invocation goes through this. sudo's default `env_reset` drops
# PIG_IMAGE, so a plain `sudo docker compose` interpolates the `pig:local`
# fallback in docker-compose.yml instead of the release tag: the pull then fails
# with "pull access denied for pig", and — worse, if the pull is ever made
# non-fatal — the migrate, the `up` and the rollback all silently run whatever
# `pig:local` happens to be while the log reports the release. Pass it on the
# command line via env(1): `sudo -E` and bare `sudo VAR=val` are both refused by
# the default sudoers policy, `sudo env VAR=val …` is not.
dc() {
sudo env PIG_IMAGE="$IMAGE_REF" docker compose -p pig "$@"
}
# Gate failures come in two kinds and the caller must be able to tell them
# apart: scripts/autodeploy.sh reports to the journal what is serving right now.
# 1 a gate failed and the previous image was restored — the old release is up
# 3 a gate failed and the release under test is STILL LIVE
EXIT_STILL_LIVE=3
# Read one key out of .env WITHOUT sourcing it. Sourcing an environment file
# executes it, and this one holds every secret the deployment has.
env_value() {
[ -f .env ] || return 0
sed -n "s/^[[:space:]]*$1=//p" .env | tail -n1 | sed -e 's/^"\(.*\)"$/\1/' -e "s/^'\(.*\)'\$/\1/"
}
if [ -n "${PIG_IMAGE:-}" ]; then
echo "==> Deploying published image $PIG_IMAGE"
# No git sync. The caller has already detached this checkout at the tag the
# image was built from; resetting to origin/main here would silently deploy a
# compose file and a migration set from a different commit than the image.
AFTER=$(git rev-parse --short HEAD)
echo " tree at $AFTER"
else
echo "==> Fetching"
git fetch -q origin
BEFORE=$(git rev-parse --short HEAD)
git reset --hard -q origin/main
AFTER=$(git rev-parse --short HEAD)
if [ "$BEFORE" = "$AFTER" ]; then
echo " Already at $AFTER"
else
echo " $BEFORE -> $AFTER"
git --no-pager log --oneline "$BEFORE..$AFTER" | sed 's/^/ /'
fi
fi
echo "==> Backing up the database first"
# Cheap insurance. A migration that goes wrong on a database holding real deal
# data is not something to discover without a dump in hand.
mkdir -p backups
BACKUP="backups/pig-$(date +%Y%m%d-%H%M%S).sql.gz"
dc exec -T db pg_dump -U pig pig | gzip > "$BACKUP"
echo " $BACKUP ($(du -h "$BACKUP" | cut -f1))"
if [ -n "${PIG_IMAGE:-}" ]; then
echo "==> Pulling"
dc pull app
else
echo "==> Building"
dc build app
fi
echo "==> Starting the database"
dc up -d db
for _ in $(seq 1 60); do
if dc exec -T db pg_isready -U pig -d pig > /dev/null; then break; fi
sleep 1
done
if ! dc exec -T db pg_isready -U pig -d pig > /dev/null; then
echo "ERROR: database did not become ready within 60 seconds" >&2
exit 1
fi
echo "==> Migrating before the schema-dependent app starts"
# A release may query a newly introduced table during startup. Running the
# migration from a one-off container prevents that app from crash-looping
# before an `exec`-based migration can reach it.
dc run --rm --no-deps app pnpm exec tsx packages/db/src/migrate.ts
# The image the current container is running, captured before it is replaced.
# Without this a failed release stays live: both gates below used to exit 1 and
# leave the broken version serving, which is tolerable when a human is reading
# the terminal and an outage when the poller ran this at 04:00.
PREVIOUS_IMAGE=""
PREVIOUS_CID=$(dc ps -q app 2>/dev/null | head -n1 || true)
if [ -n "$PREVIOUS_CID" ]; then
PREVIOUS_IMAGE=$(sudo docker inspect -f '{{.Image}}' "$PREVIOUS_CID" 2>/dev/null || true)
fi
# Restore the previous image and exit non-zero. Deliberately NOT a database
# rollback: migrations are additive, so the previous code runs against the new
# schema, and the dump taken above is the escape hatch for the case where it
# does not. The re-tag makes $IMAGE_REF point back at the old image locally; the
# next successful pull moves it forward again.
roll_back() {
echo " !! $1"
dc logs app --tail 40 || true
if [ -z "$PREVIOUS_IMAGE" ]; then
echo " NO PREVIOUS IMAGE TO ROLL BACK TO — the broken release is live" >&2
exit "$EXIT_STILL_LIVE"
fi
echo "==> Rolling back to $PREVIOUS_IMAGE"
sudo docker tag "$PREVIOUS_IMAGE" "$IMAGE_REF"
dc up -d --no-build app
for _ in $(seq 1 60); do
if curl -sf http://127.0.0.1:8920/api/health > /dev/null; then break; fi
sleep 1
done
if curl -sf http://127.0.0.1:8920/api/health | grep -q '"ok":true'; then
echo " rolled back, previous release is healthy"
exit 1
fi
# The restore ran but did not come up. Reporting "rolled back" here would tell
# the on-call the previous release is serving when nothing is.
echo " ROLLBACK ALSO UNHEALTHY — this host needs a human" >&2
exit "$EXIT_STILL_LIVE"
}
echo "==> Starting the app"
dc up -d app
echo "==> Waiting for health"
for _ in $(seq 1 60); do
if curl -sf http://127.0.0.1:8920/api/health > /dev/null; then break; fi
sleep 1
done
echo "==> Verifying"
if curl -sf http://127.0.0.1:8920/api/health | grep -q '"ok":true'; then
echo " health ok"
else
echo " HEALTH CHECK FAILED"
roll_back "health check failed"
fi
# Authentication must be enforced. A deploy that accidentally serves the CRM
# unauthenticated is the one failure worth blocking on.
CODE=$(curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:8920/api/dashboard)
if [ "$CODE" != "401" ]; then
echo " UNAUTHENTICATED REQUEST RETURNED $CODE, EXPECTED 401"
roll_back "unauthenticated request returned $CODE, expected 401"
fi
echo " auth enforced"
# Everything above proves the container is well. It proves nothing about what
# the public actually gets, and there is a failure on this host that every
# check so far passes: a Caddy site block missing `bind 10.0.0.2` lands in a
# separate server on :443, wins for traffic arriving on that address, and
# answers with a valid certificate, HTTP 200 and an EMPTY body. The app is
# fine; the request never reached it. So ask the origin, from outside the
# compose network, and insist on seeing something the app actually renders.
PUBLIC_URL="${PIG_DEPLOY_PUBLIC_URL:-$(env_value PIG_PUBLIC_URL)}"
PUBLIC_URL="${PUBLIC_URL:-https://primeintellectgrowth.com}"
PUBLIC_MARKER="${PIG_DEPLOY_PUBLIC_MARKER:-<div id=\"root\">}"
echo "==> Verifying the public origin ($PUBLIC_URL)"
# Three outcomes, and conflating any two of them has already cost a deploy.
# `|| true` used to fold "curl never got an answer" into the empty string, so a
# host where the public name does not hairpin, or where egress to 443 is
# filtered, failed a completely healthy release — and autodeploy.sh then
# blacklists that digest for good. Keep the transport result separate from the
# body.
PUBLIC_ERR=$(mktemp)
trap 'rm -f "$PUBLIC_ERR"' EXIT
if PUBLIC_BODY=$(curl -sS --max-time 20 --location "${PUBLIC_URL%/}/" 2>"$PUBLIC_ERR"); then
CURL_STATUS=0
else
CURL_STATUS=$?
fi
PUBLIC_BYTES=${#PUBLIC_BODY}
if [ "$CURL_STATUS" -ne 0 ]; then
# Could not connect at all: DNS, TLS, no hairpin, filtered egress. This says
# nothing about the release, and nothing about the proxy either — from this
# host the two are indistinguishable. Do not fail a deploy that passed every
# check that can actually be trusted; set PIG_DEPLOY_REQUIRE_PUBLIC=1 on a
# host where the public name IS reachable from here and this must block.
echo " could not reach the public origin (curl exit $CURL_STATUS)" >&2
sed 's/^/ /' "$PUBLIC_ERR" >&2
if [ "${PIG_DEPLOY_REQUIRE_PUBLIC:-0}" = '1' ]; then
echo " PIG_DEPLOY_REQUIRE_PUBLIC=1, so this is fatal." >&2
exit "$EXIT_STILL_LIVE"
fi
echo " Not treated as a failure: the container passed every local gate," >&2
echo " and this host may simply not be able to route to its own name." >&2
echo " Verify $PUBLIC_URL from somewhere else." >&2
elif [ "$PUBLIC_BYTES" -eq 0 ]; then
echo " PUBLIC ORIGIN ANSWERED WITH AN EMPTY BODY" >&2
echo " Something terminated TLS and answered, so this is the proxy, not" >&2
echo " the release. Check that every Caddy site block on this host has" >&2
echo " 'bind 10.0.0.2'." >&2
# Deliberately no rollback: the previous image would fail this same check.
# Rolling back would churn the service and not fix a proxy fault. The release
# under test therefore stays live, which is what EXIT_STILL_LIVE tells the
# poller to say.
exit "$EXIT_STILL_LIVE"
elif ! printf '%s' "$PUBLIC_BODY" | grep -qF "$PUBLIC_MARKER"; then
echo " PUBLIC ORIGIN ANSWERED $PUBLIC_BYTES BYTES WITHOUT '$PUBLIC_MARKER'" >&2
# This one DOES roll back, unlike the empty-body case above. A proxy fault
# cannot serve a wrong-but-non-empty page for this hostname; a bad release
# can — a changed Vite `base`, a Dockerfile that stopped copying
# apps/web/dist. Every earlier gate passes in that state (the API is well,
# /api/dashboard still 401s) and only this one fires, so if it is the release
# then refusing to roll back leaves the broken release serving the public.
# Restoring the previous image is right when it is, and harmless when it is
# not: the marker check is the last gate, so the rollback costs one restart.
roll_back "public origin answered without '$PUBLIC_MARKER'"
else
echo " public origin serving PIG ($PUBLIC_BYTES bytes)"
fi
echo "==> Deployed $AFTER"