# Continuous integration. # # Runs on every push and pull request. The job is deliberately one sequence # rather than a fan-out: this is a small project, the whole thing takes a # couple of minutes, and a single log is easier to read than five. # # What it actually proves, in order of how likely each is to catch something: # # 1. Every package typechecks. # 2. The migration chain applies to a REAL, empty Postgres. This has already # caught one migration that Drizzle generated but Postgres refused # (a jsonb -> integer cast with no USING clause). # 3. The seed is idempotent — running it twice leaves the same row counts. # This caught a seed that silently duplicated 27 contacts. # 4. The unit tests pass. # 5. The server boots against that database and answers. # 6. Piggy boots against that same database, answers /internal/health, and # the API — wired to it through the environment, not through a stub — # reports it enabled. The relay's own tests inject a resolver, so they # stay green whether or not the real wiring exists; only this step reads # it. A crash on boot and an unset PIGGY_INTERNAL_URL look identical from # the browser: the dock simply never appears. # 7. The front end builds, and the CSP hash for the inline theme script still # matches what the proxy is configured to allow. Editing that script # changes its hash, and the failure mode is a silent white flash for # dark-mode users rather than an error. # 8. docker-compose.yml renders, and the piggy service is passed every # environment key the worker's schema requires. That is the one failure # nothing else here can see, because it lives between two files that are # each individually correct. # # THE CSP HASH IS DUPLICATED IN THREE PLACES: the `expected` constant below, # `deploy/Caddyfile.example`, and the LIVE Caddyfile on cloud-2. Only the first # two are checked by anything. The live one is the copy that actually decides # whether a browser runs the script, and nothing in this repository can see it, # so changing the script means editing all three by hand — see deploy/README.md. # # Shipping is a two-step, and the second step is a human: # # push to main -> `verify` only. Nothing is published, nothing deploys. # tag release-* -> `verify`, then `publish` pushes the image to the Gitea # registry. The production host notices it and deploys. # # So the tag IS the ship decision. No credential on this runner can reach # cloud-2; the host pulls, the runner never pushes to it. name: CI on: push: branches: [main] # A tag push runs the same verification and then, and only then, publishes. tags: ['release-*'] pull_request: jobs: verify: runs-on: ubuntu-latest # Postgres is started as a step rather than through `services:`, because # this runner is configured with `container.network: host`. # # That single setting explains three failed attempts, and is worth writing # down so nobody repeats them: # # `services:` Service containers are not resolvable by # name from a host-networked job, giving # "getaddrinfo EAI_AGAIN postgres". # `--network container:$HOSTNAME` /etc/hostname reports the HOST's name, # not a container id, so the namespace # join finds no such container. # default-gateway addressing The wrong idea entirely: with host # networking the default route is the # real router, not a docker bridge. # # Because the job shares the host's network namespace, a published port is # simply on 127.0.0.1. The port is derived from the run id so concurrent # runs cannot collide. env: PG_CONTAINER: pig-ci-pg-${{ github.run_id }} steps: - uses: actions/checkout@v4 - uses: actions/setup-node@v4 with: node-version: '22' # Corepack ships with Node and installs the exact pnpm pinned by # `packageManager` in package.json, so CI, the image and a laptop all run # the same version. `--activate` puts it on PATH; the download prompt is # disabled because a non-interactive runner cannot answer it and would # otherwise hang until the job times out. - name: Enable pnpm env: COREPACK_ENABLE_DOWNLOAD_PROMPT: '0' run: | corepack enable corepack prepare --activate pnpm --version - name: Start Postgres run: | PG_PORT=$(( 45000 + (${{ github.run_id }} % 15000) )) echo "Publishing Postgres on 127.0.0.1:${PG_PORT}" docker rm -f "$PG_CONTAINER" 2>/dev/null || true docker run -d --name "$PG_CONTAINER" \ -p "127.0.0.1:${PG_PORT}:5432" \ -e POSTGRES_USER=pig -e POSTGRES_PASSWORD=pig -e POSTGRES_DB=pig \ postgres:16-alpine # pg_isready inside the container only proves the server started. # What matters is that THIS job can reach it through the published # port, so the readiness check is made from here, over TCP. for i in $(seq 1 60); do if node -e " const net=require('net'); const s=net.connect(${PG_PORT},'127.0.0.1'); s.on('connect',()=>{s.end();process.exit(0)}); s.on('error',()=>process.exit(1)); " 2>/dev/null; then echo "Reachable after ${i}s" echo "DATABASE_URL=postgres://pig:pig@127.0.0.1:${PG_PORT}/pig" >> "$GITHUB_ENV" exit 0 fi sleep 1 done echo "Postgres never became reachable on 127.0.0.1:${PG_PORT}" docker logs "$PG_CONTAINER" 2>&1 | tail -30 exit 1 - name: Install # --frozen-lockfile fails rather than quietly resolving a different # tree when the lockfile and manifests disagree. That is the whole # point of committing a lockfile, and it is the default in CI anyway — # stated here so it survives someone running this locally. run: pnpm install --frozen-lockfile - name: Typecheck every package run: pnpm run typecheck - name: Unit tests run: pnpm run test - name: Migrations apply to a real Postgres run: pnpm exec tsx packages/db/src/migrate.ts - name: Migrations are re-runnable run: pnpm exec tsx packages/db/src/migrate.ts - name: Seed is idempotent # A seed that duplicates on a second run corrupts any database it is # pointed at twice, and nobody notices until the counts look odd. # # Two tables are counted, not one. This gate only ever watched # `contacts`, and `contacts` is idempotent by an explicit existence # check — so it could not see the class of regression it exists to # catch, which is `onConflictDoNothing` firing at a constraint that is # no longer there. The Motion starter library relies on exactly that # clause against `motion_templates_slug_version_key`, so it is counted # here too. Any table whose idempotency rests on a conflict target # belongs in this list. run: | pnpm exec tsx packages/db/src/seed/index.ts > /dev/null count() { docker exec "$PG_CONTAINER" psql -U pig -d pig -tAc "select count(*) from $1"; } CONTACTS_BEFORE=$(count contacts) TEMPLATES_BEFORE=$(count motion_templates) pnpm exec tsx packages/db/src/seed/index.ts > /dev/null CONTACTS_AFTER=$(count contacts) TEMPLATES_AFTER=$(count motion_templates) echo "contacts: $CONTACTS_BEFORE -> $CONTACTS_AFTER" echo "motion_templates: $TEMPLATES_BEFORE -> $TEMPLATES_AFTER" test "$CONTACTS_BEFORE" = "$CONTACTS_AFTER" || { echo "SEED IS NOT IDEMPOTENT"; exit 1; } test "$TEMPLATES_BEFORE" = "$TEMPLATES_AFTER" || { echo "SEED IS NOT IDEMPOTENT"; exit 1; } - name: Critical path E2E against Postgres and Hono run: pnpm run test:e2e - name: Server boots and answers run: | NODE_ENV=development PIG_PORT=8930 pnpm exec tsx apps/api/src/server.ts & for i in $(seq 1 30); do curl -sf http://127.0.0.1:8930/api/health && break sleep 1 done curl -sf http://127.0.0.1:8930/api/health | grep -q '"ok":true' - name: Piggy boots, and the API reports it enabled # apps/api's own comment admits the gap this closes: its tests inject a # resolver, so they pass whether or not the process is really wired to a # Piggy. Here the relay is given nothing but environment variables and # has to reach a Piggy that actually booted. # # It runs after the seed on purpose: with no identity provider every # request is the development user, and that user is a seeded row. run: | LOGS=$(mktemp -d) # Derived from the run id for the same reason Postgres's port is: this # job shares the host's network namespace, so a fixed port belongs to # the whole machine and two concurrent runs would fight over it. PIGGY_PORT=$(( 30000 + (${{ github.run_id }} % 5000) )) API_PORT=$(( 36000 + (${{ github.run_id }} % 5000) )) # Worthless, and long enough for the schema's 32-character minimum. INTERNAL_TOKEN='piggy-ci-internal-token-0123456789' PIGGY_PID='' API_PID='' # There are two processes between the job's pid and the server that # holds the port — pnpm launches tsx, tsx launches node — so the whole # descendant tree has to go. Verified by watching a plain `kill` leave # a Piggy behind, still holding its Postgres connections. # # SIGKILL, not the polite signal: nothing here needs a clean shutdown, # and a server still listening when the next step runs is worse than # an abrupt one. stop() { for pid in "$@"; do [ -n "$pid" ] || continue for child in $(pgrep -P "$pid" 2>/dev/null); do stop "$child"; done kill -9 "$pid" 2>/dev/null || true done } trap 'stop "$PIGGY_PID" "$API_PID"' EXIT # Nothing here calls a model: the task queue is empty and the status # route never reaches one. The inference base points at the discard # port so that a future version which DID call out would fail loudly # rather than quietly billing somebody's real endpoint. PIGGY_INFERENCE_API_KEY=ci-stub-key \ PIGGY_INFERENCE_BASE=http://127.0.0.1:9/v1 \ PIGGY_INTERNAL_TOKEN="$INTERNAL_TOKEN" \ PIGGY_CHAT_HOST=127.0.0.1 \ PIGGY_CHAT_PORT="$PIGGY_PORT" \ pnpm exec tsx apps/piggy/src/main.ts > "$LOGS/piggy.log" 2>&1 & PIGGY_PID=$! for i in $(seq 1 30); do curl -sf "http://127.0.0.1:${PIGGY_PORT}/internal/health" >/dev/null && break sleep 1 done HEALTH=$(curl -sf "http://127.0.0.1:${PIGGY_PORT}/internal/health" || true) echo "GET /internal/health -> ${HEALTH:-}" case "$HEALTH" in *'"ok":true'*) ;; *) echo 'Piggy never answered. Its configuration schema rejects an incomplete environment on start, so the reason is usually the last line here:' tail -30 "$LOGS/piggy.log" exit 1 ;; esac NODE_ENV=development PIG_PORT="$API_PORT" \ PIGGY_ENABLED=true \ PIGGY_INTERNAL_URL="http://127.0.0.1:${PIGGY_PORT}" \ PIGGY_INTERNAL_TOKEN="$INTERNAL_TOKEN" \ pnpm exec tsx apps/api/src/server.ts > "$LOGS/api.log" 2>&1 & API_PID=$! for i in $(seq 1 30); do curl -sf "http://127.0.0.1:${API_PORT}/api/health" >/dev/null && break sleep 1 done # The stored admin toggle is the inner gate, and an earlier step in # this job has already created the settings row with Piggy off — the # insert is ON CONFLICT DO NOTHING, so booting with PIGGY_ENABLED=true # cannot correct it. Flip it here: what is under test is the wiring, # not the switch. docker exec "$PG_CONTAINER" psql -U pig -d pig \ -c 'update platform_settings set piggy_enabled = true' >/dev/null STATUS=$(curl -sf "http://127.0.0.1:${API_PORT}/api/piggy/status" || true) echo "GET /api/piggy/status -> ${STATUS:-}" case "$STATUS" in *'"enabled":true'*) ;; *) echo 'The API does not consider Piggy available, which is what the browser sees as a dock that never appears. PIGGY_ENABLED, PIGGY_INTERNAL_URL and PIGGY_INTERNAL_TOKEN are all read where the routes are composed; one of them is no longer reaching them.' tail -30 "$LOGS/api.log" exit 1 ;; esac - name: Front end builds run: pnpm -F @pig/web run build - name: Inline theme script still matches the deployed CSP hash # The proxy allows exactly one inline script by hash. If the script # changes and the CSP is not updated, dark-mode users get a white flash # on every load and nothing anywhere reports an error. run: | node -e " const fs=require('fs'), crypto=require('crypto'); const html=fs.readFileSync('apps/web/dist/index.html','utf8'); const m=html.match(/