Files
pig/.gitea/workflows/ci.yml
T
karti 13dec6b4b8
CI / verify (push) Successful in 3m45s
CI / publish (push) Has been skipped
Rebuild the shell, add Calendar and Learn, and govern reads
Seven parallel agents and an adversarial verification pass. The three things
worth knowing before reading the diff:

RBAC WAS ALREADY BUILT. docs/build-plan.md marks F2 and F3 outstanding and is
stale — packages/core/src/permissions.ts and lib/mutation.ts shipped long ago.
So this does not rebuild them; it closes the gaps an audit found. The big one
is that reads were entirely ungoverned: every GET was "any authenticated
member", so a junior demand rep and a research contractor could both pull
per-block supplier cost and break-even prices from /api/capacity/margin, and
every contract's negotiated terms. For a company whose margin is the business,
that was the hole that mattered. Adds book:read / economics:read / team:read,
a readGuard middleware, and a `viewer` role below member.

THE BUTTON AND THE 403 DISAGREED — the exact thing F3 said must never happen.
Contracts.tsx never called can() at all, so its save button was always enabled
against a server requiring contract:sign; Capacity.tsx gated commitment
creation on deal:write/demand while the server wanted commitment:write/supply.

POST /api/activities was the one write bypassing executeMutation: no capability
check, and any member could mutate accounts.lastActivityAt as a side effect.
It is now a proper mutation() behind activity:write.

The shell becomes three panes — a collapsible shadcn sidebar with an account
switcher on the Piggy accent, a header with real search, and Piggy docked to
the right, page-aware and persistent across navigation. The phone keeps its
bottom tab bar, which is the thing this product already beat trycompai/crm on,
and gains the sidebar as a sheet.

Calendar is a projection over thirteen dated sources rather than a new table,
because a table would duplicate dates that already live on contracts, deals and
commitments and would drift — and one ledger answering the question is the
whole argument. It surfaces export_authorizations and compliance_artifacts,
which had indexed expires_at columns, schema comments saying they must be
alerted on, and no read endpoint or UI anywhere.

Learn carries two tracks. Concepts are members-only; the platform track can be
opened with a share code by someone with no account. The code mints a scoped
learn-only token and never a Principal — every route here resolves a principal
and then checks capabilities, so a principal-minting code would be one missing
check away from leaking the book. "Only platform-track rows may be code-visible"
is a database CHECK constraint as well as a write-path rule, and a test asserts
a valid learn token still gets 401 on /api/dashboard, /api/accounts and
/api/contracts — the same invariant scripts/deploy.sh refuses to ship without.

CD becomes tag-to-ship. CI publishes an image to the Gitea registry on a
release-* tag and cloud-2 pulls it, so no credential on the shared runner can
execute anything on production — by construction rather than by policy. Both
halves of deploy.sh's original rule survive: nothing on the runner reaches the
host, and a human still decides when it ships. deploy.sh gains a rollback and a
public-origin check, and PIG_IMAGE now reaches compose through `sudo env`,
without which sudo's env_reset silently resolved every release to pig:local.

Tests 141 -> 261.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 15:02:48 -07:00

271 lines
12 KiB
YAML

# Continuous integration.
#
# Runs on every push and pull request. The job is deliberately one sequence
# rather than a fan-out: this is a small project, the whole thing takes a
# couple of minutes, and a single log is easier to read than five.
#
# What it actually proves, in order of how likely each is to catch something:
#
# 1. Every package typechecks.
# 2. The migration chain applies to a REAL, empty Postgres. This has already
# caught one migration that Drizzle generated but Postgres refused
# (a jsonb -> integer cast with no USING clause).
# 3. The seed is idempotent — running it twice leaves the same row counts.
# This caught a seed that silently duplicated 27 contacts.
# 4. The unit tests pass.
# 5. The server boots against that database and answers.
# 6. The front end builds, and the CSP hash for the inline theme script still
# matches what the proxy is configured to allow. Editing that script
# changes its hash, and the failure mode is a silent white flash for
# dark-mode users rather than an error.
#
# THE CSP HASH IS DUPLICATED IN THREE PLACES: the `expected` constant below,
# `deploy/Caddyfile.example`, and the LIVE Caddyfile on cloud-2. Only the first
# two are checked by anything. The live one is the copy that actually decides
# whether a browser runs the script, and nothing in this repository can see it,
# so changing the script means editing all three by hand — see deploy/README.md.
#
# Shipping is a two-step, and the second step is a human:
#
# push to main -> `verify` only. Nothing is published, nothing deploys.
# tag release-* -> `verify`, then `publish` pushes the image to the Gitea
# registry. The production host notices it and deploys.
#
# So the tag IS the ship decision. No credential on this runner can reach
# cloud-2; the host pulls, the runner never pushes to it.
name: CI
on:
push:
branches: [main]
# A tag push runs the same verification and then, and only then, publishes.
tags: ['release-*']
pull_request:
jobs:
verify:
runs-on: ubuntu-latest
# Postgres is started as a step rather than through `services:`, because
# this runner is configured with `container.network: host`.
#
# That single setting explains three failed attempts, and is worth writing
# down so nobody repeats them:
#
# `services:` Service containers are not resolvable by
# name from a host-networked job, giving
# "getaddrinfo EAI_AGAIN postgres".
# `--network container:$HOSTNAME` /etc/hostname reports the HOST's name,
# not a container id, so the namespace
# join finds no such container.
# default-gateway addressing The wrong idea entirely: with host
# networking the default route is the
# real router, not a docker bridge.
#
# Because the job shares the host's network namespace, a published port is
# simply on 127.0.0.1. The port is derived from the run id so concurrent
# runs cannot collide.
env:
PG_CONTAINER: pig-ci-pg-${{ github.run_id }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: '22'
# Corepack ships with Node and installs the exact pnpm pinned by
# `packageManager` in package.json, so CI, the image and a laptop all run
# the same version. `--activate` puts it on PATH; the download prompt is
# disabled because a non-interactive runner cannot answer it and would
# otherwise hang until the job times out.
- name: Enable pnpm
env:
COREPACK_ENABLE_DOWNLOAD_PROMPT: '0'
run: |
corepack enable
corepack prepare --activate
pnpm --version
- name: Start Postgres
run: |
PG_PORT=$(( 45000 + (${{ github.run_id }} % 15000) ))
echo "Publishing Postgres on 127.0.0.1:${PG_PORT}"
docker rm -f "$PG_CONTAINER" 2>/dev/null || true
docker run -d --name "$PG_CONTAINER" \
-p "127.0.0.1:${PG_PORT}:5432" \
-e POSTGRES_USER=pig -e POSTGRES_PASSWORD=pig -e POSTGRES_DB=pig \
postgres:16-alpine
# pg_isready inside the container only proves the server started.
# What matters is that THIS job can reach it through the published
# port, so the readiness check is made from here, over TCP.
for i in $(seq 1 60); do
if node -e "
const net=require('net');
const s=net.connect(${PG_PORT},'127.0.0.1');
s.on('connect',()=>{s.end();process.exit(0)});
s.on('error',()=>process.exit(1));
" 2>/dev/null; then
echo "Reachable after ${i}s"
echo "DATABASE_URL=postgres://pig:pig@127.0.0.1:${PG_PORT}/pig" >> "$GITHUB_ENV"
exit 0
fi
sleep 1
done
echo "Postgres never became reachable on 127.0.0.1:${PG_PORT}"
docker logs "$PG_CONTAINER" 2>&1 | tail -30
exit 1
- name: Install
# --frozen-lockfile fails rather than quietly resolving a different
# tree when the lockfile and manifests disagree. That is the whole
# point of committing a lockfile, and it is the default in CI anyway —
# stated here so it survives someone running this locally.
run: pnpm install --frozen-lockfile
- name: Typecheck every package
run: pnpm run typecheck
- name: Unit tests
run: pnpm run test
- name: Migrations apply to a real Postgres
run: pnpm exec tsx packages/db/src/migrate.ts
- name: Migrations are re-runnable
run: pnpm exec tsx packages/db/src/migrate.ts
- name: Seed is idempotent
# A seed that duplicates on a second run corrupts any database it is
# pointed at twice, and nobody notices until the counts look odd.
run: |
pnpm exec tsx packages/db/src/seed/index.ts > /dev/null
count() { docker exec "$PG_CONTAINER" psql -U pig -d pig -tAc "select count(*) from contacts"; }
BEFORE=$(count)
pnpm exec tsx packages/db/src/seed/index.ts > /dev/null
AFTER=$(count)
echo "contacts: $BEFORE -> $AFTER"
test "$BEFORE" = "$AFTER" || { echo "SEED IS NOT IDEMPOTENT"; exit 1; }
- name: Critical path E2E against Postgres and Hono
run: pnpm run test:e2e
- name: Server boots and answers
run: |
NODE_ENV=development PIG_PORT=8930 pnpm exec tsx apps/api/src/server.ts &
for i in $(seq 1 30); do
curl -sf http://127.0.0.1:8930/api/health && break
sleep 1
done
curl -sf http://127.0.0.1:8930/api/health | grep -q '"ok":true'
- name: Front end builds
run: pnpm -F @pig/web run build
- name: Inline theme script still matches the deployed CSP hash
# The proxy allows exactly one inline script by hash. If the script
# changes and the CSP is not updated, dark-mode users get a white flash
# on every load and nothing anywhere reports an error.
run: |
node -e "
const fs=require('fs'), crypto=require('crypto');
const html=fs.readFileSync('apps/web/dist/index.html','utf8');
const m=html.match(/<script>([\s\S]*?)<\/script>/);
if(!m){ console.error('No inline script found in index.html'); process.exit(1); }
const hash='sha256-'+crypto.createHash('sha256').update(m[1]).digest('base64');
const expected='sha256-1tTDwCq+TCEyPDSZeYqW5HbmP+unUg8hrgRiZBiH/IU=';
if(hash!==expected){
console.error('Inline script hash changed.');
console.error(' now: '+hash);
console.error(' expected: '+expected);
console.error('Update the CSP in deploy/Caddyfile.example AND on the server,');
console.error('then update the expected hash in this workflow.');
process.exit(1);
}
console.log('CSP hash unchanged: '+hash);
"
- name: Docker image builds
run: docker build -t pig:ci .
- name: Stop Postgres
if: always()
run: docker rm -f "$PG_CONTAINER" 2>/dev/null || true
# Publish the image that production will run.
#
# Only on a `release-*` tag. A push to main proves the commit is sound and
# stops there; tagging is the deliberate, human act that says "ship this".
# The production host polls the registry for the newest release tag and
# deploys it (scripts/autodeploy.sh) — which is how this gets automated
# WITHOUT the thing scripts/deploy.sh refuses to do. Nothing here holds a
# credential for cloud-2, and nothing here can execute anything on cloud-2.
#
# `gitea.ref` and `github.ref` are the same object in Gitea Actions; the
# gitea-prefixed spelling is used for the ref test because that is the one
# documented for tag conditions, and github.* elsewhere to match the job
# above.
publish:
needs: verify
if: startsWith(gitea.ref, 'refs/tags/release-')
runs-on: ubuntu-latest
env:
REGISTRY: git.karti.ai
# Gitea namespaces packages under the lowercased owner, so PIG/pig is
# published as pig/pig.
IMAGE: git.karti.ai/pig/pig
# THE POINT OF THIS VARIABLE: the shared act_runner on cloud-1 runs with
# `container.network: host`, and its docker config is visible to jobs
# from every other repository on that host. A plain `docker login` would
# leave a credential in ~/.docker/config.json that any of them could
# read. Pointing DOCKER_CONFIG at a per-run directory keeps the token out
# of the shared file entirely; the logout step below is the second belt.
DOCKER_CONFIG: /tmp/pig-docker-${{ github.run_id }}
steps:
- uses: actions/checkout@v4
- name: Log in to the Gitea registry
# The per-run Actions token, not a long-lived secret: it is minted for
# this run and dies with it. --password-stdin because an argument is
# visible in the runner's process list to anything else on that host.
run: |
mkdir -p "$DOCKER_CONFIG"
printf '%s' '${{ secrets.GITHUB_TOKEN }}' \
| docker login "$REGISTRY" -u '${{ github.actor }}' --password-stdin
- name: Build and push
# Both cloud-1 and cloud-2 are aarch64, so this is a native build and
# needs no --platform. The layer cache from the `verify` job's
# `docker build` is warm on this same daemon, so the rebuild is cheap.
#
# Two tags, always pushed together: the tag is what a human asked for,
# the short sha is what is unambiguous a year later when tags have been
# moved or deleted.
run: |
TAG="${GITHUB_REF#refs/tags/}"
SHORT_SHA=$(printf '%s' "${{ github.sha }}" | cut -c1-7)
echo "Publishing $IMAGE:$TAG and $IMAGE:$SHORT_SHA"
docker build -t "$IMAGE:$TAG" -t "$IMAGE:$SHORT_SHA" .
docker push "$IMAGE:$TAG"
docker push "$IMAGE:$SHORT_SHA"
# Print the digest: it is what the host poller compares against, and
# the only identifier that cannot be reassigned.
docker image inspect "$IMAGE:$TAG" \
--format '{{range .RepoDigests}}{{println .}}{{end}}'
- name: Log out
if: always()
# Runs even when the build failed, because a failed job that left a
# credential behind is exactly the leak this is guarding against.
run: |
docker logout "$REGISTRY" || true
rm -rf "$DOCKER_CONFIG"