thermograph/infra/deploy/forgejo
Emi Griffith fc455a0473
All checks were successful
PR build (required check) / changes (pull_request) Successful in 9s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
secrets-guard / encrypted (pull_request) Successful in 7s
shell-lint / shellcheck (pull_request) Successful in 9s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
ci: collapse the eight deploy and build-push workflows into two
The six deploy workflows were one file written six times, differing only in a
branch name, a paths filter, a concurrency group, the service name, its
*_IMAGE_TAG variable and a secret prefix. The two build-push workflows were the
same file twice, differing only in the domain string. The contract into
infra/deploy/deploy.sh (SERVICE + BACKEND_IMAGE_TAG/FRONTEND_IMAGE_TAG) was
already fully parameterised, so the duplication bought nothing and cost eight
files to keep in step.

deploy.yml: branch selects the environment (main -> beta, release -> prod), a
matrix covers backend and frontend, and each leg decides whether this push
actually touched its domain before rolling anything. The workflow-level paths
filter only says backend OR frontend moved; without the per-leg refinement a
backend-only push would also roll the frontend and lose the independent-deploy
property the FE/BE split exists for.

build-push.yml: the same shape for images. Images stay separate and
independently deployable; nothing about the published artefacts changes.

The two *-deploy-dev.yml workflows are deleted rather than ported. They were
already documented as inert -- they call a monorepo path on the LAN box whose
~/thermograph-dev is still a split-era thermograph-infra checkout. LAN dev is a
local `make dev-up` concern, not a CI environment.

Deliberately boring expressions throughout. No dynamic matrix (fromJSON), and
the *_IMAGE_TAG selection is done in shell rather than with a
`matrix.service == 'x' && a || b` ternary. Those are GitHub idioms a Forgejo/act
runner may evaluate differently, and the failure mode is silent: an empty
*_IMAGE_TAG makes deploy.sh fall back to the tag already running, so the job
goes green having deployed nothing. Beta and prod get two explicit, mutually
exclusive steps rather than a ternary over secrets, where an empty host would be
worse still.

Everything load-bearing is preserved: fetch-depth 0 and the domain-keyed 12-hex
tag (the branch tip is often another domain's commit), per-service per-ref
concurrency with cancel-in-progress false, separate PROD_SSH_* credentials, the
v*.*.* both-images exception, and appleboy/ssh-action by full URL.

pr-build.yml is untouched. Its workflow name and `gate` job are the required
status check branch protection is configured against, and renaming that context
makes every PR unmergeable with fully green CI. secrets-guard.yml and
shell-lint.yml are likewise left alone: they are the only other consolidation
candidates that run on pull_request, and the branch-protection API needs
credentials this could not read, so merging them was not worth the risk for two
files.

Logic verified by replaying the plan step against real history: an infra-only
push rolls nothing, a backend-only push rolls backend and leaves frontend
alone, and a run with no usable before-sha defaults to deploying rather than
silently skipping. The computed frontend tag for a backend-only push is
sha-ca84e0ce95f0 -- exactly the tag live on beta -- so the derivation matches
what the old workflows produced.

15 workflows -> 9, 1365 lines -> 1019. Docs naming the deleted files are
updated in the same commit, per the rule the root CLAUDE.md now carries.
2026-07-25 00:47:41 -07:00
..
ci-runner web/worker: add a process-level liveness heartbeat (#80) 2026-07-25 04:13:47 +00:00
caddy-git.conf Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
docker-stack.yml web/worker: add a process-level liveness heartbeat (#80) 2026-07-25 04:13:47 +00:00
README.md ci: collapse the eight deploy and build-push workflows into two 2026-07-25 00:47:41 -07:00
register-lan-runner.sh infra/forgejo: give the LAN runner's PATH ~/.local/bin (#84) 2026-07-25 06:59:21 +00:00

Forgejo on the Swarm cluster

Runs as deploy/forgejo/docker-stack.yml — the only Swarm-scheduled workload this cluster carries (the Thermograph app itself stays on the Terraform-managed docker compose deploys; see terraform/README.md). Pinned to the beta node (old VPS) via the role=forge label from deploy/swarm/label-forge-node.sh.

The Actions runner is deliberately not part of this stack — it runs on the desktop as a plain systemd service (register-lan-runner.sh below), per thermograph-docs/runbooks/implementation-handoff.md Track B step 5. That's the canonical placement; an earlier revision of this stack ran the runner as a Swarm-scheduled Docker-in-Docker sidecar pinned to beta, which is gone now.

Prerequisites

  1. All three nodes have joined the swarm (deploy/swarm/) and beta is labeled role=forge.
  2. docker node ls (from the manager) shows all three Ready.

One-time setup: Swarm secret

One secret the stack expects to already exist (a Swarm secret, not a file — external: true in the stack file, so docker stack deploy never creates or sees the value, only references it):

# A strong random password for Forgejo's own Postgres (NOT related to
# Thermograph's app database — entirely separate instance/network).
openssl rand -base64 32 | docker secret create forgejo_db_password -

Deploy / update

docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo

Re-running is safe — Swarm only touches services whose spec actually changed. Do this before the DNS + Caddy step below — Caddy's reverse_proxy target (127.0.0.1:3080) needs the forgejo service actually listening first, or its first health check just fails harmlessly until it is.

This stack has no auto-deploy trigger — nothing in .forgejo/workflows/ redeploys it on push. A change to docker-stack.yml only takes effect once someone re-runs docker stack deploy by hand on the manager (prod).

db/forgejo both carry resources.limits (defaults: db 1 CPU/1g, forgejo 2 CPU/2g — several times observed steady-state usage), overridable with FORGEJO_DB_CPUS/FORGEJO_DB_MEMORY/FORGEJO_CPUS/FORGEJO_MEMORY env vars before docker stack deploy, same convention as the app stack.

DNS + TLS: reusing beta's existing Caddy, not a second reverse proxy

Forgejo is pinned to beta (role=forge) — but beta is also today's live thermograph.org host, and its Caddy already owns ports 80/443 (/etc/caddy/Caddyfile on that box). A second ingress (Traefik) trying to bind the same ports would collide with it. So there's no Traefik in this stack: forgejo's web port publishes to 127.0.0.1:3080 only (host-local), and beta's existing Caddy gets one more site block reverse-proxying to it — same pattern as its thermograph.org block, same automatic-HTTPS.

  1. Point the Forgejo domain (default git.thermograph.org; override with FORGEJO_DOMAIN=... before docker stack deploy) at beta's public IP — that's where the task actually runs, not prod's or the desktop's.
  2. Append deploy/forgejo/caddy-git.conf to beta's /etc/caddy/Caddyfile, adjusting the domain if you didn't use the default, then systemctl reload caddy.
  3. That file also resolves the registry-exposure hazard (#15 in thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md): /v2/* (the registry API) is blocked to everything except the WireGuard mesh CIDR (10.10.0.0/24); the git/web UI stays public. CI runners and Swarm nodes reach the registry over the mesh, not the public internet — see "Registry access from mesh clients" below.

Registry access from mesh clients

Any node that needs docker login/push/pull against the registry (the desktop's CI runner building/pushing images, later any Swarm node pulling them) must reach git.thermograph.org over the WireGuard tunnel, not beta's public IP — otherwise Caddy's /v2/* block above refuses the connection. Public DNS resolves the domain to beta's public IP, so add a /etc/hosts override on each such node pinning it to beta's WireGuard address instead:

echo "10.10.0.2 git.thermograph.org" | sudo tee -a /etc/hosts

(10.10.0.2 is beta's WG address per deploy/swarm/README.md's peer numbering — adjust if you assigned it differently.) The git/web UI keeps working normally for everyone else since only /v2/* is restricted.

Register the Actions runner (on the desktop, not through Swarm)

Once Forgejo answers at its domain:

# On the desktop:
#   Forgejo web UI -> Site Administration -> Actions -> Runners -> Create new Runner
#   (or, for a repo-scoped runner: <repo> -> Settings -> Actions -> Runners)
#   copy the registration token, then:
bash deploy/forgejo/register-lan-runner.sh https://<forgejo-domain> <token>

See that script's header for exactly what it replaces (the pre-Forgejo GitHub self-hosted runner on this same machine) and why it registers with two labels where there used to be two separate runners.

config.yaml's runner.capacity is raised from the tool's default of 1 to 8 (override with CAPACITY=) — a single PR push fires pr-build, secrets-guard, and shell-lint simultaneously (3 independent workflows, no needs: between them), so capacity 1 serializes work that could run in parallel, and even capacity 3 (an earlier, undocumented hand-tune) leaves build-backend/build-frontend/validate-observability queued behind those three before they get a slot. The desktop has 16 cores / 34GB free today; 8 concurrent jobs is comfortable headroom without starving LAN dev's own compose stack.

Adding more runner capacity should mean raising this number, or adding a second runner on the desktop itself — not putting a runner on prod or beta. container.docker_host: automount gives job containers the host's Docker socket; on prod or beta that would mean any CI job has root-equivalent access to whatever else is running there (the live app stack, or Forgejo itself). An earlier revision of this stack actually did run the runner as a Swarm-hosted container on beta and was deliberately reverted to the desktop for this reason — see the note at the top of docker-stack.yml.

Custom CI job image (ci-runner/)

ci-runner/Dockerfile still bases on node:20-bookworm — Node is a hard requirement, not leftover: Forgejo's runner executes actions/checkout@v4 (used by every workflow) as node dist/index.js inside the job container, regardless of whether the workflow itself uses npm/node. (v1 of this image tried a Node-free Debian-slim base and broke every job's checkout step — node: executable file not found — within a minute of going live; reverted immediately.) What it actually fixes: every docker-labeled build-push job currently re-installs the Docker CLI on each run (apt-get install docker.io), which pulls in the classic builder rather than BuildKit (the classic builder mishandles COPY --chown=<name> group resolution — a real bug hit during the frontend Go rewrite). ci-runner adds docker-ce-cli + docker-buildx-plugin (BuildKit) on top of the same Node base, plus git/python3/python3-yaml for the other jobs that need them (shell-lint, observability-validate).

Current tag: git.thermograph.org/emi/thermograph/ci-runner:v2 (v1 is broken — do not register any runner against it). Rebuild/push (requires a PAT with write:package scope — the embedded git-remote token lacks it, same requirement documented in build-push.yml):

docker build -t git.thermograph.org/emi/thermograph/ci-runner:vN deploy/forgejo/ci-runner
docker push git.thermograph.org/emi/thermograph/ci-runner:vN

register-lan-runner.sh's LABELS default points at the current tag, so fresh registrations pick it up automatically. The live runner is cut over by editing the labels array in ~/forgejo-runner/.runner on the desktop (same runner id/token, no re-registration needed) and restarting the service — verify a real job runs green under the new image before relying on it, same way v1's break was caught. Only after that verification should the now-redundant apt-get install docker.io / python3-yaml steps be removed from the workflows that had them — removing them first would break every job still running on the stock node:20-bookworm image.

Why Postgres here and not the Thermograph app's TimescaleDB

Separate instance, separate network (forgejo_net, not the app's compose network), separate volume. Forgejo is a distinct product with its own schema and its own backup/restore lifecycle — sharing a database with the app would couple two things that should be able to fail, migrate, and restore independently.

Verifying

docker service ls                       # forgejo_db, forgejo_forgejo both Running, 1/1
curl -I https://git.thermograph.org/    # 200, valid cert (Caddy's, not a new one)
curl -I https://git.thermograph.org/v2/ # 403 from anywhere off the WireGuard mesh
# On the desktop, after registering the runner:
systemctl --user status forgejo-runner   # active, both labels registered

Rollback / removal

docker stack rm forgejo
# volumes (forgejo_data, forgejo_db, ...) survive a stack rm — remove them
# explicitly only if you actually want to destroy the Forgejo instance's data.