thermograph/infra/deploy/forgejo/README.md
Emi Griffith 396b80cf91 forgejo: serve dev.jinemi.com alongside git.thermograph.org
Both hostnames go on a single Caddy site block. That is load-bearing rather
than cosmetic: the /v2/* mesh-only matcher that keeps the OCI registry API
off the public internet (hazard #15) is scoped per block, so a separate
block for the new name would serve the same Forgejo with the registry open.

git.thermograph.org stays canonical — Forgejo has one ROOT_URL and builds
every absolute URL from it, so clone URLs, the OAuth callback and post-login
redirects keep naming the .org host. Flipping FORGEJO_DOMAIN additionally
requires the Google OAuth redirect URI, the registry host baked into image
names and runner labels, and the mesh /etc/hosts pins; README documents
those prerequisites.

Config verified with `caddy validate` and by asserting against the adapted
JSON that the /v2 403 guard matches both hostnames.
2026-08-01 09:01:59 -07:00

12 KiB

Forgejo on the Swarm cluster

Runs as deploy/forgejo/docker-stack.yml — the only Swarm-scheduled workload this cluster carries (prod and beta's app stacks are separate docker stack deploys that happen to run on the manager node, vps2 — see deploy/stack/README context in DEPLOY.md). Pinned to the vps1 node (75.119.132.91) via the role=forge label from deploy/swarm/label-forge-node.sh.

The Actions runner is deliberately not part of this stack: a Swarm-scheduled runner cannot redeploy the Swarm that schedules it, so CI would be unavailable exactly when the cluster is what needs repairing. An earlier revision did run it as a Swarm-scheduled Docker-in-Docker sidecar pinned to the Forgejo node; that is gone.

Runners live in two places, and only one of them counts:

  • runner-vps2/ — the always-on runner, a plain restart: always container on vps2. This is the one CI depends on.
  • register-lan-runner.sh — the desktop's systemd runner. Extra capacity. Nothing may assume it is up. It was the only registered runner in the estate until 2026-08-01, and its 21-hour outage on 2026-07-31 froze every merge and deploy and stopped the nightly backup from firing.

Prerequisites

  1. All three nodes have joined the swarm (deploy/swarm/) and vps1 is labeled role=forge.
  2. docker node ls (from the manager, vps2) shows all three Ready.

One-time setup: Swarm secret

One secret the stack expects to already exist (a Swarm secret, not a file — external: true in the stack file, so docker stack deploy never creates or sees the value, only references it):

# A strong random password for Forgejo's own Postgres (NOT related to
# Thermograph's app database — entirely separate instance/network).
openssl rand -base64 32 | docker secret create forgejo_db_password -

Deploy / update

docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo

Re-running is safe — Swarm only touches services whose spec actually changed. Do this before the DNS + Caddy step below — Caddy's reverse_proxy target (127.0.0.1:3080) needs the forgejo service actually listening first, or its first health check just fails harmlessly until it is.

This stack has no auto-deploy trigger — nothing in .forgejo/workflows/ redeploys it on push. A change to docker-stack.yml only takes effect once someone re-runs docker stack deploy by hand on the manager (vps2).

db/forgejo both carry resources.limits (defaults: db 1 CPU/1g, forgejo 2 CPU/2g — several times observed steady-state usage), overridable with FORGEJO_DB_CPUS/FORGEJO_DB_MEMORY/FORGEJO_CPUS/FORGEJO_MEMORY env vars before docker stack deploy, same convention as the app stack.

DNS + TLS: reusing vps1's existing Caddy, not a second reverse proxy

Forgejo is pinned to vps1 (role=forge) — the same box that also runs Grafana/Loki/Alloy and the emigriffith.dev portfolio site, each fronted by that host's own Caddy. A second ingress (Traefik) trying to bind the same ports would collide with it. So there's no Traefik in this stack: forgejo's web port publishes to 127.0.0.1:3080 only (host-local), and vps1's existing Caddy gets one more site block reverse-proxying to it — same pattern as its other site blocks, same automatic-HTTPS.

  1. Point the Forgejo domain (default git.thermograph.org; override with FORGEJO_DOMAIN=... before docker stack deploy) at vps1's public IP — that's where the task actually runs, not vps2's or the desktop's.
  2. Append deploy/forgejo/caddy-git.conf to vps1's /etc/caddy/Caddyfile, adjusting the domain if you didn't use the default, then systemctl reload caddy.
  3. That file also resolves the registry-exposure hazard (#15 in thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md): /v2/* (the registry API) is blocked to everything except the WireGuard mesh CIDR (10.10.0.0/24); the git/web UI stays public. CI runners and Swarm nodes reach the registry over the mesh, not the public internet — see "Registry access from mesh clients" below.

Migrating the domain (git.thermograph.org -> dev.jinemi.com)

In progress. Both hostnames are served; git.thermograph.org is still the canonical one.

caddy-git.conf lists both names on a single site block. That is deliberate, not cosmetic: the /v2/* mesh-only matcher is per-block, so a second block for dev.jinemi.com would re-expose the registry API publicly and undo hazard #15. Whatever else changes, keep the two names in one block.

Serving a second hostname is the cheap half. Forgejo has exactly one ROOT_URL, and it builds every absolute URL from it, so on the non-canonical host today:

  • clone URLs on repo pages read git.thermograph.org,
  • the OAuth callback and post-login redirect land on git.thermograph.org (login is Google-SSO-only — see docker-stack.yml),
  • webhook payload URLs and notification mail links use git.thermograph.org.

Browsing, git-over-HTTP and the API all work on dev.jinemi.com regardless.

Flipping canonical (FORGEJO_DOMAIN=dev.jinemi.com, then docker stack deploy) needs these first, or it breaks login and CI rather than just renaming things:

  1. Google OAuth — add https://dev.jinemi.com/user/oauth2/<name>/callback as an authorized redirect URI in the Google Cloud console before the flip. Not in this repo; nothing here can verify it. Miss it and no one can log in, including to undo the flip.
  2. The registry host in image names — images are named by registry host, and the git.thermograph.org/ prefix is baked into the runner labels (register-lan-runner.sh, runner-vps2/README.md), the CI-runner image (ci-runner/Dockerfile), REGISTRY_HOST in infra/.env.example, and every already-pushed tag. Renaming the registry host is a separate migration from renaming the web UI; the two need not happen together, and doing the web UI alone is much the smaller change.
  3. Mesh /etc/hosts pins — every node in "Registry access from mesh clients" below pins git.thermograph.org to 10.10.0.2. A node pulling under a new hostname needs its own pin, or Caddy's /v2/* matcher sees a public source IP and returns 403.
  4. Runner instance URL — registered runners hold the instance URL they were registered with. Existing runners keep working against the old name as long as it resolves; a re-registration must use whichever name is canonical then.

Until all four are handled, leaving ROOT_URL on git.thermograph.org is the correct state, not an unfinished one.

Registry access from mesh clients

Any node that needs docker login/push/pull against the registry (the CI runner building/pushing images, any Swarm node pulling them, prod or beta on vps2 pulling app images) must reach git.thermograph.org over the WireGuard tunnel, not vps1's public IP — otherwise Caddy's /v2/* block above refuses the connection. Public DNS resolves the domain to vps1's public IP, so add a /etc/hosts override on each such node pinning it to vps1's WireGuard address instead:

echo "10.10.0.2 git.thermograph.org" | sudo tee -a /etc/hosts

(10.10.0.2 is vps1's WG address per deploy/swarm/README.md's peer numbering — adjust if you assigned it differently.) The git/web UI keeps working normally for everyone else since only /v2/* is restricted.

Register the Actions runner

Once Forgejo answers at its domain:

# On the runner host (see DEPLOY-DEV.md for where that is today):
#   Forgejo web UI -> Site Administration -> Actions -> Runners -> Create new Runner
#   (or, for a repo-scoped runner: <repo> -> Settings -> Actions -> Runners)
#   copy the registration token, then:
bash deploy/forgejo/register-lan-runner.sh https://<forgejo-domain> <token>

See that script's header for exactly what it replaces (the pre-Forgejo GitHub self-hosted runner) and why it registers with two labels where there used to be two separate runners.

config.yaml's runner.capacity is raised from the tool's default of 1 to 8 (override with CAPACITY=) — a single PR push fires pr-build, secrets-guard, and shell-lint simultaneously (3 independent workflows, no needs: between them), so capacity 1 serializes work that could run in parallel, and even capacity 3 (an earlier, undocumented hand-tune) leaves build-backend/build-frontend/validate-observability queued behind those three before they get a slot.

Adding more runner capacity should mean raising this number, or adding a second runner alongside it — not putting a runner on prod or beta (vps2). container.docker_host: automount gives job containers the host's Docker socket; on vps2 that would mean any CI job has root-equivalent access to both the prod and beta stacks running there. An earlier revision of this stack actually ran the runner as a Swarm-hosted container on the box Forgejo was pinned to and was deliberately reverted for this same class of reason — see the note at the top of docker-stack.yml.

Custom CI job image (ci-runner/)

ci-runner/Dockerfile still bases on node:20-bookworm — Node is a hard requirement, not leftover: Forgejo's runner executes actions/checkout@v4 (used by every workflow) as node dist/index.js inside the job container, regardless of whether the workflow itself uses npm/node. (v1 of this image tried a Node-free Debian-slim base and broke every job's checkout step — node: executable file not found — within a minute of going live; reverted immediately.) What it actually fixes: every docker-labeled build-push job currently re-installs the Docker CLI on each run (apt-get install docker.io), which pulls in the classic builder rather than BuildKit (the classic builder mishandles COPY --chown=<name> group resolution — a real bug hit during the frontend Go rewrite). ci-runner adds docker-ce-cli + docker-buildx-plugin (BuildKit) on top of the same Node base, plus git/python3/python3-yaml for the other jobs that need them (shell-lint, observability-validate).

Current tag: git.thermograph.org/emi/thermograph/ci-runner:v2 (v1 is broken — do not register any runner against it). Rebuild/push (requires a PAT with write:package scope — the embedded git-remote token lacks it, same requirement documented in build-push.yml):

docker build -t git.thermograph.org/emi/thermograph/ci-runner:vN deploy/forgejo/ci-runner
docker push git.thermograph.org/emi/thermograph/ci-runner:vN

register-lan-runner.sh's LABELS default points at the current tag, so fresh registrations pick it up automatically. The live runner is cut over by editing the labels array in ~/forgejo-runner/.runner on the runner host (same runner id/token, no re-registration needed) and restarting the service — verify a real job runs green under the new image before relying on it, same way v1's break was caught. Only after that verification should the now-redundant apt-get install docker.io / python3-yaml steps be removed from the workflows that had them — removing them first would break every job still running on the stock node:20-bookworm image.

Why Postgres here and not the Thermograph app's TimescaleDB

Separate instance, separate network (forgejo_net, not the app's overlay network), separate volume. Forgejo is a distinct product with its own schema and its own backup/restore lifecycle — sharing a database with the app would couple two things that should be able to fail, migrate, and restore independently.

Verifying

docker service ls                       # forgejo_db, forgejo_forgejo both Running, 1/1
curl -I https://git.thermograph.org/    # 200, valid cert (Caddy's, not a new one)
curl -I https://git.thermograph.org/v2/ # 403 from anywhere off the WireGuard mesh
# On the runner host, after registering the runner:
systemctl --user status forgejo-runner   # active, both labels registered

Rollback / removal

docker stack rm forgejo
# volumes (forgejo_data, forgejo_db, ...) survive a stack rm — remove them
# explicitly only if you actually want to destroy the Forgejo instance's data.