Bring the Swarm+Forgejo layer in line with the canonical topology in docs/runbooks/implementation-handoff.md: three nodes (prod, beta, and the desktop LAN dev machine) instead of two, with the Forgejo Actions runner registered on the desktop as a plain systemd service rather than a Swarm-scheduled Docker-in-Docker sidecar. - setup-wireguard.sh: full N-peer mesh instead of point-to-point - docker-stack.yml: drop the runner/runner-dind services, volumes, and runner-token secret; only forgejo_db_password remains - register-lan-runner.sh: register under both docker and thermograph-lan labels, since one runner now covers both job types - init-swarm.sh, join-swarm.sh, swarm/README.md, forgejo/README.md, INFRA.md: updated node lists, order of operations, and access-state table for three nodes - deploy.yml: renamed to "Deploy to beta VPS" and reconciled the main-vs-release branch question against terraform.tfvars.example (this workflow already targets beta; a release-triggered prod deploy doesn't exist yet and isn't invented here)
3.4 KiB
Forgejo on the Swarm cluster
Runs as deploy/forgejo/docker-stack.yml — the only Swarm-scheduled workload
this cluster carries (the Thermograph app itself stays on the
Terraform-managed docker compose deploys; see terraform/README.md).
Pinned to the beta node (old VPS) via the role=forge label from
deploy/swarm/label-forge-node.sh.
The Actions runner is deliberately not part of this stack — it runs on
the desktop as a plain systemd service (register-lan-runner.sh below),
per docs/runbooks/implementation-handoff.md Track B step 5. That's the
canonical placement; an earlier revision of this stack ran the runner as a
Swarm-scheduled Docker-in-Docker sidecar pinned to beta, which is gone now.
Prerequisites
- All three nodes have joined the swarm (
deploy/swarm/) and beta is labeledrole=forge. docker node ls(from the manager) shows all threeReady.
One-time setup: Swarm secret
One secret the stack expects to already exist (a Swarm secret, not a file —
external: true in the stack file, so docker stack deploy never creates or
sees the value, only references it):
# A strong random password for Forgejo's own Postgres (NOT related to
# Thermograph's app database — entirely separate instance/network).
openssl rand -base64 32 | docker secret create forgejo_db_password -
Deploy / update
docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo
Re-running is safe — Swarm only touches services whose spec actually changed.
Register the Actions runner (on the desktop, not through Swarm)
Once Forgejo answers at its domain:
# On the desktop:
# Forgejo web UI -> Site Administration -> Actions -> Runners -> Create new Runner
# (or, for a repo-scoped runner: <repo> -> Settings -> Actions -> Runners)
# copy the registration token, then:
bash deploy/forgejo/register-lan-runner.sh https://<forgejo-domain> <token>
See that script's header for exactly what it replaces (the pre-Forgejo GitHub self-hosted runner on this same machine) and why it registers with two labels where there used to be two separate runners.
DNS
Point the Forgejo domain (default git.thermograph.org; override with
FORGEJO_DOMAIN=... before docker stack deploy, Swarm reads it from the
deploying shell's environment) at prod's or beta's public IP — not the
desktop's, which isn't a stable publicly-reachable address. The routing mesh
forwards published ports to wherever the task actually landed, so either VPS
IP works.
Why Postgres here and not the Thermograph app's TimescaleDB
Separate instance, separate network (forgejo_net, not the app's compose
network), separate volume. Forgejo is a distinct product with its own schema
and its own backup/restore lifecycle — sharing a database with the app would
couple two things that should be able to fail, migrate, and restore
independently.
Verifying
docker service ls # forgejo_db, forgejo_forgejo, forgejo_traefik all Running, 1/1
curl -I https://git.thermograph.org/ # 200, valid cert
# On the desktop, after registering the runner:
systemctl --user status forgejo-runner # active, both labels registered
Rollback / removal
docker stack rm forgejo
# volumes (forgejo_data, forgejo_db, ...) survive a stack rm — remove them
# explicitly only if you actually want to destroy the Forgejo instance's data.