thermograph/deploy/swarm
Emi Griffith 0d8fc9f4d0 Add agent VPS access, a 2-node Docker Swarm, and Forgejo CI/CD (#234)
Three additive infrastructure layers on top of the two VPS boxes Terraform
already provisions (prod: new 48 GB/12-core box, thermograph.org; beta: old
VPS, 75.119.132.91). None of this touches backend/, Dockerfile,
docker-compose*.yml, terraform/, or deploy/db/ — that stays owned by the
app-containerization work in flight elsewhere; this is strictly the layer on
top. See INFRA.md for the full runbook and order of operations.

- deploy/provision-agent-access.sh: a dedicated, auditable full-sudo login
  (not raw root) for agent-driven ops — passwordless sudo under a distinct
  username, sshd hardened to key-only auth, auditd logging every
  root-effective command. One line to revoke.

- deploy/swarm/: a 2-node Swarm (prod=manager, beta=worker) joined over a
  WireGuard tunnel rather than trusting the public internet with the
  overlay data plane, which Docker's own guidance says should never face it
  directly. Swarm ports locked to the tunnel interface once joined. This
  cluster's only workload is Forgejo — it does not orchestrate the
  Terraform-managed app deploys, so nothing here can strand the app's
  single-writer database.

- deploy/forgejo/ + .forgejo/workflows/: Forgejo + Traefik + a
  Docker-in-Docker-sandboxed runner as a Swarm stack pinned to beta, plus
  Forgejo Actions workflows mirroring .github/workflows/*.yml. The custom
  auto-merge workflow step is dropped — it existed only to work around
  GitHub's paywalled branch protection on private free-tier repos, which
  Forgejo has no such tier for; native "auto merge when checks succeed"
  replaces it, and as a real git push (unlike GitHub's non-triggering
  token-merge) it fires the LAN deploy naturally with no double-trigger
  logic needed. appleboy/ssh-action is referenced by full URL (not mirrored
  on Forgejo's default action registry); actions/checkout and
  actions/setup-python resolve unchanged.

Migration is mirror-first: the GitHub repo import and workflow files land
here, but cutting deploy secrets over and retiring GitHub happens only after
verification (INFRA.md 3d) — GitHub stays live as a fallback throughout.

One flagged, unresolved mismatch: deploy.yml still triggers on `main`, but
terraform.tfvars.example names prod's deploy branch `release`. Left as a
faithful mirror rather than guessed at — reconcile with whoever's driving
Terraform/deploy.
2026-07-21 00:36:39 +00:00
..
firewall-swarm.sh Add agent VPS access, a 2-node Docker Swarm, and Forgejo CI/CD (#234) 2026-07-21 00:36:39 +00:00
init-swarm.sh Add agent VPS access, a 2-node Docker Swarm, and Forgejo CI/CD (#234) 2026-07-21 00:36:39 +00:00
join-swarm.sh Add agent VPS access, a 2-node Docker Swarm, and Forgejo CI/CD (#234) 2026-07-21 00:36:39 +00:00
label-forge-node.sh Add agent VPS access, a 2-node Docker Swarm, and Forgejo CI/CD (#234) 2026-07-21 00:36:39 +00:00
README.md Add agent VPS access, a 2-node Docker Swarm, and Forgejo CI/CD (#234) 2026-07-21 00:36:39 +00:00
setup-wireguard.sh Add agent VPS access, a 2-node Docker Swarm, and Forgejo CI/CD (#234) 2026-07-21 00:36:39 +00:00

2-node Docker Swarm (prod + beta), for hosting Forgejo

This Swarm cluster's only job is to run Forgejo (deploy/forgejo/) — it does not orchestrate the Thermograph app itself, which stays on the Terraform-managed docker compose deploys on each box independently (see terraform/README.md). Keeping those separate means nothing here can strand or interfere with the app's already-working, single-writer Postgres/TimescaleDB deploys.

Nodes:

  • manager — prod, the new 48 GB / 12-core box (more headroom).
  • worker — beta, the old VPS (75.119.132.91).

1 manager + 1 worker, not 2 managers: Raft needs 3 nodes for real quorum-based HA, so a second manager here would add complexity (split-brain risk) without adding actual failover. If the manager goes down, the worker keeps running whatever was already scheduled on it (Forgejo, since it's pinned there) but the cluster can't reschedule anything until the manager's back — acceptable for a 2-box hobby/small-team cluster whose only job is CI/CD.

Order of operations

  1. Agent access first (deploy/provision-agent-access.sh) on both boxes — everything below is run through that access.
  2. WireGuard tunnel (setup-wireguard.sh) — run on both boxes; see the script's header for the two-pass key-exchange dance. Verify with ping <peer_wg_ip> before continuing.
  3. Swarm init (init-swarm.sh <manager_wg_ip>) on the manager (prod) only.
  4. Swarm join (join-swarm.sh <manager_wg_ip> <token>) on the worker (beta) only, using the token init-swarm.sh printed.
  5. Firewall lockdown (firewall-swarm.sh) on both boxes — closes 2377/7946/4789 to everything except the WireGuard interface. Do this after joining is confirmed working, not before (locking the ports first would make the join itself fail).
  6. Label beta (label-forge-node.sh <beta-node-name>) on the manager — docker node ls shows beta's node name/ID.
  7. Deploy Forgejo: see deploy/forgejo/README.md.

Why WireGuard instead of relying on Swarm's built-in TLS alone

Swarm's control plane (port 2377) is TLS-encrypted and mutually authenticated by default. Its overlay data plane (VXLAN, port 4789) is not encrypted by default, and Docker's own guidance is that port must never face the public internet — these two boxes are on different providers' public IPs, not a private LAN, so the tunnel is the network boundary the Swarm ports advertise into, rather than trusting the public internet directly.

Verifying

# On the manager:
docker node ls                      # both nodes Ready
docker node inspect <beta-node> --format '{{.Spec.Labels}}'   # role:forge

# From outside the WireGuard interface (e.g. your own machine), confirm the
# Swarm ports are NOT reachable on the public IP:
nc -zv -w2 <public_ip> 2377   # should fail/timeout
nc -zvu -w2 <public_ip> 4789  # should fail/timeout

Adding a node label back out (undo)

docker node update --label-rm role <beta-node>