thermograph/INFRA.md
Emi Griffith 7ff3873661 Realign Swarm/Forgejo infra to the 3-node design (#243)
Bring the Swarm+Forgejo layer in line with the canonical topology in
docs/runbooks/implementation-handoff.md: three nodes (prod, beta, and
the desktop LAN dev machine) instead of two, with the Forgejo Actions
runner registered on the desktop as a plain systemd service rather than
a Swarm-scheduled Docker-in-Docker sidecar.

- setup-wireguard.sh: full N-peer mesh instead of point-to-point
- docker-stack.yml: drop the runner/runner-dind services, volumes, and
  runner-token secret; only forgejo_db_password remains
- register-lan-runner.sh: register under both docker and
  thermograph-lan labels, since one runner now covers both job types
- init-swarm.sh, join-swarm.sh, swarm/README.md, forgejo/README.md,
  INFRA.md: updated node lists, order of operations, and access-state
  table for three nodes
- deploy.yml: renamed to "Deploy to beta VPS" and reconciled the
  main-vs-release branch question against terraform.tfvars.example
  (this workflow already targets beta; a release-triggered prod
  deploy doesn't exist yet and isn't invented here)
2026-07-21 01:43:19 +00:00

9.1 KiB
Raw Blame History

Agent access, Docker Swarm, and Forgejo CI/CD

This covers three additive infrastructure layers on top of the app containerization / Postgres / Terraform work described in docs/runbooks/implementation-handoff.md (read that first — it's the canonical Track A/Track B plan this whole effort implements) and docs/architecture/repo-topology-and-infrastructure.md. Terraform provisions prod (the new 48 GB / 12-core box, thermograph.org) and beta (the old VPS, 75.119.132.91) — see terraform/README.md / terraform.tfvars.example. None of this touches backend/, Dockerfile, docker-compose*.yml, terraform/, or deploy/db/ — those stay owned by the app-containerization work; this is strictly the layer above it.

Track 1: deploy/provision-agent-access.sh   — a dedicated full-root login for me
Track 2: deploy/swarm/                      — Swarm cluster spanning THREE nodes
Track 3: deploy/forgejo/ + .forgejo/         — Forgejo + Forgejo Actions, replacing GitHub

Three nodes, not two: prod, beta, and the desktop (the LAN dev machine — same box that already runs the pre-Forgejo GitHub self-hosted runner). This matches the canonical Track B design in the handoff doc; an earlier revision of this whole effort covered just prod+beta and has been realigned.

Track 1 — Agent access

Run sudo bash deploy/provision-agent-access.sh on prod and beta (not the desktop — that's wherever you're already working from, no separate access-provisioning step needed there). Creates a dedicated agent user (not raw root login — a distinct name gives a clean audit trail) with passwordless sudo, installs the agent's public key, disables SSH password auth fleet-wide, and turns on auditd logging of every root-effective command. See the script's header comment for the full rationale and revocation steps (one line to delete, or delete the whole account — your own access is never affected).

Current access state (live)

Both VPS boxes are provisioned and reachable as of this writing:

Host Role Public IP Login
prod Swarm manager, Thermograph's eventual home 169.58.46.181 agent
beta Swarm worker, old VPS 75.119.132.91 agent
desktop Swarm worker, LAN dev machine + Actions runner (local) (already have access)
ssh -i ~/.ssh/thermograph_agent_ed25519 agent@169.58.46.181   # prod
ssh -i ~/.ssh/thermograph_agent_ed25519 agent@75.119.132.91    # beta

The private half of the agent's dedicated keypair lives at ~/.ssh/thermograph_agent_ed25519 on the operator's own machine (never in this repo, never in a CI secret) — same convention DEPLOY.md already uses for the separate CI deploy key (~/.ssh/thermograph_ci). If you're setting this up fresh elsewhere, generate a new pair the same way (ssh-keygen -t ed25519 -a 100 -C "claude-agent@thermograph-infra" -f ~/.ssh/thermograph_agent_ed25519 -N ""), embed the new public half in AGENT_PUBKEY at the top of provision-agent-access.sh, and re-run it on both boxes — that replaces the authorized key wholesale rather than appending, so an old key stops working the moment the new one lands.

Keep the private key somewhere that survives — not a scratch/temp directory that gets cleaned up, not anywhere-in-repo. Losing it just means regenerating and re-running the provisioning script; it does not lock either box, since your own login is untouched.

As of this writing, neither VPS has Docker installed yet and Terraform has not been applied to either (prod is a bare OS install; beta has an old /opt/thermograph checkout but no Docker) — Track 2 is paused on that. The desktop already has Docker (used for the LAN dev containerized stack) but hasn't joined anything yet (docker info shows Swarm: inactive).

Track 2 — Swarm

See deploy/swarm/README.md for the exact order of operations across all three nodes (WireGuard mesh first, then swarm init/join, then lock the Swarm ports down to the tunnel interface, then label beta for Forgejo placement). This cluster's only job is hosting Forgejo — it does not orchestrate the Terraform-managed app deploys.

Track 3 — Forgejo, replacing GitHub

3a. Stand up Forgejo

See deploy/forgejo/README.md — deploys the stack (pinned to beta), then walks through registering the Actions runner on the desktop as a plain systemd service (register-lan-runner.sh), not as a Swarm-scheduled container. Only one Swarm secret to mint now (forgejo_db_password) — the runner token is no longer a Swarm secret, since the runner isn't a Swarm service anymore.

3b. Migrate the repo (mirror first, verify, then cut over)

  1. In Forgejo: + New Migration → GitHub. Point it at https://github.com/griffemi/thermograph, pull code + issues + PRs + releases + labels. This is non-destructive to GitHub — do this as a mirror and leave GitHub live and untouched until you're confident.
  2. Spot-check: branch list matches, a few PRs/issues render correctly, LFS (if any) came across.
  3. Don't retarget any secret or disable a GitHub workflow yet — see 3d.

3c. Workflows

.forgejo/workflows/{build,pr-build,deploy-dev,deploy}.yml mirror .github/workflows/*.yml, copied with the mechanical changes Forgejo needs:

  • runs-on: ubuntu-latestruns-on: docker (one of the two labels the desktop's runner registers under — see 3a; general CI/deploy jobs get a fresh container, the LAN-deploy job in deploy-dev.yml runs bare/host-native under the thermograph-lan label instead, since it needs real filesystem access).
  • appleboy/ssh-action referenced by full GitHub URL (not mirrored in Forgejo's default action registry, unlike actions/checkout / actions/setup-python, which resolve unchanged).
  • The custom auto-merge workflow step is gone. It existed only because GitHub's native branch protection/auto-merge is paywalled on private free-tier repos. Forgejo has no such tier — configure it directly: Settings → Branches → protect dev → Enable Status Check → list build → merge style: squash. Then a ready PR gets a real "Auto merge when checks succeed" button.
  • One knock-on simplification: a native auto-merge is an ordinary git push (unlike GitHub's token-based merge commit, which doesn't re-trigger push), so deploy-dev.yml's push trigger fires naturally after auto-merge — no more "direct push vs. PR-merge" double-trigger logic to maintain.
  • .forgejo/workflows/deploy.yml ("Deploy to beta VPS") triggers on main, same as it always did. This was flagged as a possible mismatch against terraform.tfvars.example's release branch for prod, but turned out not to be one: this workflow predates the prod/beta split, so its secrets almost certainly already point at beta (main is beta's branch per Terraform) — see the file's own header comment. A separate release → prod deploy path doesn't exist anywhere yet; don't invent one speculatively, confirm with whoever owns Terraform whether prod even wants a push-triggered deploy or is meant to go through terraform apply only.

A second, independent set of Forgejo workflows (.forgejo/workflows/{ci,build-push,ops-cron}.yml) was added by a parallel effort implementing docs/runbooks/implementation-handoff.md Track A chunk 7 — see that doc and its own reconciliation PR (which already fixed its runner-label and ssh-action guesses to match the real infrastructure here). ci.yml duplicated this PR's build.yml exactly and was dropped in that reconciliation; build-push.yml (image → Forgejo's built-in registry) and ops-cron.yml (backup + IndexNow) are novel and still present.

3d. Cut over

Only after 3a3c are verified:

  1. Move the deploy secrets (SSH_HOST, SSH_USER, SSH_KEY, SSH_PORT) into Forgejo's repo Secrets (Settings → Actions → Secrets).
  2. Re-point the LAN dev runner: bash deploy/forgejo/register-lan-runner.sh <forgejo_url> <token> — stops/disables the old GitHub runner service on that same machine and registers forgejo-runner in its place, same systemd --user pattern as before (see DEPLOY-DEV.md for the original; this supersedes it for the runner specifically).
  3. Open a test PR against dev in Forgejo, confirm the required check runs, auto-merge fires, and the LAN deploy lands.
  4. Only once that's solid: disable/archive the .github/workflows/*.yml triggers (or the whole GitHub repo) — GitHub was the fallback during this whole migration; retire it deliberately, not by accident.

Verification checklist

  • ssh agent@<prod> and ssh agent@<beta> both work; sudo whoamiroot
  • Password SSH auth confirmed dead on both boxes (tested with an actual failed login attempt, not just config inspection)
  • docker node ls (from prod) shows both nodes Ready
  • Swarm ports (2377/7946/4789) unreachable from outside the WireGuard tunnel
  • https://git.thermograph.org (or your chosen domain) serves Forgejo over TLS
  • A test PR: required check runs → auto-merges → LAN dev deploy lands
  • GitHub retired only after all of the above hold