# Agent access, Docker Swarm, and Forgejo CI/CD This covers three additive infrastructure layers on top of the app containerization / Postgres / Terraform work described in **`docs/runbooks/implementation-handoff.md`** (read that first — it's the canonical Track A/Track B plan this whole effort implements) and `docs/architecture/repo-topology-and-infrastructure.md`. Terraform provisions **prod** (the new 48 GB / 12-core box, `thermograph.org`) and **beta** (the old VPS, `75.119.132.91`) — see `terraform/README.md` / `terraform.tfvars.example`. **None of this touches `backend/`, `Dockerfile`, `docker-compose*.yml`, `terraform/`, or `deploy/db/`** — those stay owned by the app-containerization work; this is strictly the layer above it. ``` Track 1: deploy/provision-agent-access.sh — a dedicated full-root login for me Track 2: deploy/swarm/ — Swarm cluster spanning THREE nodes Track 3: deploy/forgejo/ + .forgejo/ — Forgejo + Forgejo Actions, replacing GitHub ``` **Three nodes, not two**: prod, beta, and **the desktop** (the LAN dev machine — same box that already runs the pre-Forgejo GitHub self-hosted runner). This matches the canonical Track B design in the handoff doc; an earlier revision of this whole effort covered just prod+beta and has been realigned. ## Track 1 — Agent access Run `sudo bash deploy/provision-agent-access.sh` on **prod and beta** (not the desktop — that's wherever you're already working from, no separate access-provisioning step needed there). Creates a dedicated `agent` user (not raw root login — a distinct name gives a clean audit trail) with passwordless sudo, installs the agent's public key, disables SSH password auth fleet-wide, and turns on `auditd` logging of every root-effective command. See the script's header comment for the full rationale and revocation steps (one line to delete, or delete the whole account — your own access is never affected). ### Current access state (live) Both VPS boxes are provisioned and reachable as of this writing: | Host | Role | Public IP | Login | |------|------|-----------|-------| | prod | Swarm manager, Thermograph's eventual home | `169.58.46.181` | `agent` | | beta | Swarm worker, old VPS | `75.119.132.91` | `agent` | | desktop | Swarm worker, LAN dev machine + Actions runner | (local) | (already have access) | ``` ssh -i ~/.ssh/thermograph_agent_ed25519 agent@169.58.46.181 # prod ssh -i ~/.ssh/thermograph_agent_ed25519 agent@75.119.132.91 # beta ``` The private half of the agent's dedicated keypair lives at `~/.ssh/thermograph_agent_ed25519` on the operator's own machine (never in this repo, never in a CI secret) — same convention `DEPLOY.md` already uses for the separate CI deploy key (`~/.ssh/thermograph_ci`). If you're setting this up fresh elsewhere, generate a new pair the same way (`ssh-keygen -t ed25519 -a 100 -C "claude-agent@thermograph-infra" -f ~/.ssh/thermograph_agent_ed25519 -N ""`), embed the new public half in `AGENT_PUBKEY` at the top of `provision-agent-access.sh`, and re-run it on both boxes — that replaces the authorized key wholesale rather than appending, so an old key stops working the moment the new one lands. **Keep the private key somewhere that survives** — not a scratch/temp directory that gets cleaned up, not anywhere-in-repo. Losing it just means regenerating and re-running the provisioning script; it does not lock either box, since your own login is untouched. As of this writing, neither VPS has Docker installed yet and Terraform has not been applied to either (`prod` is a bare OS install; `beta` has an old `/opt/thermograph` checkout but no Docker) — Track 2 is paused on that. The desktop already has Docker (used for the LAN dev containerized stack) but hasn't joined anything yet (`docker info` shows `Swarm: inactive`). ## Track 2 — Swarm See `deploy/swarm/README.md` for the exact order of operations across all **three** nodes (WireGuard mesh first, then swarm init/join, then lock the Swarm ports down to the tunnel interface, then label beta for Forgejo placement). This cluster's only job is hosting Forgejo — it does not orchestrate the Terraform-managed app deploys. ## Track 3 — Forgejo, replacing GitHub ### 3a. Stand up Forgejo See `deploy/forgejo/README.md` — deploys the stack (pinned to beta), then walks through registering the Actions runner **on the desktop** as a plain systemd service (`register-lan-runner.sh`), not as a Swarm-scheduled container. Only one Swarm secret to mint now (`forgejo_db_password`) — the runner token is no longer a Swarm secret, since the runner isn't a Swarm service anymore. ### 3b. Migrate the repo (mirror first, verify, then cut over) 1. In Forgejo: **+ New Migration → GitHub**. Point it at `https://github.com/griffemi/thermograph`, pull code + issues + PRs + releases + labels. This is non-destructive to GitHub — do this as a mirror and leave GitHub live and untouched until you're confident. 2. Spot-check: branch list matches, a few PRs/issues render correctly, LFS (if any) came across. 3. **Don't retarget any secret or disable a GitHub workflow yet** — see 3d. ### 3c. Workflows `.forgejo/workflows/{build,pr-build,deploy-dev,deploy}.yml` mirror `.github/workflows/*.yml`, copied with the mechanical changes Forgejo needs: - `runs-on: ubuntu-latest` → `runs-on: docker` (one of the two labels the desktop's runner registers under — see 3a; general CI/deploy jobs get a fresh container, the LAN-deploy job in `deploy-dev.yml` runs bare/host-native under the `thermograph-lan` label instead, since it needs real filesystem access). - `appleboy/ssh-action` referenced by full GitHub URL (not mirrored in Forgejo's default action registry, unlike `actions/checkout` / `actions/setup-python`, which resolve unchanged). - **The custom auto-merge workflow step is gone.** It existed only because GitHub's native branch protection/auto-merge is paywalled on private free-tier repos. Forgejo has no such tier — configure it directly: **Settings → Branches → protect `dev` → Enable Status Check → list `build` → merge style: squash.** Then a ready PR gets a real "Auto merge when checks succeed" button. - One knock-on simplification: a native auto-merge is an ordinary git push (unlike GitHub's token-based merge commit, which doesn't re-trigger `push`), so `deploy-dev.yml`'s push trigger fires naturally after auto-merge — no more "direct push vs. PR-merge" double-trigger logic to maintain. - `.forgejo/workflows/deploy.yml` ("Deploy to beta VPS") triggers on `main`, same as it always did. This was flagged as a possible mismatch against `terraform.tfvars.example`'s `release` branch for prod, but turned out not to be one: this workflow predates the prod/beta split, so its secrets almost certainly already point at beta (`main` is beta's branch per Terraform) — see the file's own header comment. A separate `release` → prod deploy path doesn't exist anywhere yet; don't invent one speculatively, confirm with whoever owns Terraform whether prod even wants a push-triggered deploy or is meant to go through `terraform apply` only. A second, independent set of Forgejo workflows (`.forgejo/workflows/{ci,build-push,ops-cron}.yml`) was added by a parallel effort implementing `docs/runbooks/implementation-handoff.md` Track A chunk 7 — see that doc and its own reconciliation PR (which already fixed its runner-label and ssh-action guesses to match the real infrastructure here). `ci.yml` duplicated this PR's `build.yml` exactly and was dropped in that reconciliation; `build-push.yml` (image → Forgejo's built-in registry) and `ops-cron.yml` (backup + IndexNow) are novel and still present. ### 3d. Cut over Only after 3a–3c are verified: 1. Move the deploy secrets (`SSH_HOST`, `SSH_USER`, `SSH_KEY`, `SSH_PORT`) into Forgejo's repo Secrets (Settings → Actions → Secrets). 2. Re-point the LAN dev runner: `bash deploy/forgejo/register-lan-runner.sh ` — stops/disables the old GitHub runner service on that same machine and registers `forgejo-runner` in its place, same `systemd --user` pattern as before (see `DEPLOY-DEV.md` for the original; this supersedes it for the runner specifically). 3. Open a test PR against `dev` in Forgejo, confirm the required check runs, auto-merge fires, and the LAN deploy lands. 4. Only once that's solid: disable/archive the `.github/workflows/*.yml` triggers (or the whole GitHub repo) — GitHub was the fallback during this whole migration; retire it deliberately, not by accident. ## Verification checklist - [x] `ssh agent@` and `ssh agent@` both work; `sudo whoami` → `root` - [x] Password SSH auth confirmed dead on both boxes (tested with an actual failed login attempt, not just config inspection) - [ ] `docker node ls` (from prod) shows both nodes `Ready` - [ ] Swarm ports (2377/7946/4789) unreachable from outside the WireGuard tunnel - [ ] `https://git.thermograph.org` (or your chosen domain) serves Forgejo over TLS - [ ] A test PR: required check runs → auto-merges → LAN dev deploy lands - [ ] GitHub retired only after all of the above hold