Bring the Swarm+Forgejo layer in line with the canonical topology in docs/runbooks/implementation-handoff.md: three nodes (prod, beta, and the desktop LAN dev machine) instead of two, with the Forgejo Actions runner registered on the desktop as a plain systemd service rather than a Swarm-scheduled Docker-in-Docker sidecar. - setup-wireguard.sh: full N-peer mesh instead of point-to-point - docker-stack.yml: drop the runner/runner-dind services, volumes, and runner-token secret; only forgejo_db_password remains - register-lan-runner.sh: register under both docker and thermograph-lan labels, since one runner now covers both job types - init-swarm.sh, join-swarm.sh, swarm/README.md, forgejo/README.md, INFRA.md: updated node lists, order of operations, and access-state table for three nodes - deploy.yml: renamed to "Deploy to beta VPS" and reconciled the main-vs-release branch question against terraform.tfvars.example (this workflow already targets beta; a release-triggered prod deploy doesn't exist yet and isn't invented here)
165 lines
9.1 KiB
Markdown
165 lines
9.1 KiB
Markdown
# Agent access, Docker Swarm, and Forgejo CI/CD
|
||
|
||
This covers three additive infrastructure layers on top of the app
|
||
containerization / Postgres / Terraform work described in
|
||
**`docs/runbooks/implementation-handoff.md`** (read that first — it's the
|
||
canonical Track A/Track B plan this whole effort implements) and
|
||
`docs/architecture/repo-topology-and-infrastructure.md`. Terraform provisions
|
||
**prod** (the new 48 GB / 12-core box, `thermograph.org`) and **beta** (the
|
||
old VPS, `75.119.132.91`) — see `terraform/README.md` /
|
||
`terraform.tfvars.example`. **None of this touches `backend/`, `Dockerfile`,
|
||
`docker-compose*.yml`, `terraform/`, or `deploy/db/`** — those stay owned by
|
||
the app-containerization work; this is strictly the layer above it.
|
||
|
||
```
|
||
Track 1: deploy/provision-agent-access.sh — a dedicated full-root login for me
|
||
Track 2: deploy/swarm/ — Swarm cluster spanning THREE nodes
|
||
Track 3: deploy/forgejo/ + .forgejo/ — Forgejo + Forgejo Actions, replacing GitHub
|
||
```
|
||
|
||
**Three nodes, not two**: prod, beta, and **the desktop** (the LAN dev
|
||
machine — same box that already runs the pre-Forgejo GitHub self-hosted
|
||
runner). This matches the canonical Track B design in the handoff doc; an
|
||
earlier revision of this whole effort covered just prod+beta and has been
|
||
realigned.
|
||
|
||
## Track 1 — Agent access
|
||
|
||
Run `sudo bash deploy/provision-agent-access.sh` on **prod and beta** (not
|
||
the desktop — that's wherever you're already working from, no separate
|
||
access-provisioning step needed there). Creates a dedicated `agent` user (not
|
||
raw root login — a distinct name gives a clean audit trail) with passwordless
|
||
sudo, installs the agent's public key, disables SSH password auth
|
||
fleet-wide, and turns on `auditd` logging of every root-effective command.
|
||
See the script's header comment for the full rationale and revocation steps
|
||
(one line to delete, or delete the whole account — your own access is never
|
||
affected).
|
||
|
||
### Current access state (live)
|
||
|
||
Both VPS boxes are provisioned and reachable as of this writing:
|
||
|
||
| Host | Role | Public IP | Login |
|
||
|------|------|-----------|-------|
|
||
| prod | Swarm manager, Thermograph's eventual home | `169.58.46.181` | `agent` |
|
||
| beta | Swarm worker, old VPS | `75.119.132.91` | `agent` |
|
||
| desktop | Swarm worker, LAN dev machine + Actions runner | (local) | (already have access) |
|
||
|
||
```
|
||
ssh -i ~/.ssh/thermograph_agent_ed25519 agent@169.58.46.181 # prod
|
||
ssh -i ~/.ssh/thermograph_agent_ed25519 agent@75.119.132.91 # beta
|
||
```
|
||
|
||
The private half of the agent's dedicated keypair lives at
|
||
`~/.ssh/thermograph_agent_ed25519` on the operator's own machine (never in
|
||
this repo, never in a CI secret) — same convention `DEPLOY.md` already uses
|
||
for the separate CI deploy key (`~/.ssh/thermograph_ci`). If you're setting
|
||
this up fresh elsewhere, generate a new pair the same way
|
||
(`ssh-keygen -t ed25519 -a 100 -C "claude-agent@thermograph-infra" -f
|
||
~/.ssh/thermograph_agent_ed25519 -N ""`), embed the new public half in
|
||
`AGENT_PUBKEY` at the top of `provision-agent-access.sh`, and re-run it on
|
||
both boxes — that replaces the authorized key wholesale rather than
|
||
appending, so an old key stops working the moment the new one lands.
|
||
|
||
**Keep the private key somewhere that survives** — not a scratch/temp
|
||
directory that gets cleaned up, not anywhere-in-repo. Losing it just means
|
||
regenerating and re-running the provisioning script; it does not lock either
|
||
box, since your own login is untouched.
|
||
|
||
As of this writing, neither VPS has Docker installed yet and Terraform has
|
||
not been applied to either (`prod` is a bare OS install; `beta` has an old
|
||
`/opt/thermograph` checkout but no Docker) — Track 2 is paused on that. The
|
||
desktop already has Docker (used for the LAN dev containerized stack) but
|
||
hasn't joined anything yet (`docker info` shows `Swarm: inactive`).
|
||
|
||
## Track 2 — Swarm
|
||
|
||
See `deploy/swarm/README.md` for the exact order of operations across all
|
||
**three** nodes (WireGuard mesh first, then swarm init/join, then lock the
|
||
Swarm ports down to the tunnel interface, then label beta for Forgejo
|
||
placement). This cluster's only job is hosting Forgejo — it does not
|
||
orchestrate the Terraform-managed app deploys.
|
||
|
||
## Track 3 — Forgejo, replacing GitHub
|
||
|
||
### 3a. Stand up Forgejo
|
||
See `deploy/forgejo/README.md` — deploys the stack (pinned to beta), then
|
||
walks through registering the Actions runner **on the desktop** as a plain
|
||
systemd service (`register-lan-runner.sh`), not as a Swarm-scheduled
|
||
container. Only one Swarm secret to mint now (`forgejo_db_password`) — the
|
||
runner token is no longer a Swarm secret, since the runner isn't a Swarm
|
||
service anymore.
|
||
|
||
### 3b. Migrate the repo (mirror first, verify, then cut over)
|
||
1. In Forgejo: **+ New Migration → GitHub**. Point it at
|
||
`https://github.com/griffemi/thermograph`, pull code + issues + PRs +
|
||
releases + labels. This is non-destructive to GitHub — do this as a mirror
|
||
and leave GitHub live and untouched until you're confident.
|
||
2. Spot-check: branch list matches, a few PRs/issues render correctly, LFS (if
|
||
any) came across.
|
||
3. **Don't retarget any secret or disable a GitHub workflow yet** — see 3d.
|
||
|
||
### 3c. Workflows
|
||
`.forgejo/workflows/{build,pr-build,deploy-dev,deploy}.yml` mirror
|
||
`.github/workflows/*.yml`, copied with the mechanical changes Forgejo needs:
|
||
- `runs-on: ubuntu-latest` → `runs-on: docker` (one of the two labels the
|
||
desktop's runner registers under — see 3a; general CI/deploy jobs get a
|
||
fresh container, the LAN-deploy job in `deploy-dev.yml` runs bare/host-native
|
||
under the `thermograph-lan` label instead, since it needs real filesystem
|
||
access).
|
||
- `appleboy/ssh-action` referenced by full GitHub URL (not mirrored in
|
||
Forgejo's default action registry, unlike `actions/checkout` /
|
||
`actions/setup-python`, which resolve unchanged).
|
||
- **The custom auto-merge workflow step is gone.** It existed only because
|
||
GitHub's native branch protection/auto-merge is paywalled on private
|
||
free-tier repos. Forgejo has no such tier — configure it directly:
|
||
**Settings → Branches → protect `dev` → Enable Status Check → list `build`
|
||
→ merge style: squash.** Then a ready PR gets a real "Auto merge when checks
|
||
succeed" button.
|
||
- One knock-on simplification: a native auto-merge is an ordinary git push
|
||
(unlike GitHub's token-based merge commit, which doesn't re-trigger `push`),
|
||
so `deploy-dev.yml`'s push trigger fires naturally after auto-merge — no
|
||
more "direct push vs. PR-merge" double-trigger logic to maintain.
|
||
- `.forgejo/workflows/deploy.yml` ("Deploy to beta VPS") triggers on `main`,
|
||
same as it always did. This was flagged as a possible mismatch against
|
||
`terraform.tfvars.example`'s `release` branch for prod, but turned out not
|
||
to be one: this workflow predates the prod/beta split, so its secrets almost
|
||
certainly already point at beta (`main` is beta's branch per Terraform) —
|
||
see the file's own header comment. A separate `release` → prod deploy path
|
||
doesn't exist anywhere yet; don't invent one speculatively, confirm with
|
||
whoever owns Terraform whether prod even wants a push-triggered deploy or
|
||
is meant to go through `terraform apply` only.
|
||
|
||
A second, independent set of Forgejo workflows
|
||
(`.forgejo/workflows/{ci,build-push,ops-cron}.yml`) was added by a parallel
|
||
effort implementing `docs/runbooks/implementation-handoff.md` Track A chunk
|
||
7 — see that doc and its own reconciliation PR (which already fixed its
|
||
runner-label and ssh-action guesses to match the real infrastructure here).
|
||
`ci.yml` duplicated this PR's `build.yml` exactly and was dropped in that
|
||
reconciliation; `build-push.yml` (image → Forgejo's built-in registry) and
|
||
`ops-cron.yml` (backup + IndexNow) are novel and still present.
|
||
|
||
### 3d. Cut over
|
||
Only after 3a–3c are verified:
|
||
1. Move the deploy secrets (`SSH_HOST`, `SSH_USER`, `SSH_KEY`, `SSH_PORT`) into
|
||
Forgejo's repo Secrets (Settings → Actions → Secrets).
|
||
2. Re-point the LAN dev runner: `bash deploy/forgejo/register-lan-runner.sh
|
||
<forgejo_url> <token>` — stops/disables the old GitHub runner service on
|
||
that same machine and registers `forgejo-runner` in its place, same
|
||
`systemd --user` pattern as before (see `DEPLOY-DEV.md` for the original;
|
||
this supersedes it for the runner specifically).
|
||
3. Open a test PR against `dev` in Forgejo, confirm the required check runs,
|
||
auto-merge fires, and the LAN deploy lands.
|
||
4. Only once that's solid: disable/archive the `.github/workflows/*.yml`
|
||
triggers (or the whole GitHub repo) — GitHub was the fallback during this
|
||
whole migration; retire it deliberately, not by accident.
|
||
|
||
## Verification checklist
|
||
- [x] `ssh agent@<prod>` and `ssh agent@<beta>` both work; `sudo whoami` → `root`
|
||
- [x] Password SSH auth confirmed dead on both boxes (tested with an actual
|
||
failed login attempt, not just config inspection)
|
||
- [ ] `docker node ls` (from prod) shows both nodes `Ready`
|
||
- [ ] Swarm ports (2377/7946/4789) unreachable from outside the WireGuard tunnel
|
||
- [ ] `https://git.thermograph.org` (or your chosen domain) serves Forgejo over TLS
|
||
- [ ] A test PR: required check runs → auto-merges → LAN dev deploy lands
|
||
- [ ] GitHub retired only after all of the above hold
|