Three additive infrastructure layers on top of the two VPS boxes Terraform already provisions (prod: new 48 GB/12-core box, thermograph.org; beta: old VPS, 75.119.132.91). None of this touches backend/, Dockerfile, docker-compose*.yml, terraform/, or deploy/db/ — that stays owned by the app-containerization work in flight elsewhere; this is strictly the layer on top. See INFRA.md for the full runbook and order of operations. - deploy/provision-agent-access.sh: a dedicated, auditable full-sudo login (not raw root) for agent-driven ops — passwordless sudo under a distinct username, sshd hardened to key-only auth, auditd logging every root-effective command. One line to revoke. - deploy/swarm/: a 2-node Swarm (prod=manager, beta=worker) joined over a WireGuard tunnel rather than trusting the public internet with the overlay data plane, which Docker's own guidance says should never face it directly. Swarm ports locked to the tunnel interface once joined. This cluster's only workload is Forgejo — it does not orchestrate the Terraform-managed app deploys, so nothing here can strand the app's single-writer database. - deploy/forgejo/ + .forgejo/workflows/: Forgejo + Traefik + a Docker-in-Docker-sandboxed runner as a Swarm stack pinned to beta, plus Forgejo Actions workflows mirroring .github/workflows/*.yml. The custom auto-merge workflow step is dropped — it existed only to work around GitHub's paywalled branch protection on private free-tier repos, which Forgejo has no such tier for; native "auto merge when checks succeed" replaces it, and as a real git push (unlike GitHub's non-triggering token-merge) it fires the LAN deploy naturally with no double-trigger logic needed. appleboy/ssh-action is referenced by full URL (not mirrored on Forgejo's default action registry); actions/checkout and actions/setup-python resolve unchanged. Migration is mirror-first: the GitHub repo import and workflow files land here, but cutting deploy secrets over and retiring GitHub happens only after verification (INFRA.md 3d) — GitHub stays live as a fallback throughout. One flagged, unresolved mismatch: deploy.yml still triggers on `main`, but terraform.tfvars.example names prod's deploy branch `release`. Left as a faithful mirror rather than guessed at — reconcile with whoever's driving Terraform/deploy.
65 lines
3 KiB
Markdown
65 lines
3 KiB
Markdown
# 2-node Docker Swarm (prod + beta), for hosting Forgejo
|
|
|
|
This Swarm cluster's only job is to run Forgejo (`deploy/forgejo/`) — it does
|
|
**not** orchestrate the Thermograph app itself, which stays on the
|
|
Terraform-managed `docker compose` deploys on each box independently (see
|
|
`terraform/README.md`). Keeping those separate means nothing here can strand
|
|
or interfere with the app's already-working, single-writer Postgres/TimescaleDB
|
|
deploys.
|
|
|
|
**Nodes:**
|
|
- **manager** — prod, the new 48 GB / 12-core box (more headroom).
|
|
- **worker** — beta, the old VPS (`75.119.132.91`).
|
|
|
|
1 manager + 1 worker, not 2 managers: Raft needs 3 nodes for real quorum-based
|
|
HA, so a second manager here would add complexity (split-brain risk) without
|
|
adding actual failover. If the manager goes down, the worker keeps running
|
|
whatever was already scheduled on it (Forgejo, since it's pinned there) but
|
|
the cluster can't reschedule anything until the manager's back — acceptable
|
|
for a 2-box hobby/small-team cluster whose only job is CI/CD.
|
|
|
|
## Order of operations
|
|
|
|
1. **Agent access first** (`deploy/provision-agent-access.sh`) on both boxes —
|
|
everything below is run through that access.
|
|
2. **WireGuard tunnel** (`setup-wireguard.sh`) — run on both boxes; see the
|
|
script's header for the two-pass key-exchange dance. Verify with
|
|
`ping <peer_wg_ip>` before continuing.
|
|
3. **Swarm init** (`init-swarm.sh <manager_wg_ip>`) on the manager (prod) only.
|
|
4. **Swarm join** (`join-swarm.sh <manager_wg_ip> <token>`) on the worker
|
|
(beta) only, using the token `init-swarm.sh` printed.
|
|
5. **Firewall lockdown** (`firewall-swarm.sh`) on **both** boxes — closes
|
|
2377/7946/4789 to everything except the WireGuard interface. Do this
|
|
*after* joining is confirmed working, not before (locking the ports first
|
|
would make the join itself fail).
|
|
6. **Label beta** (`label-forge-node.sh <beta-node-name>`) on the manager —
|
|
`docker node ls` shows beta's node name/ID.
|
|
7. Deploy Forgejo: see `deploy/forgejo/README.md`.
|
|
|
|
## Why WireGuard instead of relying on Swarm's built-in TLS alone
|
|
|
|
Swarm's control plane (port 2377) is TLS-encrypted and mutually authenticated
|
|
by default. Its overlay data plane (VXLAN, port 4789) is **not** encrypted by
|
|
default, and Docker's own guidance is that port must never face the public
|
|
internet — these two boxes are on different providers' public IPs, not a
|
|
private LAN, so the tunnel is the network boundary the Swarm ports advertise
|
|
into, rather than trusting the public internet directly.
|
|
|
|
## Verifying
|
|
|
|
```bash
|
|
# On the manager:
|
|
docker node ls # both nodes Ready
|
|
docker node inspect <beta-node> --format '{{.Spec.Labels}}' # role:forge
|
|
|
|
# From outside the WireGuard interface (e.g. your own machine), confirm the
|
|
# Swarm ports are NOT reachable on the public IP:
|
|
nc -zv -w2 <public_ip> 2377 # should fail/timeout
|
|
nc -zvu -w2 <public_ip> 4789 # should fail/timeout
|
|
```
|
|
|
|
## Adding a node label back out (undo)
|
|
|
|
```bash
|
|
docker node update --label-rm role <beta-node>
|
|
```
|