66 lines
3 KiB
Markdown
66 lines
3 KiB
Markdown
|
|
# 2-node Docker Swarm (prod + beta), for hosting Forgejo
|
||
|
|
|
||
|
|
This Swarm cluster's only job is to run Forgejo (`deploy/forgejo/`) — it does
|
||
|
|
**not** orchestrate the Thermograph app itself, which stays on the
|
||
|
|
Terraform-managed `docker compose` deploys on each box independently (see
|
||
|
|
`terraform/README.md`). Keeping those separate means nothing here can strand
|
||
|
|
or interfere with the app's already-working, single-writer Postgres/TimescaleDB
|
||
|
|
deploys.
|
||
|
|
|
||
|
|
**Nodes:**
|
||
|
|
- **manager** — prod, the new 48 GB / 12-core box (more headroom).
|
||
|
|
- **worker** — beta, the old VPS (`75.119.132.91`).
|
||
|
|
|
||
|
|
1 manager + 1 worker, not 2 managers: Raft needs 3 nodes for real quorum-based
|
||
|
|
HA, so a second manager here would add complexity (split-brain risk) without
|
||
|
|
adding actual failover. If the manager goes down, the worker keeps running
|
||
|
|
whatever was already scheduled on it (Forgejo, since it's pinned there) but
|
||
|
|
the cluster can't reschedule anything until the manager's back — acceptable
|
||
|
|
for a 2-box hobby/small-team cluster whose only job is CI/CD.
|
||
|
|
|
||
|
|
## Order of operations
|
||
|
|
|
||
|
|
1. **Agent access first** (`deploy/provision-agent-access.sh`) on both boxes —
|
||
|
|
everything below is run through that access.
|
||
|
|
2. **WireGuard tunnel** (`setup-wireguard.sh`) — run on both boxes; see the
|
||
|
|
script's header for the two-pass key-exchange dance. Verify with
|
||
|
|
`ping <peer_wg_ip>` before continuing.
|
||
|
|
3. **Swarm init** (`init-swarm.sh <manager_wg_ip>`) on the manager (prod) only.
|
||
|
|
4. **Swarm join** (`join-swarm.sh <manager_wg_ip> <token>`) on the worker
|
||
|
|
(beta) only, using the token `init-swarm.sh` printed.
|
||
|
|
5. **Firewall lockdown** (`firewall-swarm.sh`) on **both** boxes — closes
|
||
|
|
2377/7946/4789 to everything except the WireGuard interface. Do this
|
||
|
|
*after* joining is confirmed working, not before (locking the ports first
|
||
|
|
would make the join itself fail).
|
||
|
|
6. **Label beta** (`label-forge-node.sh <beta-node-name>`) on the manager —
|
||
|
|
`docker node ls` shows beta's node name/ID.
|
||
|
|
7. Deploy Forgejo: see `deploy/forgejo/README.md`.
|
||
|
|
|
||
|
|
## Why WireGuard instead of relying on Swarm's built-in TLS alone
|
||
|
|
|
||
|
|
Swarm's control plane (port 2377) is TLS-encrypted and mutually authenticated
|
||
|
|
by default. Its overlay data plane (VXLAN, port 4789) is **not** encrypted by
|
||
|
|
default, and Docker's own guidance is that port must never face the public
|
||
|
|
internet — these two boxes are on different providers' public IPs, not a
|
||
|
|
private LAN, so the tunnel is the network boundary the Swarm ports advertise
|
||
|
|
into, rather than trusting the public internet directly.
|
||
|
|
|
||
|
|
## Verifying
|
||
|
|
|
||
|
|
```bash
|
||
|
|
# On the manager:
|
||
|
|
docker node ls # both nodes Ready
|
||
|
|
docker node inspect <beta-node> --format '{{.Spec.Labels}}' # role:forge
|
||
|
|
|
||
|
|
# From outside the WireGuard interface (e.g. your own machine), confirm the
|
||
|
|
# Swarm ports are NOT reachable on the public IP:
|
||
|
|
nc -zv -w2 <public_ip> 2377 # should fail/timeout
|
||
|
|
nc -zvu -w2 <public_ip> 4789 # should fail/timeout
|
||
|
|
```
|
||
|
|
|
||
|
|
## Adding a node label back out (undo)
|
||
|
|
|
||
|
|
```bash
|
||
|
|
docker node update --label-rm role <beta-node>
|
||
|
|
```
|