Realign Swarm/Forgejo infra to the 3-node design (#243)
Bring the Swarm+Forgejo layer in line with the canonical topology in docs/runbooks/implementation-handoff.md: three nodes (prod, beta, and the desktop LAN dev machine) instead of two, with the Forgejo Actions runner registered on the desktop as a plain systemd service rather than a Swarm-scheduled Docker-in-Docker sidecar. - setup-wireguard.sh: full N-peer mesh instead of point-to-point - docker-stack.yml: drop the runner/runner-dind services, volumes, and runner-token secret; only forgejo_db_password remains - register-lan-runner.sh: register under both docker and thermograph-lan labels, since one runner now covers both job types - init-swarm.sh, join-swarm.sh, swarm/README.md, forgejo/README.md, INFRA.md: updated node lists, order of operations, and access-state table for three nodes - deploy.yml: renamed to "Deploy to beta VPS" and reconciled the main-vs-release branch question against terraform.tfvars.example (this workflow already targets beta; a release-triggered prod deploy doesn't exist yet and isn't invented here)
This commit is contained in:
parent
9611a7be8b
commit
7ff3873661
8 changed files with 281 additions and 203 deletions
102
INFRA.md
102
INFRA.md
|
|
@ -1,37 +1,49 @@
|
|||
# Agent access, Docker Swarm, and Forgejo CI/CD
|
||||
|
||||
This covers three additive infrastructure layers on top of the two VPS boxes
|
||||
Terraform already provisions (see `terraform/README.md` and
|
||||
`terraform.tfvars.example`, where they're named **prod** — the new 48 GB /
|
||||
12-core box, `thermograph.org` — and **beta** — the old VPS,
|
||||
`75.119.132.91`). **None of this touches `backend/`, `Dockerfile`,
|
||||
This covers three additive infrastructure layers on top of the app
|
||||
containerization / Postgres / Terraform work described in
|
||||
**`docs/runbooks/implementation-handoff.md`** (read that first — it's the
|
||||
canonical Track A/Track B plan this whole effort implements) and
|
||||
`docs/architecture/repo-topology-and-infrastructure.md`. Terraform provisions
|
||||
**prod** (the new 48 GB / 12-core box, `thermograph.org`) and **beta** (the
|
||||
old VPS, `75.119.132.91`) — see `terraform/README.md` /
|
||||
`terraform.tfvars.example`. **None of this touches `backend/`, `Dockerfile`,
|
||||
`docker-compose*.yml`, `terraform/`, or `deploy/db/`** — those stay owned by
|
||||
the app-containerization work; this is strictly the layer above it.
|
||||
|
||||
```
|
||||
Track 1: deploy/provision-agent-access.sh — a dedicated full-root login for me
|
||||
Track 2: deploy/swarm/ — Swarm cluster spanning both boxes
|
||||
Track 2: deploy/swarm/ — Swarm cluster spanning THREE nodes
|
||||
Track 3: deploy/forgejo/ + .forgejo/ — Forgejo + Forgejo Actions, replacing GitHub
|
||||
```
|
||||
|
||||
**Three nodes, not two**: prod, beta, and **the desktop** (the LAN dev
|
||||
machine — same box that already runs the pre-Forgejo GitHub self-hosted
|
||||
runner). This matches the canonical Track B design in the handoff doc; an
|
||||
earlier revision of this whole effort covered just prod+beta and has been
|
||||
realigned.
|
||||
|
||||
## Track 1 — Agent access
|
||||
|
||||
Run `sudo bash deploy/provision-agent-access.sh` on **both** boxes. Creates a
|
||||
dedicated `agent` user (not raw root login — a distinct name gives a clean
|
||||
audit trail) with passwordless sudo, installs the agent's public key, disables
|
||||
SSH password auth fleet-wide, and turns on `auditd` logging of every
|
||||
root-effective command. See the script's header comment for the full
|
||||
rationale and revocation steps (one line to delete, or delete the whole
|
||||
account — your own access is never affected).
|
||||
Run `sudo bash deploy/provision-agent-access.sh` on **prod and beta** (not
|
||||
the desktop — that's wherever you're already working from, no separate
|
||||
access-provisioning step needed there). Creates a dedicated `agent` user (not
|
||||
raw root login — a distinct name gives a clean audit trail) with passwordless
|
||||
sudo, installs the agent's public key, disables SSH password auth
|
||||
fleet-wide, and turns on `auditd` logging of every root-effective command.
|
||||
See the script's header comment for the full rationale and revocation steps
|
||||
(one line to delete, or delete the whole account — your own access is never
|
||||
affected).
|
||||
|
||||
### Current access state (live)
|
||||
|
||||
Both boxes are provisioned and reachable as of this writing:
|
||||
Both VPS boxes are provisioned and reachable as of this writing:
|
||||
|
||||
| Host | Role | Public IP | Login |
|
||||
|------|------|-----------|-------|
|
||||
| prod | manager (Swarm), Thermograph's eventual home | `169.58.46.181` | `agent` |
|
||||
| beta | worker (Swarm), old VPS | `75.119.132.91` | `agent` |
|
||||
| prod | Swarm manager, Thermograph's eventual home | `169.58.46.181` | `agent` |
|
||||
| beta | Swarm worker, old VPS | `75.119.132.91` | `agent` |
|
||||
| desktop | Swarm worker, LAN dev machine + Actions runner | (local) | (already have access) |
|
||||
|
||||
```
|
||||
ssh -i ~/.ssh/thermograph_agent_ed25519 agent@169.58.46.181 # prod
|
||||
|
|
@ -54,23 +66,29 @@ directory that gets cleaned up, not anywhere-in-repo. Losing it just means
|
|||
regenerating and re-running the provisioning script; it does not lock either
|
||||
box, since your own login is untouched.
|
||||
|
||||
As of this writing, neither box has Docker installed yet and Terraform has
|
||||
As of this writing, neither VPS has Docker installed yet and Terraform has
|
||||
not been applied to either (`prod` is a bare OS install; `beta` has an old
|
||||
`/opt/thermograph` checkout but no Docker) — Track 2 is paused on that.
|
||||
`/opt/thermograph` checkout but no Docker) — Track 2 is paused on that. The
|
||||
desktop already has Docker (used for the LAN dev containerized stack) but
|
||||
hasn't joined anything yet (`docker info` shows `Swarm: inactive`).
|
||||
|
||||
## Track 2 — Swarm
|
||||
|
||||
See `deploy/swarm/README.md` for the exact order of operations (WireGuard
|
||||
tunnel first, then swarm init/join, then lock the Swarm ports down to the
|
||||
tunnel interface, then label beta for Forgejo placement). This cluster's only
|
||||
job is hosting Forgejo — it does not orchestrate the Terraform-managed app
|
||||
deploys.
|
||||
See `deploy/swarm/README.md` for the exact order of operations across all
|
||||
**three** nodes (WireGuard mesh first, then swarm init/join, then lock the
|
||||
Swarm ports down to the tunnel interface, then label beta for Forgejo
|
||||
placement). This cluster's only job is hosting Forgejo — it does not
|
||||
orchestrate the Terraform-managed app deploys.
|
||||
|
||||
## Track 3 — Forgejo, replacing GitHub
|
||||
|
||||
### 3a. Stand up Forgejo
|
||||
See `deploy/forgejo/README.md` — deploys the stack, walks through minting the
|
||||
two Swarm secrets it needs (`forgejo_db_password`, `forgejo_runner_token`).
|
||||
See `deploy/forgejo/README.md` — deploys the stack (pinned to beta), then
|
||||
walks through registering the Actions runner **on the desktop** as a plain
|
||||
systemd service (`register-lan-runner.sh`), not as a Swarm-scheduled
|
||||
container. Only one Swarm secret to mint now (`forgejo_db_password`) — the
|
||||
runner token is no longer a Swarm secret, since the runner isn't a Swarm
|
||||
service anymore.
|
||||
|
||||
### 3b. Migrate the repo (mirror first, verify, then cut over)
|
||||
1. In Forgejo: **+ New Migration → GitHub**. Point it at
|
||||
|
|
@ -82,10 +100,13 @@ two Swarm secrets it needs (`forgejo_db_password`, `forgejo_runner_token`).
|
|||
3. **Don't retarget any secret or disable a GitHub workflow yet** — see 3d.
|
||||
|
||||
### 3c. Workflows
|
||||
`.forgejo/workflows/*.yml` already mirror `.github/workflows/*.yml` in this
|
||||
PR — copied with the mechanical changes Forgejo needs:
|
||||
- `runs-on: ubuntu-latest` → `runs-on: docker` (the label the Swarm-hosted
|
||||
runner in `deploy/forgejo/docker-stack.yml` registers under).
|
||||
`.forgejo/workflows/{build,pr-build,deploy-dev,deploy}.yml` mirror
|
||||
`.github/workflows/*.yml`, copied with the mechanical changes Forgejo needs:
|
||||
- `runs-on: ubuntu-latest` → `runs-on: docker` (one of the two labels the
|
||||
desktop's runner registers under — see 3a; general CI/deploy jobs get a
|
||||
fresh container, the LAN-deploy job in `deploy-dev.yml` runs bare/host-native
|
||||
under the `thermograph-lan` label instead, since it needs real filesystem
|
||||
access).
|
||||
- `appleboy/ssh-action` referenced by full GitHub URL (not mirrored in
|
||||
Forgejo's default action registry, unlike `actions/checkout` /
|
||||
`actions/setup-python`, which resolve unchanged).
|
||||
|
|
@ -99,11 +120,24 @@ PR — copied with the mechanical changes Forgejo needs:
|
|||
(unlike GitHub's token-based merge commit, which doesn't re-trigger `push`),
|
||||
so `deploy-dev.yml`'s push trigger fires naturally after auto-merge — no
|
||||
more "direct push vs. PR-merge" double-trigger logic to maintain.
|
||||
- `.forgejo/workflows/deploy.yml` still targets `main` faithfully mirroring
|
||||
the current GitHub workflow, but `terraform.tfvars.example` describes prod's
|
||||
branch as `release`. That mismatch is flagged in the file's own header
|
||||
comment — reconcile with whoever's driving the Terraform/deploy side rather
|
||||
than guessing here.
|
||||
- `.forgejo/workflows/deploy.yml` ("Deploy to beta VPS") triggers on `main`,
|
||||
same as it always did. This was flagged as a possible mismatch against
|
||||
`terraform.tfvars.example`'s `release` branch for prod, but turned out not
|
||||
to be one: this workflow predates the prod/beta split, so its secrets almost
|
||||
certainly already point at beta (`main` is beta's branch per Terraform) —
|
||||
see the file's own header comment. A separate `release` → prod deploy path
|
||||
doesn't exist anywhere yet; don't invent one speculatively, confirm with
|
||||
whoever owns Terraform whether prod even wants a push-triggered deploy or
|
||||
is meant to go through `terraform apply` only.
|
||||
|
||||
A second, independent set of Forgejo workflows
|
||||
(`.forgejo/workflows/{ci,build-push,ops-cron}.yml`) was added by a parallel
|
||||
effort implementing `docs/runbooks/implementation-handoff.md` Track A chunk
|
||||
7 — see that doc and its own reconciliation PR (which already fixed its
|
||||
runner-label and ssh-action guesses to match the real infrastructure here).
|
||||
`ci.yml` duplicated this PR's `build.yml` exactly and was dropped in that
|
||||
reconciliation; `build-push.yml` (image → Forgejo's built-in registry) and
|
||||
`ops-cron.yml` (backup + IndexNow) are novel and still present.
|
||||
|
||||
### 3d. Cut over
|
||||
Only after 3a–3c are verified:
|
||||
|
|
|
|||
|
|
@ -1,37 +1,33 @@
|
|||
# Forgejo on the Swarm cluster
|
||||
|
||||
Runs as `deploy/forgejo/docker-stack.yml` — the only workload this Swarm
|
||||
cluster carries (the Thermograph app itself stays on the Terraform-managed
|
||||
`docker compose` deploys; see `terraform/README.md`). Pinned to the **beta**
|
||||
node (old VPS) via the `role=forge` label from
|
||||
Runs as `deploy/forgejo/docker-stack.yml` — the only Swarm-scheduled workload
|
||||
this cluster carries (the Thermograph app itself stays on the
|
||||
Terraform-managed `docker compose` deploys; see `terraform/README.md`).
|
||||
Pinned to the **beta** node (old VPS) via the `role=forge` label from
|
||||
`deploy/swarm/label-forge-node.sh`.
|
||||
|
||||
The Actions **runner** is deliberately *not* part of this stack — it runs on
|
||||
the **desktop** as a plain systemd service (`register-lan-runner.sh` below),
|
||||
per `docs/runbooks/implementation-handoff.md` Track B step 5. That's the
|
||||
canonical placement; an earlier revision of this stack ran the runner as a
|
||||
Swarm-scheduled Docker-in-Docker sidecar pinned to beta, which is gone now.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
1. Both boxes have joined the swarm (`deploy/swarm/`) and beta is labeled
|
||||
`role=forge`.
|
||||
2. `docker node ls` (from the manager) shows both `Ready`.
|
||||
1. All three nodes have joined the swarm (`deploy/swarm/`) and beta is
|
||||
labeled `role=forge`.
|
||||
2. `docker node ls` (from the manager) shows all three `Ready`.
|
||||
|
||||
## One-time setup: Swarm secrets
|
||||
## One-time setup: Swarm secret
|
||||
|
||||
Two secrets the stack expects to already exist (Swarm secrets, not files —
|
||||
One secret the stack expects to already exist (a Swarm secret, not a file —
|
||||
`external: true` in the stack file, so `docker stack deploy` never creates or
|
||||
sees the values, only references them):
|
||||
sees the value, only references it):
|
||||
|
||||
```bash
|
||||
# A strong random password for Forgejo's own Postgres (NOT related to
|
||||
# Thermograph's app database — entirely separate instance/network).
|
||||
openssl rand -base64 32 | docker secret create forgejo_db_password -
|
||||
|
||||
# The runner registration token. Forgejo can't issue one before it's running,
|
||||
# so this is a two-step dance the first time:
|
||||
docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo # 1. bring Forgejo up (runner will crashloop briefly — expected)
|
||||
# 2. once Forgejo answers at the domain, log in, go to
|
||||
# Site Administration -> Actions -> Runners -> Create new Runner
|
||||
# (or, for a repo-scoped runner: <repo> -> Settings -> Actions -> Runners),
|
||||
# copy the token, then:
|
||||
echo -n "PASTE_TOKEN_HERE" | docker secret create forgejo_runner_token -
|
||||
docker service update --force forgejo_runner # picks up the new secret and registers
|
||||
```
|
||||
|
||||
## Deploy / update
|
||||
|
|
@ -42,12 +38,30 @@ docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo
|
|||
|
||||
Re-running is safe — Swarm only touches services whose spec actually changed.
|
||||
|
||||
## Register the Actions runner (on the desktop, not through Swarm)
|
||||
|
||||
Once Forgejo answers at its domain:
|
||||
|
||||
```bash
|
||||
# On the desktop:
|
||||
# Forgejo web UI -> Site Administration -> Actions -> Runners -> Create new Runner
|
||||
# (or, for a repo-scoped runner: <repo> -> Settings -> Actions -> Runners)
|
||||
# copy the registration token, then:
|
||||
bash deploy/forgejo/register-lan-runner.sh https://<forgejo-domain> <token>
|
||||
```
|
||||
|
||||
See that script's header for exactly what it replaces (the pre-Forgejo GitHub
|
||||
self-hosted runner on this same machine) and why it registers with two
|
||||
labels where there used to be two separate runners.
|
||||
|
||||
## DNS
|
||||
|
||||
Point the Forgejo domain (default `git.thermograph.org`; override with
|
||||
`FORGEJO_DOMAIN=...` before `docker stack deploy`, Swarm reads it from the
|
||||
deploying shell's environment) at **either** node's public IP — the routing
|
||||
mesh forwards published ports to wherever the task actually landed.
|
||||
deploying shell's environment) at **prod's or beta's** public IP — not the
|
||||
desktop's, which isn't a stable publicly-reachable address. The routing mesh
|
||||
forwards published ports to wherever the task actually landed, so either VPS
|
||||
IP works.
|
||||
|
||||
## Why Postgres here and not the Thermograph app's TimescaleDB
|
||||
|
||||
|
|
@ -60,9 +74,10 @@ independently.
|
|||
## Verifying
|
||||
|
||||
```bash
|
||||
docker service ls # all forgejo_* services Running, 1/1
|
||||
docker service ls # forgejo_db, forgejo_forgejo, forgejo_traefik all Running, 1/1
|
||||
curl -I https://git.thermograph.org/ # 200, valid cert
|
||||
docker service logs forgejo_runner --tail 50 # "runner: successfully registered" then idle
|
||||
# On the desktop, after registering the runner:
|
||||
systemctl --user status forgejo-runner # active, both labels registered
|
||||
```
|
||||
|
||||
## Rollback / removal
|
||||
|
|
|
|||
|
|
@ -1,20 +1,38 @@
|
|||
# Forgejo (self-hosted Git + CI/CD) as a Docker Swarm stack — the only workload
|
||||
# this Swarm cluster runs (see deploy/swarm/README.md). Deliberately separate
|
||||
# from the Terraform-managed docker-compose.yml that runs the Thermograph app
|
||||
# itself: this stack's only job is Forgejo, its runner, and the ingress in
|
||||
# front of it.
|
||||
# itself: this stack's only job is Forgejo, its container registry, and the
|
||||
# ingress in front of it.
|
||||
#
|
||||
# Deploy from the manager node (prod), after both boxes have joined the swarm
|
||||
# and beta is labeled role=forge:
|
||||
# The Actions RUNNER is deliberately NOT a service in this stack. Per
|
||||
# docs/runbooks/implementation-handoff.md (Track B step 5), the canonical
|
||||
# design registers the runner on the desktop node as a plain systemd service
|
||||
# (deploy/forgejo/register-lan-runner.sh) — the same place the pre-Forgejo
|
||||
# GitHub self-hosted runner already lived, not a Swarm-scheduled container. An
|
||||
# earlier revision of this file ran the runner + a Docker-in-Docker sidecar as
|
||||
# Swarm services pinned to beta; that's gone now in favor of the canonical
|
||||
# placement.
|
||||
#
|
||||
# Deploy from the manager node (prod), after all three nodes (prod, beta,
|
||||
# desktop) have joined the swarm and beta is labeled role=forge:
|
||||
#
|
||||
# docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo
|
||||
#
|
||||
# Requires these Swarm secrets to exist first (see deploy/forgejo/README.md):
|
||||
# forgejo_db_password, forgejo_runner_token
|
||||
# Requires this Swarm secret to exist first (see deploy/forgejo/README.md):
|
||||
# forgejo_db_password
|
||||
# (forgejo_runner_token is no longer a Swarm secret — see register-lan-runner.sh,
|
||||
# which takes the registration token as a plain argument on the desktop instead.)
|
||||
#
|
||||
# DNS: point git.thermograph.org (or your chosen domain — see plan) at EITHER
|
||||
# node's public IP. Swarm's routing mesh accepts published ports on every node
|
||||
# and forwards to wherever the task actually runs, so either IP works.
|
||||
# DNS: point the Forgejo domain at prod's or beta's public IP (not the
|
||||
# desktop's — it isn't a stable, publicly-reachable address). Swarm's routing
|
||||
# mesh accepts published ports on every node and forwards to wherever the
|
||||
# task actually runs, so either of the two VPS IPs works.
|
||||
#
|
||||
# Registry exposure (hazard #15 in docs/runbooks/hop1-forgejo-registry-cutover.md)
|
||||
# is still an open call, not resolved here: this file exposes Forgejo (and its
|
||||
# built-in registry) publicly via Traefik/Let's Encrypt. The alternative —
|
||||
# mesh-only registry + DNS-01 ACME + an outbound pull-mirror — is a deliberate
|
||||
# design decision for whoever executes Track B, not something to guess at here.
|
||||
|
||||
services:
|
||||
db:
|
||||
|
|
@ -75,54 +93,6 @@ services:
|
|||
- traefik.http.routers.forgejo.tls.certresolver=le
|
||||
- traefik.http.services.forgejo.loadbalancer.server.port=3000
|
||||
|
||||
# Docker-in-Docker sidecar for the runner. Chosen over mounting the host
|
||||
# docker.sock (root-equivalent host access, zero isolation) — DinD is the
|
||||
# documented mitigation for a trusted single-maintainer runner that only
|
||||
# ever executes this repo's own workflows.
|
||||
runner-dind:
|
||||
image: docker:27-dind
|
||||
privileged: true
|
||||
environment:
|
||||
DOCKER_TLS_CERTDIR: ""
|
||||
volumes:
|
||||
- runner_dind_certs:/certs
|
||||
networks: [forgejo_net]
|
||||
deploy:
|
||||
placement:
|
||||
constraints: [node.labels.role == forge]
|
||||
restart_policy:
|
||||
condition: on-failure
|
||||
|
||||
runner:
|
||||
image: code.forgejo.org/forgejo/runner:6
|
||||
depends_on: [runner-dind, forgejo]
|
||||
environment:
|
||||
DOCKER_HOST: tcp://runner-dind:2375
|
||||
FORGEJO_INSTANCE_URL: "https://${FORGEJO_DOMAIN:-git.thermograph.org}"
|
||||
secrets:
|
||||
- source: forgejo_runner_token
|
||||
target: forgejo_runner_token
|
||||
# Registers on first start using the token secret (see README for how to
|
||||
# mint one), then just executes jobs — no docker.sock mount.
|
||||
entrypoint: >
|
||||
sh -c '
|
||||
if [ ! -f /data/.runner ]; then
|
||||
forgejo-runner register --no-interactive
|
||||
--instance "$$FORGEJO_INSTANCE_URL"
|
||||
--token "$$(cat /run/secrets/forgejo_runner_token)"
|
||||
--name swarm-runner --labels docker:docker://node:20-bookworm;
|
||||
fi;
|
||||
forgejo-runner daemon
|
||||
'
|
||||
volumes:
|
||||
- runner_data:/data
|
||||
networks: [forgejo_net]
|
||||
deploy:
|
||||
placement:
|
||||
constraints: [node.labels.role == forge]
|
||||
restart_policy:
|
||||
condition: on-failure
|
||||
|
||||
traefik:
|
||||
image: traefik:v3.1
|
||||
command:
|
||||
|
|
@ -163,12 +133,8 @@ networks:
|
|||
volumes:
|
||||
forgejo_db:
|
||||
forgejo_data:
|
||||
runner_data:
|
||||
runner_dind_certs:
|
||||
traefik_certs:
|
||||
|
||||
secrets:
|
||||
forgejo_db_password:
|
||||
external: true
|
||||
forgejo_runner_token:
|
||||
external: true
|
||||
|
|
|
|||
|
|
@ -4,6 +4,17 @@
|
|||
# runner (see DEPLOY-DEV.md) — this replaces that runner, it doesn't add a
|
||||
# second one. Sudo-free, systemd --user, same pattern as the app service.
|
||||
#
|
||||
# This is THE runner (docs/runbooks/implementation-handoff.md Track B step 5
|
||||
# — canonical placement is the desktop, not a Swarm-hosted container), so it
|
||||
# registers with BOTH labels the workflows need, where one runner used to be
|
||||
# two: general CI/build/deploy.yml jobs (`docker`, containerized via this
|
||||
# machine's own already-installed Docker — no Docker-in-Docker sidecar needed,
|
||||
# unlike the Swarm-hosted design this replaces, since a real host with a real
|
||||
# Docker install needs no such indirection) and the LAN-deploy job
|
||||
# (`thermograph-lan`, bare/host-native — it writes to ~/thermograph-dev and
|
||||
# restarts a systemd --user service, which only works running directly on the
|
||||
# host, not inside a container).
|
||||
#
|
||||
# bash deploy/forgejo/register-lan-runner.sh <forgejo_url> <registration_token>
|
||||
#
|
||||
# Get <registration_token> from the Forgejo web UI:
|
||||
|
|
@ -15,12 +26,18 @@ set -euo pipefail
|
|||
FORGEJO_URL="${1:?usage: $0 <forgejo_url> <registration_token>}"
|
||||
TOKEN="${2:?}"
|
||||
RUNNER_DIR="${RUNNER_DIR:-$HOME/forgejo-runner}"
|
||||
LABELS="${LABELS:-thermograph-lan}"
|
||||
LABELS="${LABELS:-docker:docker://node:20-bookworm,thermograph-lan}"
|
||||
|
||||
echo "==> Stopping and disabling the old GitHub Actions runner service, if present"
|
||||
systemctl --user stop github-actions-runner 2>/dev/null || true
|
||||
systemctl --user disable github-actions-runner 2>/dev/null || true
|
||||
|
||||
if ! id -nG "$USER" 2>/dev/null | grep -qw docker; then
|
||||
echo "WARNING: $USER is not in the 'docker' group — the 'docker:docker://...'"
|
||||
echo "labeled jobs (general CI) will fail to start a container until this is"
|
||||
echo "fixed: sudo usermod -aG docker $USER && (log out and back in)."
|
||||
fi
|
||||
|
||||
echo "==> Installing forgejo-runner into $RUNNER_DIR"
|
||||
mkdir -p "$RUNNER_DIR"
|
||||
cd "$RUNNER_DIR"
|
||||
|
|
@ -49,7 +66,7 @@ echo "==> systemd --user unit"
|
|||
mkdir -p "$HOME/.config/systemd/user"
|
||||
cat > "$HOME/.config/systemd/user/forgejo-runner.service" <<EOF
|
||||
[Unit]
|
||||
Description=Forgejo Actions runner (thermograph-lan)
|
||||
Description=Forgejo Actions runner (docker + thermograph-lan)
|
||||
After=network-online.target
|
||||
|
||||
[Service]
|
||||
|
|
@ -68,7 +85,9 @@ loginctl enable-linger "$USER" 2>/dev/null || true
|
|||
|
||||
cat <<EOF
|
||||
|
||||
Done. The runner now serves Forgejo, not GitHub.
|
||||
Done. The runner now serves Forgejo, not GitHub — one runner, both labels
|
||||
($LABELS), replacing what used to be a separate Swarm-hosted runner for
|
||||
general CI plus this machine's own GitHub runner for LAN deploys.
|
||||
status: systemctl --user status forgejo-runner
|
||||
logs: journalctl --user -u forgejo-runner -f
|
||||
restart: systemctl --user restart forgejo-runner
|
||||
|
|
|
|||
|
|
@ -1,61 +1,79 @@
|
|||
# 2-node Docker Swarm (prod + beta), for hosting Forgejo
|
||||
# 3-node Docker Swarm (prod + beta + desktop), for hosting Forgejo
|
||||
|
||||
This Swarm cluster's only job is to run Forgejo (`deploy/forgejo/`) — it does
|
||||
**not** orchestrate the Thermograph app itself, which stays on the
|
||||
Terraform-managed `docker compose` deploys on each box independently (see
|
||||
Terraform-managed `docker compose` deploys on prod/beta independently (see
|
||||
`terraform/README.md`). Keeping those separate means nothing here can strand
|
||||
or interfere with the app's already-working, single-writer Postgres/TimescaleDB
|
||||
deploys.
|
||||
|
||||
This is the canonical topology from
|
||||
`docs/runbooks/implementation-handoff.md` (Track B steps 2-3) — three nodes,
|
||||
not two. An earlier revision of this doc/scripts covered just prod+beta;
|
||||
the desktop (this LAN dev machine) joins too.
|
||||
|
||||
**Nodes:**
|
||||
- **manager** — prod, the new 48 GB / 12-core box (more headroom).
|
||||
- **worker** — beta, the old VPS (`75.119.132.91`).
|
||||
- **worker** — desktop, this LAN dev machine (also runs the Forgejo Actions
|
||||
runner as a plain systemd service — see `deploy/forgejo/README.md` — not as
|
||||
a Swarm-scheduled container).
|
||||
|
||||
1 manager + 1 worker, not 2 managers: Raft needs 3 nodes for real quorum-based
|
||||
HA, so a second manager here would add complexity (split-brain risk) without
|
||||
adding actual failover. If the manager goes down, the worker keeps running
|
||||
whatever was already scheduled on it (Forgejo, since it's pinned there) but
|
||||
the cluster can't reschedule anything until the manager's back — acceptable
|
||||
for a 2-box hobby/small-team cluster whose only job is CI/CD.
|
||||
One manager, not more: Raft needs 3 nodes for real quorum-based HA, and this
|
||||
cluster only has 3 nodes total, so making even one more of them a manager
|
||||
would still fall short of real HA while adding split-brain risk. If the
|
||||
manager (prod) goes down, the workers keep running whatever was already
|
||||
scheduled on them (Forgejo, pinned to beta) but the cluster can't reschedule
|
||||
anything until prod's back — acceptable for a small cluster whose only job is
|
||||
CI/CD.
|
||||
|
||||
## Order of operations
|
||||
|
||||
1. **Agent access first** (`deploy/provision-agent-access.sh`) on both boxes —
|
||||
everything below is run through that access.
|
||||
2. **WireGuard tunnel** (`setup-wireguard.sh`) — run on both boxes; see the
|
||||
script's header for the two-pass key-exchange dance. Verify with
|
||||
`ping <peer_wg_ip>` before continuing.
|
||||
1. **Agent access first** (`deploy/provision-agent-access.sh`) on prod and
|
||||
beta — everything below on those two boxes is run through that access. The
|
||||
desktop is wherever you're already working from; no separate access step
|
||||
needed there.
|
||||
2. **WireGuard mesh** (`setup-wireguard.sh <my_wg_ip> <peers_file>`) — run on
|
||||
**all three** nodes. See the script's header for the peer-list format and
|
||||
the two-pass key-exchange dance (pubkeys aren't known until every node has
|
||||
run it once). Verify with `ping <peer_wg_ip>` to each of the other two
|
||||
before continuing.
|
||||
3. **Swarm init** (`init-swarm.sh <manager_wg_ip>`) on the manager (prod) only.
|
||||
4. **Swarm join** (`join-swarm.sh <manager_wg_ip> <token>`) on the worker
|
||||
(beta) only, using the token `init-swarm.sh` printed.
|
||||
5. **Firewall lockdown** (`firewall-swarm.sh`) on **both** boxes — closes
|
||||
4. **Swarm join** (`join-swarm.sh <manager_wg_ip> <token>`) on **each** of the
|
||||
two workers (beta, desktop) — same token for both.
|
||||
5. **Firewall lockdown** (`firewall-swarm.sh`) on **all three** nodes — closes
|
||||
2377/7946/4789 to everything except the WireGuard interface. Do this
|
||||
*after* joining is confirmed working, not before (locking the ports first
|
||||
would make the join itself fail).
|
||||
*after* joining is confirmed working on all three, not before (locking the
|
||||
ports first would make the join itself fail).
|
||||
6. **Label beta** (`label-forge-node.sh <beta-node-name>`) on the manager —
|
||||
`docker node ls` shows beta's node name/ID.
|
||||
`docker node ls` shows each node's name/ID. Only beta gets `role=forge`;
|
||||
the desktop and prod don't need a Swarm label for anything in this setup.
|
||||
7. Deploy Forgejo: see `deploy/forgejo/README.md`.
|
||||
8. Register the Actions runner **on the desktop** (not through Swarm):
|
||||
`deploy/forgejo/register-lan-runner.sh`.
|
||||
|
||||
## Why WireGuard instead of relying on Swarm's built-in TLS alone
|
||||
|
||||
Swarm's control plane (port 2377) is TLS-encrypted and mutually authenticated
|
||||
by default. Its overlay data plane (VXLAN, port 4789) is **not** encrypted by
|
||||
default, and Docker's own guidance is that port must never face the public
|
||||
internet — these two boxes are on different providers' public IPs, not a
|
||||
private LAN, so the tunnel is the network boundary the Swarm ports advertise
|
||||
into, rather than trusting the public internet directly.
|
||||
internet — these nodes are on different networks (two separate providers'
|
||||
public IPs, plus a home/LAN connection for the desktop), not one private LAN,
|
||||
so the tunnel is the network boundary the Swarm ports advertise into, rather
|
||||
than trusting the public internet (or the desktop's home network) directly.
|
||||
|
||||
## Verifying
|
||||
|
||||
```bash
|
||||
# On the manager:
|
||||
docker node ls # both nodes Ready
|
||||
docker node ls # all three nodes Ready
|
||||
docker node inspect <beta-node> --format '{{.Spec.Labels}}' # role:forge
|
||||
|
||||
# From outside the WireGuard interface (e.g. your own machine), confirm the
|
||||
# Swarm ports are NOT reachable on the public IP:
|
||||
nc -zv -w2 <public_ip> 2377 # should fail/timeout
|
||||
nc -zvu -w2 <public_ip> 4789 # should fail/timeout
|
||||
# From a FOURTH machine outside the mesh entirely, confirm the Swarm ports
|
||||
# are NOT reachable on either VPS's public IP (the desktop has no public IP
|
||||
# to check this way):
|
||||
nc -zv -w2 <prod_or_beta_public_ip> 2377 # should fail/timeout
|
||||
nc -zvu -w2 <prod_or_beta_public_ip> 4789 # should fail/timeout
|
||||
```
|
||||
|
||||
## Adding a node label back out (undo)
|
||||
|
|
|
|||
|
|
@ -1,7 +1,9 @@
|
|||
#!/usr/bin/env bash
|
||||
# Run ONCE, on the manager node only (prod — the new 48 GB box). Initializes
|
||||
# the Swarm advertising the WireGuard address, so cluster traffic never
|
||||
# touches the public interface. Run setup-wireguard.sh on both boxes first.
|
||||
# touches the public interface. Run setup-wireguard.sh on all THREE nodes
|
||||
# first (prod, beta, and the desktop — see
|
||||
# docs/runbooks/implementation-handoff.md Track B steps 2-3).
|
||||
set -euo pipefail
|
||||
|
||||
MY_WG_IP="${1:?usage: $0 <my_wg_ip>}"
|
||||
|
|
@ -19,8 +21,9 @@ echo "==> docker swarm init, advertising ${MY_WG_IP} (the WireGuard address, not
|
|||
docker swarm init --advertise-addr "$MY_WG_IP" --listen-addr "${MY_WG_IP}:2377"
|
||||
|
||||
echo
|
||||
echo "==> Worker join command (run this on the worker/beta node):"
|
||||
echo "==> Worker join command (run this on EACH of the two workers — beta and"
|
||||
echo " the desktop; the same token works for both):"
|
||||
TOKEN="$(docker swarm join-token -q worker)"
|
||||
echo " docker swarm join --token ${TOKEN} ${MY_WG_IP}:2377"
|
||||
echo
|
||||
echo "(Or run deploy/swarm/join-swarm.sh <manager_wg_ip> <token> on the worker.)"
|
||||
echo "(Or run deploy/swarm/join-swarm.sh <manager_wg_ip> <token> on each worker.)"
|
||||
|
|
|
|||
|
|
@ -1,6 +1,7 @@
|
|||
#!/usr/bin/env bash
|
||||
# Run ONCE, on the worker node only (beta — the old VPS). Joins the Swarm
|
||||
# initialized by init-swarm.sh on the manager, over the WireGuard tunnel.
|
||||
# Run ONCE on EACH worker node (beta, and the desktop — not the manager,
|
||||
# prod). Joins the Swarm initialized by init-swarm.sh, over the WireGuard
|
||||
# tunnel. The join token is the same for both workers.
|
||||
set -euo pipefail
|
||||
|
||||
MANAGER_WG_IP="${1:?usage: $0 <manager_wg_ip> <join_token>}"
|
||||
|
|
|
|||
|
|
@ -1,28 +1,33 @@
|
|||
#!/usr/bin/env bash
|
||||
# Sets up a point-to-point WireGuard tunnel between the two Swarm hosts. Docker
|
||||
# Swarm's control plane (2377/tcp) is TLS-encrypted by default, but the overlay
|
||||
# data plane (VXLAN, 4789/udp) is NOT — and it should never face the public
|
||||
# internet. Swarm is joined over this tunnel's private IPs instead.
|
||||
# Sets up this node's side of a full-mesh WireGuard tunnel across all of
|
||||
# prod, beta, and the desktop (three nodes, not two — see
|
||||
# docs/runbooks/implementation-handoff.md Track B step 2). Docker Swarm's
|
||||
# control plane (2377/tcp) is TLS-encrypted by default, but the overlay data
|
||||
# plane (VXLAN, 4789/udp) is NOT — and it should never face the public
|
||||
# internet. Swarm joins over these private WireGuard IPs instead.
|
||||
#
|
||||
# Run on EACH box, in either order. Each run generates that box's own WireGuard
|
||||
# keypair (if missing) and prints its public key — paste it into the OTHER
|
||||
# box's invocation. Two runs total, one per box:
|
||||
# A full mesh means every node peers directly with every other node (not a
|
||||
# hub-and-spoke through one box) — with three nodes that's 2 [Peer] blocks per
|
||||
# node. Run this on EACH of the three nodes, in any order, using a small peer
|
||||
# list file so the script isn't hardcoded to any particular node count.
|
||||
#
|
||||
# box A (prod): bash setup-wireguard.sh 10.10.0.1 10.10.0.2 <PROD_PUBLIC_IP> <BETA_PUBLIC_IP> <BETA_WG_PUBKEY_or_empty>
|
||||
# box B (beta): bash setup-wireguard.sh 10.10.0.2 10.10.0.1 <BETA_PUBLIC_IP> <PROD_PUBLIC_IP> <PROD_WG_PUBKEY>
|
||||
# --- Peer list format --------------------------------------------------
|
||||
# A text file, one line per OTHER node (not including the one you're running
|
||||
# on), each line: <wg_ip> <public_ip> <pubkey-or-dash>
|
||||
# 10.10.0.1 169.58.46.181 <prod's pubkey, or - if not yet known>
|
||||
# 10.10.0.2 75.119.132.91 <beta's pubkey, or ->
|
||||
# 10.10.0.3 <desktop's public/reachable IP or a DDNS name> <desktop's pubkey, or ->
|
||||
#
|
||||
# First run on either box: leave the last argument empty, note the printed
|
||||
# pubkey, then run the second box with that value, then re-run the first box
|
||||
# once more with the second box's now-known pubkey. (A local key never
|
||||
# changes across re-runs — only the peer section updates.)
|
||||
# First pass on any node: peer pubkeys you don't have yet are "-". Run this
|
||||
# script on all three nodes once (prints each node's own pubkey), fill those
|
||||
# into everyone's peer-list files, then re-run on all three once more to
|
||||
# actually establish the tunnels. Re-running is safe (idempotent) at any point.
|
||||
#
|
||||
# bash setup-wireguard.sh <my_wg_ip> <peers_file>
|
||||
set -euo pipefail
|
||||
|
||||
MY_WG_IP="${1:?usage: $0 <my_wg_ip> <peer_wg_ip> <my_public_ip> <peer_public_ip> [peer_wg_pubkey]}"
|
||||
PEER_WG_IP="${2:?}"
|
||||
MY_PUBLIC_IP="${3:?}"
|
||||
PEER_PUBLIC_IP="${4:?}"
|
||||
PEER_PUBKEY="${5:-}"
|
||||
|
||||
MY_WG_IP="${1:?usage: $0 <my_wg_ip> <peers_file>}"
|
||||
PEERS_FILE="${2:?}"
|
||||
WG_PORT="${WG_PORT:-51820}"
|
||||
WG_DIR=/etc/wireguard
|
||||
|
||||
|
|
@ -30,13 +35,17 @@ if [ "$(id -u)" -ne 0 ]; then
|
|||
echo "Run as root (or via the agent user's sudo)." >&2
|
||||
exit 1
|
||||
fi
|
||||
if [ ! -f "$PEERS_FILE" ]; then
|
||||
echo "Peer list not found: $PEERS_FILE (see this script's header for the format)" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "==> Installing WireGuard if needed"
|
||||
command -v wg >/dev/null 2>&1 || { apt-get update -y -q && apt-get install -y -q wireguard; }
|
||||
|
||||
install -d -m 700 "$WG_DIR"
|
||||
if [ ! -f "$WG_DIR/privatekey" ]; then
|
||||
echo "==> Generating this box's WireGuard keypair"
|
||||
echo "==> Generating this node's WireGuard keypair"
|
||||
umask 077
|
||||
wg genkey | tee "$WG_DIR/privatekey" | wg pubkey > "$WG_DIR/publickey"
|
||||
fi
|
||||
|
|
@ -44,50 +53,63 @@ MY_PRIVKEY="$(cat "$WG_DIR/privatekey")"
|
|||
MY_PUBKEY="$(cat "$WG_DIR/publickey")"
|
||||
|
||||
echo
|
||||
echo "==> This box's WireGuard public key (give this to the OTHER box's run):"
|
||||
echo "==> This node's WireGuard public key (put this in the OTHER two nodes'"
|
||||
echo " peer-list files, replacing the '-' for this node's line):"
|
||||
echo " $MY_PUBKEY"
|
||||
echo
|
||||
|
||||
if [ -z "$PEER_PUBKEY" ]; then
|
||||
echo "No peer public key supplied yet — writing a config with no [Peer] section."
|
||||
echo "Run this same command again once you have it (from the other box's output above)."
|
||||
cat > "$WG_DIR/wg0.conf" <<EOF
|
||||
[Interface]
|
||||
Address = ${MY_WG_IP}/24
|
||||
PrivateKey = ${MY_PRIVKEY}
|
||||
ListenPort = ${WG_PORT}
|
||||
EOF
|
||||
else
|
||||
cat > "$WG_DIR/wg0.conf" <<EOF
|
||||
[Interface]
|
||||
Address = ${MY_WG_IP}/24
|
||||
PrivateKey = ${MY_PRIVKEY}
|
||||
ListenPort = ${WG_PORT}
|
||||
missing=0
|
||||
{
|
||||
echo "[Interface]"
|
||||
echo "Address = ${MY_WG_IP}/24"
|
||||
echo "PrivateKey = ${MY_PRIVKEY}"
|
||||
echo "ListenPort = ${WG_PORT}"
|
||||
echo
|
||||
|
||||
[Peer]
|
||||
PublicKey = ${PEER_PUBKEY}
|
||||
Endpoint = ${PEER_PUBLIC_IP}:${WG_PORT}
|
||||
AllowedIPs = ${PEER_WG_IP}/32
|
||||
PersistentKeepalive = 25
|
||||
EOF
|
||||
fi
|
||||
while read -r peer_wg_ip peer_public_ip peer_pubkey; do
|
||||
[ -z "${peer_wg_ip:-}" ] && continue
|
||||
case "$peer_wg_ip" in \#*) continue ;; esac
|
||||
if [ "$peer_pubkey" = "-" ] || [ -z "$peer_pubkey" ]; then
|
||||
echo "# SKIPPED (no pubkey yet) — peer at ${peer_public_ip}" >&2
|
||||
missing=$((missing + 1))
|
||||
continue
|
||||
fi
|
||||
echo "[Peer]"
|
||||
echo "PublicKey = ${peer_pubkey}"
|
||||
echo "Endpoint = ${peer_public_ip}:${WG_PORT}"
|
||||
echo "AllowedIPs = ${peer_wg_ip}/32"
|
||||
echo "PersistentKeepalive = 25"
|
||||
echo
|
||||
done < "$PEERS_FILE"
|
||||
} > "$WG_DIR/wg0.conf"
|
||||
chmod 600 "$WG_DIR/wg0.conf"
|
||||
|
||||
echo "==> Bringing up wg0"
|
||||
systemctl enable --now wg-quick@wg0 2>/dev/null || (wg-quick down wg0 2>/dev/null; wg-quick up wg0)
|
||||
# Re-apply if the config changed on an already-up interface.
|
||||
wg syncconf wg0 <(wg-quick strip wg0) 2>/dev/null || true
|
||||
|
||||
echo "==> Firewall: only let the OTHER box's public IP reach the WireGuard port"
|
||||
echo "==> Firewall: only let listed peers' public IPs reach the WireGuard port"
|
||||
if command -v ufw >/dev/null 2>&1; then
|
||||
ufw allow from "$PEER_PUBLIC_IP" to any port "$WG_PORT" proto udp comment "wireguard peer"
|
||||
while read -r _wg_ip peer_public_ip _pubkey; do
|
||||
[ -z "${peer_public_ip:-}" ] && continue
|
||||
case "$peer_public_ip" in \#*) continue ;; esac
|
||||
ufw allow from "$peer_public_ip" to any port "$WG_PORT" proto udp comment "wireguard peer"
|
||||
done < "$PEERS_FILE"
|
||||
fi
|
||||
|
||||
echo
|
||||
echo "wg0 status:"
|
||||
wg show wg0 || true
|
||||
echo
|
||||
if [ -n "$PEER_PUBKEY" ]; then
|
||||
echo "==> Verify from here once BOTH boxes have run with each other's pubkey:"
|
||||
echo " ping -c2 ${PEER_WG_IP}"
|
||||
if [ "$missing" -gt 0 ]; then
|
||||
echo "NOTE: $missing peer(s) skipped (no pubkey yet in $PEERS_FILE). Fill them in"
|
||||
echo "and re-run on this node once all three nodes have printed their pubkey."
|
||||
else
|
||||
echo "==> All peers configured. Verify once every node has run this:"
|
||||
while read -r peer_wg_ip _peer_public_ip peer_pubkey; do
|
||||
[ -z "${peer_wg_ip:-}" ] && continue
|
||||
case "$peer_wg_ip" in \#*) continue ;; esac
|
||||
[ "$peer_pubkey" = "-" ] && continue
|
||||
echo " ping -c2 ${peer_wg_ip}"
|
||||
done < "$PEERS_FILE"
|
||||
fi
|
||||
|
|
|
|||
Loading…
Reference in a new issue