# Forgejo on the Swarm cluster Runs as `deploy/forgejo/docker-stack.yml` — the only Swarm-scheduled workload this cluster carries (prod and beta's app stacks are separate `docker stack deploy`s that happen to run on the manager node, vps2 — see `deploy/stack/README` context in `DEPLOY.md`). Pinned to the **vps1** node (`75.119.132.91`) via the `role=forge` label from `deploy/swarm/label-forge-node.sh`. The Actions **runner** is deliberately *not* part of this stack — see `DEPLOY-DEV.md` and the note in `register-lan-runner.sh`'s own header for where it runs today. That's the canonical placement per `thermograph-docs/runbooks/implementation-handoff.md` Track B step 5; an earlier revision of this stack ran the runner as a Swarm-scheduled Docker-in-Docker sidecar pinned to the Forgejo node, which is gone now. ## Prerequisites 1. All three nodes have joined the swarm (`deploy/swarm/`) and vps1 is labeled `role=forge`. 2. `docker node ls` (from the manager, vps2) shows all three `Ready`. ## One-time setup: Swarm secret One secret the stack expects to already exist (a Swarm secret, not a file — `external: true` in the stack file, so `docker stack deploy` never creates or sees the value, only references it): ```bash # A strong random password for Forgejo's own Postgres (NOT related to # Thermograph's app database — entirely separate instance/network). openssl rand -base64 32 | docker secret create forgejo_db_password - ``` ## Deploy / update ```bash docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo ``` Re-running is safe — Swarm only touches services whose spec actually changed. Do this **before** the DNS + Caddy step below — Caddy's reverse_proxy target (`127.0.0.1:3080`) needs the `forgejo` service actually listening first, or its first health check just fails harmlessly until it is. This stack has **no auto-deploy trigger** — nothing in `.forgejo/workflows/` redeploys it on push. A change to `docker-stack.yml` only takes effect once someone re-runs `docker stack deploy` by hand on the manager (vps2). `db`/`forgejo` both carry `resources.limits` (defaults: db 1 CPU/1g, forgejo 2 CPU/2g — several times observed steady-state usage), overridable with `FORGEJO_DB_CPUS`/`FORGEJO_DB_MEMORY`/`FORGEJO_CPUS`/`FORGEJO_MEMORY` env vars before `docker stack deploy`, same convention as the app stack. ## DNS + TLS: reusing vps1's existing Caddy, not a second reverse proxy Forgejo is pinned to vps1 (`role=forge`) — the same box that also runs Grafana/Loki/Alloy and the `emigriffith.dev` portfolio site, each fronted by that host's own Caddy. A second ingress (Traefik) trying to bind the same ports would collide with it. So there's no Traefik in this stack: `forgejo`'s web port publishes to `127.0.0.1:3080` only (host-local), and vps1's *existing* Caddy gets one more site block reverse-proxying to it — same pattern as its other site blocks, same automatic-HTTPS. 1. Point the Forgejo domain (default `git.thermograph.org`; override with `FORGEJO_DOMAIN=...` before `docker stack deploy`) at **vps1's** public IP — that's where the task actually runs, not vps2's or the desktop's. 2. Append `deploy/forgejo/caddy-git.conf` to vps1's `/etc/caddy/Caddyfile`, adjusting the domain if you didn't use the default, then `systemctl reload caddy`. 3. That file also resolves the registry-exposure hazard (#15 in `thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md`): `/v2/*` (the registry API) is blocked to everything except the WireGuard mesh CIDR (`10.10.0.0/24`); the git/web UI stays public. CI runners and Swarm nodes reach the registry over the mesh, not the public internet — see "Registry access from mesh clients" below. ## Registry access from mesh clients Any node that needs `docker login`/push/pull against the registry (the CI runner building/pushing images, any Swarm node pulling them, prod or beta on vps2 pulling app images) must reach `git.thermograph.org` **over the WireGuard tunnel**, not vps1's public IP — otherwise Caddy's `/v2/*` block above refuses the connection. Public DNS resolves the domain to vps1's public IP, so add a `/etc/hosts` override on each such node pinning it to vps1's WireGuard address instead: ``` echo "10.10.0.2 git.thermograph.org" | sudo tee -a /etc/hosts ``` (`10.10.0.2` is vps1's WG address per `deploy/swarm/README.md`'s peer numbering — adjust if you assigned it differently.) The git/web UI keeps working normally for everyone else since only `/v2/*` is restricted. ## Register the Actions runner Once Forgejo answers at its domain: ```bash # On the runner host (see DEPLOY-DEV.md for where that is today): # Forgejo web UI -> Site Administration -> Actions -> Runners -> Create new Runner # (or, for a repo-scoped runner: -> Settings -> Actions -> Runners) # copy the registration token, then: bash deploy/forgejo/register-lan-runner.sh https:// ``` See that script's header for exactly what it replaces (the pre-Forgejo GitHub self-hosted runner) and why it registers with two labels where there used to be two separate runners. `config.yaml`'s `runner.capacity` is raised from the tool's default of 1 to 8 (override with `CAPACITY=`) — a single PR push fires `pr-build`, `secrets-guard`, and `shell-lint` simultaneously (3 independent workflows, no `needs:` between them), so capacity 1 serializes work that could run in parallel, and even capacity 3 (an earlier, undocumented hand-tune) leaves `build-backend`/`build-frontend`/`validate-observability` queued behind those three before they get a slot. **Adding more runner capacity should mean raising this number, or adding a second runner alongside it — not putting a runner on prod or beta (vps2).** `container.docker_host: automount` gives job containers the *host's* Docker socket; on vps2 that would mean any CI job has root-equivalent access to both the prod and beta stacks running there. An earlier revision of this stack actually ran the runner as a Swarm-hosted container on the box Forgejo was pinned to and was deliberately reverted for this same class of reason — see the note at the top of `docker-stack.yml`. ## Custom CI job image (`ci-runner/`) `ci-runner/Dockerfile` still bases on `node:20-bookworm` — Node is a hard requirement, not leftover: Forgejo's runner executes `actions/checkout@v4` (used by every workflow) as `node dist/index.js` *inside the job container*, regardless of whether the workflow itself uses npm/node. (v1 of this image tried a Node-free Debian-slim base and broke every job's checkout step — `node: executable file not found` — within a minute of going live; reverted immediately.) What it actually fixes: every `docker`-labeled build-push job currently re-installs the Docker CLI on each run (`apt-get install docker.io`), which pulls in the classic builder rather than BuildKit (the classic builder mishandles `COPY --chown=` group resolution — a real bug hit during the frontend Go rewrite). `ci-runner` adds `docker-ce-cli` + `docker-buildx-plugin` (BuildKit) on top of the same Node base, plus `git`/`python3`/`python3-yaml` for the other jobs that need them (`shell-lint`, `observability-validate`). Current tag: `git.thermograph.org/emi/thermograph/ci-runner:v2` (`v1` is broken — do not register any runner against it). Rebuild/push (requires a PAT with `write:package` scope — the embedded git-remote token lacks it, same requirement documented in `build-push.yml`): ```bash docker build -t git.thermograph.org/emi/thermograph/ci-runner:vN deploy/forgejo/ci-runner docker push git.thermograph.org/emi/thermograph/ci-runner:vN ``` `register-lan-runner.sh`'s `LABELS` default points at the current tag, so fresh registrations pick it up automatically. The live runner is cut over by editing the `labels` array in `~/forgejo-runner/.runner` on the runner host (same runner id/token, no re-registration needed) and restarting the service — **verify a real job runs green under the new image before relying on it**, same way v1's break was caught. Only after that verification should the now-redundant `apt-get install docker.io` / `python3-yaml` steps be removed from the workflows that had them — removing them first would break every job still running on the stock `node:20-bookworm` image. ## Why Postgres here and not the Thermograph app's TimescaleDB Separate instance, separate network (`forgejo_net`, not the app's overlay network), separate volume. Forgejo is a distinct product with its own schema and its own backup/restore lifecycle — sharing a database with the app would couple two things that should be able to fail, migrate, and restore independently. ## Verifying ```bash docker service ls # forgejo_db, forgejo_forgejo both Running, 1/1 curl -I https://git.thermograph.org/ # 200, valid cert (Caddy's, not a new one) curl -I https://git.thermograph.org/v2/ # 403 from anywhere off the WireGuard mesh # On the runner host, after registering the runner: systemctl --user status forgejo-runner # active, both labels registered ``` ## Rollback / removal ```bash docker stack rm forgejo # volumes (forgejo_data, forgejo_db, ...) survive a stack rm — remove them # explicitly only if you actually want to destroy the Forgejo instance's data. ```