thermograph/infra/deploy/forgejo/README.md
Emi Griffith c966ac4801
All checks were successful
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / changes (pull_request) Successful in 9s
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
infra/forgejo: add resource limits and a leaner CI job image
docker-stack.yml previously had no cpu/memory limits on db or forgejo,
unlike every service in the app stack. Add generous limits (several
times observed steady-state usage) as a backstop, overridable via env
vars matching the app stack's convention.

Also add ci-runner/Dockerfile: a slim Debian base with the Docker CLI
+ buildx preinstalled, replacing node:20-bookworm (leftover from when
the frontend was Python/Jinja; nothing in CI uses npm/node anymore).
Closes a COPY --chown group-resolution bug in the classic Docker
builder that apt-get install docker.io currently pulls in on every
build-push job, and drops that install step's cost once cut over.
Built and verified locally; not yet pushed to the registry (needs a
write:package-scoped token) or wired into the live runner.
2026-07-24 00:21:29 -07:00

167 lines
7.7 KiB
Markdown

# Forgejo on the Swarm cluster
Runs as `deploy/forgejo/docker-stack.yml` — the only Swarm-scheduled workload
this cluster carries (the Thermograph app itself stays on the
Terraform-managed `docker compose` deploys; see `terraform/README.md`).
Pinned to the **beta** node (old VPS) via the `role=forge` label from
`deploy/swarm/label-forge-node.sh`.
The Actions **runner** is deliberately *not* part of this stack — it runs on
the **desktop** as a plain systemd service (`register-lan-runner.sh` below),
per `thermograph-docs/runbooks/implementation-handoff.md` Track B step 5. That's the
canonical placement; an earlier revision of this stack ran the runner as a
Swarm-scheduled Docker-in-Docker sidecar pinned to beta, which is gone now.
## Prerequisites
1. All three nodes have joined the swarm (`deploy/swarm/`) and beta is
labeled `role=forge`.
2. `docker node ls` (from the manager) shows all three `Ready`.
## One-time setup: Swarm secret
One secret the stack expects to already exist (a Swarm secret, not a file —
`external: true` in the stack file, so `docker stack deploy` never creates or
sees the value, only references it):
```bash
# A strong random password for Forgejo's own Postgres (NOT related to
# Thermograph's app database — entirely separate instance/network).
openssl rand -base64 32 | docker secret create forgejo_db_password -
```
## Deploy / update
```bash
docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo
```
Re-running is safe — Swarm only touches services whose spec actually changed.
Do this **before** the DNS + Caddy step below — Caddy's reverse_proxy target
(`127.0.0.1:3080`) needs the `forgejo` service actually listening first, or
its first health check just fails harmlessly until it is.
This stack has **no auto-deploy trigger** — nothing in `.forgejo/workflows/`
redeploys it on push. A change to `docker-stack.yml` only takes effect once
someone re-runs `docker stack deploy` by hand on the manager (prod).
`db`/`forgejo` both carry `resources.limits` (defaults: db 1 CPU/1g, forgejo 2
CPU/2g — several times observed steady-state usage), overridable with
`FORGEJO_DB_CPUS`/`FORGEJO_DB_MEMORY`/`FORGEJO_CPUS`/`FORGEJO_MEMORY` env vars
before `docker stack deploy`, same convention as the app stack.
## DNS + TLS: reusing beta's existing Caddy, not a second reverse proxy
Forgejo is pinned to beta (`role=forge`) — but beta is also **today's live
thermograph.org host**, and its Caddy already owns ports 80/443
(`/etc/caddy/Caddyfile` on that box). A second ingress (Traefik) trying to
bind the same ports would collide with it. So there's no Traefik in this
stack: `forgejo`'s web port publishes to `127.0.0.1:3080` only (host-local),
and beta's *existing* Caddy gets one more site block reverse-proxying to it —
same pattern as its `thermograph.org` block, same automatic-HTTPS.
1. Point the Forgejo domain (default `git.thermograph.org`; override with
`FORGEJO_DOMAIN=...` before `docker stack deploy`) at **beta's** public IP
— that's where the task actually runs, not prod's or the desktop's.
2. Append `deploy/forgejo/caddy-git.conf` to beta's `/etc/caddy/Caddyfile`,
adjusting the domain if you didn't use the default, then `systemctl reload
caddy`.
3. That file also resolves the registry-exposure hazard (#15 in
`thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md`): `/v2/*` (the registry
API) is blocked to everything except the WireGuard mesh CIDR
(`10.10.0.0/24`); the git/web UI stays public. CI runners and Swarm nodes
reach the registry over the mesh, not the public internet — see "Registry
access from mesh clients" below.
## Registry access from mesh clients
Any node that needs `docker login`/push/pull against the registry (the
desktop's CI runner building/pushing images, later any Swarm node pulling
them) must reach `git.thermograph.org` **over the WireGuard tunnel**, not
beta's public IP — otherwise Caddy's `/v2/*` block above refuses the
connection. Public DNS resolves the domain to beta's public IP, so add a
`/etc/hosts` override on each such node pinning it to beta's WireGuard
address instead:
```
echo "10.10.0.2 git.thermograph.org" | sudo tee -a /etc/hosts
```
(`10.10.0.2` is beta's WG address per `deploy/swarm/README.md`'s peer
numbering — adjust if you assigned it differently.) The git/web UI keeps
working normally for everyone else since only `/v2/*` is restricted.
## Register the Actions runner (on the desktop, not through Swarm)
Once Forgejo answers at its domain:
```bash
# On the desktop:
# Forgejo web UI -> Site Administration -> Actions -> Runners -> Create new Runner
# (or, for a repo-scoped runner: <repo> -> Settings -> Actions -> Runners)
# copy the registration token, then:
bash deploy/forgejo/register-lan-runner.sh https://<forgejo-domain> <token>
```
See that script's header for exactly what it replaces (the pre-Forgejo GitHub
self-hosted runner on this same machine) and why it registers with two
labels where there used to be two separate runners.
## Custom CI job image (`ci-runner/`)
`ci-runner/Dockerfile` replaces the stock `node:20-bookworm` image most
`docker`-labeled workflow jobs run in. That base is leftover from when the
frontend was Python/Jinja and the old GitHub runner also built JS assets for
it — the frontend is Go now, nothing in `.forgejo/workflows/` uses npm/node.
Every `docker`-labeled build-push job also currently re-installs the Docker
CLI on each run (`apt-get install docker.io`), which pulls in the classic
builder rather than BuildKit (the classic builder mishandles `COPY
--chown=<name>` group resolution — a real bug hit during the frontend Go
rewrite). `ci-runner` is a slim Debian base with `docker-ce-cli` +
`docker-buildx-plugin` (BuildKit) preinstalled, plus `git`/`python3`/
`python3-yaml` for the other jobs that need them (`shell-lint`,
`observability-validate`).
Built and verified locally; **not yet pushed** — pushing to
`git.thermograph.org/emi/thermograph/ci-runner` needs a PAT with
`write:package` scope (the embedded git-remote token lacks it, same
requirement documented in `backend-build-push.yml`). Once pushed:
```bash
docker build -t git.thermograph.org/emi/thermograph/ci-runner:v1 deploy/forgejo/ci-runner
docker push git.thermograph.org/emi/thermograph/ci-runner:v1
```
then cut the live runner over by editing the `labels` array in
`~/forgejo-runner/.runner` on the desktop (same runner id/token, no
re-registration) and updating `register-lan-runner.sh`'s `LABELS` default so
future re-provisioning matches. Only after that cutover is verified working
should the now-redundant `apt-get install docker.io` / `python3-yaml` steps be
removed from the workflows that had them — removing them first would break
every job still running on `node:20-bookworm`.
## Why Postgres here and not the Thermograph app's TimescaleDB
Separate instance, separate network (`forgejo_net`, not the app's compose
network), separate volume. Forgejo is a distinct product with its own schema
and its own backup/restore lifecycle — sharing a database with the app would
couple two things that should be able to fail, migrate, and restore
independently.
## Verifying
```bash
docker service ls # forgejo_db, forgejo_forgejo both Running, 1/1
curl -I https://git.thermograph.org/ # 200, valid cert (Caddy's, not a new one)
curl -I https://git.thermograph.org/v2/ # 403 from anywhere off the WireGuard mesh
# On the desktop, after registering the runner:
systemctl --user status forgejo-runner # active, both labels registered
```
## Rollback / removal
```bash
docker stack rm forgejo
# volumes (forgejo_data, forgejo_db, ...) survive a stack rm — remove them
# explicitly only if you actually want to destroy the Forgejo instance's data.
```