thermograph/infra/deploy/forgejo/README.md
Emi Griffith 396b80cf91 forgejo: serve dev.jinemi.com alongside git.thermograph.org
Both hostnames go on a single Caddy site block. That is load-bearing rather
than cosmetic: the /v2/* mesh-only matcher that keeps the OCI registry API
off the public internet (hazard #15) is scoped per block, so a separate
block for the new name would serve the same Forgejo with the registry open.

git.thermograph.org stays canonical — Forgejo has one ROOT_URL and builds
every absolute URL from it, so clone URLs, the OAuth callback and post-login
redirects keep naming the .org host. Flipping FORGEJO_DOMAIN additionally
requires the Google OAuth redirect URI, the registry host baked into image
names and runner labels, and the mesh /etc/hosts pins; README documents
those prerequisites.

Config verified with `caddy validate` and by asserting against the adapted
JSON that the /v2 403 guard matches both hostnames.
2026-08-01 09:01:59 -07:00

245 lines
12 KiB
Markdown

# Forgejo on the Swarm cluster
Runs as `deploy/forgejo/docker-stack.yml` — the only Swarm-scheduled workload
this cluster carries (prod and beta's app stacks are separate `docker stack
deploy`s that happen to run on the manager node, vps2 — see
`deploy/stack/README` context in `DEPLOY.md`). Pinned to the **vps1** node
(`75.119.132.91`) via the `role=forge` label from
`deploy/swarm/label-forge-node.sh`.
The Actions **runner** is deliberately *not* part of this stack: a
Swarm-scheduled runner cannot redeploy the Swarm that schedules it, so CI would
be unavailable exactly when the cluster is what needs repairing. An earlier
revision did run it as a Swarm-scheduled Docker-in-Docker sidecar pinned to the
Forgejo node; that is gone.
Runners live in two places, and only one of them counts:
- **`runner-vps2/`** — the always-on runner, a plain `restart: always` container
on vps2. This is the one CI depends on.
- **`register-lan-runner.sh`** — the desktop's systemd runner. Extra capacity.
**Nothing may assume it is up.** It was the only registered runner in the
estate until 2026-08-01, and its 21-hour outage on 2026-07-31 froze every
merge and deploy and stopped the nightly backup from firing.
## Prerequisites
1. All three nodes have joined the swarm (`deploy/swarm/`) and vps1 is
labeled `role=forge`.
2. `docker node ls` (from the manager, vps2) shows all three `Ready`.
## One-time setup: Swarm secret
One secret the stack expects to already exist (a Swarm secret, not a file —
`external: true` in the stack file, so `docker stack deploy` never creates or
sees the value, only references it):
```bash
# A strong random password for Forgejo's own Postgres (NOT related to
# Thermograph's app database — entirely separate instance/network).
openssl rand -base64 32 | docker secret create forgejo_db_password -
```
## Deploy / update
```bash
docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo
```
Re-running is safe — Swarm only touches services whose spec actually changed.
Do this **before** the DNS + Caddy step below — Caddy's reverse_proxy target
(`127.0.0.1:3080`) needs the `forgejo` service actually listening first, or
its first health check just fails harmlessly until it is.
This stack has **no auto-deploy trigger** — nothing in `.forgejo/workflows/`
redeploys it on push. A change to `docker-stack.yml` only takes effect once
someone re-runs `docker stack deploy` by hand on the manager (vps2).
`db`/`forgejo` both carry `resources.limits` (defaults: db 1 CPU/1g, forgejo 2
CPU/2g — several times observed steady-state usage), overridable with
`FORGEJO_DB_CPUS`/`FORGEJO_DB_MEMORY`/`FORGEJO_CPUS`/`FORGEJO_MEMORY` env vars
before `docker stack deploy`, same convention as the app stack.
## DNS + TLS: reusing vps1's existing Caddy, not a second reverse proxy
Forgejo is pinned to vps1 (`role=forge`) — the same box that also runs
Grafana/Loki/Alloy and the `emigriffith.dev` portfolio site, each fronted by
that host's own Caddy. A second ingress (Traefik) trying to bind the same
ports would collide with it. So there's no Traefik in this stack: `forgejo`'s
web port publishes to `127.0.0.1:3080` only (host-local), and vps1's
*existing* Caddy gets one more site block reverse-proxying to it — same
pattern as its other site blocks, same automatic-HTTPS.
1. Point the Forgejo domain (default `git.thermograph.org`; override with
`FORGEJO_DOMAIN=...` before `docker stack deploy`) at **vps1's** public IP
— that's where the task actually runs, not vps2's or the desktop's.
2. Append `deploy/forgejo/caddy-git.conf` to vps1's `/etc/caddy/Caddyfile`,
adjusting the domain if you didn't use the default, then `systemctl reload
caddy`.
3. That file also resolves the registry-exposure hazard (#15 in
`thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md`): `/v2/*` (the registry
API) is blocked to everything except the WireGuard mesh CIDR
(`10.10.0.0/24`); the git/web UI stays public. CI runners and Swarm nodes
reach the registry over the mesh, not the public internet — see "Registry
access from mesh clients" below.
## Migrating the domain (git.thermograph.org -> dev.jinemi.com)
In progress. **Both hostnames are served**; `git.thermograph.org` is still the
canonical one.
`caddy-git.conf` lists both names on a *single* site block. That is deliberate,
not cosmetic: the `/v2/*` mesh-only matcher is per-block, so a second block for
`dev.jinemi.com` would re-expose the registry API publicly and undo hazard #15.
Whatever else changes, keep the two names in one block.
Serving a second hostname is the cheap half. Forgejo has exactly one
`ROOT_URL`, and it builds every absolute URL from it, so on the non-canonical
host today:
- clone URLs on repo pages read `git.thermograph.org`,
- the OAuth callback and post-login redirect land on `git.thermograph.org`
(login is Google-SSO-only — see `docker-stack.yml`),
- webhook payload URLs and notification mail links use `git.thermograph.org`.
Browsing, git-over-HTTP and the API all work on `dev.jinemi.com` regardless.
Flipping canonical (`FORGEJO_DOMAIN=dev.jinemi.com`, then `docker stack deploy`)
needs these **first**, or it breaks login and CI rather than just renaming
things:
1. **Google OAuth** — add `https://dev.jinemi.com/user/oauth2/<name>/callback`
as an authorized redirect URI in the Google Cloud console *before* the flip.
Not in this repo; nothing here can verify it. Miss it and no one can log in,
including to undo the flip.
2. **The registry host in image names** — images are named by registry host, and
the `git.thermograph.org/` prefix is baked into the runner labels
(`register-lan-runner.sh`, `runner-vps2/README.md`), the CI-runner image
(`ci-runner/Dockerfile`), `REGISTRY_HOST` in `infra/.env.example`, and every
already-pushed tag. Renaming the registry host is a separate migration from
renaming the web UI; the two need not happen together, and doing the web UI
alone is much the smaller change.
3. **Mesh `/etc/hosts` pins** — every node in "Registry access from mesh
clients" below pins `git.thermograph.org` to `10.10.0.2`. A node pulling
under a new hostname needs its own pin, or Caddy's `/v2/*` matcher sees a
public source IP and returns 403.
4. **Runner instance URL** — registered runners hold the instance URL they were
registered with. Existing runners keep working against the old name as long
as it resolves; a *re*-registration must use whichever name is canonical then.
Until all four are handled, leaving `ROOT_URL` on `git.thermograph.org` is the
correct state, not an unfinished one.
## Registry access from mesh clients
Any node that needs `docker login`/push/pull against the registry (the CI
runner building/pushing images, any Swarm node pulling them, prod or beta on
vps2 pulling app images) must reach `git.thermograph.org` **over the
WireGuard tunnel**, not vps1's public IP — otherwise Caddy's `/v2/*` block
above refuses the connection. Public DNS resolves the domain to vps1's public
IP, so add a `/etc/hosts` override on each such node pinning it to vps1's
WireGuard address instead:
```
echo "10.10.0.2 git.thermograph.org" | sudo tee -a /etc/hosts
```
(`10.10.0.2` is vps1's WG address per `deploy/swarm/README.md`'s peer
numbering — adjust if you assigned it differently.) The git/web UI keeps
working normally for everyone else since only `/v2/*` is restricted.
## Register the Actions runner
Once Forgejo answers at its domain:
```bash
# On the runner host (see DEPLOY-DEV.md for where that is today):
# Forgejo web UI -> Site Administration -> Actions -> Runners -> Create new Runner
# (or, for a repo-scoped runner: <repo> -> Settings -> Actions -> Runners)
# copy the registration token, then:
bash deploy/forgejo/register-lan-runner.sh https://<forgejo-domain> <token>
```
See that script's header for exactly what it replaces (the pre-Forgejo GitHub
self-hosted runner) and why it registers with two labels where there used to
be two separate runners.
`config.yaml`'s `runner.capacity` is raised from the tool's default of 1 to 8
(override with `CAPACITY=`) — a single PR push fires `pr-build`,
`secrets-guard`, and `shell-lint` simultaneously (3 independent workflows,
no `needs:` between them), so capacity 1 serializes work that could run in
parallel, and even capacity 3 (an earlier, undocumented hand-tune) leaves
`build-backend`/`build-frontend`/`validate-observability` queued behind
those three before they get a slot.
**Adding more runner capacity should mean raising this number, or adding a
second runner alongside it — not putting a runner on prod or beta (vps2).**
`container.docker_host: automount` gives job containers the *host's* Docker
socket; on vps2 that would mean any CI job has root-equivalent access to both
the prod and beta stacks running there. An earlier revision of this stack
actually ran the runner as a Swarm-hosted container on the box Forgejo was
pinned to and was deliberately reverted for this same class of reason — see
the note at the top of `docker-stack.yml`.
## Custom CI job image (`ci-runner/`)
`ci-runner/Dockerfile` still bases on `node:20-bookworm` — Node is a hard
requirement, not leftover: Forgejo's runner executes `actions/checkout@v4`
(used by every workflow) as `node dist/index.js` *inside the job container*,
regardless of whether the workflow itself uses npm/node. (v1 of this image
tried a Node-free Debian-slim base and broke every job's checkout step —
`node: executable file not found` — within a minute of going live; reverted
immediately.) What it actually fixes: every `docker`-labeled build-push job
currently re-installs the Docker CLI on each run (`apt-get install
docker.io`), which pulls in the classic builder rather than BuildKit (the
classic builder mishandles `COPY --chown=<name>` group resolution — a real
bug hit during the frontend Go rewrite). `ci-runner` adds `docker-ce-cli` +
`docker-buildx-plugin` (BuildKit) on top of the same Node base, plus
`git`/`python3`/`python3-yaml` for the other jobs that need them
(`shell-lint`, `observability-validate`).
Current tag: `git.thermograph.org/emi/thermograph/ci-runner:v2` (`v1` is
broken — do not register any runner against it). Rebuild/push (requires a PAT
with `write:package` scope — the embedded git-remote token lacks it, same
requirement documented in `build-push.yml`):
```bash
docker build -t git.thermograph.org/emi/thermograph/ci-runner:vN deploy/forgejo/ci-runner
docker push git.thermograph.org/emi/thermograph/ci-runner:vN
```
`register-lan-runner.sh`'s `LABELS` default points at the current tag, so
fresh registrations pick it up automatically. The live runner is cut over by
editing the `labels` array in `~/forgejo-runner/.runner` on the runner host
(same runner id/token, no re-registration needed) and restarting the service —
**verify a real job runs green under the new image before relying on it**,
same way v1's break was caught. Only after that verification should the
now-redundant `apt-get install docker.io` / `python3-yaml` steps be removed
from the workflows that had them — removing them first would break every job
still running on the stock `node:20-bookworm` image.
## Why Postgres here and not the Thermograph app's TimescaleDB
Separate instance, separate network (`forgejo_net`, not the app's overlay
network), separate volume. Forgejo is a distinct product with its own schema
and its own backup/restore lifecycle — sharing a database with the app would
couple two things that should be able to fail, migrate, and restore
independently.
## Verifying
```bash
docker service ls # forgejo_db, forgejo_forgejo both Running, 1/1
curl -I https://git.thermograph.org/ # 200, valid cert (Caddy's, not a new one)
curl -I https://git.thermograph.org/v2/ # 403 from anywhere off the WireGuard mesh
# On the runner host, after registering the runner:
systemctl --user status forgejo-runner # active, both labels registered
```
## Rollback / removal
```bash
docker stack rm forgejo
# volumes (forgejo_data, forgejo_db, ...) survive a stack rm — remove them
# explicitly only if you actually want to destroy the Forgejo instance's data.
```