thermograph/infra/deploy/forgejo/README.md
Emi Griffith 53e2eb7c84
All checks were successful
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-centralis (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 9s
secrets-guard / encrypted (push) Successful in 5s
Validate observability stack / validate (push) Successful in 12s
shell-lint / shellcheck (push) Successful in 8s
Build + push images (Forgejo registry) / build-push (frontend) (push) Successful in 26s
Build + push images (Forgejo registry) / build-push (backend) (push) Successful in 1m10s
Deploy / deploy (backend) (push) Successful in 1m41s
Deploy / deploy (frontend) (push) Successful in 1m48s
secrets-guard / encrypted (pull_request) Successful in 8s
shell-lint / shellcheck (pull_request) Successful in 9s
PR build (required check) / changes (pull_request) Successful in 20s
PR build (required check) / validate-observability (pull_request) Successful in 15s
PR build (required check) / build-frontend (pull_request) Successful in 17s
PR build (required check) / build-backend (pull_request) Successful in 2m26s
PR build (required check) / gate (pull_request) Successful in 1s
registry: move the registry host to dev.jinemi.com, MCP to mcp.jinemi.com
Completes the Forgejo domain migration. ROOT_URL moved to dev.jinemi.com
earlier; the registry half was deliberately deferred. Every image name,
REGISTRY_HOST default, runner label and --add-host pin now names
dev.jinemi.com, so the registry host and the bearer-token realm agree again.

Both names address the same Forgejo, so no image needs re-pushing and a
rollback to a tag pushed under the old prefix still resolves.
git.thermograph.org therefore stays served off the same Caddy site block --
one block, so the /v2/* mesh-only matcher keeps covering both names -- for
pre-migration tags and for runners holding it as their registered instance
URL.

Mesh clients now pin both names in /etc/hosts: the new one as registry host
and token realm, the old one for pre-migration tags. runner-vps2/config.yaml
carries both --add-host entries for the same reason.

Also renames Centralis' endpoint to mcp.jinemi.com in the two places this
repo names it; Centralis itself is provisioned outside this repo.

Host-side steps this cannot do (documented in deploy/forgejo/README.md,
"Host-side steps"): the Forgejo Actions variable REGISTRY_HOST, docker login
against the new host, and the /etc/hosts pins.
2026-08-01 16:06:22 -07:00

337 lines
17 KiB
Markdown

# Forgejo on the Swarm cluster
Runs as `deploy/forgejo/docker-stack.yml` — the only Swarm-scheduled workload
this cluster carries (prod and beta's app stacks are separate `docker stack
deploy`s that happen to run on the manager node, vps2 — see
`DEPLOY.md`). Pinned to the **vps1** node
(`75.119.132.91`) via the `role=forge` label from
`deploy/swarm/label-forge-node.sh`.
The Actions **runner** is deliberately *not* part of this stack: a
Swarm-scheduled runner cannot redeploy the Swarm that schedules it, so CI would
be unavailable exactly when the cluster is what needs repairing. An earlier
revision did run it as a Swarm-scheduled Docker-in-Docker sidecar pinned to the
Forgejo node; that is gone.
Runners live in two places, and only one of them counts:
- **`runner-vps2/`** — the always-on runner, a plain `restart: always` container
on vps2. This is the one CI depends on.
- **`register-lan-runner.sh`** — the desktop's systemd runner. Extra capacity.
**Nothing may assume it is up.** It was the only registered runner in the
estate until 2026-08-01, and its 21-hour outage on 2026-07-31 froze every
merge and deploy and stopped the nightly backup from firing.
## Prerequisites
1. All three nodes have joined the swarm (`deploy/swarm/`) and vps1 is
labeled `role=forge`.
2. `docker node ls` (from the manager, vps2) shows all three `Ready`.
## One-time setup: Swarm secret
One secret the stack expects to already exist (a Swarm secret, not a file —
`external: true` in the stack file, so `docker stack deploy` never creates or
sees the value, only references it):
```bash
# A strong random password for Forgejo's own Postgres (NOT related to
# Thermograph's app database — entirely separate instance/network).
openssl rand -base64 32 | docker secret create forgejo_db_password -
```
## Deploy / update
```bash
docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo
```
Re-running is safe — Swarm only touches services whose spec actually changed.
Do this **before** the DNS + Caddy step below — Caddy's reverse_proxy target
(`127.0.0.1:3080`) needs the `forgejo` service actually listening first, or
its first health check just fails harmlessly until it is.
This stack has **no auto-deploy trigger** — nothing in `.forgejo/workflows/`
redeploys it on push. A change to `docker-stack.yml` only takes effect once
someone re-runs `docker stack deploy` by hand on the manager (vps2).
`db`/`forgejo` both carry `resources.limits` (defaults: db 1 CPU/1g, forgejo 2
CPU/2g — several times observed steady-state usage), overridable with
`FORGEJO_DB_CPUS`/`FORGEJO_DB_MEMORY`/`FORGEJO_CPUS`/`FORGEJO_MEMORY` env vars
before `docker stack deploy`, same convention as the app stack.
## DNS + TLS: reusing vps1's existing Caddy, not a second reverse proxy
Forgejo is pinned to vps1 (`role=forge`) — the same box that also runs
Grafana/Loki/Alloy and the `emigriffith.dev` portfolio site, each fronted by
that host's own Caddy. A second ingress (Traefik) trying to bind the same
ports would collide with it. So there's no Traefik in this stack: `forgejo`'s
web port publishes to `127.0.0.1:3080` only (host-local), and vps1's
*existing* Caddy gets one more site block reverse-proxying to it — same
pattern as its other site blocks, same automatic-HTTPS.
1. Point the Forgejo domain (default `dev.jinemi.com`; override with
`FORGEJO_DOMAIN=...` before `docker stack deploy`) at **vps1's** public IP
— that's where the task actually runs, not vps2's or the desktop's.
2. Append `deploy/forgejo/caddy-git.conf` to vps1's `/etc/caddy/Caddyfile`,
adjusting the domain if you didn't use the default, then `systemctl reload
caddy`.
3. That file also resolves the registry-exposure hazard (#15 in
`thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md`): `/v2/*` (the registry
API) is blocked to everything except the WireGuard mesh CIDR
(`10.10.0.0/24`); the git/web UI stays public. CI runners and Swarm nodes
reach the registry over the mesh, not the public internet — see "Registry
access from mesh clients" below.
## Migrating the domain (git.thermograph.org -> dev.jinemi.com)
**Web UI and OAuth: done.** `dev.jinemi.com` is Forgejo's `ROOT_URL`
(`FORGEJO_DOMAIN` in `docker-stack.yml`), so it is the name Forgejo generates in
clone URLs, the Google OAuth callback, webhook payload URLs and mail links.
**Registry: done in this repo, and it needs host-side steps to match.** Every
image name, `REGISTRY_HOST` default, runner label and `--add-host` pin under
version control now says `dev.jinemi.com`. None of that is self-applying — see
"Host-side steps" below before expecting a build to push.
Both names address the *same* Forgejo and therefore the same packages. A tag
pushed as `git.thermograph.org/jinemi/thermograph/backend:sha-abc` is pullable
as `dev.jinemi.com/jinemi/thermograph/backend:sha-abc` — the host part is only
how you reach the registry, not part of the package's identity. That is what
makes this migration safe to do in one step instead of a dual-push period: no
image needs re-pushing, and a rollback to an old tag still resolves.
`caddy-git.conf` lists both names on a *single* site block. That is deliberate,
not cosmetic: the `/v2/*` mesh-only matcher is per-block, so a second block for
either name would re-expose the registry API publicly and undo hazard #15.
Whatever else changes, keep the two names in one block.
**`git.thermograph.org` must stay served.** The list of reasons is shorter than
it was, but it is not empty, and neither remaining item is a bookmark:
- image tags already pushed under the old prefix name that host, so a rollback
to one asks for `git.thermograph.org/...` even though nothing builds that
name any more;
- registered runners hold the instance URL they registered with
(`https://git.thermograph.org`); they keep working while that name resolves,
and stop the moment it doesn't. Re-registering them is the only way to retire
the name, and it is not required for this migration.
### Host-side steps
None of these live in the repo, and CI will not do them for you. In order:
1. **`/etc/hosts` on every mesh client** — pin `dev.jinemi.com` as well as
`git.thermograph.org` (both, for the reasons under "Registry access from
mesh clients" below). A client that resolves the new name publicly gets a
403 from the `/v2/*` matcher, which reads exactly like a bad credential.
2. **`docker login dev.jinemi.com`** on each host that pushes or pulls — the
CI runners, vps2 (prod and beta), the desktop. A login against the old name
does **not** carry over: docker keys stored credentials by registry host.
3. **The Forgejo Actions variable `REGISTRY_HOST`**`dev.jinemi.com`
(Site Administration → Actions → Variables, or the repo's own). This is what
`build-push.yml` reads; the defaults in this repo are fallbacks for by-hand
runs, so leaving the variable stale silently keeps CI on the old name.
4. **Re-register the runners** against `https://dev.jinemi.com`, *optionally*
see above. Until you do, leave `git.thermograph.org` served.
Verify with a push before assuming: the failure mode for a missed step 1 or 2
is `unauthorized: reqPackageAccess`, which names neither.
### Before changing ROOT_URL again
Login is Google-SSO-only, so the Google OAuth client must already carry a
matching redirect URI, or nobody can reach the UI — including to undo the
change. Both of these are registered on client
`337247886376-4360kepi20f24pue62ptg0lnqhcht46l.apps.googleusercontent.com`:
https://dev.jinemi.com/user/oauth2/google/callback
https://git.thermograph.org/user/oauth2/google/callback
Verify rather than assume. Google answers this with no credentials: a
registered URI serves the sign-in page, an unregistered one returns
`redirect_uri_mismatch`. Substitute the host you intend to make canonical:
```sh
CID=337247886376-4360kepi20f24pue62ptg0lnqhcht46l.apps.googleusercontent.com
curl -sL "https://accounts.google.com/o/oauth2/v2/auth?client_id=$CID&redirect_uri=https%3A%2F%2F<host>%2Fuser%2Foauth2%2Fgoogle%2Fcallback&response_type=code&scope=openid+email+profile&state=probe" \
| grep -q redirect_uri_mismatch && echo REJECTED || echo ACCEPTED
```
The stack has **no auto-deploy**: a `FORGEJO_DOMAIN` change here only takes
effect when someone re-runs `docker stack deploy` by hand on the manager (vps2)
— see "Deploy / update" above. Until then the running service keeps whatever
`ROOT_URL` it was last deployed with, regardless of what this file says.
## Registry access from mesh clients
Any node that needs `docker login`/push/pull against the registry (the CI
runner building/pushing images, any Swarm node pulling them, prod or beta on
vps2 pulling app images) must reach Forgejo **over the WireGuard tunnel**, not
vps1's public IP — otherwise Caddy's `/v2/*` block above refuses the
connection. Public DNS resolves both names to vps1's public IP, so add an
`/etc/hosts` override on each such node pinning them to vps1's WireGuard
address instead:
```
echo "10.10.0.2 dev.jinemi.com git.thermograph.org" | sudo tee -a /etc/hosts
```
(`10.10.0.2` is vps1's WG address per `deploy/swarm/README.md`'s peer
numbering — adjust if you assigned it differently.) The git/web UI keeps
working normally for everyone else since only `/v2/*` is restricted.
**Both names, not just the one in the image.** `dev.jinemi.com` now carries two
independent jobs and needs the pin for each: it is the host in every image
name, *and* it is `ROOT_URL`, from which Forgejo derives the registry's
bearer-token realm — so a client is sent to `https://<ROOT_URL host>/v2/token`
to collect a token no matter which name it dialled. `git.thermograph.org` stays
pinned for a third reason: image tags pushed before the migration name it, and
a rollback to one of those dials it directly.
This is not hypothetical. When `ROOT_URL` first became `dev.jinemi.com` while
images were still named `git.thermograph.org`, the token request took the
public route, the `/v2/*` matcher answered 403, and docker fell back to
anonymous — every push failed with `unauthorized: reqPackageAccess`, which
reads exactly like a revoked token or a missing scope. It was neither. Pinning
one name and not the other reproduces it in either direction.
The realm moves whenever `ROOT_URL` does, so read it rather than assuming:
```sh
curl -sI https://dev.jinemi.com/v2/ | grep -i www-authenticate
```
These `/etc/hosts` pins are host state. Nothing in this repo writes them, so a
reprovisioned node has to be pinned again before it can pull.
## Register the Actions runner
Once Forgejo answers at its domain:
```bash
# On the runner host (see DEPLOY-DEV.md for where that is today):
# Forgejo web UI -> Site Administration -> Actions -> Runners -> Create new Runner
# (or, for a repo-scoped runner: <repo> -> Settings -> Actions -> Runners)
# copy the registration token, then:
bash deploy/forgejo/register-lan-runner.sh https://<forgejo-domain> <token>
```
See that script's header for exactly what it replaces (the pre-Forgejo GitHub
self-hosted runner) and why it registers with two labels where there used to
be two separate runners.
`config.yaml`'s `runner.capacity` is raised from the tool's default of 1 to 8
(override with `CAPACITY=`) — a single PR push fires `pr-build`,
`secrets-guard`, and `shell-lint` simultaneously (3 independent workflows,
no `needs:` between them), so capacity 1 serializes work that could run in
parallel, and even capacity 3 (an earlier, undocumented hand-tune) leaves
`build-backend`/`build-frontend`/`validate-observability` queued behind
those three before they get a slot.
**This paragraph used to say a runner must never go on vps2. That was reversed
deliberately on 2026-08-01, and the reasoning is worth keeping rather than
quietly deleting.**
The original objection stands on its merits: `container.docker_host: automount`
gives job containers the *host's* Docker socket, so a runner on vps2 means any
CI job has root-equivalent access to both the prod and beta stacks running
there. An earlier revision of this stack ran the runner as a Swarm-hosted
container on the Forgejo node and was reverted for the same class of reason.
What changed is that the alternative turned out to be worse. The desktop was
the estate's **only** registered runner, and when it went offline on
2026-07-31 nothing merged, nothing deployed, and the nightly backup did not
fire for 21 hours — on a repo where every branch is protected and every change
is a PR, a single absent runner freezes the whole estate. Availability of CI
is not a convenience here; it is what the backup and the deploy path both
hang off.
So the runner on vps2 (`runner-vps2/`) is a considered trade, not an
oversight, and it is bounded rather than assumed benign:
- `capacity: 1` — one job at a time, so a CI burst cannot contend with
`thermograph_web` for the box's six cores.
- `--cpus=2 --memory=4g` on job containers, set in `container.options`. Job
containers are siblings of the runner rather than children, so the runner's
own compose limits do not reach them; that option is the only lever that
does.
- `valid_volumes: []` — a job cannot bind-mount a host path, which is the
difference between "a job can read `/etc/thermograph.env`" and "a job
cannot", on the box where that file is prod's.
- The compose project joins no `thermograph` network.
Be honest about what that buys: it is defence against **accident**, not
against a hostile workflow author. Forgejo's own database on vps1 already
stores `VPS2_SSH_KEY`, which is root on vps2, so the material was already
reachable — what changed is that it is now reachable by a *job* rather than
only at rest. On a two-person estate where both people can already SSH to that
box as root, that is the honest boundary.
**If you are adding capacity rather than redundancy, raise `capacity` or add a
runner on a box that hosts nothing.** The vps2 runner exists so that CI has a
second home, not because prod is a good place to run CI.
## Custom CI job image (`ci-runner/`)
`ci-runner/Dockerfile` still bases on `node:20-bookworm` — Node is a hard
requirement, not leftover: Forgejo's runner executes `actions/checkout@v4`
(used by every workflow) as `node dist/index.js` *inside the job container*,
regardless of whether the workflow itself uses npm/node. (v1 of this image
tried a Node-free Debian-slim base and broke every job's checkout step —
`node: executable file not found` — within a minute of going live; reverted
immediately.) What it actually fixes: every `docker`-labeled build-push job
currently re-installs the Docker CLI on each run (`apt-get install
docker.io`), which pulls in the classic builder rather than BuildKit (the
classic builder mishandles `COPY --chown=<name>` group resolution — a real
bug hit during the frontend Go rewrite). `ci-runner` adds `docker-ce-cli` +
`docker-buildx-plugin` (BuildKit) on top of the same Node base, plus
`git`/`python3`/`python3-yaml` for the other jobs that need them
(`shell-lint`, `observability-validate`).
Current tag: `dev.jinemi.com/jinemi/thermograph/ci-runner:v2` (`v1` is
broken — do not register any runner against it). Rebuild/push (requires a PAT
with `write:package` scope — the embedded git-remote token lacks it, same
requirement documented in `build-push.yml`):
```bash
docker build -t dev.jinemi.com/jinemi/thermograph/ci-runner:vN deploy/forgejo/ci-runner
docker push dev.jinemi.com/jinemi/thermograph/ci-runner:vN
```
`register-lan-runner.sh`'s `LABELS` default points at the current tag, so
fresh registrations pick it up automatically. The live runner is cut over by
editing the `labels` array in `~/forgejo-runner/.runner` on the runner host
(same runner id/token, no re-registration needed) and restarting the service —
**verify a real job runs green under the new image before relying on it**,
same way v1's break was caught. Only after that verification should the
now-redundant `apt-get install docker.io` / `python3-yaml` steps be removed
from the workflows that had them — removing them first would break every job
still running on the stock `node:20-bookworm` image.
## Why Postgres here and not the Thermograph app's TimescaleDB
Separate instance, separate network (`forgejo_net`, not the app's overlay
network), separate volume. Forgejo is a distinct product with its own schema
and its own backup/restore lifecycle — sharing a database with the app would
couple two things that should be able to fail, migrate, and restore
independently.
## Verifying
```bash
docker service ls # forgejo_db, forgejo_forgejo both Running, 1/1
curl -I https://dev.jinemi.com/ # 200, valid cert (Caddy's, not a new one)
curl -I https://dev.jinemi.com/v2/ # 403 from anywhere off the WireGuard mesh
curl -I https://git.thermograph.org/ # 200 too — the alias must stay served
# On the runner host, after registering the runner:
systemctl --user status forgejo-runner # active, both labels registered
```
## Rollback / removal
```bash
docker stack rm forgejo
# volumes (forgejo_data, forgejo_db, ...) survive a stack rm — remove them
# explicitly only if you actually want to destroy the Forgejo instance's data.
```