All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
secrets-guard / encrypted (pull_request) Successful in 9s
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 1s
Forgejo derives every absolute URL from its single ROOT_URL, so this is what moves clone URLs, the Google OAuth callback, webhook payload URLs and mail links onto the new domain. Both redirect URIs are registered on the Google OAuth client and were verified against Google before the flip — login is SSO-only, so a mismatch locks everyone out of the UI including out of undoing it. README documents the probe so the check is repeatable rather than assumed. git.thermograph.org stays on the same Caddy site block and stays load-bearing: the registry host baked into image names, the mesh /etc/hosts pins that make the /v2/* matcher see a mesh source IP, and the runners' registered instance URL all still use it. Renaming the registry is a separate migration; the docs now say so where someone would otherwise retire the name as dead weight. The stack has no auto-deploy, so this takes effect on a manual `docker stack deploy` from the manager.
250 lines
12 KiB
Markdown
250 lines
12 KiB
Markdown
# Forgejo on the Swarm cluster
|
|
|
|
Runs as `deploy/forgejo/docker-stack.yml` — the only Swarm-scheduled workload
|
|
this cluster carries (prod and beta's app stacks are separate `docker stack
|
|
deploy`s that happen to run on the manager node, vps2 — see
|
|
`deploy/stack/README` context in `DEPLOY.md`). Pinned to the **vps1** node
|
|
(`75.119.132.91`) via the `role=forge` label from
|
|
`deploy/swarm/label-forge-node.sh`.
|
|
|
|
The Actions **runner** is deliberately *not* part of this stack: a
|
|
Swarm-scheduled runner cannot redeploy the Swarm that schedules it, so CI would
|
|
be unavailable exactly when the cluster is what needs repairing. An earlier
|
|
revision did run it as a Swarm-scheduled Docker-in-Docker sidecar pinned to the
|
|
Forgejo node; that is gone.
|
|
|
|
Runners live in two places, and only one of them counts:
|
|
|
|
- **`runner-vps2/`** — the always-on runner, a plain `restart: always` container
|
|
on vps2. This is the one CI depends on.
|
|
- **`register-lan-runner.sh`** — the desktop's systemd runner. Extra capacity.
|
|
**Nothing may assume it is up.** It was the only registered runner in the
|
|
estate until 2026-08-01, and its 21-hour outage on 2026-07-31 froze every
|
|
merge and deploy and stopped the nightly backup from firing.
|
|
|
|
## Prerequisites
|
|
|
|
1. All three nodes have joined the swarm (`deploy/swarm/`) and vps1 is
|
|
labeled `role=forge`.
|
|
2. `docker node ls` (from the manager, vps2) shows all three `Ready`.
|
|
|
|
## One-time setup: Swarm secret
|
|
|
|
One secret the stack expects to already exist (a Swarm secret, not a file —
|
|
`external: true` in the stack file, so `docker stack deploy` never creates or
|
|
sees the value, only references it):
|
|
|
|
```bash
|
|
# A strong random password for Forgejo's own Postgres (NOT related to
|
|
# Thermograph's app database — entirely separate instance/network).
|
|
openssl rand -base64 32 | docker secret create forgejo_db_password -
|
|
```
|
|
|
|
## Deploy / update
|
|
|
|
```bash
|
|
docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo
|
|
```
|
|
|
|
Re-running is safe — Swarm only touches services whose spec actually changed.
|
|
Do this **before** the DNS + Caddy step below — Caddy's reverse_proxy target
|
|
(`127.0.0.1:3080`) needs the `forgejo` service actually listening first, or
|
|
its first health check just fails harmlessly until it is.
|
|
|
|
This stack has **no auto-deploy trigger** — nothing in `.forgejo/workflows/`
|
|
redeploys it on push. A change to `docker-stack.yml` only takes effect once
|
|
someone re-runs `docker stack deploy` by hand on the manager (vps2).
|
|
|
|
`db`/`forgejo` both carry `resources.limits` (defaults: db 1 CPU/1g, forgejo 2
|
|
CPU/2g — several times observed steady-state usage), overridable with
|
|
`FORGEJO_DB_CPUS`/`FORGEJO_DB_MEMORY`/`FORGEJO_CPUS`/`FORGEJO_MEMORY` env vars
|
|
before `docker stack deploy`, same convention as the app stack.
|
|
|
|
## DNS + TLS: reusing vps1's existing Caddy, not a second reverse proxy
|
|
|
|
Forgejo is pinned to vps1 (`role=forge`) — the same box that also runs
|
|
Grafana/Loki/Alloy and the `emigriffith.dev` portfolio site, each fronted by
|
|
that host's own Caddy. A second ingress (Traefik) trying to bind the same
|
|
ports would collide with it. So there's no Traefik in this stack: `forgejo`'s
|
|
web port publishes to `127.0.0.1:3080` only (host-local), and vps1's
|
|
*existing* Caddy gets one more site block reverse-proxying to it — same
|
|
pattern as its other site blocks, same automatic-HTTPS.
|
|
|
|
1. Point the Forgejo domain (default `git.thermograph.org`; override with
|
|
`FORGEJO_DOMAIN=...` before `docker stack deploy`) at **vps1's** public IP
|
|
— that's where the task actually runs, not vps2's or the desktop's.
|
|
2. Append `deploy/forgejo/caddy-git.conf` to vps1's `/etc/caddy/Caddyfile`,
|
|
adjusting the domain if you didn't use the default, then `systemctl reload
|
|
caddy`.
|
|
3. That file also resolves the registry-exposure hazard (#15 in
|
|
`thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md`): `/v2/*` (the registry
|
|
API) is blocked to everything except the WireGuard mesh CIDR
|
|
(`10.10.0.0/24`); the git/web UI stays public. CI runners and Swarm nodes
|
|
reach the registry over the mesh, not the public internet — see "Registry
|
|
access from mesh clients" below.
|
|
|
|
## Migrating the domain (git.thermograph.org -> dev.jinemi.com)
|
|
|
|
**Web UI and OAuth: done.** `dev.jinemi.com` is Forgejo's `ROOT_URL`
|
|
(`FORGEJO_DOMAIN` in `docker-stack.yml`), so it is the name Forgejo generates in
|
|
clone URLs, the Google OAuth callback, webhook payload URLs and mail links.
|
|
|
|
**Registry: deliberately not done.** Renaming the registry host is a separate
|
|
migration from renaming the web UI, and the two need not happen together.
|
|
|
|
`caddy-git.conf` lists both names on a *single* site block. That is deliberate,
|
|
not cosmetic: the `/v2/*` mesh-only matcher is per-block, so a second block for
|
|
either name would re-expose the registry API publicly and undo hazard #15.
|
|
Whatever else changes, keep the two names in one block.
|
|
|
|
**`git.thermograph.org` must stay served.** It is not a courtesy redirect for
|
|
old bookmarks — CI resolves it:
|
|
|
|
- images are named by registry host, and the `git.thermograph.org/` prefix is
|
|
baked into the runner labels (`register-lan-runner.sh`,
|
|
`runner-vps2/README.md`), the CI-runner image (`ci-runner/Dockerfile`),
|
|
`REGISTRY_HOST` in `infra/.env.example`, and every already-pushed tag;
|
|
- every mesh client in "Registry access from mesh clients" below pins
|
|
`git.thermograph.org` to `10.10.0.2` in `/etc/hosts`, which is what makes
|
|
Caddy's `/v2/*` matcher see a mesh source IP instead of returning 403;
|
|
- registered runners hold the instance URL they registered with
|
|
(`https://git.thermograph.org`); they keep working while that name resolves.
|
|
|
|
### Before changing ROOT_URL again
|
|
|
|
Login is Google-SSO-only, so the Google OAuth client must already carry a
|
|
matching redirect URI, or nobody can reach the UI — including to undo the
|
|
change. Both of these are registered on client
|
|
`337247886376-4360kepi20f24pue62ptg0lnqhcht46l.apps.googleusercontent.com`:
|
|
|
|
https://dev.jinemi.com/user/oauth2/google/callback
|
|
https://git.thermograph.org/user/oauth2/google/callback
|
|
|
|
Verify rather than assume. Google answers this with no credentials: a
|
|
registered URI serves the sign-in page, an unregistered one returns
|
|
`redirect_uri_mismatch`. Substitute the host you intend to make canonical:
|
|
|
|
```sh
|
|
CID=337247886376-4360kepi20f24pue62ptg0lnqhcht46l.apps.googleusercontent.com
|
|
curl -sL "https://accounts.google.com/o/oauth2/v2/auth?client_id=$CID&redirect_uri=https%3A%2F%2F<host>%2Fuser%2Foauth2%2Fgoogle%2Fcallback&response_type=code&scope=openid+email+profile&state=probe" \
|
|
| grep -q redirect_uri_mismatch && echo REJECTED || echo ACCEPTED
|
|
```
|
|
|
|
The stack has **no auto-deploy**: a `FORGEJO_DOMAIN` change here only takes
|
|
effect when someone re-runs `docker stack deploy` by hand on the manager (vps2)
|
|
— see "Deploy / update" above. Until then the running service keeps whatever
|
|
`ROOT_URL` it was last deployed with, regardless of what this file says.
|
|
|
|
## Registry access from mesh clients
|
|
|
|
Any node that needs `docker login`/push/pull against the registry (the CI
|
|
runner building/pushing images, any Swarm node pulling them, prod or beta on
|
|
vps2 pulling app images) must reach `git.thermograph.org` **over the
|
|
WireGuard tunnel**, not vps1's public IP — otherwise Caddy's `/v2/*` block
|
|
above refuses the connection. Public DNS resolves the domain to vps1's public
|
|
IP, so add a `/etc/hosts` override on each such node pinning it to vps1's
|
|
WireGuard address instead:
|
|
|
|
```
|
|
echo "10.10.0.2 git.thermograph.org" | sudo tee -a /etc/hosts
|
|
```
|
|
|
|
(`10.10.0.2` is vps1's WG address per `deploy/swarm/README.md`'s peer
|
|
numbering — adjust if you assigned it differently.) The git/web UI keeps
|
|
working normally for everyone else since only `/v2/*` is restricted.
|
|
|
|
## Register the Actions runner
|
|
|
|
Once Forgejo answers at its domain:
|
|
|
|
```bash
|
|
# On the runner host (see DEPLOY-DEV.md for where that is today):
|
|
# Forgejo web UI -> Site Administration -> Actions -> Runners -> Create new Runner
|
|
# (or, for a repo-scoped runner: <repo> -> Settings -> Actions -> Runners)
|
|
# copy the registration token, then:
|
|
bash deploy/forgejo/register-lan-runner.sh https://<forgejo-domain> <token>
|
|
```
|
|
|
|
See that script's header for exactly what it replaces (the pre-Forgejo GitHub
|
|
self-hosted runner) and why it registers with two labels where there used to
|
|
be two separate runners.
|
|
|
|
`config.yaml`'s `runner.capacity` is raised from the tool's default of 1 to 8
|
|
(override with `CAPACITY=`) — a single PR push fires `pr-build`,
|
|
`secrets-guard`, and `shell-lint` simultaneously (3 independent workflows,
|
|
no `needs:` between them), so capacity 1 serializes work that could run in
|
|
parallel, and even capacity 3 (an earlier, undocumented hand-tune) leaves
|
|
`build-backend`/`build-frontend`/`validate-observability` queued behind
|
|
those three before they get a slot.
|
|
|
|
**Adding more runner capacity should mean raising this number, or adding a
|
|
second runner alongside it — not putting a runner on prod or beta (vps2).**
|
|
`container.docker_host: automount` gives job containers the *host's* Docker
|
|
socket; on vps2 that would mean any CI job has root-equivalent access to both
|
|
the prod and beta stacks running there. An earlier revision of this stack
|
|
actually ran the runner as a Swarm-hosted container on the box Forgejo was
|
|
pinned to and was deliberately reverted for this same class of reason — see
|
|
the note at the top of `docker-stack.yml`.
|
|
|
|
## Custom CI job image (`ci-runner/`)
|
|
|
|
`ci-runner/Dockerfile` still bases on `node:20-bookworm` — Node is a hard
|
|
requirement, not leftover: Forgejo's runner executes `actions/checkout@v4`
|
|
(used by every workflow) as `node dist/index.js` *inside the job container*,
|
|
regardless of whether the workflow itself uses npm/node. (v1 of this image
|
|
tried a Node-free Debian-slim base and broke every job's checkout step —
|
|
`node: executable file not found` — within a minute of going live; reverted
|
|
immediately.) What it actually fixes: every `docker`-labeled build-push job
|
|
currently re-installs the Docker CLI on each run (`apt-get install
|
|
docker.io`), which pulls in the classic builder rather than BuildKit (the
|
|
classic builder mishandles `COPY --chown=<name>` group resolution — a real
|
|
bug hit during the frontend Go rewrite). `ci-runner` adds `docker-ce-cli` +
|
|
`docker-buildx-plugin` (BuildKit) on top of the same Node base, plus
|
|
`git`/`python3`/`python3-yaml` for the other jobs that need them
|
|
(`shell-lint`, `observability-validate`).
|
|
|
|
Current tag: `git.thermograph.org/emi/thermograph/ci-runner:v2` (`v1` is
|
|
broken — do not register any runner against it). Rebuild/push (requires a PAT
|
|
with `write:package` scope — the embedded git-remote token lacks it, same
|
|
requirement documented in `build-push.yml`):
|
|
|
|
```bash
|
|
docker build -t git.thermograph.org/emi/thermograph/ci-runner:vN deploy/forgejo/ci-runner
|
|
docker push git.thermograph.org/emi/thermograph/ci-runner:vN
|
|
```
|
|
|
|
`register-lan-runner.sh`'s `LABELS` default points at the current tag, so
|
|
fresh registrations pick it up automatically. The live runner is cut over by
|
|
editing the `labels` array in `~/forgejo-runner/.runner` on the runner host
|
|
(same runner id/token, no re-registration needed) and restarting the service —
|
|
**verify a real job runs green under the new image before relying on it**,
|
|
same way v1's break was caught. Only after that verification should the
|
|
now-redundant `apt-get install docker.io` / `python3-yaml` steps be removed
|
|
from the workflows that had them — removing them first would break every job
|
|
still running on the stock `node:20-bookworm` image.
|
|
|
|
## Why Postgres here and not the Thermograph app's TimescaleDB
|
|
|
|
Separate instance, separate network (`forgejo_net`, not the app's overlay
|
|
network), separate volume. Forgejo is a distinct product with its own schema
|
|
and its own backup/restore lifecycle — sharing a database with the app would
|
|
couple two things that should be able to fail, migrate, and restore
|
|
independently.
|
|
|
|
## Verifying
|
|
|
|
```bash
|
|
docker service ls # forgejo_db, forgejo_forgejo both Running, 1/1
|
|
curl -I https://git.thermograph.org/ # 200, valid cert (Caddy's, not a new one)
|
|
curl -I https://git.thermograph.org/v2/ # 403 from anywhere off the WireGuard mesh
|
|
# On the runner host, after registering the runner:
|
|
systemctl --user status forgejo-runner # active, both labels registered
|
|
```
|
|
|
|
## Rollback / removal
|
|
|
|
```bash
|
|
docker stack rm forgejo
|
|
# volumes (forgejo_data, forgejo_db, ...) survive a stack rm — remove them
|
|
# explicitly only if you actually want to destroy the Forgejo instance's data.
|
|
```
|