All checks were successful
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-centralis (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 9s
secrets-guard / encrypted (push) Successful in 5s
Validate observability stack / validate (push) Successful in 12s
shell-lint / shellcheck (push) Successful in 8s
Build + push images (Forgejo registry) / build-push (frontend) (push) Successful in 26s
Build + push images (Forgejo registry) / build-push (backend) (push) Successful in 1m10s
Deploy / deploy (backend) (push) Successful in 1m41s
Deploy / deploy (frontend) (push) Successful in 1m48s
secrets-guard / encrypted (pull_request) Successful in 8s
shell-lint / shellcheck (pull_request) Successful in 9s
PR build (required check) / changes (pull_request) Successful in 20s
PR build (required check) / validate-observability (pull_request) Successful in 15s
PR build (required check) / build-frontend (pull_request) Successful in 17s
PR build (required check) / build-backend (pull_request) Successful in 2m26s
PR build (required check) / gate (pull_request) Successful in 1s
Completes the Forgejo domain migration. ROOT_URL moved to dev.jinemi.com earlier; the registry half was deliberately deferred. Every image name, REGISTRY_HOST default, runner label and --add-host pin now names dev.jinemi.com, so the registry host and the bearer-token realm agree again. Both names address the same Forgejo, so no image needs re-pushing and a rollback to a tag pushed under the old prefix still resolves. git.thermograph.org therefore stays served off the same Caddy site block -- one block, so the /v2/* mesh-only matcher keeps covering both names -- for pre-migration tags and for runners holding it as their registered instance URL. Mesh clients now pin both names in /etc/hosts: the new one as registry host and token realm, the old one for pre-migration tags. runner-vps2/config.yaml carries both --add-host entries for the same reason. Also renames Centralis' endpoint to mcp.jinemi.com in the two places this repo names it; Centralis itself is provisioned outside this repo. Host-side steps this cannot do (documented in deploy/forgejo/README.md, "Host-side steps"): the Forgejo Actions variable REGISTRY_HOST, docker login against the new host, and the /etc/hosts pins.
337 lines
17 KiB
Markdown
337 lines
17 KiB
Markdown
# Forgejo on the Swarm cluster
|
|
|
|
Runs as `deploy/forgejo/docker-stack.yml` — the only Swarm-scheduled workload
|
|
this cluster carries (prod and beta's app stacks are separate `docker stack
|
|
deploy`s that happen to run on the manager node, vps2 — see
|
|
`DEPLOY.md`). Pinned to the **vps1** node
|
|
(`75.119.132.91`) via the `role=forge` label from
|
|
`deploy/swarm/label-forge-node.sh`.
|
|
|
|
The Actions **runner** is deliberately *not* part of this stack: a
|
|
Swarm-scheduled runner cannot redeploy the Swarm that schedules it, so CI would
|
|
be unavailable exactly when the cluster is what needs repairing. An earlier
|
|
revision did run it as a Swarm-scheduled Docker-in-Docker sidecar pinned to the
|
|
Forgejo node; that is gone.
|
|
|
|
Runners live in two places, and only one of them counts:
|
|
|
|
- **`runner-vps2/`** — the always-on runner, a plain `restart: always` container
|
|
on vps2. This is the one CI depends on.
|
|
- **`register-lan-runner.sh`** — the desktop's systemd runner. Extra capacity.
|
|
**Nothing may assume it is up.** It was the only registered runner in the
|
|
estate until 2026-08-01, and its 21-hour outage on 2026-07-31 froze every
|
|
merge and deploy and stopped the nightly backup from firing.
|
|
|
|
## Prerequisites
|
|
|
|
1. All three nodes have joined the swarm (`deploy/swarm/`) and vps1 is
|
|
labeled `role=forge`.
|
|
2. `docker node ls` (from the manager, vps2) shows all three `Ready`.
|
|
|
|
## One-time setup: Swarm secret
|
|
|
|
One secret the stack expects to already exist (a Swarm secret, not a file —
|
|
`external: true` in the stack file, so `docker stack deploy` never creates or
|
|
sees the value, only references it):
|
|
|
|
```bash
|
|
# A strong random password for Forgejo's own Postgres (NOT related to
|
|
# Thermograph's app database — entirely separate instance/network).
|
|
openssl rand -base64 32 | docker secret create forgejo_db_password -
|
|
```
|
|
|
|
## Deploy / update
|
|
|
|
```bash
|
|
docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo
|
|
```
|
|
|
|
Re-running is safe — Swarm only touches services whose spec actually changed.
|
|
Do this **before** the DNS + Caddy step below — Caddy's reverse_proxy target
|
|
(`127.0.0.1:3080`) needs the `forgejo` service actually listening first, or
|
|
its first health check just fails harmlessly until it is.
|
|
|
|
This stack has **no auto-deploy trigger** — nothing in `.forgejo/workflows/`
|
|
redeploys it on push. A change to `docker-stack.yml` only takes effect once
|
|
someone re-runs `docker stack deploy` by hand on the manager (vps2).
|
|
|
|
`db`/`forgejo` both carry `resources.limits` (defaults: db 1 CPU/1g, forgejo 2
|
|
CPU/2g — several times observed steady-state usage), overridable with
|
|
`FORGEJO_DB_CPUS`/`FORGEJO_DB_MEMORY`/`FORGEJO_CPUS`/`FORGEJO_MEMORY` env vars
|
|
before `docker stack deploy`, same convention as the app stack.
|
|
|
|
## DNS + TLS: reusing vps1's existing Caddy, not a second reverse proxy
|
|
|
|
Forgejo is pinned to vps1 (`role=forge`) — the same box that also runs
|
|
Grafana/Loki/Alloy and the `emigriffith.dev` portfolio site, each fronted by
|
|
that host's own Caddy. A second ingress (Traefik) trying to bind the same
|
|
ports would collide with it. So there's no Traefik in this stack: `forgejo`'s
|
|
web port publishes to `127.0.0.1:3080` only (host-local), and vps1's
|
|
*existing* Caddy gets one more site block reverse-proxying to it — same
|
|
pattern as its other site blocks, same automatic-HTTPS.
|
|
|
|
1. Point the Forgejo domain (default `dev.jinemi.com`; override with
|
|
`FORGEJO_DOMAIN=...` before `docker stack deploy`) at **vps1's** public IP
|
|
— that's where the task actually runs, not vps2's or the desktop's.
|
|
2. Append `deploy/forgejo/caddy-git.conf` to vps1's `/etc/caddy/Caddyfile`,
|
|
adjusting the domain if you didn't use the default, then `systemctl reload
|
|
caddy`.
|
|
3. That file also resolves the registry-exposure hazard (#15 in
|
|
`thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md`): `/v2/*` (the registry
|
|
API) is blocked to everything except the WireGuard mesh CIDR
|
|
(`10.10.0.0/24`); the git/web UI stays public. CI runners and Swarm nodes
|
|
reach the registry over the mesh, not the public internet — see "Registry
|
|
access from mesh clients" below.
|
|
|
|
## Migrating the domain (git.thermograph.org -> dev.jinemi.com)
|
|
|
|
**Web UI and OAuth: done.** `dev.jinemi.com` is Forgejo's `ROOT_URL`
|
|
(`FORGEJO_DOMAIN` in `docker-stack.yml`), so it is the name Forgejo generates in
|
|
clone URLs, the Google OAuth callback, webhook payload URLs and mail links.
|
|
|
|
**Registry: done in this repo, and it needs host-side steps to match.** Every
|
|
image name, `REGISTRY_HOST` default, runner label and `--add-host` pin under
|
|
version control now says `dev.jinemi.com`. None of that is self-applying — see
|
|
"Host-side steps" below before expecting a build to push.
|
|
|
|
Both names address the *same* Forgejo and therefore the same packages. A tag
|
|
pushed as `git.thermograph.org/jinemi/thermograph/backend:sha-abc` is pullable
|
|
as `dev.jinemi.com/jinemi/thermograph/backend:sha-abc` — the host part is only
|
|
how you reach the registry, not part of the package's identity. That is what
|
|
makes this migration safe to do in one step instead of a dual-push period: no
|
|
image needs re-pushing, and a rollback to an old tag still resolves.
|
|
|
|
`caddy-git.conf` lists both names on a *single* site block. That is deliberate,
|
|
not cosmetic: the `/v2/*` mesh-only matcher is per-block, so a second block for
|
|
either name would re-expose the registry API publicly and undo hazard #15.
|
|
Whatever else changes, keep the two names in one block.
|
|
|
|
**`git.thermograph.org` must stay served.** The list of reasons is shorter than
|
|
it was, but it is not empty, and neither remaining item is a bookmark:
|
|
|
|
- image tags already pushed under the old prefix name that host, so a rollback
|
|
to one asks for `git.thermograph.org/...` even though nothing builds that
|
|
name any more;
|
|
- registered runners hold the instance URL they registered with
|
|
(`https://git.thermograph.org`); they keep working while that name resolves,
|
|
and stop the moment it doesn't. Re-registering them is the only way to retire
|
|
the name, and it is not required for this migration.
|
|
|
|
### Host-side steps
|
|
|
|
None of these live in the repo, and CI will not do them for you. In order:
|
|
|
|
1. **`/etc/hosts` on every mesh client** — pin `dev.jinemi.com` as well as
|
|
`git.thermograph.org` (both, for the reasons under "Registry access from
|
|
mesh clients" below). A client that resolves the new name publicly gets a
|
|
403 from the `/v2/*` matcher, which reads exactly like a bad credential.
|
|
2. **`docker login dev.jinemi.com`** on each host that pushes or pulls — the
|
|
CI runners, vps2 (prod and beta), the desktop. A login against the old name
|
|
does **not** carry over: docker keys stored credentials by registry host.
|
|
3. **The Forgejo Actions variable `REGISTRY_HOST`** → `dev.jinemi.com`
|
|
(Site Administration → Actions → Variables, or the repo's own). This is what
|
|
`build-push.yml` reads; the defaults in this repo are fallbacks for by-hand
|
|
runs, so leaving the variable stale silently keeps CI on the old name.
|
|
4. **Re-register the runners** against `https://dev.jinemi.com`, *optionally* —
|
|
see above. Until you do, leave `git.thermograph.org` served.
|
|
|
|
Verify with a push before assuming: the failure mode for a missed step 1 or 2
|
|
is `unauthorized: reqPackageAccess`, which names neither.
|
|
|
|
### Before changing ROOT_URL again
|
|
|
|
Login is Google-SSO-only, so the Google OAuth client must already carry a
|
|
matching redirect URI, or nobody can reach the UI — including to undo the
|
|
change. Both of these are registered on client
|
|
`337247886376-4360kepi20f24pue62ptg0lnqhcht46l.apps.googleusercontent.com`:
|
|
|
|
https://dev.jinemi.com/user/oauth2/google/callback
|
|
https://git.thermograph.org/user/oauth2/google/callback
|
|
|
|
Verify rather than assume. Google answers this with no credentials: a
|
|
registered URI serves the sign-in page, an unregistered one returns
|
|
`redirect_uri_mismatch`. Substitute the host you intend to make canonical:
|
|
|
|
```sh
|
|
CID=337247886376-4360kepi20f24pue62ptg0lnqhcht46l.apps.googleusercontent.com
|
|
curl -sL "https://accounts.google.com/o/oauth2/v2/auth?client_id=$CID&redirect_uri=https%3A%2F%2F<host>%2Fuser%2Foauth2%2Fgoogle%2Fcallback&response_type=code&scope=openid+email+profile&state=probe" \
|
|
| grep -q redirect_uri_mismatch && echo REJECTED || echo ACCEPTED
|
|
```
|
|
|
|
The stack has **no auto-deploy**: a `FORGEJO_DOMAIN` change here only takes
|
|
effect when someone re-runs `docker stack deploy` by hand on the manager (vps2)
|
|
— see "Deploy / update" above. Until then the running service keeps whatever
|
|
`ROOT_URL` it was last deployed with, regardless of what this file says.
|
|
|
|
## Registry access from mesh clients
|
|
|
|
Any node that needs `docker login`/push/pull against the registry (the CI
|
|
runner building/pushing images, any Swarm node pulling them, prod or beta on
|
|
vps2 pulling app images) must reach Forgejo **over the WireGuard tunnel**, not
|
|
vps1's public IP — otherwise Caddy's `/v2/*` block above refuses the
|
|
connection. Public DNS resolves both names to vps1's public IP, so add an
|
|
`/etc/hosts` override on each such node pinning them to vps1's WireGuard
|
|
address instead:
|
|
|
|
```
|
|
echo "10.10.0.2 dev.jinemi.com git.thermograph.org" | sudo tee -a /etc/hosts
|
|
```
|
|
|
|
(`10.10.0.2` is vps1's WG address per `deploy/swarm/README.md`'s peer
|
|
numbering — adjust if you assigned it differently.) The git/web UI keeps
|
|
working normally for everyone else since only `/v2/*` is restricted.
|
|
|
|
**Both names, not just the one in the image.** `dev.jinemi.com` now carries two
|
|
independent jobs and needs the pin for each: it is the host in every image
|
|
name, *and* it is `ROOT_URL`, from which Forgejo derives the registry's
|
|
bearer-token realm — so a client is sent to `https://<ROOT_URL host>/v2/token`
|
|
to collect a token no matter which name it dialled. `git.thermograph.org` stays
|
|
pinned for a third reason: image tags pushed before the migration name it, and
|
|
a rollback to one of those dials it directly.
|
|
|
|
This is not hypothetical. When `ROOT_URL` first became `dev.jinemi.com` while
|
|
images were still named `git.thermograph.org`, the token request took the
|
|
public route, the `/v2/*` matcher answered 403, and docker fell back to
|
|
anonymous — every push failed with `unauthorized: reqPackageAccess`, which
|
|
reads exactly like a revoked token or a missing scope. It was neither. Pinning
|
|
one name and not the other reproduces it in either direction.
|
|
|
|
The realm moves whenever `ROOT_URL` does, so read it rather than assuming:
|
|
|
|
```sh
|
|
curl -sI https://dev.jinemi.com/v2/ | grep -i www-authenticate
|
|
```
|
|
|
|
These `/etc/hosts` pins are host state. Nothing in this repo writes them, so a
|
|
reprovisioned node has to be pinned again before it can pull.
|
|
|
|
## Register the Actions runner
|
|
|
|
Once Forgejo answers at its domain:
|
|
|
|
```bash
|
|
# On the runner host (see DEPLOY-DEV.md for where that is today):
|
|
# Forgejo web UI -> Site Administration -> Actions -> Runners -> Create new Runner
|
|
# (or, for a repo-scoped runner: <repo> -> Settings -> Actions -> Runners)
|
|
# copy the registration token, then:
|
|
bash deploy/forgejo/register-lan-runner.sh https://<forgejo-domain> <token>
|
|
```
|
|
|
|
See that script's header for exactly what it replaces (the pre-Forgejo GitHub
|
|
self-hosted runner) and why it registers with two labels where there used to
|
|
be two separate runners.
|
|
|
|
`config.yaml`'s `runner.capacity` is raised from the tool's default of 1 to 8
|
|
(override with `CAPACITY=`) — a single PR push fires `pr-build`,
|
|
`secrets-guard`, and `shell-lint` simultaneously (3 independent workflows,
|
|
no `needs:` between them), so capacity 1 serializes work that could run in
|
|
parallel, and even capacity 3 (an earlier, undocumented hand-tune) leaves
|
|
`build-backend`/`build-frontend`/`validate-observability` queued behind
|
|
those three before they get a slot.
|
|
|
|
**This paragraph used to say a runner must never go on vps2. That was reversed
|
|
deliberately on 2026-08-01, and the reasoning is worth keeping rather than
|
|
quietly deleting.**
|
|
|
|
The original objection stands on its merits: `container.docker_host: automount`
|
|
gives job containers the *host's* Docker socket, so a runner on vps2 means any
|
|
CI job has root-equivalent access to both the prod and beta stacks running
|
|
there. An earlier revision of this stack ran the runner as a Swarm-hosted
|
|
container on the Forgejo node and was reverted for the same class of reason.
|
|
|
|
What changed is that the alternative turned out to be worse. The desktop was
|
|
the estate's **only** registered runner, and when it went offline on
|
|
2026-07-31 nothing merged, nothing deployed, and the nightly backup did not
|
|
fire for 21 hours — on a repo where every branch is protected and every change
|
|
is a PR, a single absent runner freezes the whole estate. Availability of CI
|
|
is not a convenience here; it is what the backup and the deploy path both
|
|
hang off.
|
|
|
|
So the runner on vps2 (`runner-vps2/`) is a considered trade, not an
|
|
oversight, and it is bounded rather than assumed benign:
|
|
|
|
- `capacity: 1` — one job at a time, so a CI burst cannot contend with
|
|
`thermograph_web` for the box's six cores.
|
|
- `--cpus=2 --memory=4g` on job containers, set in `container.options`. Job
|
|
containers are siblings of the runner rather than children, so the runner's
|
|
own compose limits do not reach them; that option is the only lever that
|
|
does.
|
|
- `valid_volumes: []` — a job cannot bind-mount a host path, which is the
|
|
difference between "a job can read `/etc/thermograph.env`" and "a job
|
|
cannot", on the box where that file is prod's.
|
|
- The compose project joins no `thermograph` network.
|
|
|
|
Be honest about what that buys: it is defence against **accident**, not
|
|
against a hostile workflow author. Forgejo's own database on vps1 already
|
|
stores `VPS2_SSH_KEY`, which is root on vps2, so the material was already
|
|
reachable — what changed is that it is now reachable by a *job* rather than
|
|
only at rest. On a two-person estate where both people can already SSH to that
|
|
box as root, that is the honest boundary.
|
|
|
|
**If you are adding capacity rather than redundancy, raise `capacity` or add a
|
|
runner on a box that hosts nothing.** The vps2 runner exists so that CI has a
|
|
second home, not because prod is a good place to run CI.
|
|
|
|
## Custom CI job image (`ci-runner/`)
|
|
|
|
`ci-runner/Dockerfile` still bases on `node:20-bookworm` — Node is a hard
|
|
requirement, not leftover: Forgejo's runner executes `actions/checkout@v4`
|
|
(used by every workflow) as `node dist/index.js` *inside the job container*,
|
|
regardless of whether the workflow itself uses npm/node. (v1 of this image
|
|
tried a Node-free Debian-slim base and broke every job's checkout step —
|
|
`node: executable file not found` — within a minute of going live; reverted
|
|
immediately.) What it actually fixes: every `docker`-labeled build-push job
|
|
currently re-installs the Docker CLI on each run (`apt-get install
|
|
docker.io`), which pulls in the classic builder rather than BuildKit (the
|
|
classic builder mishandles `COPY --chown=<name>` group resolution — a real
|
|
bug hit during the frontend Go rewrite). `ci-runner` adds `docker-ce-cli` +
|
|
`docker-buildx-plugin` (BuildKit) on top of the same Node base, plus
|
|
`git`/`python3`/`python3-yaml` for the other jobs that need them
|
|
(`shell-lint`, `observability-validate`).
|
|
|
|
Current tag: `dev.jinemi.com/jinemi/thermograph/ci-runner:v2` (`v1` is
|
|
broken — do not register any runner against it). Rebuild/push (requires a PAT
|
|
with `write:package` scope — the embedded git-remote token lacks it, same
|
|
requirement documented in `build-push.yml`):
|
|
|
|
```bash
|
|
docker build -t dev.jinemi.com/jinemi/thermograph/ci-runner:vN deploy/forgejo/ci-runner
|
|
docker push dev.jinemi.com/jinemi/thermograph/ci-runner:vN
|
|
```
|
|
|
|
`register-lan-runner.sh`'s `LABELS` default points at the current tag, so
|
|
fresh registrations pick it up automatically. The live runner is cut over by
|
|
editing the `labels` array in `~/forgejo-runner/.runner` on the runner host
|
|
(same runner id/token, no re-registration needed) and restarting the service —
|
|
**verify a real job runs green under the new image before relying on it**,
|
|
same way v1's break was caught. Only after that verification should the
|
|
now-redundant `apt-get install docker.io` / `python3-yaml` steps be removed
|
|
from the workflows that had them — removing them first would break every job
|
|
still running on the stock `node:20-bookworm` image.
|
|
|
|
## Why Postgres here and not the Thermograph app's TimescaleDB
|
|
|
|
Separate instance, separate network (`forgejo_net`, not the app's overlay
|
|
network), separate volume. Forgejo is a distinct product with its own schema
|
|
and its own backup/restore lifecycle — sharing a database with the app would
|
|
couple two things that should be able to fail, migrate, and restore
|
|
independently.
|
|
|
|
## Verifying
|
|
|
|
```bash
|
|
docker service ls # forgejo_db, forgejo_forgejo both Running, 1/1
|
|
curl -I https://dev.jinemi.com/ # 200, valid cert (Caddy's, not a new one)
|
|
curl -I https://dev.jinemi.com/v2/ # 403 from anywhere off the WireGuard mesh
|
|
curl -I https://git.thermograph.org/ # 200 too — the alias must stay served
|
|
# On the runner host, after registering the runner:
|
|
systemctl --user status forgejo-runner # active, both labels registered
|
|
```
|
|
|
|
## Rollback / removal
|
|
|
|
```bash
|
|
docker stack rm forgejo
|
|
# volumes (forgejo_data, forgejo_db, ...) survive a stack rm — remove them
|
|
# explicitly only if you actually want to destroy the Forgejo instance's data.
|
|
```
|