All checks were successful
PR build (required check) / changes (pull_request) Successful in 9s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
secrets-guard / encrypted (pull_request) Successful in 7s
shell-lint / shellcheck (pull_request) Successful in 9s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
The six deploy workflows were one file written six times, differing only in a branch name, a paths filter, a concurrency group, the service name, its *_IMAGE_TAG variable and a secret prefix. The two build-push workflows were the same file twice, differing only in the domain string. The contract into infra/deploy/deploy.sh (SERVICE + BACKEND_IMAGE_TAG/FRONTEND_IMAGE_TAG) was already fully parameterised, so the duplication bought nothing and cost eight files to keep in step. deploy.yml: branch selects the environment (main -> beta, release -> prod), a matrix covers backend and frontend, and each leg decides whether this push actually touched its domain before rolling anything. The workflow-level paths filter only says backend OR frontend moved; without the per-leg refinement a backend-only push would also roll the frontend and lose the independent-deploy property the FE/BE split exists for. build-push.yml: the same shape for images. Images stay separate and independently deployable; nothing about the published artefacts changes. The two *-deploy-dev.yml workflows are deleted rather than ported. They were already documented as inert -- they call a monorepo path on the LAN box whose ~/thermograph-dev is still a split-era thermograph-infra checkout. LAN dev is a local `make dev-up` concern, not a CI environment. Deliberately boring expressions throughout. No dynamic matrix (fromJSON), and the *_IMAGE_TAG selection is done in shell rather than with a `matrix.service == 'x' && a || b` ternary. Those are GitHub idioms a Forgejo/act runner may evaluate differently, and the failure mode is silent: an empty *_IMAGE_TAG makes deploy.sh fall back to the tag already running, so the job goes green having deployed nothing. Beta and prod get two explicit, mutually exclusive steps rather than a ternary over secrets, where an empty host would be worse still. Everything load-bearing is preserved: fetch-depth 0 and the domain-keyed 12-hex tag (the branch tip is often another domain's commit), per-service per-ref concurrency with cancel-in-progress false, separate PROD_SSH_* credentials, the v*.*.* both-images exception, and appleboy/ssh-action by full URL. pr-build.yml is untouched. Its workflow name and `gate` job are the required status check branch protection is configured against, and renaming that context makes every PR unmergeable with fully green CI. secrets-guard.yml and shell-lint.yml are likewise left alone: they are the only other consolidation candidates that run on pull_request, and the branch-protection API needs credentials this could not read, so merging them was not worth the risk for two files. Logic verified by replaying the plan step against real history: an infra-only push rolls nothing, a backend-only push rolls backend and leaves frontend alone, and a run with no usable before-sha defaults to deploying rather than silently skipping. The computed frontend tag for a backend-only push is sha-ca84e0ce95f0 -- exactly the tag live on beta -- so the derivation matches what the old workflows produced. 15 workflows -> 9, 1365 lines -> 1019. Docs naming the deleted files are updated in the same commit, per the rule the root CLAUDE.md now carries.
190 lines
9.2 KiB
Markdown
190 lines
9.2 KiB
Markdown
# Forgejo on the Swarm cluster
|
|
|
|
Runs as `deploy/forgejo/docker-stack.yml` — the only Swarm-scheduled workload
|
|
this cluster carries (the Thermograph app itself stays on the
|
|
Terraform-managed `docker compose` deploys; see `terraform/README.md`).
|
|
Pinned to the **beta** node (old VPS) via the `role=forge` label from
|
|
`deploy/swarm/label-forge-node.sh`.
|
|
|
|
The Actions **runner** is deliberately *not* part of this stack — it runs on
|
|
the **desktop** as a plain systemd service (`register-lan-runner.sh` below),
|
|
per `thermograph-docs/runbooks/implementation-handoff.md` Track B step 5. That's the
|
|
canonical placement; an earlier revision of this stack ran the runner as a
|
|
Swarm-scheduled Docker-in-Docker sidecar pinned to beta, which is gone now.
|
|
|
|
## Prerequisites
|
|
|
|
1. All three nodes have joined the swarm (`deploy/swarm/`) and beta is
|
|
labeled `role=forge`.
|
|
2. `docker node ls` (from the manager) shows all three `Ready`.
|
|
|
|
## One-time setup: Swarm secret
|
|
|
|
One secret the stack expects to already exist (a Swarm secret, not a file —
|
|
`external: true` in the stack file, so `docker stack deploy` never creates or
|
|
sees the value, only references it):
|
|
|
|
```bash
|
|
# A strong random password for Forgejo's own Postgres (NOT related to
|
|
# Thermograph's app database — entirely separate instance/network).
|
|
openssl rand -base64 32 | docker secret create forgejo_db_password -
|
|
```
|
|
|
|
## Deploy / update
|
|
|
|
```bash
|
|
docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo
|
|
```
|
|
|
|
Re-running is safe — Swarm only touches services whose spec actually changed.
|
|
Do this **before** the DNS + Caddy step below — Caddy's reverse_proxy target
|
|
(`127.0.0.1:3080`) needs the `forgejo` service actually listening first, or
|
|
its first health check just fails harmlessly until it is.
|
|
|
|
This stack has **no auto-deploy trigger** — nothing in `.forgejo/workflows/`
|
|
redeploys it on push. A change to `docker-stack.yml` only takes effect once
|
|
someone re-runs `docker stack deploy` by hand on the manager (prod).
|
|
|
|
`db`/`forgejo` both carry `resources.limits` (defaults: db 1 CPU/1g, forgejo 2
|
|
CPU/2g — several times observed steady-state usage), overridable with
|
|
`FORGEJO_DB_CPUS`/`FORGEJO_DB_MEMORY`/`FORGEJO_CPUS`/`FORGEJO_MEMORY` env vars
|
|
before `docker stack deploy`, same convention as the app stack.
|
|
|
|
## DNS + TLS: reusing beta's existing Caddy, not a second reverse proxy
|
|
|
|
Forgejo is pinned to beta (`role=forge`) — but beta is also **today's live
|
|
thermograph.org host**, and its Caddy already owns ports 80/443
|
|
(`/etc/caddy/Caddyfile` on that box). A second ingress (Traefik) trying to
|
|
bind the same ports would collide with it. So there's no Traefik in this
|
|
stack: `forgejo`'s web port publishes to `127.0.0.1:3080` only (host-local),
|
|
and beta's *existing* Caddy gets one more site block reverse-proxying to it —
|
|
same pattern as its `thermograph.org` block, same automatic-HTTPS.
|
|
|
|
1. Point the Forgejo domain (default `git.thermograph.org`; override with
|
|
`FORGEJO_DOMAIN=...` before `docker stack deploy`) at **beta's** public IP
|
|
— that's where the task actually runs, not prod's or the desktop's.
|
|
2. Append `deploy/forgejo/caddy-git.conf` to beta's `/etc/caddy/Caddyfile`,
|
|
adjusting the domain if you didn't use the default, then `systemctl reload
|
|
caddy`.
|
|
3. That file also resolves the registry-exposure hazard (#15 in
|
|
`thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md`): `/v2/*` (the registry
|
|
API) is blocked to everything except the WireGuard mesh CIDR
|
|
(`10.10.0.0/24`); the git/web UI stays public. CI runners and Swarm nodes
|
|
reach the registry over the mesh, not the public internet — see "Registry
|
|
access from mesh clients" below.
|
|
|
|
## Registry access from mesh clients
|
|
|
|
Any node that needs `docker login`/push/pull against the registry (the
|
|
desktop's CI runner building/pushing images, later any Swarm node pulling
|
|
them) must reach `git.thermograph.org` **over the WireGuard tunnel**, not
|
|
beta's public IP — otherwise Caddy's `/v2/*` block above refuses the
|
|
connection. Public DNS resolves the domain to beta's public IP, so add a
|
|
`/etc/hosts` override on each such node pinning it to beta's WireGuard
|
|
address instead:
|
|
|
|
```
|
|
echo "10.10.0.2 git.thermograph.org" | sudo tee -a /etc/hosts
|
|
```
|
|
|
|
(`10.10.0.2` is beta's WG address per `deploy/swarm/README.md`'s peer
|
|
numbering — adjust if you assigned it differently.) The git/web UI keeps
|
|
working normally for everyone else since only `/v2/*` is restricted.
|
|
|
|
## Register the Actions runner (on the desktop, not through Swarm)
|
|
|
|
Once Forgejo answers at its domain:
|
|
|
|
```bash
|
|
# On the desktop:
|
|
# Forgejo web UI -> Site Administration -> Actions -> Runners -> Create new Runner
|
|
# (or, for a repo-scoped runner: <repo> -> Settings -> Actions -> Runners)
|
|
# copy the registration token, then:
|
|
bash deploy/forgejo/register-lan-runner.sh https://<forgejo-domain> <token>
|
|
```
|
|
|
|
See that script's header for exactly what it replaces (the pre-Forgejo GitHub
|
|
self-hosted runner on this same machine) and why it registers with two
|
|
labels where there used to be two separate runners.
|
|
|
|
`config.yaml`'s `runner.capacity` is raised from the tool's default of 1 to 8
|
|
(override with `CAPACITY=`) — a single PR push fires `pr-build`,
|
|
`secrets-guard`, and `shell-lint` simultaneously (3 independent workflows,
|
|
no `needs:` between them), so capacity 1 serializes work that could run in
|
|
parallel, and even capacity 3 (an earlier, undocumented hand-tune) leaves
|
|
`build-backend`/`build-frontend`/`validate-observability` queued behind
|
|
those three before they get a slot. The desktop has 16 cores / 34GB free
|
|
today; 8 concurrent jobs is comfortable headroom without starving LAN dev's
|
|
own compose stack.
|
|
|
|
**Adding more runner capacity should mean raising this number, or adding a
|
|
second runner on the desktop itself — not putting a runner on prod or
|
|
beta.** `container.docker_host: automount` gives job containers the *host's*
|
|
Docker socket; on prod or beta that would mean any CI job has root-equivalent
|
|
access to whatever else is running there (the live app stack, or Forgejo
|
|
itself). An earlier revision of this stack actually did run the runner as a
|
|
Swarm-hosted container on beta and was deliberately reverted to the desktop
|
|
for this reason — see the note at the top of `docker-stack.yml`.
|
|
|
|
## Custom CI job image (`ci-runner/`)
|
|
|
|
`ci-runner/Dockerfile` still bases on `node:20-bookworm` — Node is a hard
|
|
requirement, not leftover: Forgejo's runner executes `actions/checkout@v4`
|
|
(used by every workflow) as `node dist/index.js` *inside the job container*,
|
|
regardless of whether the workflow itself uses npm/node. (v1 of this image
|
|
tried a Node-free Debian-slim base and broke every job's checkout step —
|
|
`node: executable file not found` — within a minute of going live; reverted
|
|
immediately.) What it actually fixes: every `docker`-labeled build-push job
|
|
currently re-installs the Docker CLI on each run (`apt-get install
|
|
docker.io`), which pulls in the classic builder rather than BuildKit (the
|
|
classic builder mishandles `COPY --chown=<name>` group resolution — a real
|
|
bug hit during the frontend Go rewrite). `ci-runner` adds `docker-ce-cli` +
|
|
`docker-buildx-plugin` (BuildKit) on top of the same Node base, plus
|
|
`git`/`python3`/`python3-yaml` for the other jobs that need them
|
|
(`shell-lint`, `observability-validate`).
|
|
|
|
Current tag: `git.thermograph.org/emi/thermograph/ci-runner:v2` (`v1` is
|
|
broken — do not register any runner against it). Rebuild/push (requires a PAT
|
|
with `write:package` scope — the embedded git-remote token lacks it, same
|
|
requirement documented in `build-push.yml`):
|
|
|
|
```bash
|
|
docker build -t git.thermograph.org/emi/thermograph/ci-runner:vN deploy/forgejo/ci-runner
|
|
docker push git.thermograph.org/emi/thermograph/ci-runner:vN
|
|
```
|
|
|
|
`register-lan-runner.sh`'s `LABELS` default points at the current tag, so
|
|
fresh registrations pick it up automatically. The live runner is cut over by
|
|
editing the `labels` array in `~/forgejo-runner/.runner` on the desktop (same
|
|
runner id/token, no re-registration needed) and restarting the service —
|
|
**verify a real job runs green under the new image before relying on it**,
|
|
same way v1's break was caught. Only after that verification should the
|
|
now-redundant `apt-get install docker.io` / `python3-yaml` steps be removed
|
|
from the workflows that had them — removing them first would break every job
|
|
still running on the stock `node:20-bookworm` image.
|
|
|
|
## Why Postgres here and not the Thermograph app's TimescaleDB
|
|
|
|
Separate instance, separate network (`forgejo_net`, not the app's compose
|
|
network), separate volume. Forgejo is a distinct product with its own schema
|
|
and its own backup/restore lifecycle — sharing a database with the app would
|
|
couple two things that should be able to fail, migrate, and restore
|
|
independently.
|
|
|
|
## Verifying
|
|
|
|
```bash
|
|
docker service ls # forgejo_db, forgejo_forgejo both Running, 1/1
|
|
curl -I https://git.thermograph.org/ # 200, valid cert (Caddy's, not a new one)
|
|
curl -I https://git.thermograph.org/v2/ # 403 from anywhere off the WireGuard mesh
|
|
# On the desktop, after registering the runner:
|
|
systemctl --user status forgejo-runner # active, both labels registered
|
|
```
|
|
|
|
## Rollback / removal
|
|
|
|
```bash
|
|
docker stack rm forgejo
|
|
# volumes (forgejo_data, forgejo_db, ...) survive a stack rm — remove them
|
|
# explicitly only if you actually want to destroy the Forgejo instance's data.
|
|
```
|