All checks were successful
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / changes (pull_request) Successful in 13s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 7s
secrets-guard / encrypted (push) Successful in 7s
shell-lint / shellcheck (push) Successful in 7s
The estate had exactly one registered runner, on the desktop. It went offline at 2026-07-31 16:31Z; for the next 21 hours no PR could satisfy a required check, no deploy could run, and the 03:00Z ops-cron -- the only backup for both application databases and for Forgejo -- did not fire. Forgejo queued that scheduled run rather than dropping it, so it completed on reconnect and nothing was lost. A longer outage would have meant real gaps. Three files claimed an "always-on Swarm-hosted runner" existed and that the estate therefore no longer depended on the desktop. It did not exist: an early revision of docker-stack.yml ran one as a Docker-in-Docker sidecar and it was removed. That claim is why a single point of failure sat unnoticed. Corrected in docker-stack.yml, forgejo/README.md and register-lan-runner.sh. The new runner is a plain restart:always container, not a Swarm service: a Swarm-scheduled runner cannot redeploy the Swarm that schedules it, so CI would be gone exactly when the cluster is what is broken. It runs from /opt/forgejo-runner rather than in place, because the checkout is reset on every prod deploy and one `git clean -fdx` there would destroy the registration. vps2 runs prod, so the socket mount is bounded rather than assumed benign: capacity 1, --cpus=2/--memory=4g on job containers, valid_volumes empty so no job can bind-mount /etc/thermograph.env, and no thermograph network joined. This defends against accident, not against a hostile workflow author -- stated plainly in the compose header rather than implied. The desktop runner stays registered as extra capacity. Nothing may assume it is up.
174 lines
8.4 KiB
YAML
174 lines
8.4 KiB
YAML
# Forgejo (self-hosted Git + CI/CD) as a Docker Swarm stack — the only workload
|
|
# this Swarm cluster runs (see deploy/swarm/README.md). Deliberately separate
|
|
# from the Terraform-managed docker-compose.yml that runs the Thermograph app
|
|
# itself: this stack's only job is Forgejo and its container registry.
|
|
#
|
|
# The Actions RUNNER is deliberately NOT a service in this stack, and that is a
|
|
# reasoned choice rather than an omission: a Swarm-scheduled runner cannot
|
|
# redeploy the Swarm that schedules it, so CI would be gone exactly when the
|
|
# cluster is the thing that is broken.
|
|
#
|
|
# The runner that CI depends on runs on vps2 as a plain restart:always container
|
|
# — deploy/forgejo/runner-vps2/. The desktop's systemd runner
|
|
# (deploy/forgejo/register-lan-runner.sh) stays registered as extra capacity.
|
|
# Nothing may assume the desktop is up: it was the ONLY runner until 2026-08-01,
|
|
# and its 21-hour outage on 2026-07-31 froze every merge, deploy and backup.
|
|
#
|
|
# No Traefik here. Forgejo is pinned to vps1 (node.labels.role == forge)
|
|
# because that's what was chosen, and vps1 is ALSO the emigriffith.dev portfolio
|
|
# host — Caddy already owns its ports 80/443 (see /etc/caddy/Caddyfile on that
|
|
# box). A second reverse proxy binding those same ports would either fail to
|
|
# start or fight Caddy. Instead: forgejo's web port publishes to
|
|
# 127.0.0.1:3080 only (host-local, mode: host), and vps1's existing Caddy gets
|
|
# a new site block reverse-proxying git.thermograph.org -> 127.0.0.1:3080,
|
|
# same pattern as its other blocks. TLS is Caddy's existing automatic-HTTPS
|
|
# (HTTP-01), not a second ACME flow.
|
|
#
|
|
# Deploy from the manager node (vps2, which runs prod), after all three nodes
|
|
# (vps1, vps2, desktop) have joined the swarm and vps1 is labeled role=forge:
|
|
#
|
|
# docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo
|
|
#
|
|
# Requires this Swarm secret to exist first (see deploy/forgejo/README.md):
|
|
# forgejo_db_password
|
|
#
|
|
# Registry exposure (hazard #15 in thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md):
|
|
# resolved as the runbook's first listed option — serve Git/UI publicly but
|
|
# firewall the /v2/ registry API paths so only WireGuard-mesh clients can
|
|
# reach them. That's implemented in the Caddy site block (see
|
|
# deploy/forgejo/README.md), not here — nothing in this stack file is
|
|
# registry-specific, the restriction lives entirely in Caddy's config on vps1.
|
|
|
|
services:
|
|
db:
|
|
image: postgres:16-alpine
|
|
environment:
|
|
POSTGRES_USER: forgejo
|
|
POSTGRES_DB: forgejo
|
|
POSTGRES_PASSWORD_FILE: /run/secrets/forgejo_db_password
|
|
secrets: [forgejo_db_password]
|
|
volumes:
|
|
- forgejo_db:/var/lib/postgresql/data
|
|
networks: [forgejo_net]
|
|
deploy:
|
|
placement:
|
|
constraints: [node.labels.role == forge]
|
|
resources:
|
|
limits:
|
|
cpus: "${FORGEJO_DB_CPUS:-1}"
|
|
memory: ${FORGEJO_DB_MEMORY:-1g}
|
|
restart_policy:
|
|
# `any`, NOT `on-failure`. This took Forgejo down for 27 hours on 2026-07-29:
|
|
# Postgres hit an invalid data-directory lock file ("could not open file
|
|
# postmaster.pid ... performing immediate shutdown") and exited **0**. A clean
|
|
# exit is not a failure, so Swarm considered the task Complete, dropped the
|
|
# service to 0/1 replicas, and never rescheduled it. Forgejo itself stayed Up
|
|
# and served its homepage while every repo page, the API and all CI returned
|
|
# 500 with `dial tcp: lookup db ... no such host`.
|
|
#
|
|
# `on-failure` is the wrong policy for any always-on stateful service: it
|
|
# cannot distinguish "finished successfully" from "shut itself down and should
|
|
# be restarted", and Postgres does the latter with status 0 on several paths.
|
|
condition: any
|
|
delay: 5s
|
|
|
|
forgejo:
|
|
image: codeberg.org/forgejo/forgejo:9-rootless
|
|
depends_on: [db]
|
|
environment:
|
|
FORGEJO__database__DB_TYPE: postgres
|
|
FORGEJO__database__HOST: db:5432
|
|
FORGEJO__database__NAME: forgejo
|
|
FORGEJO__database__USER: forgejo
|
|
FORGEJO__server__DOMAIN: "${FORGEJO_DOMAIN:-git.thermograph.org}"
|
|
FORGEJO__server__ROOT_URL: "https://${FORGEJO_DOMAIN:-git.thermograph.org}/"
|
|
FORGEJO__server__SSH_PORT: "2222"
|
|
# Without this, Forgejo starts no SSH server at all despite SSH_PORT
|
|
# being set — git@ clones get "connection refused", not a slow failure.
|
|
FORGEJO__server__START_SSH_SERVER: "true"
|
|
# Config is fully supplied via env, so lock the install wizard rather
|
|
# than leave it open on the public internet waiting for someone to
|
|
# complete it first (DB config alone does NOT imply this).
|
|
FORGEJO__security__INSTALL_LOCK: "true"
|
|
# Actions on: repo/org/user-level runners register against this instance.
|
|
FORGEJO__actions__ENABLED: "true"
|
|
# --- Access model: OAuth-only login, multi-user via an approval gate ---
|
|
# Login is Google SSO only. Forgejo 9.0.3 has no setting to hide the
|
|
# password sign-in form, so "OAuth-only" is enforced operationally: every
|
|
# account is given an unusable (random) password and no NEW local account
|
|
# can be created (ALLOW_ONLY_EXTERNAL_REGISTRATION). ENABLE_PASSWORD_SIGNIN_FORM
|
|
# is set anyway — ignored on 9.0.3, effective if the instance is upgraded.
|
|
# Auto-registration stays ON so a Google login auto-links (by verified email)
|
|
# to a PRE-CREATED account (`forgejo admin user create --username X --email
|
|
# X@gmail.com --random-password`). A single required-claim-value can't
|
|
# allowlist more than one email, so instead of pinning one address the
|
|
# required-claim-value is CLEARED (a DB change on the auth source, persisted
|
|
# in the forgejo_data volume — not re-applied from here) and every NEW signup
|
|
# is gated: a stranger who signs in with Google lands INACTIVE
|
|
# (REGISTER_MANUAL_CONFIRM) and RESTRICTED (DEFAULT_USER_IS_RESTRICTED)
|
|
# pending admin approval, so an un-provisioned Google user gets zero access.
|
|
FORGEJO__oauth2_client__ACCOUNT_LINKING: "auto"
|
|
FORGEJO__oauth2_client__ENABLE_AUTO_REGISTRATION: "true"
|
|
FORGEJO__service__ALLOW_ONLY_EXTERNAL_REGISTRATION: "true"
|
|
FORGEJO__service__ENABLE_PASSWORD_SIGNIN_FORM: "false"
|
|
FORGEJO__service__REGISTER_MANUAL_CONFIRM: "true"
|
|
FORGEJO__service__DEFAULT_USER_IS_RESTRICTED: "true"
|
|
# Outbound mail via vps2's Postfix null client (prod's box) over the
|
|
# WireGuard mesh (10.10.0.1:25 — mynetworks permits vps1, 10.10.0.2,
|
|
# where this Forgejo runs); self-signed cert on :25, so trust it.
|
|
# Send-only; used for admin/approval and notification mail.
|
|
FORGEJO__mailer__ENABLED: "true"
|
|
FORGEJO__mailer__PROTOCOL: "smtp"
|
|
FORGEJO__mailer__SMTP_ADDR: "10.10.0.1"
|
|
FORGEJO__mailer__SMTP_PORT: "25"
|
|
FORGEJO__mailer__FORCE_TRUST_SERVER_CERT: "true"
|
|
FORGEJO__mailer__FROM: "Thermograph Git <no-reply@thermograph.org>"
|
|
# Gitea/Forgejo's app.ini env-mapping honors a __FILE suffix to read a
|
|
# value from a file instead of the literal env var — same convention as
|
|
# the official Postgres image's POSTGRES_PASSWORD_FILE above.
|
|
FORGEJO__database__PASSWD__FILE: /run/secrets/forgejo_db_password
|
|
secrets:
|
|
- source: forgejo_db_password
|
|
target: forgejo_db_password
|
|
volumes:
|
|
- forgejo_data:/var/lib/gitea
|
|
networks: [forgejo_net]
|
|
ports:
|
|
# SSH for git@ clones — published on whichever node the task lands on
|
|
# (pinned to vps1 by the placement constraint below, so effectively
|
|
# always vps1's public IP, port 2222).
|
|
- target: 2222
|
|
published: 2222
|
|
protocol: tcp
|
|
mode: host
|
|
# Web UI/API. Swarm's port schema has no host_ip scoping, so this binds
|
|
# 0.0.0.0:3080 on vps1 — NOT actually localhost-only by itself. A
|
|
# DOCKER-USER iptables rule (see deploy/forgejo/README.md) is what
|
|
# actually restricts it, since Docker's own iptables rules bypass ufw
|
|
# for published ports. vps1's Caddy reverse-proxies to 127.0.0.1:3080.
|
|
- target: 3000
|
|
published: 3080
|
|
protocol: tcp
|
|
mode: host
|
|
deploy:
|
|
placement:
|
|
constraints: [node.labels.role == forge]
|
|
resources:
|
|
limits:
|
|
cpus: "${FORGEJO_CPUS:-2}"
|
|
memory: ${FORGEJO_MEMORY:-2g}
|
|
restart_policy:
|
|
condition: on-failure
|
|
|
|
networks:
|
|
forgejo_net:
|
|
driver: overlay
|
|
attachable: false
|
|
|
|
volumes:
|
|
forgejo_db:
|
|
forgejo_data:
|
|
|
|
secrets:
|
|
forgejo_db_password:
|
|
external: true
|