Thermograph monorepo: graded-climate API + SSR frontend + infra, domain-specific containerized deploys
Find a file
Emi Griffith 001e6b1365
All checks were successful
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / changes (pull_request) Successful in 13s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 7s
secrets-guard / encrypted (push) Successful in 7s
shell-lint / shellcheck (push) Successful in 7s
ci: add an always-on Actions runner on vps2
The estate had exactly one registered runner, on the desktop. It went offline
at 2026-07-31 16:31Z; for the next 21 hours no PR could satisfy a required
check, no deploy could run, and the 03:00Z ops-cron -- the only backup for both
application databases and for Forgejo -- did not fire. Forgejo queued that
scheduled run rather than dropping it, so it completed on reconnect and nothing
was lost. A longer outage would have meant real gaps.

Three files claimed an "always-on Swarm-hosted runner" existed and that the
estate therefore no longer depended on the desktop. It did not exist: an early
revision of docker-stack.yml ran one as a Docker-in-Docker sidecar and it was
removed. That claim is why a single point of failure sat unnoticed. Corrected
in docker-stack.yml, forgejo/README.md and register-lan-runner.sh.

The new runner is a plain restart:always container, not a Swarm service: a
Swarm-scheduled runner cannot redeploy the Swarm that schedules it, so CI would
be gone exactly when the cluster is what is broken. It runs from
/opt/forgejo-runner rather than in place, because the checkout is reset on
every prod deploy and one `git clean -fdx` there would destroy the
registration.

vps2 runs prod, so the socket mount is bounded rather than assumed benign:
capacity 1, --cpus=2/--memory=4g on job containers, valid_volumes empty so no
job can bind-mount /etc/thermograph.env, and no thermograph network joined.
This defends against accident, not against a hostile workflow author -- stated
plainly in the compose header rather than implied.

The desktop runner stays registered as extra capacity. Nothing may assume it
is up.
2026-08-01 07:54:28 -07:00
.claude BRANCHING: the escape hatch, now that apply_to_admins binds the owner too 2026-07-26 15:05:26 -07:00
.forgejo/workflows ci: add an always-on Actions runner on vps2 2026-08-01 07:54:28 -07:00
backend accounts: sign in with Google, on a shared provider-agnostic OAuth engine (#122) 2026-07-27 00:56:43 +00:00
docs/onboarding infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00
frontend accounts: sign in with Google, on a shared provider-agnostic OAuth engine (#122) 2026-07-27 00:56:43 +00:00
infra ci: add an always-on Actions runner on vps2 2026-08-01 07:54:28 -07:00
observability infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00
.gitignore Drop accidentally-committed worktrees; ignore .claude/worktrees 2026-07-24 15:57:34 -07:00
CLAUDE.md docs: the orchestrator runbook and the module map 2026-07-26 12:53:33 -07:00
CUTOVER-NOTES.md infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00
README.md infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00

thermograph

The Thermograph monorepo — the split repos reunified (2026-07-22) with full history via subtree merges, while keeping everything the split was actually for: per-domain images, per-domain deploys, and an async FE/BE contract.

Domains

Dir What CI
backend/ FastAPI graded-climate API, accounts, notifications (Discord bot, push, mail), data pipeline build-push → image emi/thermograph/backend; deploy
frontend/ Public client: static JS/CSS + SSR pages same build-push / deploy workflows, matrixed by domain; image emi/thermograph/frontend
infra/ Compose (dev, on vps1) + two Swarm stacks co-resident on vps2 (beta, prod), deploy scripts, terraform, SOPS secrets vault, ops cron infra-sync (host checkout + secrets render), secrets-guard, ops-cron
observability/ Loki + Grafana + Alloy stack observability-validate

thermograph-docs deliberately stays its own repo (ADRs + runbooks, no build artifacts, different change cadence).

New here?

docs/onboarding/ is the developer onboarding path: orientation, verified local-setup recipes, a per-domain deep dive, the cross-service contracts that break silently, the release flow, and a list of which docs in this repo are currently stale.

How CI stays decoupled

Every workflow in .forgejo/workflows/ is path-filtered to its domain: a push touching only frontend/** builds/deploys nothing else. Images stay separate (emi/thermograph/backend, emi/thermograph/frontend, each tagged sha-<12hex>), deploys stay per-service (infra/deploy/deploy.sh SERVICE=backend|frontend|all), and the API version contract (GET /api/version, PAYLOAD_VER) still lets FE and BE ship out of lockstep. The one intentionally coupled piece is pr-build.yml: a single always-running gate required check that builds only the domains a PR touches (a path-filtered required check would deadlock auto-merge).

Branch model (unchanged from the split era): PRs → dev, main → beta, release → prod. Infra isn't environment-staged the same way app images are: beta's and prod's checkouts (both on vps2) track main; dev's checkout (on vps1) tracks dev itself, since it's the one environment that isn't a rehearsal for something downstream.

Before pointing anything live at this repo, read CUTOVER-NOTES.md.