Thermograph monorepo: graded-climate API + SSR frontend + infra, domain-specific containerized deploys
Find a file
Emi Griffith 8a2c838663
All checks were successful
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / gate (pull_request) Successful in 1s
secrets-guard / encrypted (pull_request) Successful in 4s
PR build (required check) / changes (pull_request) Successful in 8s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 5s
secrets-guard / encrypted (push) Successful in 4s
shell-lint / shellcheck (push) Successful in 6s
forgejo: restart the db service on a clean exit, not just on failure
forgejo_db has been at 0/1 replicas since 2026-07-29 01:24 UTC. Postgres hit
an invalid data-directory lock file ("could not open file postmaster.pid ...
performing immediate shutdown because data directory lock file is invalid")
and exited 0. With restart_policy.condition=on-failure, Swarm read the zero
status as successful completion, marked the task Complete, and never
rescheduled it.

The forgejo service itself stayed Up and kept serving its homepage, so the
outage presented as every repository page, the whole API and all CI returning
500 with "dial tcp: lookup db on 127.0.0.11:53: no such host" — including the
auth path, which is why API calls reported "user does not exist [uid: 0]"
rather than a database error.

on-failure cannot distinguish "finished successfully" from "shut itself down
and should be restarted", and Postgres exits 0 on several such paths, so it is
the wrong policy for an always-on stateful service.

This is the durable fix; it does not restart the currently stopped task.
2026-07-30 06:30:30 -07:00
.claude BRANCHING: the escape hatch, now that apply_to_admins binds the owner too 2026-07-26 15:05:26 -07:00
.forgejo/workflows infra-sync: refuse to render dev's vault onto a host that is not dev (#107) 2026-07-26 07:05:09 +00:00
backend accounts: sign in with Google, on a shared provider-agnostic OAuth engine (#122) 2026-07-27 00:56:43 +00:00
docs/onboarding infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00
frontend accounts: sign in with Google, on a shared provider-agnostic OAuth engine (#122) 2026-07-27 00:56:43 +00:00
infra forgejo: restart the db service on a clean exit, not just on failure 2026-07-30 06:30:30 -07:00
observability infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00
.gitignore Drop accidentally-committed worktrees; ignore .claude/worktrees 2026-07-24 15:57:34 -07:00
CLAUDE.md docs: the orchestrator runbook and the module map 2026-07-26 12:53:33 -07:00
CUTOVER-NOTES.md infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00
README.md infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00

thermograph

The Thermograph monorepo — the split repos reunified (2026-07-22) with full history via subtree merges, while keeping everything the split was actually for: per-domain images, per-domain deploys, and an async FE/BE contract.

Domains

Dir What CI
backend/ FastAPI graded-climate API, accounts, notifications (Discord bot, push, mail), data pipeline build-push → image emi/thermograph/backend; deploy
frontend/ Public client: static JS/CSS + SSR pages same build-push / deploy workflows, matrixed by domain; image emi/thermograph/frontend
infra/ Compose (dev, on vps1) + two Swarm stacks co-resident on vps2 (beta, prod), deploy scripts, terraform, SOPS secrets vault, ops cron infra-sync (host checkout + secrets render), secrets-guard, ops-cron
observability/ Loki + Grafana + Alloy stack observability-validate

thermograph-docs deliberately stays its own repo (ADRs + runbooks, no build artifacts, different change cadence).

New here?

docs/onboarding/ is the developer onboarding path: orientation, verified local-setup recipes, a per-domain deep dive, the cross-service contracts that break silently, the release flow, and a list of which docs in this repo are currently stale.

How CI stays decoupled

Every workflow in .forgejo/workflows/ is path-filtered to its domain: a push touching only frontend/** builds/deploys nothing else. Images stay separate (emi/thermograph/backend, emi/thermograph/frontend, each tagged sha-<12hex>), deploys stay per-service (infra/deploy/deploy.sh SERVICE=backend|frontend|all), and the API version contract (GET /api/version, PAYLOAD_VER) still lets FE and BE ship out of lockstep. The one intentionally coupled piece is pr-build.yml: a single always-running gate required check that builds only the domains a PR touches (a path-filtered required check would deadlock auto-merge).

Branch model (unchanged from the split era): PRs → dev, main → beta, release → prod. Infra isn't environment-staged the same way app images are: beta's and prod's checkouts (both on vps2) track main; dev's checkout (on vps1) tracks dev itself, since it's the one environment that isn't a rehearsal for something downstream.

Before pointing anything live at this repo, read CUTOVER-NOTES.md.