Thermograph monorepo: graded-climate API + SSR frontend + infra, domain-specific containerized deploys
Find a file
Emi Griffith b2077b5294
All checks were successful
PR build (required check) / changes (pull_request) Successful in 5s
secrets-guard / encrypted (pull_request) Successful in 4s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 1s
secrets: make the two-host age key quorum a checked policy
The copies of /etc/thermograph/age.key on vps1 and vps2 have always read as
provisioning residue — a side effect of provision-secrets.sh rather than a
decision. They are in fact the estate's recovery quorum: that key is the single
recovery root for all five SOPS vaults AND for every off-box backup, since
ops-cron.yml:142,197 pipe those dumps through `age -r` to the same recipient.

On 2026-07-30 the operator's copy was shredded from the desktop before it had been
copied to its replacement machine, leaving zero operator-side copies. These two were
the only surviving material and the key was restored from vps2. Documenting the
ordering rule that would have prevented it, and adding the check that makes the
quorum verifiable rather than assumed.

verify-age-quorum.sh asserts presence, permissions and hash equality across both
hosts. Divergence is treated as seriously as absence: two hosts holding different
keys means one renders secrets the other cannot read, and a restore silently picks
the wrong one. Prints SHA-256 digests and modes only — a digest of a 32-byte random
key is not reversible, so it is safe in a CI log, and the key is never read or
copied anywhere. Exit status makes it usable as a cron gate.

Also replaces the "wasn't re-verified for this pass" hedge in the README with what
was actually measured: vps1 0440 root:deploy, vps2 0400 root:root, digests equal.
The dev render does not need vps1's group bit — it runs as `agent`, which is not in
group `deploy` and reaches the key through passwordless sudo, with
/opt/thermograph-dev owned agent:agent. Tightening vps1 to 0400 root:root is
therefore expected to be safe, and is called out as the follow-up rather than done
here because it changes a live host on the deploy path.

The check reports that group read as an advisory rather than a failure. It is the
current provisioned state, and a check that fails on day one is a check that is
ignored by day three.
2026-07-31 05:12:58 +00:00
.claude BRANCHING: the escape hatch, now that apply_to_admins binds the owner too 2026-07-26 15:05:26 -07:00
.forgejo/workflows infra-sync: refuse to render dev's vault onto a host that is not dev (#107) 2026-07-26 07:05:09 +00:00
backend accounts: sign in with Google, on a shared provider-agnostic OAuth engine (#122) 2026-07-27 00:56:43 +00:00
docs/onboarding infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00
frontend accounts: sign in with Google, on a shared provider-agnostic OAuth engine (#122) 2026-07-27 00:56:43 +00:00
infra secrets: make the two-host age key quorum a checked policy 2026-07-31 05:12:58 +00:00
observability infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00
.gitignore Drop accidentally-committed worktrees; ignore .claude/worktrees 2026-07-24 15:57:34 -07:00
CLAUDE.md docs: the orchestrator runbook and the module map 2026-07-26 12:53:33 -07:00
CUTOVER-NOTES.md infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00
README.md infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00

thermograph

The Thermograph monorepo — the split repos reunified (2026-07-22) with full history via subtree merges, while keeping everything the split was actually for: per-domain images, per-domain deploys, and an async FE/BE contract.

Domains

Dir What CI
backend/ FastAPI graded-climate API, accounts, notifications (Discord bot, push, mail), data pipeline build-push → image emi/thermograph/backend; deploy
frontend/ Public client: static JS/CSS + SSR pages same build-push / deploy workflows, matrixed by domain; image emi/thermograph/frontend
infra/ Compose (dev, on vps1) + two Swarm stacks co-resident on vps2 (beta, prod), deploy scripts, terraform, SOPS secrets vault, ops cron infra-sync (host checkout + secrets render), secrets-guard, ops-cron
observability/ Loki + Grafana + Alloy stack observability-validate

thermograph-docs deliberately stays its own repo (ADRs + runbooks, no build artifacts, different change cadence).

New here?

docs/onboarding/ is the developer onboarding path: orientation, verified local-setup recipes, a per-domain deep dive, the cross-service contracts that break silently, the release flow, and a list of which docs in this repo are currently stale.

How CI stays decoupled

Every workflow in .forgejo/workflows/ is path-filtered to its domain: a push touching only frontend/** builds/deploys nothing else. Images stay separate (emi/thermograph/backend, emi/thermograph/frontend, each tagged sha-<12hex>), deploys stay per-service (infra/deploy/deploy.sh SERVICE=backend|frontend|all), and the API version contract (GET /api/version, PAYLOAD_VER) still lets FE and BE ship out of lockstep. The one intentionally coupled piece is pr-build.yml: a single always-running gate required check that builds only the domains a PR touches (a path-filtered required check would deadlock auto-merge).

Branch model (unchanged from the split era): PRs → dev, main → beta, release → prod. Infra isn't environment-staged the same way app images are: beta's and prod's checkouts (both on vps2) track main; dev's checkout (on vps1) tracks dev itself, since it's the one environment that isn't a rehearsal for something downstream.

Before pointing anything live at this repo, read CUTOVER-NOTES.md.