|
All checks were successful
PR build (required check) / changes (pull_request) Successful in 5s
secrets-guard / encrypted (pull_request) Successful in 4s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 1s
The copies of /etc/thermograph/age.key on vps1 and vps2 have always read as provisioning residue — a side effect of provision-secrets.sh rather than a decision. They are in fact the estate's recovery quorum: that key is the single recovery root for all five SOPS vaults AND for every off-box backup, since ops-cron.yml:142,197 pipe those dumps through `age -r` to the same recipient. On 2026-07-30 the operator's copy was shredded from the desktop before it had been copied to its replacement machine, leaving zero operator-side copies. These two were the only surviving material and the key was restored from vps2. Documenting the ordering rule that would have prevented it, and adding the check that makes the quorum verifiable rather than assumed. verify-age-quorum.sh asserts presence, permissions and hash equality across both hosts. Divergence is treated as seriously as absence: two hosts holding different keys means one renders secrets the other cannot read, and a restore silently picks the wrong one. Prints SHA-256 digests and modes only — a digest of a 32-byte random key is not reversible, so it is safe in a CI log, and the key is never read or copied anywhere. Exit status makes it usable as a cron gate. Also replaces the "wasn't re-verified for this pass" hedge in the README with what was actually measured: vps1 0440 root:deploy, vps2 0400 root:root, digests equal. The dev render does not need vps1's group bit — it runs as `agent`, which is not in group `deploy` and reaches the key through passwordless sudo, with /opt/thermograph-dev owned agent:agent. Tightening vps1 to 0400 root:root is therefore expected to be safe, and is called out as the follow-up rather than done here because it changes a live host on the deploy path. The check reports that group read as an advisory rather than a failure. It is the current provisioned state, and a check that fails on day one is a check that is ignored by day three. |
||
|---|---|---|
| .claude | ||
| .forgejo/workflows | ||
| backend | ||
| docs/onboarding | ||
| frontend | ||
| infra | ||
| observability | ||
| .gitignore | ||
| CLAUDE.md | ||
| CUTOVER-NOTES.md | ||
| README.md | ||
thermograph
The Thermograph monorepo — the split repos reunified (2026-07-22) with full history via subtree merges, while keeping everything the split was actually for: per-domain images, per-domain deploys, and an async FE/BE contract.
Domains
| Dir | What | CI |
|---|---|---|
backend/ |
FastAPI graded-climate API, accounts, notifications (Discord bot, push, mail), data pipeline | build-push → image emi/thermograph/backend; deploy |
frontend/ |
Public client: static JS/CSS + SSR pages | same build-push / deploy workflows, matrixed by domain; image emi/thermograph/frontend |
infra/ |
Compose (dev, on vps1) + two Swarm stacks co-resident on vps2 (beta, prod), deploy scripts, terraform, SOPS secrets vault, ops cron | infra-sync (host checkout + secrets render), secrets-guard, ops-cron |
observability/ |
Loki + Grafana + Alloy stack | observability-validate |
thermograph-docs deliberately stays its own repo (ADRs + runbooks, no
build artifacts, different change cadence).
New here?
docs/onboarding/ is the developer onboarding
path: orientation, verified local-setup recipes, a per-domain deep dive, the
cross-service contracts that break silently, the release flow, and a list of
which docs in this repo are currently stale.
How CI stays decoupled
Every workflow in .forgejo/workflows/ is path-filtered to its domain: a
push touching only frontend/** builds/deploys nothing else. Images stay
separate (emi/thermograph/backend, emi/thermograph/frontend, each tagged
sha-<12hex>), deploys stay per-service (infra/deploy/deploy.sh SERVICE=backend|frontend|all), and the API version contract
(GET /api/version, PAYLOAD_VER) still lets FE and BE ship out of lockstep.
The one intentionally coupled piece is pr-build.yml: a single always-running
gate required check that builds only the domains a PR touches (a
path-filtered required check would deadlock auto-merge).
Branch model (unchanged from the split era): PRs → dev, main → beta,
release → prod. Infra isn't environment-staged the same way app images are:
beta's and prod's checkouts (both on vps2) track main; dev's checkout (on
vps1) tracks dev itself, since it's the one environment that isn't a
rehearsal for something downstream.
Before pointing anything live at this repo, read CUTOVER-NOTES.md.