Thermograph monorepo: graded-climate API + SSR frontend + infra, domain-specific containerized deploys
Find a file
Emi Griffith cb69a22c0f
All checks were successful
PR build (required check) / changes (pull_request) Successful in 7s
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Successful in 1m11s
PR build (required check) / build-backend (pull_request) Successful in 1m25s
PR build (required check) / gate (pull_request) Successful in 2s
secrets: factor shared values into common.yaml; stop the renderer failing open
Two changes to how credentials are stored and distributed.

1. common.yaml now exists.

The renderer has always concatenated common.yaml then <env>.yaml, host winning,
and the README has always documented common.yaml with a checkmark. The file was
never created. So every value shared by prod and beta was duplicated in both
vaults, free to drift apart with nothing to detect it.

16 values move to common.yaml: VAPID keypair, metrics token, IndexNow key,
REGISTRY_TOKEN, the S3 endpoint/bucket and both S3 keypairs, plus shared config.
5 stay per-host because their values genuinely differ (APP_CPUS, DB_CPUS,
DB_MEMORY, WORKERS, THERMOGRAPH_BASE_URL). 8 exist only on prod — the Discord
and mail credentials, which beta does not have at all.

Three more are held back deliberately despite being identical today:
POSTGRES_PASSWORD, THERMOGRAPH_AUTH_SECRET, THERMOGRAPH_DATABASE_URL. These are
the credentials that let one environment act as another, and beta is the more
exposed box — it serves public Forgejo and Grafana. They match only because beta
was seeded from prod. Keeping them per-host costs one line each and preserves
the ability to diverge; putting them in common.yaml would encode the equivalence
as intentional and make breaking it a migration rather than an edit. Reasoning
recorded in the vault README, since a future reader will otherwise "fix" it.

Done without any plaintext leaving prod: ciphertext was shipped up, decrypted
against the host's age key, recombined, re-encrypted, and shipped back. The
consolidation refuses to write unless it has proved merged(common + env) equals
the original env vault exactly — same keys, same values — and re-verifies from
the written files afterwards. Both checks passed for prod and beta.

2. render_thermograph_secrets no longer fails open.

One `return 0` covered two different situations: "this host is not configured
for SOPS" (true of the LAN dev box, and correct to succeed) and "this host IS
configured but its vault is not where we looked". The second is a failure, and
returning 0 made it a silent no-op wearing a success code — a deploy would
report success having rendered nothing, and the host would keep serving whatever
/etc/thermograph.env already held, including after a rotation.

The likely cause is passing the wrong root: the function wants the directory
containing deploy/secrets, which on the hosts is /opt/thermograph/infra, not
/opt/thermograph. That mistake looked exactly like "not configured here", which
is why it went unnoticed. It now exits 1 and names the probable cause.

Claude-Session: https://claude.ai/code/session_0182KTMrsTHJc3TcewCatJFY
2026-07-24 17:40:26 -07:00
.forgejo/workflows observability: add the estate's first alerting; supervise Postfix 2026-07-24 13:19:19 -07:00
backend data/climate: fix geocode_nominatim's NameError on every call (P0, live) (#51) 2026-07-24 19:45:49 +00:00
frontend shell: add shellcheck CI guard and drive the tree to zero findings (#19) 2026-07-23 22:26:05 +00:00
infra secrets: factor shared values into common.yaml; stop the renderer failing open 2026-07-24 17:40:26 -07:00
observability observability: add the estate's first alerting; supervise Postfix 2026-07-24 13:19:19 -07:00
.gitignore Drop accidentally-committed worktrees; ignore .claude/worktrees 2026-07-24 15:57:34 -07:00
CLAUDE.md docs: monorepo README, cutover runbook, root agent instructions 2026-07-22 22:11:33 -07:00
CUTOVER-NOTES.md docs: record the 2026-07-22 branch-migration sweep in cutover notes 2026-07-22 22:27:09 -07:00
README.md docs: monorepo README, cutover runbook, root agent instructions 2026-07-22 22:11:33 -07:00

thermograph

The Thermograph monorepo — the split repos reunified (2026-07-22) with full history via subtree merges, while keeping everything the split was actually for: per-domain images, per-domain deploys, and an async FE/BE contract.

Domains

Dir What CI
backend/ FastAPI graded-climate API, accounts, notifications (Discord bot, push, mail), data pipeline backend-build-push → image emi/thermograph/backend; backend-deploy[-prod|-dev]
frontend/ Public client: static JS/CSS + SSR pages frontend-* mirrors of the above; image emi/thermograph/frontend
infra/ Compose, deploy scripts, terraform, SOPS secrets vault, ops cron infra-sync (host checkout + secrets render), secrets-guard, ops-cron
observability/ Loki + Grafana + Alloy stack observability-validate

thermograph-docs deliberately stays its own repo (ADRs + runbooks, no build artifacts, different change cadence).

How CI stays decoupled

Every workflow in .forgejo/workflows/ is path-filtered to its domain: a push touching only frontend/** builds/deploys nothing else. Images stay separate (emi/thermograph/backend, emi/thermograph/frontend, each tagged sha-<12hex>), deploys stay per-service (infra/deploy/deploy.sh SERVICE=backend|frontend|all), and the API version contract (GET /api/version, PAYLOAD_VER) still lets FE and BE ship out of lockstep. The one intentionally coupled piece is pr-build.yml: a single always-running gate required check that builds only the domains a PR touches (a path-filtered required check would deadlock auto-merge).

Branch model (unchanged from the split era): PRs → dev, main → beta, release → prod; infra tracked via main on all hosts.

Before pointing anything live at this repo, read CUTOVER-NOTES.md.