Thermograph monorepo: graded-climate API + SSR frontend + infra, domain-specific containerized deploys
Find a file
Emi Griffith 486a194086
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 1s
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 5s
Sync infra to hosts / sync-centralis (push) Has been skipped
secrets-guard / encrypted (push) Successful in 4s
shell-lint / shellcheck (push) Successful in 7s
forgejo: record why the runner is on vps2, rather than leaving a contradiction
This README said, in bold, that a runner must never go on prod or beta. A
runner has been registered and running on vps2 since 2026-08-01. Both
statements cannot stand: an instruction the estate visibly ignores teaches
readers to ignore the next one too.

The original objection is kept rather than deleted, because it is correct on
its merits -- docker_host: automount hands job containers the host's Docker
socket, which on vps2 is root over both the prod and beta stacks. What changed
is that the alternative proved worse. The desktop was the ONLY registered
runner in the estate; when it dropped on 2026-07-31 nothing merged, nothing
deployed and the nightly backup did not fire for 21 hours. On a repo where
every branch is protected and every change is a PR, one absent runner freezes
everything, and the backup hangs off the same path.

So the section now records the trade and the bounds actually applied --
capacity 1, --cpus=2/--memory=4g on job containers, valid_volumes empty, no
thermograph network -- and states plainly that this is defence against
accident, not against a hostile workflow author. Forgejo's database already
stored VPS2_SSH_KEY, which is root on that box; what changed is that the
material is now reachable by a job rather than only at rest.

It also keeps the original advice for the case it was actually written for:
if you are adding CAPACITY rather than REDUNDANCY, raise capacity or use a box
that hosts nothing.
2026-08-01 09:59:14 -07:00
.claude BRANCHING: the escape hatch, now that apply_to_admins binds the owner too 2026-07-26 15:05:26 -07:00
.forgejo/workflows registry: move image and repo references to the Jinemi namespace 2026-08-01 09:25:02 -07:00
backend registry: move image and repo references to the Jinemi namespace 2026-08-01 09:25:02 -07:00
docs/onboarding registry: move image and repo references to the Jinemi namespace 2026-08-01 09:25:02 -07:00
frontend registry: move image and repo references to the Jinemi namespace 2026-08-01 09:25:02 -07:00
infra forgejo: record why the runner is on vps2, rather than leaving a contradiction 2026-08-01 09:59:14 -07:00
observability registry: move image and repo references to the Jinemi namespace 2026-08-01 09:25:02 -07:00
.gitignore gitignore: ignore .claude/settings.local.json 2026-08-01 09:16:06 -07:00
CLAUDE.md registry: move image and repo references to the Jinemi namespace 2026-08-01 09:25:02 -07:00
CUTOVER-NOTES.md registry: move image and repo references to the Jinemi namespace 2026-08-01 09:25:02 -07:00
README.md registry: move image and repo references to the Jinemi namespace 2026-08-01 09:25:02 -07:00

thermograph

The Thermograph monorepo — the split repos reunified (2026-07-22) with full history via subtree merges, while keeping everything the split was actually for: per-domain images, per-domain deploys, and an async FE/BE contract.

Domains

Dir What CI
backend/ FastAPI graded-climate API, accounts, notifications (Discord bot, push, mail), data pipeline build-push → image jinemi/thermograph/backend; deploy
frontend/ Public client: static JS/CSS + SSR pages same build-push / deploy workflows, matrixed by domain; image jinemi/thermograph/frontend
infra/ Compose (dev, on vps1) + two Swarm stacks co-resident on vps2 (beta, prod), deploy scripts, terraform, SOPS secrets vault, ops cron infra-sync (host checkout + secrets render), secrets-guard, ops-cron
observability/ Loki + Grafana + Alloy stack observability-validate

thermograph-docs deliberately stays its own repo (ADRs + runbooks, no build artifacts, different change cadence).

New here?

docs/onboarding/ is the developer onboarding path: orientation, verified local-setup recipes, a per-domain deep dive, the cross-service contracts that break silently, the release flow, and a list of which docs in this repo are currently stale.

How CI stays decoupled

Every workflow in .forgejo/workflows/ is path-filtered to its domain: a push touching only frontend/** builds/deploys nothing else. Images stay separate (jinemi/thermograph/backend, jinemi/thermograph/frontend, each tagged sha-<12hex>), deploys stay per-service (infra/deploy/deploy.sh SERVICE=backend|frontend|all), and the API version contract (GET /api/version, PAYLOAD_VER) still lets FE and BE ship out of lockstep. The one intentionally coupled piece is pr-build.yml: a single always-running gate required check that builds only the domains a PR touches (a path-filtered required check would deadlock auto-merge).

Branch model (unchanged from the split era): PRs → dev, main → beta, release → prod. Infra isn't environment-staged the same way app images are: beta's and prod's checkouts (both on vps2) track main; dev's checkout (on vps1) tracks dev itself, since it's the one environment that isn't a rehearsal for something downstream.

Before pointing anything live at this repo, read CUTOVER-NOTES.md.