thermograph/infra
Emi Griffith f25466ca76
All checks were successful
PR build (required check) / changes (pull_request) Successful in 9s
shell-lint / shellcheck (pull_request) Successful in 9s
secrets-guard / encrypted (pull_request) Successful in 11s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
forgejo: make the desktop runner unit restart-always, fix the vps2 config path
Both runners exist and are up; these are the two defects found while verifying
that.

The desktop's systemd --user unit had Restart=on-failure. That is the same
distinction that took Forgejo down for 27 hours on 2026-07-29 (docker-stack.yml
records it for the db service): a daemon that exits 0 is not "finished
successfully", and on-failure cannot tell that from a clean shutdown. The
failure mode is silent — unit inactive (dead), runner offline in the UI, and
every protected-branch merge blocked on a check nothing will produce.

The StartLimit directives that bound that retry loop were in [Service], where
systemd accepts them without complaint and ignores them; the live unit was
running the 10s default rather than the intended 300s. Moved to [Unit], which
is where they are read.

runner-vps2/README told you to copy config.yaml next to docker-compose.yml, but
the compose file mounts ./data:/data and loads --config /data/config.yaml, so a
config there is invisible to the container — the daemon starts on defaults with
no --add-host, and every registry push then fails as if the credential were
wrong. vps2 had a stray copy at the documented path proving the instruction had
been followed.
2026-08-01 11:50:42 -07:00
..
.claude/skills/key-gaps infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00
deploy forgejo: make the desktop runner unit restart-always, fix the vps2 config path 2026-08-01 11:50:42 -07:00
lake-iceberg Iceberg conversion container for the ERA5 lake (infra/lake-iceberg) (#24) 2026-07-23 22:56:27 +00:00
openbao openbao: make verify-parity --all host-aware 2026-08-01 09:18:05 -07:00
ops registry: move image and repo references to the Jinemi namespace 2026-08-01 09:25:02 -07:00
terraform registry: move image and repo references to the Jinemi namespace 2026-08-01 09:25:02 -07:00
.env.example registry: move image and repo references to the Jinemi namespace 2026-08-01 09:25:02 -07:00
.gitignore infra: mirror LAN dev secrets under $HOME for snap-confined Docker (#85) 2026-07-25 07:16:08 +00:00
.sops.yaml Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
ACCESS.md forgejo: make dev.jinemi.com canonical (ROOT_URL and OAuth callback) 2026-08-01 09:32:12 -07:00
CLAUDE.md forgejo: make dev.jinemi.com canonical (ROOT_URL and OAuth callback) 2026-08-01 09:32:12 -07:00
DEPLOY-DEV.md infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00
DEPLOY.md infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00
docker-compose.dev.yml dev: serve at dev.thermograph.org behind basic auth, bound to loopback (#111) 2026-07-26 07:49:43 +00:00
docker-compose.openmeteo.yml Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
docker-compose.yml registry: move image and repo references to the Jinemi namespace 2026-08-01 09:25:02 -07:00
Makefile infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103) 2026-07-26 06:56:38 +00:00
README.md registry: move image and repo references to the Jinemi namespace 2026-08-01 09:25:02 -07:00

infra/

Infrastructure for Thermograph: Terraform host provisioning, the SOPS+age secrets vault, WireGuard/Swarm networking, Forgejo, Caddy, mail, and the deploy scripts that run the already-built app images on each host. This is a domain of the Jinemi/thermograph monorepo — a host's checkout (/opt/thermograph, /opt/thermograph-beta, or /opt/thermograph-dev) is a checkout of the whole monorepo, and infra/ never builds app source; it only runs published images.

  • terraform/ — provisions/configures hosts (SSH-driven by default; an optional GCP-creating module is scaffolded, no live resources yet). See terraform/README.md. No tfstate is persisted anywhere — treat apply as executable documentation, not a routine operation.
  • deploy/secrets/ — the git-native SOPS+age secrets vault (every app secret, encrypted at rest, rendered at deploy time). See deploy/secrets/README.md.
  • deploy/swarm/, deploy/forgejo/ — the WireGuard/Swarm mesh spanning vps1, vps2 and the desktop, and Forgejo (git + CI + registry), pinned to vps1. See ACCESS.md.
  • deploy/env-topology.sh — the single source of truth for where each environment (dev/beta/prod) lives: host, checkout path, branch, deploy mode, stack/compose name, env file, LB ports, DB role/database, service-name prefix. Every deploy path sources it and derives its behavior from it rather than guessing from the host it happens to be running on — necessary since vps2 alone now runs two environments.
  • deploy/deploy.sh — the single deploy entry point for dev, beta and prod. Takes SERVICE=backend|frontend|all plus BACKEND_IMAGE_TAG/FRONTEND_IMAGE_TAG (and, on vps2, THERMOGRAPH_ENV=beta|prod to say which of the two checkouts it's acting on), resets the host checkout, renders secrets, and routes to the right orchestrator per env-topology.sh.
  • deploy/deploy-dev.sh — a thin dev-specific wrapper around deploy.sh (dev compose overlay, dev's secrets policy). See DEPLOY-DEV.md.
  • deploy/stack/ — the Swarm path, live on vps2 for both prod and beta: thermograph-stack.yml (prod: db, web, worker, lake, daemon, frontend, autoscaler, autoscaler-lake) and thermograph-beta-stack.yml (beta: the same service shape minus db and the autoscalers, every service name prefixed beta-). deploy-stack.sh, autoscale.sh and lb/ are shared by both. Rolling updates are start-first, health-gated, with auto-rollback. STACK_TEST=1 rehearses the whole stack on throwaway volumes and ports.
  • docker-compose*.yml — the compose path, live only on dev (vps1) (db, backend, lake, daemon, frontend). docker-compose.dev.yml is dev's mesh-only overlay; docker-compose.openmeteo.yml is the self-hosted Open-Meteo overlay (prod only). make dev-up also runs this path locally as a laptop convenience — that is not an "environment", just a local rehearsal.

Which path an environment takes is decided by deploy/env-topology.sh (TG_DEPLOY_MODE, keyed by dev/beta/prod): dev is compose, beta and prod are both stack. The old host-wide marker /etc/thermograph/deploy-mode still exists as a fallback for a by-hand run with no explicit environment, but it cannot describe vps2, which runs two environments in two different checkouts — so it is no longer the thing that decides where files go. The workflows never need to know which mode an environment runs; they only pass THERMOGRAPH_ENV.

Branches & how changes reach each environment

  • dev — deploys to dev on vps1 (/opt/thermograph-dev).
  • main — deploys to beta on vps2 (/opt/thermograph-beta).
  • release — deploys to prod on vps2 (/opt/thermograph).

App code IS environment-staged this way (devmainrelease maps to vps1/dev → vps2/beta → vps2/prod via image tags, one Deploy workflow keyed by branch). Infra itself is not environment-staged the same way: infra-sync.yml fires on a push touching infra/** — on dev it fast-forwards vps1's dev checkout, on main it fast-forwards both of vps2's checkouts (beta and prod) — re-rendering each environment's own env file from the vault. It deliberately does not roll any service: image tags are the app domains' axis, not infra's. A compose or stack change that must recreate containers takes effect on the next app deploy, or a by-hand SERVICE=all … deploy/deploy.sh (or deploy-dev.sh) run.

Note the asymmetry this leaves: dev's infra checkout tracks dev, the same branch its app images are staged by. Beta's and prod's infra checkouts both track main — prod's app images are staged by release, but prod's infra checkout follows main, same as beta's.