|
All checks were successful
PR build (required check) / changes (pull_request) Successful in 7s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 6s
secrets-guard / encrypted (push) Successful in 5s
shell-lint / shellcheck (push) Successful in 7s
Three fixes, all on the path between "the migration looks ready" and "the migration is ready". verify-parity.sh never sourced env-topology.sh, so TG_BAO_APPROLE was unset and render-secrets-openbao.sh fell through to the bare default -- prod's credential -- for every environment. On vps2 that made `--env beta` authenticate as tg-prod and take a 403 from tg-host-prod.hcl on thermograph/data/env/beta. `--all` failed the same way, taking dev with it. deploy.sh and deploy-stack.sh always sourced it; only the verifier did not, which is the worst place for the omission: the tool whose job is to notice divergence was itself diverging from the path it verifies. It now calls thermograph_topology per environment and takes TG_SKIP_COMMON from there rather than re-deriving it. infra-sync.yml renders all three env files on every push touching infra/**, and none of its three jobs sourced env-topology.sh either. TG_SECRETS_BACKEND was therefore unset and render-secrets.sh:44 defaulted to sops. Flipping the selector would have changed deploy.sh and deploy-stack.sh but not this workflow, which would have kept re-rendering from SOPS -- silent while parity holds, and a hard failure the moment the SOPS files are retired. The one-line cutover the README describes was never sufficient on its own. ops-cron.yml gains a secrets-parity job: prod and beta from vps2, dev from vps1, nightly, failing loudly on a mismatch. README.md called this the gate for a cutover; nothing implemented it. A gate that exists only in prose gets satisfied by assertion. It grants CI no vault access -- the host renders, CI only asks it to. First measured run, immediately after the verifier fix: dev 12 keys PASS, byte-identical beta 24 keys FAIL, 3 values differ prod 32 keys FAIL, 3 values differ The same three keys on both: THERMOGRAPH_S3_SECRET_KEY, THERMOGRAPH_LAKE_S3_SECRET_KEY, THERMOGRAPH_VAPID_CONTACT. All three live in common.yaml, seeded 2026-07-31 02:30Z -- before the Contabo rotation landed. dev passes because dev never layers common. One `seed-from-sops.sh --env common` fixes beta and prod together, and the 7-day clock starts after that. |
||
|---|---|---|
| .. | ||
| .claude/skills/key-gaps | ||
| deploy | ||
| lake-iceberg | ||
| openbao | ||
| ops | ||
| terraform | ||
| .env.example | ||
| .gitignore | ||
| .sops.yaml | ||
| ACCESS.md | ||
| CLAUDE.md | ||
| DEPLOY-DEV.md | ||
| DEPLOY.md | ||
| docker-compose.dev.yml | ||
| docker-compose.openmeteo.yml | ||
| docker-compose.yml | ||
| Makefile | ||
| README.md | ||
infra/
Infrastructure for Thermograph: Terraform host
provisioning, the SOPS+age secrets vault, WireGuard/Swarm networking, Forgejo,
Caddy, mail, and the deploy scripts that run the already-built app images on each
host. This is a domain of the emi/thermograph monorepo — a host's checkout
(/opt/thermograph, /opt/thermograph-beta, or /opt/thermograph-dev) is a
checkout of the whole monorepo, and infra/ never builds app source; it only
runs published images.
terraform/— provisions/configures hosts (SSH-driven by default; an optional GCP-creating module is scaffolded, no live resources yet). Seeterraform/README.md. No tfstate is persisted anywhere — treatapplyas executable documentation, not a routine operation.deploy/secrets/— the git-native SOPS+age secrets vault (every app secret, encrypted at rest, rendered at deploy time). Seedeploy/secrets/README.md.deploy/swarm/,deploy/forgejo/— the WireGuard/Swarm mesh spanning vps1, vps2 and the desktop, and Forgejo (git + CI + registry), pinned to vps1. SeeACCESS.md.deploy/env-topology.sh— the single source of truth for where each environment (dev/beta/prod) lives: host, checkout path, branch, deploy mode, stack/compose name, env file, LB ports, DB role/database, service-name prefix. Every deploy path sources it and derives its behavior from it rather than guessing from the host it happens to be running on — necessary since vps2 alone now runs two environments.deploy/deploy.sh— the single deploy entry point for dev, beta and prod. TakesSERVICE=backend|frontend|allplusBACKEND_IMAGE_TAG/FRONTEND_IMAGE_TAG(and, on vps2,THERMOGRAPH_ENV=beta|prodto say which of the two checkouts it's acting on), resets the host checkout, renders secrets, and routes to the right orchestrator perenv-topology.sh.deploy/deploy-dev.sh— a thin dev-specific wrapper arounddeploy.sh(dev compose overlay, dev's secrets policy). SeeDEPLOY-DEV.md.deploy/stack/— the Swarm path, live on vps2 for both prod and beta:thermograph-stack.yml(prod: db, web, worker, lake, daemon, frontend, autoscaler, autoscaler-lake) andthermograph-beta-stack.yml(beta: the same service shape minusdband the autoscalers, every service name prefixedbeta-).deploy-stack.sh,autoscale.shandlb/are shared by both. Rolling updates are start-first, health-gated, with auto-rollback.STACK_TEST=1rehearses the whole stack on throwaway volumes and ports.docker-compose*.yml— the compose path, live only on dev (vps1) (db, backend, lake, daemon, frontend).docker-compose.dev.ymlis dev's mesh-only overlay;docker-compose.openmeteo.ymlis the self-hosted Open-Meteo overlay (prod only).make dev-upalso runs this path locally as a laptop convenience — that is not an "environment", just a local rehearsal.
Which path an environment takes is decided by deploy/env-topology.sh
(TG_DEPLOY_MODE, keyed by dev/beta/prod): dev is compose, beta and prod
are both stack. The old host-wide marker /etc/thermograph/deploy-mode still
exists as a fallback for a by-hand run with no explicit environment, but it
cannot describe vps2, which runs two environments in two different checkouts —
so it is no longer the thing that decides where files go. The workflows never
need to know which mode an environment runs; they only pass THERMOGRAPH_ENV.
Branches & how changes reach each environment
dev— deploys to dev on vps1 (/opt/thermograph-dev).main— deploys to beta on vps2 (/opt/thermograph-beta).release— deploys to prod on vps2 (/opt/thermograph).
App code IS environment-staged this way (dev→main→release maps to
vps1/dev → vps2/beta → vps2/prod via image tags, one Deploy workflow keyed by
branch). Infra itself is not environment-staged the same way: infra-sync.yml
fires on a push touching infra/** — on dev it fast-forwards vps1's dev
checkout, on main it fast-forwards both of vps2's checkouts (beta and
prod) — re-rendering each environment's own env file from the vault. It
deliberately does not roll any service: image tags are the app domains'
axis, not infra's. A compose or stack change that must recreate containers
takes effect on the next app deploy, or a by-hand SERVICE=all … deploy/deploy.sh (or deploy-dev.sh) run.
Note the asymmetry this leaves: dev's infra checkout tracks dev, the same
branch its app images are staged by. Beta's and prod's infra checkouts both
track main — prod's app images are staged by release, but prod's infra
checkout follows main, same as beta's.