Twelve documents under docs/onboarding/ covering orientation, local setup, the repo map, per-domain deep dives, the cross-service contracts, CI and the release flow, infra and secrets, observability, task recipes, and a list of which docs in this tree are currently stale. Every command in the setup guide was run against this checkout: the backend suite (429 passed, 8 skipped), the frontend Go suite, a venv boot of the backend, and a build+boot of the Go frontend. Two findings recorded along the way: - static/units.js's F_REGIONS is guarded by no test, despite three source comments claiming "a test asserts all three stay identical". The Go test only cross-checks the Go copy against the backend's Python. All four copies are currently identical. - backend/ and frontend/docker-compose.test.yml still default to the retired emi/thermograph-backend/app image path, and the frontend harness pins the split-era v0.0.2-split-ci tag. Claude-Session: https://claude.ai/code/session_01AfXqHrxCJLs2D7hpQkiUiJ
13 KiB
8. Infra and secrets
infra/ owns how and where the already-built images run. It never builds
app source. Hosts' /opt/thermograph is a checkout of this whole monorepo.
Read infra/CLAUDE.md alongside this.
The machines
| Host | Public | Mesh | Orchestrator | Login |
|---|---|---|---|---|
| prod | 169.58.46.181 / thermograph.org |
10.10.0.1 |
Docker Swarm | agent, passwordless sudo |
| beta | 75.119.132.91 / beta.thermograph.org |
10.10.0.2 |
compose | agent, passwordless sudo |
| desktop | — | 10.10.0.3 |
compose (LAN dev) | it's your box |
ssh -i ~/.ssh/thermograph_agent_ed25519 agent@169.58.46.181 # prod
ssh -i ~/.ssh/thermograph_agent_ed25519 agent@75.119.132.91 # beta
Beta also hosts Forgejo (git + CI + registry) and Grafana + Loki, which is why beta is guarded as strictly as prod: a destructive command there takes out git, CI and the registry at once. The LAN dev box is deliberately unguarded.
Every root-effective command on the VPSes is logged by auditd
(ausearch -k agentcmd). That's a feature — fixes are attributable.
Prefer Centralis for routine work (run_on_host, fleet_status,
deployed_version) over raw SSH.
The two orchestrators
Which path a host takes is decided by /etc/thermograph/deploy-mode: the
string stack makes deploy.sh exec deploy/stack/deploy-stack.sh; anything
else is compose. The workflows never need to know which mode a host runs.
| compose (beta, LAN dev) | Swarm (prod) | |
|---|---|---|
| File | infra/docker-compose.yml |
infra/deploy/stack/thermograph-stack.yml |
| Services | db, backend, lake, daemon, frontend |
db, web, worker, lake, daemon, frontend, autoscaler, autoscaler-lake |
| Rolling | up -d --no-deps <targets> |
start-first, health-gated, auto-rollback |
| Tag file | deploy/.image-tags.env |
deploy/.stack-image-tags.env |
Note prod splits backend into web and worker (the
THERMOGRAPH_ROLE split from 04); compose runs one backend
service in role all.
STACK_TEST=1 rehearses the entire Swarm deploy under stack name
thermograph-test on throwaway volumes and ports 18137/18080. Leftover
thermograph-test_* containers in logs are rehearsal residue — ignore them.
The volume/project-name coupling
Compose creates thermograph_pgdata / _appdata / _applogs; the Swarm stack
declares those exact names as external: true at the same mount paths. That
coupling is why name: thermograph is pinned in the compose file — running
compose from infra/ without it derives project infra, silently creating a
whole new stack with empty volumes beside the running one. deploy-dev.sh
exports COMPOSE_PROJECT_NAME=thermograph-dev (env wins over the file key) to
keep LAN dev separate on purpose. Keep both halves.
deploy/deploy.sh — read this before you touch a deploy
Single entry point for beta and prod. Contract:
SERVICE=backend BACKEND_IMAGE_TAG=sha-<12hex> /opt/thermograph/infra/deploy/deploy.sh
SERVICE=frontend FRONTEND_IMAGE_TAG=sha-<12hex> /opt/thermograph/infra/deploy/deploy.sh
SERVICE=all BACKEND_IMAGE_TAG=sha-<a> FRONTEND_IMAGE_TAG=sha-<b> …/deploy.sh
What it does, in order, and why each step is the way it is:
flockondeploy/.deploy.lock, with a self re-exec so the lock spans the whole run. Backend and frontend deploys can fire for the same push seconds apart and both SSH into one checkout; concurrent runs race the git reset, the compose project and the tag file.-w 600bounds the wait.- Render secrets from the SOPS vault into
/etc/thermograph.env, then source it, so a by-hand run interpolates the same as the systemd unit does. Guarded on the helper's existence so the very deploy that introducesrender-secrets.shis safe. git reset --hard origin/$BRANCHon the checkout root (BRANCHdefaults tomain). ⚠️ Uncommitted edits in/opt/thermographevaporate here. The one exception isdeploy/.image-tags.env, untracked on purpose.- Read
.image-tags.envso a single-service roll re-renders compose with both services' real tags and never accidentally recreates or downgrades the sibling. - Registry login only if
REGISTRY_TOKENis set. The CI paths don't pass one — the host is alreadydocker logined — and an unconditional login with an empty token would abort underset -e. - Pull, with a bounded ~5-minute retry.
build-push.ymlis a separate workflow triggered by the same push, and Forgejo'sneeds:only orders jobs within one file. A deploy racing ahead of its build and failing "not found" is confirmed-live behaviour; the retry covers a normal build and still fails loudly on a genuine problem. - Daemon capability probe. Infra tracks
mainwhile image tags are env-staged, so a host can legitimately be asked to roll a backend image older than the compose file — one built before the daemon binary existed. The scriptdocker runs the image to test for/usr/local/bin/thermograph-daemonand dropsdaemonfrom this run if it's absent, rather than leaving a container crash-looping on a missing binary. - Roll.
TARGETSarebackend → (backend, lake, daemon),frontend → (frontend),all → everything.daemonandlakeride withbackendand are never targets on their own — all three run the same image at the same tag, and rolling one without the others is exactly the version skew the/internal/*contract has no negotiation for. Single-service rolls use--no-deps; a fullalluses--remove-orphans(a renamed-away service's old container would otherwise squat its port — confirmed live). - Write
.image-tags.envafterup, so a failed pull never records a tag that isn't running. - Health check (backend 8137, frontend 8080), then detached city warming.
Rollback
Rollback is redeploying the previous image tag. But do not treat image
retention as a rollback guarantee: both deploy.sh and deploy-stack.sh
end by deleting images outside the running pair. On beta that succeeds, so
beta typically holds no local rollback target at all; on prod it usually
fails only because Swarm's stopped task containers still hold references, which
is why prod keeps a handful. Verify the target tag is actually present before
promising a rollback — Centralis's rollback_to(dry_run: true) checks both host
and registry.
docker system prune -a on a live box is forbidden for the same reason.
docker image prune -f (dangling only) is the safe form.
Secrets — the SOPS + age vault
The single source of truth is infra/deploy/secrets/*.yaml, committed
encrypted (values only — keys stay readable so diffs mean something) and
rendered into /etc/thermograph.env at deploy time by
deploy/render-secrets.sh.
Cycling a key is: edit → commit → deploy. No SSH, no hand-edited root file, no per-host duplication.
| File | Holds |
|---|---|
common.yaml |
the 16 values identical on prod and beta — VAPID keypair, metrics token, IndexNow key, REGISTRY_TOKEN, S3 endpoint/bucket and both S3 keypairs, plus shared non-secret config |
prod.yaml |
prod's own — the three held-back credentials, sizing, THERMOGRAPH_BASE_URL, Discord + mail credentials that exist nowhere else |
beta.yaml |
beta's own — the three held-back credentials, sizing, base URL |
dev.yaml |
the LAN box's own, self-contained, 12 values, no production credential |
centralis.prod.yaml |
Centralis's nine variables → /etc/centralis.env, its own renderer |
example.yaml |
format reference / CI fixture (fake values) |
../../.sops.yaml |
which age recipients files encrypt to (plaintext config) |
The renderer concatenates common.yaml then <env>.yaml, so a host value
wins (last occurrence of a duplicate key). <env> comes from
/etc/thermograph/secrets-env on the box.
Two design decisions worth understanding
Three credentials are deliberately not in common.yaml —
POSTGRES_PASSWORD, THERMOGRAPH_AUTH_SECRET, THERMOGRAPH_DATABASE_URL.
They hold identical values on prod and beta today, so by the mechanical rule
they belong there. They're kept per-host anyway because these are the
credentials that let one environment act as another: with them, a foothold on
beta (the more exposed box — public Forgejo and Grafana) is a foothold on prod's
database and prod's session signing. They match because beta was seeded from
prod, not because the two are meant to be one system. Keeping them per-host
costs one extra line and buys the ability to diverge.
dev renders dev.yaml alone — deploy-dev.sh exports
THERMOGRAPH_SECRETS_SKIP_COMMON=1. Eleven of common.yaml's sixteen values
are live production credentials, including a read-write object-storage
keypair on the bucket that also holds prod's database backups, and the VAPID
private key that signs push to real subscribers. The render is plaintext
concatenation, so layering it would put all of those into every container in the
dev stack — on the operator's desktop, which is also the CI runner whose
docker-labelled jobs get the host socket, running unreviewed dev code. An
override in dev.yaml wouldn't help: last-wins governs consumers, but the
production value is still physically a line in the file.
Dev works fine without them: no S3 keys means the lake service answers 503 and
history falls through to the Open-Meteo archive (an accelerator, never a
dependency); VAPID and IndexNow both self-generate and persist to the appdata
volume; with no metrics token /api/v2/metrics is direct-loopback-only, which
is right on a LAN box. One common.yaml value is simply wrong for dev —
THERMOGRAPH_COOKIE_SECURE=1 silently breaks login over plain HTTP.
Working with the vault
sops edit infra/deploy/secrets/beta.yaml # never a plain editor
infra/deploy/secrets/dry-run-render.sh # preview a render
- Never hand-edit
/etc/thermograph.env— it's a rendered artifact, overwritten on the next deploy. - Never paste a plaintext secret onto a box.
secrets-guardCI rejects a plaintext commit; the repo'ssecrets-guard.shhook blocks a direct Write/Edit to those files locally.deploy/secrets/seed-from-live.shreads production secrets and is not for an agent to run.- The
key-gapsskill audits which keys are missing or drifting across environments. It reads key names only, never values.
There's a cautionary tale in deploy/secrets/README.md worth reading: Centralis
was hand-configured until a hand-edit wrote CENTRALIS_TOKENS as bare JSON into
a file that gets sourced, bash stripped the quotes, and Centralis — correctly
failing closed on malformed JSON — dropped to a single identity. Nothing logged
an error. It looked exactly like the file had never been written. That's why
centralis.prod.yaml exists: the only durable fix was a human stops writing
the file.
Terraform
infra/terraform/ provisions and configures hosts, SSH-driven.
Never run
terraform applycasually. No tfstate is persisted anywhere, so an apply would attempt full re-provisioning of live hosts. Terraform here is executable documentation until state is bootstrapped.
Ops query tooling
Read-only, uniform across environments:
infra/ops/dbq.sh dev "select count(*) from climate_history"
infra/ops/dbq.sh prod -tA -c "select max(date) from climate_history"
infra/ops/iceberg.sh prod -c "select count(*) from era5_daily"
Both exec into the container. No database is exposed over TCP — each listens
only on its private docker network, and prod's is a Swarm overlay the prod
host itself cannot route to, so ssh -L works for beta and is impossible
for prod. Exec works identically everywhere with no ports, no tunnels, no infra
changes. The prod container name is a Swarm task name that changes on every
redeploy, so it's resolved at call time via docker ps --filter name=, never
hardcoded.
Queries connect as thermograph_ro — NOSUPERUSER, granted only
pg_read_all_data. Read-only is enforced by Postgres, not by convention:
$ infra/ops/dbq.sh prod -c "create table t(i int)"
ERROR: permission denied for schema public
Centralis's sql_query does the same thing with the same role. write: true
escalates to the app's superuser and takes effect immediately with no
confirmation — on prod, know what you are doing.
Backups
The nightly ops-cron.yml pg_dump into agent@prod:~/thermograph-backups/
is it. Known gaps, stated plainly: a single copy on the same box as the
database, no offsite yet, and restores must handle TimescaleDB's
continuous_agg circular-FK warning (--disable-triggers). The backups/
prefix in the object-storage bucket belongs to those jobs — don't write outside
the lake's own prefix.
Next: Observability.