Some checks failed
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Failing after 6s
secrets-guard / encrypted (push) Successful in 6s
shell-lint / shellcheck (push) Successful in 13s
Validate observability stack / validate (push) Successful in 17s
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
323 lines
18 KiB
Markdown
323 lines
18 KiB
Markdown
# 8. Infra and secrets
|
|
|
|
`infra/` owns **how and where the already-built images run**. It never builds
|
|
app source. Hosts' `/opt/thermograph` is a checkout of this whole monorepo.
|
|
|
|
Read [`infra/CLAUDE.md`](../../infra/CLAUDE.md) alongside this.
|
|
|
|
## The machines
|
|
|
|
Two VPS boxes named by **role**, not by environment, plus the operator's
|
|
desktop:
|
|
|
|
| Host | Public | Mesh | Runs | Login |
|
|
|---|---|---|---|---|
|
|
| **vps1** | `75.119.132.91` / git.thermograph.org, dashboard.thermograph.org | `10.10.0.2` | Forgejo, Grafana + Loki + Alloy, `emigriffith.dev`, and **dev** (own Postgres, compose, mesh-only) | `agent`, passwordless sudo |
|
|
| **vps2** | `169.58.46.181` / thermograph.org, beta.thermograph.org | `10.10.0.1` | **prod** and **beta**, as two co-resident Docker **Swarm** stacks; Centralis; Postfix; backups | `agent`, passwordless sudo |
|
|
| **desktop** | — | `10.10.0.3` | AI-model hosting + flex Swarm capacity. No Thermograph environment; `make dev-up` is a laptop convenience there | it's your box |
|
|
|
|
```bash
|
|
ssh -i ~/.ssh/thermograph_agent_ed25519 agent@169.58.46.181 # vps2 (prod + beta)
|
|
ssh -i ~/.ssh/thermograph_agent_ed25519 agent@75.119.132.91 # vps1 (Forgejo + Grafana + dev)
|
|
```
|
|
|
|
The mesh IPs did not move when this topology landed — vps1 is the box that used
|
|
to be called "beta" (still mesh `10.10.0.2`), vps2 the one that used to be
|
|
called "prod" (still mesh `10.10.0.1`). What moved is which environment runs on
|
|
which box: beta moved from vps1 onto vps2 to sit next to prod, and dev moved
|
|
onto vps1.
|
|
|
|
**vps1 is guarded as strictly as vps2**, not more loosely because it "used to
|
|
be beta" or because dev now lives there. It hosts Forgejo, its CI runner and
|
|
the registry, plus Grafana/Loki — a destructive command there takes out git, CI
|
|
and the registry at once. Dev being mesh-only and a public VPS's tenant, rather
|
|
than someone's desktop, is exactly why it gets the same guard as everything
|
|
else on that box: it shares a kernel with the estate's source of truth for code
|
|
and the monitoring stack, sudo there is passwordless, and it runs whatever
|
|
branch is currently in flight.
|
|
|
|
Every root-effective command on both VPSes is logged by `auditd`
|
|
(`ausearch -k agentcmd`). That's a feature — fixes are attributable.
|
|
|
|
Prefer Centralis for routine work (`run_on_host`, `fleet_status`,
|
|
`deployed_version`) over raw SSH.
|
|
|
|
## The two orchestrators
|
|
|
|
Which path an environment takes is decided by **`infra/deploy/env-topology.sh`**
|
|
(`thermograph_topology <env>`, keyed on `THERMOGRAPH_ENV=dev|beta|prod`): it
|
|
derives the checkout, branch, deploy mode, stack/compose name, env file, LB
|
|
ports and DB role for that environment, and `deploy.sh` execs
|
|
`deploy/stack/deploy-stack.sh` when the mode is `stack`, or rolls compose
|
|
otherwise. The old host-wide `/etc/thermograph/deploy-mode` marker still exists
|
|
as a **fallback** for a by-hand run on a box that predates this file, but it
|
|
cannot describe vps2 on its own — vps2 runs two environments, so "which mode
|
|
does this host use" is no longer a well-formed question there.
|
|
|
|
| | compose (dev) | Swarm (prod, beta) |
|
|
|---|---|---|
|
|
| Host | vps1 | vps2, both stacks co-resident |
|
|
| File | `infra/docker-compose.yml` | `infra/deploy/stack/thermograph-stack.yml` (prod), `thermograph-beta-stack.yml` (beta) |
|
|
| Services | `db`, `backend`, `lake`, `daemon`, `frontend` | prod: `db`, `web`, `worker`, `lake`, `daemon`, `frontend`, `autoscaler`, `autoscaler-lake`. beta: `beta-web`, `beta-worker`, `beta-lake`, `beta-daemon`, `beta-frontend` — prefixed, no `db`, no autoscalers |
|
|
| Rolling | `up -d --no-deps <targets>` | start-first, health-gated, auto-rollback |
|
|
| Tag file | `deploy/.image-tags.env` | `deploy/.stack-image-tags.env` |
|
|
|
|
Note prod splits `backend` into **`web`** and **`worker`** (the
|
|
`THERMOGRAPH_ROLE` split from [04](04-backend.md)); dev's compose stack runs one
|
|
`backend` service in role `all`. Beta's Swarm services are prefixed
|
|
(`beta-web`, not `web`) because Swarm registers a service's short name as a DNS
|
|
alias on every network it joins — two stacks both naming a service `web` on
|
|
the shared `data` network would make `web` ambiguous, and beta joins that
|
|
network (declared `external` in its stack file) purely to reach prod's `db`.
|
|
Beta keeps its own `internal` overlay for beta-to-beta traffic and has no `db`
|
|
service of its own: **one TimescaleDB instance serves both environments**, on
|
|
separate databases (`thermograph` / `thermograph_beta`) and separate roles
|
|
(beta's is `NOSUPERUSER`/`NOCREATEDB`/`NOCREATEROLE`, `CONNECT` revoked from
|
|
`PUBLIC` on prod's database).
|
|
|
|
`STACK_TEST=1` rehearses the entire Swarm deploy under stack name
|
|
`thermograph-test` on throwaway volumes and ports `18137`/`18080`. Leftover
|
|
`thermograph-test_*` containers in logs are rehearsal residue — ignore them.
|
|
|
|
### The volume/project-name coupling
|
|
|
|
Compose creates `thermograph_pgdata` / `_appdata` / `_applogs`; the Swarm stack
|
|
declares those exact names as `external: true` at the same mount paths. That
|
|
coupling is why `name: thermograph` is pinned in the compose file — running
|
|
compose from `infra/` without it derives project `infra`, silently creating a
|
|
whole new stack with empty volumes beside the running one. `deploy-dev.sh`
|
|
exports `COMPOSE_PROJECT_NAME=thermograph-dev` (env wins over the file key) to
|
|
keep dev separate on purpose. **Keep both halves.**
|
|
|
|
## `deploy/deploy.sh` — read this before you touch a deploy
|
|
|
|
Single entry point for dev, beta and prod — the same script for all three;
|
|
`infra/deploy/env-topology.sh` is what tells it which checkout, branch, stack
|
|
and ports apply. Contract, run from that environment's own checkout
|
|
(`/opt/thermograph` for prod, `/opt/thermograph-beta` for beta, both on vps2;
|
|
`/opt/thermograph-dev` on vps1 via the `deploy-dev.sh` wrapper):
|
|
|
|
```bash
|
|
SERVICE=backend BACKEND_IMAGE_TAG=sha-<12hex> /opt/thermograph/infra/deploy/deploy.sh
|
|
SERVICE=frontend FRONTEND_IMAGE_TAG=sha-<12hex> /opt/thermograph/infra/deploy/deploy.sh
|
|
SERVICE=all BACKEND_IMAGE_TAG=sha-<a> FRONTEND_IMAGE_TAG=sha-<b> …/deploy.sh
|
|
```
|
|
|
|
What it does, in order, and why each step is the way it is:
|
|
|
|
1. **`flock` on `deploy/.deploy.lock`**, with a self re-exec so the lock spans
|
|
the whole run. Backend and frontend deploys can fire for the same push
|
|
seconds apart and both SSH into one checkout; concurrent runs race the git
|
|
reset, the compose project and the tag file. `-w 600` bounds the wait.
|
|
2. **Render secrets** from the SOPS vault into that environment's own env file
|
|
(`/etc/thermograph.env` for prod, `/etc/thermograph-beta.env` for beta,
|
|
likewise `/etc/thermograph.env` on dev's separate host), then source it, so
|
|
a by-hand run interpolates the same as CI's does. Guarded on the helper's
|
|
existence so the very deploy that *introduces* `render-secrets.sh` is safe.
|
|
3. **`git reset --hard origin/$BRANCH`** on the checkout root (`BRANCH` defaults
|
|
to `main` for beta/prod, `dev` for dev). ⚠️ **Uncommitted edits in that
|
|
checkout evaporate here** — on vps2 that's `/opt/thermograph` for a prod
|
|
deploy or `/opt/thermograph-beta` for a beta one; each reset only ever
|
|
touches its own checkout, never the sibling's. The one exception is
|
|
`deploy/.image-tags.env`, untracked on purpose.
|
|
4. **Read `.image-tags.env`** so a single-service roll re-renders compose with
|
|
*both* services' real tags and never accidentally recreates or downgrades the
|
|
sibling.
|
|
5. **Registry login** only if `REGISTRY_TOKEN` is set. The CI paths don't pass
|
|
one — the host is already `docker login`ed — and an unconditional login with
|
|
an empty token would abort under `set -e`.
|
|
6. **Pull, with a bounded ~5-minute retry.** `build-push.yml` is a *separate*
|
|
workflow triggered by the same push, and Forgejo's `needs:` only orders jobs
|
|
within one file. A deploy racing ahead of its build and failing "not found"
|
|
is confirmed-live behaviour; the retry covers a normal build and still fails
|
|
loudly on a genuine problem.
|
|
7. **Daemon capability probe.** Infra tracks `main` while image tags are
|
|
env-staged, so a host can legitimately be asked to roll a backend image older
|
|
than the compose file — one built before the daemon binary existed. The
|
|
script `docker run`s the image to test for `/usr/local/bin/thermograph-daemon`
|
|
and drops `daemon` from this run if it's absent, rather than leaving a
|
|
container crash-looping on a missing binary.
|
|
8. **Roll.** `TARGETS` are `backend → (backend, lake, daemon)`,
|
|
`frontend → (frontend)`, `all → everything`. **`daemon` and `lake` ride with
|
|
`backend` and are never targets on their own** — all three run the same image
|
|
at the same tag, and rolling one without the others is exactly the version
|
|
skew the `/internal/*` contract has no negotiation for. Single-service rolls
|
|
use `--no-deps`; a full `all` uses `--remove-orphans` (a renamed-away
|
|
service's old container would otherwise squat its port — confirmed live).
|
|
9. **Write `.image-tags.env`** *after* `up`, so a failed pull never records a
|
|
tag that isn't running.
|
|
10. **Health check** (backend 8137, frontend 8080), then detached city warming.
|
|
|
|
### Rollback
|
|
|
|
Rollback is redeploying the previous image tag. But **do not treat image
|
|
retention as a rollback guarantee**: both `deploy.sh` and `deploy-stack.sh`
|
|
*end* by deleting images outside the running pair. Beta now runs Swarm exactly
|
|
like prod (it used to be compose, on the old topology, where this pruning
|
|
reliably left no rollback target at all); on both Swarm stacks it usually
|
|
fails only because Swarm's stopped task containers still hold references,
|
|
which is why prod and beta both tend to keep a handful. Dev, the one remaining
|
|
compose environment, is the one where this prune still reliably succeeds and
|
|
leaves nothing to roll back to. Verify the target tag is actually present
|
|
before promising a rollback — Centralis's `rollback_to(dry_run: true)` checks
|
|
both host and registry.
|
|
|
|
`docker system prune -a` on a live box is forbidden for the same reason.
|
|
`docker image prune -f` (dangling only) is the safe form.
|
|
|
|
## Secrets — the SOPS + age vault
|
|
|
|
**The single source of truth is `infra/deploy/secrets/*.yaml`**, committed
|
|
encrypted (values only — keys stay readable so diffs mean something) and
|
|
rendered at deploy time by `deploy/render-secrets.sh` into each environment's
|
|
own env file (`/etc/thermograph.env` for prod, `/etc/thermograph-beta.env` for
|
|
beta, `/etc/thermograph.env` again on dev's separate host).
|
|
|
|
**Cycling a key is: edit → commit → deploy.** No SSH, no hand-edited root file,
|
|
no per-host duplication.
|
|
|
|
| File | Holds |
|
|
|---|---|
|
|
| `common.yaml` | the 16 values identical on prod **and** beta — VAPID keypair, metrics token, IndexNow key, `REGISTRY_TOKEN`, S3 endpoint/bucket and both S3 keypairs, plus shared non-secret config. **Never layered under `dev`** — see below |
|
|
| `prod.yaml` | prod's own — the three held-back credentials, sizing, `THERMOGRAPH_BASE_URL`, Discord + mail credentials that exist nowhere else |
|
|
| `beta.yaml` | beta's own — the three held-back credentials, sizing, base URL |
|
|
| `dev.yaml` | dev's own, **self-contained**, 12 values, no production credential |
|
|
| `centralis.prod.yaml` | Centralis's nine variables → `/etc/centralis.env`, its own renderer |
|
|
| `example.yaml` | format reference / CI fixture (fake values) |
|
|
| `../../.sops.yaml` | which age recipients files encrypt to (plaintext config) |
|
|
|
|
The renderer concatenates `common.yaml` then `<env>.yaml`, so a **host value
|
|
wins** (last occurrence of a duplicate key). `<env>` is passed explicitly as
|
|
`THERMOGRAPH_ENV` by the deploy caller — required now that vps2 alone hosts two
|
|
environments and a host-wide marker can't tell them apart; the marker
|
|
(`/etc/thermograph/secrets-env`) survives only as the fallback for a by-hand
|
|
run on a single-environment box.
|
|
|
|
### Two design decisions worth understanding — re-derived for vps1/vps2
|
|
|
|
**Three credentials are deliberately *not* in `common.yaml`** —
|
|
`POSTGRES_PASSWORD`, `THERMOGRAPH_AUTH_SECRET`, `THERMOGRAPH_DATABASE_URL`.
|
|
They hold identical *values* on prod and beta today (beta was seeded from
|
|
prod), so by the mechanical rule they'd belong in `common.yaml`. They're kept
|
|
per-host anyway, and the reason used to be phrased as a host-isolation
|
|
argument — "a foothold on beta is a foothold on prod's database, and beta is
|
|
the more exposed box." **That framing no longer holds**: beta and prod are now
|
|
two Swarm stacks co-resident on vps2, so a foothold on the host reaches both
|
|
regardless of which file a credential lives in. What these three credentials
|
|
still buy, correctly stated, is a **database-level** and **file-level**
|
|
boundary, not a host-level one: beta's `THERMOGRAPH_DATABASE_URL` names its own
|
|
`NOSUPERUSER` role (`thermograph_beta`) against its own database on the shared
|
|
TimescaleDB instance, with `CONNECT` revoked from `PUBLIC` on prod's database,
|
|
and each environment's `POSTGRES_PASSWORD`/`THERMOGRAPH_AUTH_SECRET` live in a
|
|
separate file so rotating prod's stops implying "and beta's too." Keeping them
|
|
per-host costs one extra line and buys the ability to diverge — that part is
|
|
unchanged. What changed is what "diverge" is defending against: not a beta
|
|
compromise reaching a separate prod host, but a beta compromise reaching
|
|
prod's role and session-signing key on the host they now already share.
|
|
|
|
**The credential boundary that actually maps to hosts now is vps1-vs-vps2, not
|
|
beta-vs-prod** — which is exactly why `dev` renders `dev.yaml` **alone**,
|
|
never layering `common.yaml` (the fleet's shared production credentials):
|
|
`deploy-dev.sh` exports `THERMOGRAPH_SECRETS_SKIP_COMMON=1`. Eleven of
|
|
`common.yaml`'s sixteen values are live production credentials, including a
|
|
**read-write** object-storage keypair on the bucket that also holds prod's
|
|
database backups, and the VAPID private key that signs push to real
|
|
subscribers. The render is plaintext concatenation, so layering it would put
|
|
all of those into every container in the dev stack — on vps1, the box that
|
|
also runs Forgejo and its CI runner, whose `docker`-labelled jobs get the host
|
|
socket, running whatever branch is currently `dev` (unreviewed by definition:
|
|
that's the branch PRs land on before the gate promotes them to `main`). An
|
|
override in `dev.yaml` wouldn't help: last-wins governs *consumers*, but the
|
|
production value is still physically a line in the file. This is the load-
|
|
bearing boundary in the vault today, not a per-environment credential split
|
|
that co-residency has already dissolved at the host level.
|
|
|
|
Dev works fine without `common.yaml`'s values: no S3 keys means the `lake`
|
|
service answers 503 and history falls through to the Open-Meteo archive (an
|
|
accelerator, never a dependency); VAPID and IndexNow both self-generate and
|
|
persist to the `appdata` volume; with no metrics token `/api/v2/metrics` is
|
|
direct-loopback-only, which is right for a mesh-only box. One `common.yaml`
|
|
value is simply *wrong* for dev — `THERMOGRAPH_COOKIE_SECURE=1` silently breaks
|
|
login over dev's plain-HTTP mesh URL.
|
|
|
|
None of this makes dev unimportant or unguarded. It's a public VPS's tenant
|
|
now, sharing a box with Forgejo and the monitoring stack — not somebody's
|
|
desktop — which is precisely why it gets its own exclusion from `common.yaml`
|
|
rather than being treated as beneath the vault's concern.
|
|
|
|
### Working with the vault
|
|
|
|
```bash
|
|
sops edit infra/deploy/secrets/beta.yaml # never a plain editor
|
|
infra/deploy/secrets/dry-run-render.sh # preview a render
|
|
```
|
|
|
|
- **Never hand-edit `/etc/thermograph.env`** — it's a rendered artifact,
|
|
overwritten on the next deploy.
|
|
- **Never paste a plaintext secret onto a box.**
|
|
- `secrets-guard` CI rejects a plaintext commit; the repo's `secrets-guard.sh`
|
|
hook blocks a direct Write/Edit to those files locally.
|
|
- `deploy/secrets/seed-from-live.sh` reads production secrets and is **not**
|
|
for an agent to run.
|
|
- The **`key-gaps`** skill audits which keys are missing or drifting across
|
|
environments. It reads key *names* only, never values.
|
|
|
|
There's a cautionary tale in `deploy/secrets/README.md` worth reading: Centralis
|
|
was hand-configured until a hand-edit wrote `CENTRALIS_TOKENS` as bare JSON into
|
|
a file that gets `source`d, bash stripped the quotes, and Centralis — correctly
|
|
failing closed on malformed JSON — dropped to a single identity. Nothing logged
|
|
an error. It looked exactly like the file had never been written. That's why
|
|
`centralis.prod.yaml` exists: the only durable fix was *a human stops writing
|
|
the file*.
|
|
|
|
## Terraform
|
|
|
|
`infra/terraform/` provisions and configures hosts, SSH-driven.
|
|
|
|
> **Never run `terraform apply` casually.** No tfstate is persisted anywhere, so
|
|
> an apply would attempt full re-provisioning of live hosts. Terraform here is
|
|
> **executable documentation** until state is bootstrapped.
|
|
|
|
## Ops query tooling
|
|
|
|
Read-only, uniform across environments:
|
|
|
|
```bash
|
|
infra/ops/dbq.sh dev "select count(*) from climate_history"
|
|
infra/ops/dbq.sh prod -tA -c "select max(date) from climate_history"
|
|
infra/ops/iceberg.sh prod -c "select count(*) from era5_daily"
|
|
```
|
|
|
|
Both exec **into the container**. No database is exposed over TCP — each listens
|
|
only on its private docker network, and prod's and beta's are both Swarm
|
|
**overlays the vps2 host itself cannot route to**, so `ssh -L` is *impossible*
|
|
for either. (Dev's, on vps1, is plain compose, so `ssh -L` would actually work
|
|
there — but exec is used uniformly across environments anyway, so it doesn't
|
|
matter which one happens to allow the shortcut.) Exec works identically
|
|
everywhere with no ports, no tunnels, no infra changes. The prod and beta
|
|
container names are Swarm task names that change on every redeploy, so they're
|
|
resolved at call time via `docker ps --filter name=`, never hardcoded.
|
|
|
|
Queries connect as **`thermograph_ro`** — `NOSUPERUSER`, granted only
|
|
`pg_read_all_data`. Read-only is enforced by Postgres, not by convention:
|
|
|
|
```
|
|
$ infra/ops/dbq.sh prod -c "create table t(i int)"
|
|
ERROR: permission denied for schema public
|
|
```
|
|
|
|
Centralis's `sql_query` does the same thing with the same role. `write: true`
|
|
escalates to the app's superuser and takes effect immediately with no
|
|
confirmation — on prod, know what you are doing.
|
|
|
|
## Backups
|
|
|
|
The nightly `ops-cron.yml` `pg_dump` into `agent@vps2:~/thermograph-backups/`
|
|
(prod's data specifically — see [CI and release](07-ci-and-release.md)) is it.
|
|
Known gaps, stated plainly: **a single copy on the same box as the
|
|
database, no offsite yet**, and restores must handle TimescaleDB's
|
|
`continuous_agg` circular-FK warning (`--disable-triggers`). The `backups/`
|
|
prefix in the object-storage bucket belongs to those jobs — don't write outside
|
|
the lake's own prefix.
|
|
|
|
Next: [Observability](09-observability.md).
|