# infra/ — agent instructions How and where the already-built app images run: Terraform, the SOPS+age secrets vault, compose and Swarm stack files, deploy scripts, host provisioning, mail, and the ops cron (DB backup + IndexNow). Read the root `CLAUDE.md` first — it owns the branch model, the deploy contract, and which environment runs which orchestrator. ## The machines Two VPS boxes plus the operator's desktop, on one WireGuard mesh, named by ROLE rather than by environment: - **vps1** — `75.119.132.91`, mesh `10.10.0.2`. "Operational programs": Forgejo (git + CI + registry, `git.thermograph.org`, also served at `dev.jinemi.com` — one Caddy site block, the `.org` name still canonical; migration in progress, see `deploy/forgejo/README.md`), Grafana/Loki/Alloy (`dashboard.thermograph.org`), the `emigriffith.dev` portfolio site, and the **dev** environment (`/opt/thermograph-dev`, tracks `dev`), with its own Postgres container. Dev is mesh-only here — published on `10.10.0.2:8137`, no public DNS, no Caddy site, no TLS. - **vps2** — `169.58.46.181`, mesh `10.10.0.1`. "The deployed environment" — everything an external user can reach: **prod** (`/opt/thermograph`, tracks `main`, Swarm stack `thermograph`) AND **beta** (`/opt/thermograph-beta`, tracks `main`, Swarm stack `thermograph-beta`), Centralis, Postfix, and the backups. One shared TimescaleDB instance serves both, on separate databases/roles (`thermograph` / `thermograph_beta`). - **desktop** — mesh `10.10.0.3`. Hosts no Thermograph environment at all: AI model hosting plus flex Swarm-worker capacity. `make dev-up` still works there as a laptop-local rehearsal, but that is not an environment. `agent` user, passwordless sudo, on vps1 and vps2. A host's `/opt/thermograph*` is a checkout of **this monorepo**, not an infra-only repo — note vps2 carries **two** such checkouts side by side (`/opt/thermograph` and `/opt/thermograph-beta`), which is the one thing most deploy logic here has to get right that a single-environment host never had to. ## Deploy paths - **`deploy/env-topology.sh`** — sourced by every path below. `THERMOGRAPH_ENV` (`dev`/`beta`/`prod`) is the input; host, checkout, branch, deploy mode, stack/compose name, env file, LB ports, DB role/database and service-name prefix are all derived from it. This exists because vps2 alone runs two environments — a host-wide marker can no longer answer "which environment is this", so the caller (the deploy workflow) says so explicitly. - **`deploy/deploy.sh`** — the single entry point for dev, beta and prod. Resets the host checkout to `BRANCH`, renders secrets, then routes: beta and prod (`TG_DEPLOY_MODE=stack`) exec `deploy/stack/deploy-stack.sh`; dev (`compose`) rolls compose services. The old host-wide `/etc/thermograph/deploy-mode` marker is honoured only when `env-topology.sh` can't resolve an environment at all (a checkout that predates the split, or a by-hand run with nothing set). Per-service tags persist in untracked `deploy/.image-tags.env` (compose) and `deploy/.stack-image-tags.env` (stack) so the two never mix. - **`deploy/stack/deploy-stack.sh`** — the Swarm path, live on vps2 for both prod and beta, picking the stack file/ports/DB role/service prefix out of `env-topology.sh`. `backend` rolls web **and** worker; `frontend` rolls frontend; `all` runs a full `docker stack deploy`. Start-first, health-gated, auto-rollback. `STACK_TEST=1` rehearses under stack name `thermograph-test` on throwaway volumes and ports `18137`/`18080`. - **`deploy/deploy-dev.sh`** — thin wrapper around `deploy.sh` for dev on vps1 (dev compose overlay, `/opt/thermograph-dev`, `THERMOGRAPH_SECRETS_SKIP_COMMON=1`). Deployed like beta/prod now — a push to `dev` triggers the same `Deploy` workflow over SSH, no LAN-specific runner involved. - **`Makefile`** — compose orchestration only: `up`, `down`, `db-up`, `dev-up`, `om-up`, `om-backfill`. Volumes are the reason the compose project name is pinned: compose creates `thermograph_pgdata`/`_appdata`/`_applogs`, and the Swarm stack declares those same names as `external: true` at the same mount paths. Beta's stack has no `db` of its own — it reaches prod's `db` service over prod's `thermograph_internal` overlay (declared `external: true` in beta's stack file) and keeps a separate `internal` overlay for its own beta-to-beta traffic. ## Rules - **Never run `terraform apply` casually.** No tfstate is persisted anywhere, so an apply would attempt full re-provisioning of live hosts. Terraform here is executable documentation until state is bootstrapped. - **Secrets: SOPS vault only** (`deploy/secrets/*.yaml`, `sops edit` → commit → deploy). Never hand-edit `/etc/thermograph.env` / `/etc/thermograph-beta.env` — they are rendered artifacts. `secrets-guard` CI rejects plaintext. `deploy/secrets/seed-from-live.sh` reads production secrets and is **not** for an agent to run. - **vps1 must never hold prod credentials.** It's the box that runs Forgejo, its CI runner, and dev's unreviewed branch — the opposite of an isolation boundary. This is why dev renders `dev.yaml` **alone** and never layers `common.yaml` (the fleet's shared production credential set). See `deploy/secrets/README.md`. - **A foothold on vps2 is a foothold on both beta and prod's host** — they are co-resident by design now. What still separates them is the database (separate roles/databases, `CONNECT` revoked from `PUBLIC`) and the filesystem (separate checkouts, separate rendered env files). Don't describe beta and prod as isolated at the host or SSH-credential level; they aren't anymore. - **The ops cron (`.forgejo/workflows/ops-cron.yml`, at the repo root) is THE backup for both application databases and for Forgejo.** Secrets are keyed by host, not environment: `VPS2_SSH_*` for the `backup`/`indexnow` jobs (prod and beta's databases, both on vps2) and `VPS1_SSH_*` for `forgejo-backup` (Forgejo now lives on vps1, not co-located with beta). If you touch it, verify a dump actually lands for **both** the prod and beta databases, not just one — an earlier revision backed up only prod and silently left beta uncovered. - Shell here runs as root over SSH against live hosts with no test suite in front of it. `shell-lint` CI (pinned shellcheck) is the only guard — keep the tree at zero findings. ## Commits & PRs Concise and technical. Never mention AI, assistants or automated authorship.