thermograph/infra/CLAUDE.md

109 lines
6.3 KiB
Markdown
Raw Normal View History

# infra/ — agent instructions
How and where the already-built app images run: Terraform, the SOPS+age secrets
vault, compose and Swarm stack files, deploy scripts, host provisioning, mail,
and the ops cron (DB backup + IndexNow). Read the root `CLAUDE.md` first — it
owns the branch model, the deploy contract, and which environment runs which
orchestrator.
## The machines
Two VPS boxes plus the operator's desktop, on one WireGuard mesh, named by
ROLE rather than by environment:
- **vps1** — `75.119.132.91`, mesh `10.10.0.2`. "Operational programs": Forgejo
(git + CI + registry, `git.thermograph.org`), Grafana/Loki/Alloy
(`dashboard.thermograph.org`), the `emigriffith.dev` portfolio site, and the
**dev** environment (`/opt/thermograph-dev`, tracks `dev`), with its own
Postgres container. Dev is mesh-only here — published on `10.10.0.2:8137`,
no public DNS, no Caddy site, no TLS.
- **vps2** — `169.58.46.181`, mesh `10.10.0.1`. "The deployed environment" —
everything an external user can reach: **prod** (`/opt/thermograph`, tracks
`main`, Swarm stack `thermograph`) AND **beta** (`/opt/thermograph-beta`,
tracks `main`, Swarm stack `thermograph-beta`), Centralis, Postfix, and the
backups. One shared TimescaleDB instance serves both, on separate
databases/roles (`thermograph` / `thermograph_beta`).
- **desktop** — mesh `10.10.0.3`. Hosts no Thermograph environment at all: AI
model hosting plus flex Swarm-worker capacity. `make dev-up` still works
there as a laptop-local rehearsal, but that is not an environment.
`agent` user, passwordless sudo, on vps1 and vps2. A host's `/opt/thermograph*`
is a checkout of **this monorepo**, not an infra-only repo — note vps2 carries
**two** such checkouts side by side (`/opt/thermograph` and
`/opt/thermograph-beta`), which is the one thing most deploy logic here has to
get right that a single-environment host never had to.
## Deploy paths
- **`deploy/env-topology.sh`** — sourced by every path below. `THERMOGRAPH_ENV`
(`dev`/`beta`/`prod`) is the input; host, checkout, branch, deploy mode,
stack/compose name, env file, LB ports, DB role/database and service-name
prefix are all derived from it. This exists because vps2 alone runs two
environments — a host-wide marker can no longer answer "which environment is
this", so the caller (the deploy workflow) says so explicitly.
- **`deploy/deploy.sh`** — the single entry point for dev, beta and prod.
Resets the host checkout to `BRANCH`, renders secrets, then routes: beta and
prod (`TG_DEPLOY_MODE=stack`) exec `deploy/stack/deploy-stack.sh`; dev
(`compose`) rolls compose services. The old host-wide
`/etc/thermograph/deploy-mode` marker is honoured only when
`env-topology.sh` can't resolve an environment at all (a checkout that
predates the split, or a by-hand run with nothing set). Per-service tags
persist in untracked `deploy/.image-tags.env` (compose) and
`deploy/.stack-image-tags.env` (stack) so the two never mix.
- **`deploy/stack/deploy-stack.sh`** — the Swarm path, live on vps2 for both
prod and beta, picking the stack file/ports/DB role/service prefix out of
`env-topology.sh`. `backend` rolls web **and** worker; `frontend` rolls
frontend; `all` runs a full `docker stack deploy`. Start-first, health-gated,
auto-rollback. `STACK_TEST=1` rehearses under stack name `thermograph-test`
on throwaway volumes and ports `18137`/`18080`.
- **`deploy/deploy-dev.sh`** — thin wrapper around `deploy.sh` for dev on vps1
(dev compose overlay, `/opt/thermograph-dev`, `THERMOGRAPH_SECRETS_SKIP_COMMON=1`).
Deployed like beta/prod now — a push to `dev` triggers the same `Deploy`
workflow over SSH, no LAN-specific runner involved.
- **`Makefile`** — compose orchestration only: `up`, `down`, `db-up`, `dev-up`,
`om-up`, `om-backfill`.
Volumes are the reason the compose project name is pinned: compose creates
`thermograph_pgdata`/`_appdata`/`_applogs`, and the Swarm stack declares those
same names as `external: true` at the same mount paths. Beta's stack has no
`db` of its own — it reaches prod's `db` service over prod's `thermograph_internal`
overlay (declared `external: true` in beta's stack file) and keeps a separate
`internal` overlay for its own beta-to-beta traffic.
## Rules
- **Never run `terraform apply` casually.** No tfstate is persisted anywhere, so
an apply would attempt full re-provisioning of live hosts. Terraform here is
executable documentation until state is bootstrapped.
- **Secrets: SOPS vault only** (`deploy/secrets/*.yaml`, `sops edit` → commit →
deploy). Never hand-edit `/etc/thermograph.env` / `/etc/thermograph-beta.env`
— they are rendered artifacts. `secrets-guard` CI rejects plaintext.
`deploy/secrets/seed-from-live.sh` reads production secrets and is **not**
for an agent to run.
- **vps1 must never hold prod credentials.** It's the box that runs Forgejo,
its CI runner, and dev's unreviewed branch — the opposite of an isolation
boundary. This is why dev renders `dev.yaml` **alone** and never layers
`common.yaml` (the fleet's shared production credential set). See
`deploy/secrets/README.md`.
- **A foothold on vps2 is a foothold on both beta and prod's host** — they are
co-resident by design now. What still separates them is the database
(separate roles/databases, `CONNECT` revoked from `PUBLIC`) and the
filesystem (separate checkouts, separate rendered env files). Don't describe
beta and prod as isolated at the host or SSH-credential level; they aren't
anymore.
- **The ops cron (`.forgejo/workflows/ops-cron.yml`, at the repo root) is THE
backup for both application databases and for Forgejo.** Secrets are keyed
by host, not environment: `VPS2_SSH_*` for the `backup`/`indexnow` jobs
(prod and beta's databases, both on vps2) and `VPS1_SSH_*` for
`forgejo-backup` (Forgejo now lives on vps1, not co-located with beta). If
you touch it, verify a dump actually lands for **both** the prod and beta
databases, not just one — an earlier revision backed up only prod and
silently left beta uncovered.
- Shell here runs as root over SSH against live hosts with no test suite in
front of it. `shell-lint` CI (pinned shellcheck) is the only guard — keep the
tree at zero findings.
## Commits & PRs
Concise and technical. Never mention AI, assistants or automated authorship.