All checks were successful
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 10s
PR build (required check) / changes (pull_request) Successful in 16s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Successful in 20s
PR build (required check) / build-backend (pull_request) Successful in 45s
PR build (required check) / gate (pull_request) Successful in 5s
Audited the five CLAUDE.md files and all twenty-one README.md files against the
tree, machine-checking every in-repo path they name and verifying the testable
claims against the live hosts.
The one that matters is in the root file: dev was documented as reachable on
the mesh at 10.10.0.2:8137. It is not, and never was from anywhere but vps1 —
infra/docker-compose.yml binds the port to 127.0.0.1, and the address answers
from neither vps2 nor vps1 itself. Anyone following it gets a connection
refused with nothing to explain it.
The rest are stale paths, several from the reunification:
* assetlinks.json moved under frontend/static/ in the subtree merge; the TWA
README kept the pre-merge path in both places it names it. Following it
would put the file where nothing serves it and Android app-link
verification would fail silently.
* push.py and notify.py now live in backend/notifications/.
* INFRA.md and deploy/stack/README have never existed in this repo, in any
branch.
* the Caddyfile is at deploy/stack/lb/Caddyfile.
* three bare relative paths that resolve for a reader but not from the
directory the file sits in: units.js is the frontend's, deploy.sh is
infra's, entrypoint.sh is the backend's.
Also records why mesh clients must pin the ROOT_URL host and not only the image
host: the registry's bearer-token realm follows ROOT_URL, so pinning
git.thermograph.org alone still sends the token request out the public route,
where the /v2/* matcher returns 403 and docker falls back to anonymous. That
surfaces as `unauthorized: reqPackageAccess`, indistinguishable from a bad
credential.
Verified true and left alone: the four-domain layout, both .claude runbooks,
the absence of any domain-level .forgejo directory, the pinned compose project
name, the deploy contract, prod's eight stack services, beta's five prefixed
ones with no db of its own, dev's five, and every documented make target.
192 lines
11 KiB
Markdown
192 lines
11 KiB
Markdown
# Thermograph — Terraform (host provisioning)
|
|
|
|
Terraform that **provisions and configures VPS hosts** and hands the app off to
|
|
`docker compose`, which pulls the published app image (built + pushed by the app
|
|
repo's CI) and runs it — no app source is ever built or checked out on a host.
|
|
By default it does *not* create servers: `modules/thermograph-host` is
|
|
SSH-provisioner-driven against a host that already exists. An optional
|
|
`modules/gcp-host` can additionally *create* the VM on GCP first (see "GCP
|
|
scaffold" below) — scaffold-only today, no live resources until you opt in.
|
|
|
|
## What it manages
|
|
|
|
One reusable module (`modules/thermograph-host`) is instantiated per
|
|
**(host, environment)** pair via `for_each` (`main.tf`'s `local.all_hosts`,
|
|
merging `var.hosts` — SSH-managed, already-existing boxes — with any
|
|
`var.gcp_hosts` Terraform created itself). `var.hosts` is keyed by HOST, not by
|
|
environment, because **vps2 alone runs two environments** (prod and beta) as
|
|
separate Docker Swarm stacks on the same box — see the `hosts` variable's file
|
|
header in `variables.tf` for why that stopped being "one entry = one
|
|
environment" and what still needs to change in `main.tf` to consume the new
|
|
shape (flagged there as a `TODO(cutover)`). This config manages **two VPS
|
|
hosts**, three environments, today:
|
|
|
|
| host | environment | VPS | branch | domain | notes |
|
|
|--------|-------------|--------------------------------------------|--------|------------------------|----------------------------------------------|
|
|
| `vps1` | `dev` | `75.119.132.91` | `dev` (of the APP repos — see `backend_image_tag` / `frontend_image_tag`) | none (mesh-only) | Also hosts Forgejo, Grafana/Loki/Alloy, emigriffith.dev — not Terraform-managed |
|
|
| `vps2` | `prod` | 48 GB / 12-core box (`169.58.46.181`) | `main` | `thermograph.org` | Caddy TLS (this module owns `/etc/caddy/Caddyfile`); sized up (8/8/4/16g) |
|
|
| `vps2` | `beta` | the SAME box as prod (`169.58.46.181`) | `main` | none (see below) | Separate Swarm stack + checkout, shares prod's TimescaleDB instance on a separate database/role |
|
|
|
|
Only ONE environment per host may set a `domain`: the module installs a full
|
|
`/etc/caddy/Caddyfile`, and a second `terraform apply` with its own `domain` set
|
|
would silently overwrite the first's site instead of adding to it (there is no
|
|
merge). Prod claims that slot on vps2 today, so beta's public reverse-proxy (if
|
|
and when `beta.thermograph.org` is exposed) has to be a site block in vps2's
|
|
shared, hand-maintained Caddy config instead — the same pattern
|
|
`deploy/forgejo/docker-stack.yml` already uses to put `git.thermograph.org` in
|
|
front of Forgejo's loopback port on vps1, alongside that box's own Caddy-fronted
|
|
`emigriffith.dev`.
|
|
|
|
Each environment's own checkout on disk (`app_dir`, `git_branch`) is **this
|
|
infra repo**, not the app repo — the "branch" column above is which app-repo
|
|
tag an environment is meant to track conceptually; the actual pinned versions
|
|
are `var.hosts[*].environments[*].backend_image_tag` / `frontend_image_tag`
|
|
(e.g. `"sha-<12 hex>"` each — the app is two separately-published images,
|
|
`admin_emi/thermograph-backend/app` and `admin_emi/thermograph-frontend/app`), since the
|
|
host has no app checkout to derive a tag from. `app_dir` is required with no
|
|
default specifically so vps2's two environments can never collide on the same
|
|
checkout path (prod: `/opt/thermograph`, beta: `/opt/thermograph-beta`). The
|
|
LAN-laptop dev rehearsal (`make dev-up`) is **out of scope here** — it's a
|
|
local-only compose overlay, not this Terraform-managed `dev` environment.
|
|
|
|
Per (host, environment), over SSH provisioners, Terraform:
|
|
|
|
- installs Docker + the compose plugin if missing;
|
|
- configures a `ufw` firewall (22/80/443 always; on a host with **no** domain it also
|
|
opens the app port `8137`);
|
|
- ensures the git checkout at `app_dir` exists (clones **this repo** on a fresh box)
|
|
and resets it to the host's `git_branch`;
|
|
- renders `/etc/thermograph-topology.env` from Terraform variables (non-secret sizing
|
|
only — `WORKERS`/`APP_CPUS`/`DB_CPUS`/`DB_MEMORY`/etc.), then runs
|
|
`deploy/render-secrets.sh` (from that same freshly-synced checkout) to render
|
|
`/etc/thermograph.env` from the SOPS+age vault — **Terraform itself never sees or
|
|
carries an app secret** (see "Secrets" below);
|
|
- for a host **with** a domain, installs a rendered Caddyfile and reloads Caddy;
|
|
- brings the stack up on the explicit per-service tags (exports `BACKEND_IMAGE_TAG` /
|
|
`FRONTEND_IMAGE_TAG` for the compose file): `docker login` to the registry,
|
|
`docker compose <-f each compose file> pull backend frontend`, then
|
|
`... up -d --remove-orphans`, running docker as root with both env files sourced;
|
|
- health-checks `http://127.0.0.1:8137/`.
|
|
|
|
A change to the rendered topology env, the compose files, the branch, the app image
|
|
tag, or the sizing flips the `null_resource` trigger and re-runs the provisioners on
|
|
next apply.
|
|
|
|
## GCP scaffold (no live resources)
|
|
|
|
`modules/gcp-host` can create a GCE VM (+ a dedicated VPC, subnet, and a firewall
|
|
opening 22/80/443) instead of assuming the host already exists — see `main.tf`'s
|
|
`module.gcp_vm`. It contributes *nothing* to provisioning: its only output
|
|
(`external_ip`) feeds into the exact same `thermograph-host` module every
|
|
SSH-managed host uses, so there's one provisioning path regardless of how a
|
|
host came to exist. `var.gcp_hosts` defaults to `{}`, so by default no
|
|
`google_*` resource is ever planned and the `google` provider is never
|
|
invoked — `terraform plan`/`validate` succeed with no GCP credentials
|
|
configured at all. To actually use it: populate an entry in
|
|
`terraform.tfvars` (see the commented example in
|
|
`terraform.tfvars.example`) and authenticate via `gcloud auth
|
|
application-default login` or `GOOGLE_APPLICATION_CREDENTIALS`. This is the
|
|
same create-then-provision composition a future Proxmox module (the
|
|
architecture doc's longer-term target — see §6/§9 there) would use.
|
|
|
|
### Container sizing is env-driven
|
|
|
|
`docker-compose.yml` reads `WORKERS`, `APP_CPUS`, `DB_CPUS`, and `DB_MEMORY` from the
|
|
environment (defaults `4 / 4 / 2 / 8g`, identical to before). Terraform sets them per
|
|
host through `/etc/thermograph-topology.env`, so the big prod box can run larger caps
|
|
without a compose edit. The Postgres *internal* memory budget (`shared_buffers`,
|
|
`effective_cache_size`, `work_mem`, `maintenance_work_mem`) is derived from the same
|
|
`DB_MEMORY` by `deploy/db/init/20-tuning.sh` — so raising `db_memory` scales the
|
|
container cap and the tuning together (prod 16g → shared_buffers 4 GB). The tuning
|
|
applies on a fresh DB volume; on an existing volume re-run it by hand (see the script
|
|
header).
|
|
|
|
## Prerequisites
|
|
|
|
- Terraform >= 1.6 (v1.15 is installed).
|
|
- SSH key access to **both** hosts as a **sudo-capable** user (the live boxes use the
|
|
dedicated `agent` account, `~/.ssh/thermograph_agent_ed25519` — see
|
|
`deploy/provision-agent-access.sh`). Point `ssh_private_key_path` at that key (`~`
|
|
is expanded).
|
|
- The hosts are Debian/Ubuntu with `apt` and outbound internet (Docker/Caddy installs
|
|
pull from the network). Docker may already be present — installs are conditional.
|
|
- For the `prod` host: DNS for `thermograph.org` must point at the new box before apply,
|
|
or Caddy's first-request cert issuance will fail.
|
|
|
|
## Use
|
|
|
|
```sh
|
|
cd terraform
|
|
cp terraform.tfvars.example terraform.tfvars # then edit: real IPs + secrets
|
|
terraform init
|
|
terraform plan
|
|
terraform apply
|
|
```
|
|
|
|
Target one host with `-target='module.host["prod"]'` if you want to apply to just one
|
|
(the module's `for_each` key changes once `main.tf` is updated to flatten `var.hosts`'s
|
|
new per-host `environments` map — see the `TODO(cutover)` on the `hosts` variable in
|
|
`variables.tf` — at which point the key for beta becomes something like
|
|
`module.host["vps2-beta"]`, not `module.host["beta"]`).
|
|
|
|
## Secrets
|
|
|
|
App secrets (`POSTGRES_PASSWORD`, `THERMOGRAPH_AUTH_SECRET`, VAPID keys,
|
|
`REGISTRY_TOKEN`, Discord/SMTP credentials, …) are **not** Terraform variables
|
|
— they live in the git-native SOPS+age vault at `../deploy/secrets/*.yaml`
|
|
(committed, encrypted) and are rendered into `/etc/thermograph.env` at deploy
|
|
time by `../deploy/render-secrets.sh`, which the provisioner's deploy step
|
|
runs from the freshly-synced checkout. Terraform only renders the non-secret
|
|
`/etc/thermograph-topology.env` (sizing/routing). See
|
|
`../deploy/secrets/README.md` to rotate a secret or add a new one — it's a
|
|
`sops edit` + commit + deploy, no Terraform apply involved.
|
|
|
|
`om_rclone_conf` (object-storage bucket credentials for the self-hosted ERA5
|
|
archive) is the one exception still supplied via Terraform, in
|
|
`terraform.tfvars` — folding it into the vault too is a reasonable future
|
|
step, not done here.
|
|
|
|
## Local state
|
|
|
|
The backend is **local**: `terraform.tfstate` is written next to the config.
|
|
It's gitignored (`terraform/.gitignore` and the root `.gitignore`) — keep it
|
|
off shared disks and back it up somewhere private. `terraform.tfvars` is
|
|
likewise gitignored (it may carry a `repo_url` credential and always carries
|
|
`om_rclone_conf`); only `terraform.tfvars.example` (dummy values) is
|
|
committed. `.terraform.lock.hcl` **is** committed so provider versions are
|
|
pinned across machines.
|
|
|
|
## WARNING — applying against live prod
|
|
|
|
`terraform apply` runs `remote-exec` **on the server**: it resets the checkout to the
|
|
branch, renders topology config + secrets, and runs `docker compose pull && up -d`
|
|
(pulling the pinned `backend_image_tag` / `frontend_image_tag` and recreating
|
|
containers — a brief app restart).
|
|
Against the live production host this is a real deploy. Review the plan, apply in a
|
|
maintenance window, and prefer `-target` to touch one host at a time.
|
|
|
|
This is **separate from the Postgres data cutover** in `deploy/POSTGRES-MIGRATION.md`.
|
|
Terraform provisions the host and starts the stack; it does **not** migrate the
|
|
SQLite→Postgres accounts data. Sequence them deliberately: for a first cutover on a
|
|
host, follow the migration doc's freeze/backup/copy steps around the point where
|
|
Terraform brings the stack up — don't let Terraform recreate containers mid-migration.
|
|
|
|
## Assumptions / notes
|
|
|
|
- **beta has no public domain here, and can't without a different Caddy strategy.**
|
|
With `compose_files = ["docker-compose.yml"]` the app binds `127.0.0.1:8137`
|
|
(loopback), so opening the port in `ufw` alone does not expose it — and setting
|
|
`domain` on beta the way prod does would make this module overwrite the SAME
|
|
`/etc/caddy/Caddyfile` prod's apply just installed (see the table above). Reach
|
|
beta via an SSH tunnel, or front it with a site block in vps2's own
|
|
hand-maintained Caddy config (outside Terraform), the same way Forgejo's
|
|
`git.thermograph.org` reaches Forgejo's loopback port on vps1.
|
|
`COOKIE_SECURE` is auto-set to `0` when there's no domain (a Secure
|
|
cookie is never sent over plain HTTP) and `1` behind Caddy TLS.
|
|
- The rendered Caddyfile only reverse-proxies the app. The repo's
|
|
`deploy/stack/lb/Caddyfile`
|
|
additionally serves the `emigriffith.dev` portfolio and legacy redirects; those are
|
|
host-specific and not templated here.
|
|
- Provisioner-based by design: the hosts already exist, so this is not a
|
|
create-from-scratch cloud config. Re-applying is idempotent (installs are guarded,
|
|
`git reset --hard`, compose `up` reconciles).
|