thermograph/terraform
Emi Griffith dc4ee9a8db Scale DB tuning from DB_MEMORY; order Docker after the rclone mount (#226)
Two prod-readiness hardening changes:

DB tuning scales with the container budget. Replace the fixed 8 GB
20-tuning.sql with 20-tuning.sh, which derives shared_buffers (25%),
effective_cache_size (75%), work_mem, maintenance_work_mem and
duckdb.max_memory (50%) from the DB_MEMORY the compose db service now passes in.
The ratios reproduce the historical 8 GB tuning exactly and scale linearly, so
the 48 GB prod box (db_memory 16g) gets shared_buffers 4 GB / duckdb 8 GB with
no separate edit. Beta/local (8g default) are unchanged. Docs that told
operators to raise the tuning by hand are updated.

Boot ordering for the self-hosted archive. On an openmeteo host, install a
docker.service drop-in (Wants/After rclone-om.service) so Docker starts after
the object-storage mount is ready on every boot — the restart-policy containers
never bind an empty mount point. rclone-om is Type=notify, so After waits for
the mount to actually be ready. Cleaned up when openmeteo is toggled off.
2026-07-20 14:33:09 +00:00
..
modules/thermograph-host Scale DB tuning from DB_MEMORY; order Docker after the rclone mount (#226) 2026-07-20 14:33:09 +00:00
.gitignore Add Terraform to provision the VPS hosts (compose keeps running the app) (#223) 2026-07-20 07:42:15 +00:00
.terraform.lock.hcl Add Terraform to provision the VPS hosts (compose keeps running the app) (#223) 2026-07-20 07:42:15 +00:00
main.tf Self-host the ERA5 archive via Open-Meteo (object storage) (#224) 2026-07-20 13:16:56 +00:00
outputs.tf Add Terraform to provision the VPS hosts (compose keeps running the app) (#223) 2026-07-20 07:42:15 +00:00
README.md Scale DB tuning from DB_MEMORY; order Docker after the rclone mount (#226) 2026-07-20 14:33:09 +00:00
terraform.tfvars.example Scale DB tuning from DB_MEMORY; order Docker after the rclone mount (#226) 2026-07-20 14:33:09 +00:00
variables.tf Scale DB tuning from DB_MEMORY; order Docker after the rclone mount (#226) 2026-07-20 14:33:09 +00:00
versions.tf Add Terraform to provision the VPS hosts (compose keeps running the app) (#223) 2026-07-20 07:42:15 +00:00

Thermograph — Terraform (host provisioning)

Terraform that provisions and configures the existing VPS hosts and hands the app off to docker compose. It does not create servers (no cloud provider) and does not replace compose — it prepares each host (Docker, firewall, checkout, secrets, Caddy) and runs docker compose up.

What it manages

One reusable module (modules/thermograph-host) is instantiated per host via for_each. This config manages two VPS hosts:

key role VPS branch domain notes
prod prod NEW 48 GB / 12-core box release thermograph.org Caddy TLS; sized up (8/4/16g)
beta beta old box 75.119.132.91 main (none) testing/beta; no public domain yet

The dev branch is out of scope here — it deploys to the LAN dev server via deploy/deploy-dev.sh (a self-hosted GitHub Actions runner), not Terraform.

Per host, over SSH provisioners, Terraform:

  • installs Docker + the compose plugin if missing;
  • configures a ufw firewall (22/80/443 always; on a host with no domain it also opens the app port 8137);
  • ensures the git checkout at app_dir exists (clones on a fresh box) and resets it to the host's branch;
  • renders /etc/thermograph.env from Terraform variables (secrets injected from tfvars, pushed via provisioner content so they never touch local disk) and installs it root-owned 0640;
  • for a host with a domain, installs a rendered Caddyfile and reloads Caddy;
  • brings the stack up: docker compose <-f each compose file> up -d --build, running docker as root with /etc/thermograph.env sourced in the same shell;
  • health-checks http://127.0.0.1:8137/.

A change to the rendered env, the compose files, the branch, or the sizing flips the null_resource trigger and re-runs the provisioners on next apply.

Container sizing is env-driven

docker-compose.yml reads WORKERS, APP_CPUS, DB_CPUS, and DB_MEMORY from the environment (defaults 4 / 4 / 2 / 8g, identical to before). Terraform sets them per host through /etc/thermograph.env, so the big prod box can run larger caps without a compose edit. The Postgres internal memory budget (shared_buffers, effective_cache_size, work_mem, duckdb.max_memory) is derived from the same DB_MEMORY by deploy/db/init/20-tuning.sh — so raising db_memory scales the container cap and the tuning together (prod 16g → shared_buffers 4 GB, duckdb 8 GB). The tuning applies on a fresh DB volume; on an existing volume re-run it by hand (see the script header).

Prerequisites

  • Terraform >= 1.6 (v1.15 is installed).
  • SSH key access to both hosts as a sudo-capable user (default deploy). Point ssh_private_key_path at that key (~ is expanded).
  • The hosts are Debian/Ubuntu with apt and outbound internet (Docker/Caddy installs pull from the network). Docker may already be present — installs are conditional.
  • For the prod host: DNS for thermograph.org must point at the new box before apply, or Caddy's first-request cert issuance will fail.

Use

cd terraform
cp terraform.tfvars.example terraform.tfvars   # then edit: real IPs + secrets
terraform init
terraform plan
terraform apply

Target one host with -target='module.host["beta"]' if you want to apply to just one.

Local state + secrets caveat

The backend is local: terraform.tfstate is written next to the config and holds every secret in cleartext (the rendered env, VAPID keys, DB password, …). It is gitignored (terraform/.gitignore and the root .gitignore). Keep it off shared disks and back it up somewhere private. terraform.tfvars is likewise gitignored; only terraform.tfvars.example (dummy values) is committed. .terraform.lock.hcl is committed so provider versions are pinned across machines.

WARNING — applying against live prod

terraform apply runs remote-exec on the server: it resets the checkout to the branch, rewrites /etc/thermograph.env, and runs docker compose up -d --build (rebuilding images and recreating containers — a brief app restart). Against the live production host this is a real deploy. Review the plan, apply in a maintenance window, and prefer -target to touch one host at a time.

This is separate from the Postgres data cutover in deploy/POSTGRES-MIGRATION.md. Terraform provisions the host and starts the stack; it does not migrate the SQLite→Postgres accounts data. Sequence them deliberately: for a first cutover on a host, follow the migration doc's freeze/backup/copy steps around the point where Terraform brings the stack up — don't let Terraform recreate containers mid-migration.

Assumptions / notes

  • beta has no public domain by default. With compose_files = ["docker-compose.yml"] the app binds 127.0.0.1:8137 (loopback), so opening the port in ufw alone does not expose it. Reach beta via an SSH tunnel, or set domain = "beta.thermograph.org" (adds Caddy TLS) — or add the 0.0.0.0-publishing dev overlay to compose_files — to make it reachable. COOKIE_SECURE is auto-set to 0 when there's no domain (a Secure cookie is never sent over plain HTTP) and 1 behind Caddy TLS.
  • The rendered Caddyfile only reverse-proxies the app. The repo's deploy/Caddyfile additionally serves the emigriffith.dev portfolio and legacy redirects; those are host-specific and not templated here.
  • Provisioner-based by design: the hosts already exist, so this is not a create-from-scratch cloud config. Re-applying is idempotent (installs are guarded, git reset --hard, compose up reconciles).