thermograph/terraform
Emi Griffith c41ade74b1 Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235)
* Split web/worker duties with THERMOGRAPH_ROLE

Background work (the subscription notifier) is welded to the same process
that serves requests, so scaling the web tier to N replicas would also scale
notifier instances unless something restricts it further than leader
election alone.

Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process
behavior). Every replica runs the same image; ROLE only gates whether a
process is allowed to own the notifier at all, layered on top of the
existing leader election: web replicas never start it even if they'd win
leader election, worker replicas start it if they win. The decision is
pulled into _should_run_notifier() so it's unit-testable without booting the
full app (DB init, places index, neighbor warmer).

Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE)
so a worker replica - which serves no real traffic - still has something
Swarm can health-check.

* Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate

Three changes toward the hop-1 interim cutover, all inert until Track B
stands up the platform:

docker-stack.yml: the Swarm stack file for the interim cutover, distinct
from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a
pre-built image (IMAGE_TAG) instead of building in place; app/worker publish
no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing
mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted),
reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard
needs a smaller MTU or large payloads silently stall); db is placement-
pinned to a labelled node; app/worker skip inline migrations
(RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing
that ever runs Alembic; secrets are real Swarm secrets mounted at
/run/secrets, read by the entrypoint shim rather than plain env vars.

TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default
latest-pg18, today's behavior unchanged), wired through Terraform
(timescaledb_tag, default "latest-pg18") so it can actually be pinned to an
exact minor without hand-editing the host - required before any host of the
stack could replicate with another (a floating tag risks mismatched
extension minors, which blocks a physical replica and risks compressed-
chunk corruption on restore).

Caddy active health-gate: both the Terraform-rendered Caddyfile and the live
deploy/Caddyfile now health-check the app on the same cheap /healthz route
its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage,
so it's cheap enough for a tight interval and works identically for a
worker replica, which serves no public traffic at all) - Caddy won't
forward into a container that's still booting or unhealthy.

Verified live: built and booted the real image via docker compose - both
containers report healthy via the new /healthz-based HEALTHCHECK, and GET /
still renders the full SSR homepage unchanged. Both Caddyfiles validated
with the real caddy binary. docker-stack.yml validated with docker compose
config (required-var guards fire with clear messages; secrets correctly
mount at /run/secrets/<name>, matching the entrypoint shim's mapping).
docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside
the existing openmeteo overlay. terraform validate + fmt clean.
2026-07-21 00:39:48 +00:00
..
modules/thermograph-host Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235) 2026-07-21 00:39:48 +00:00
.gitignore Add Terraform to provision the VPS hosts (compose keeps running the app) (#223) 2026-07-20 07:42:15 +00:00
.terraform.lock.hcl Add Terraform to provision the VPS hosts (compose keeps running the app) (#223) 2026-07-20 07:42:15 +00:00
main.tf Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235) 2026-07-21 00:39:48 +00:00
outputs.tf Add Terraform to provision the VPS hosts (compose keeps running the app) (#223) 2026-07-20 07:42:15 +00:00
README.md Move the climate record from parquet to TimescaleDB hypertables (#227) 2026-07-20 20:15:55 +00:00
terraform.tfvars.example Scale DB tuning from DB_MEMORY; order Docker after the rclone mount (#226) 2026-07-20 14:33:09 +00:00
variables.tf Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235) 2026-07-21 00:39:48 +00:00
versions.tf Add Terraform to provision the VPS hosts (compose keeps running the app) (#223) 2026-07-20 07:42:15 +00:00

Thermograph — Terraform (host provisioning)

Terraform that provisions and configures the existing VPS hosts and hands the app off to docker compose. It does not create servers (no cloud provider) and does not replace compose — it prepares each host (Docker, firewall, checkout, secrets, Caddy) and runs docker compose up.

What it manages

One reusable module (modules/thermograph-host) is instantiated per host via for_each. This config manages two VPS hosts:

key role VPS branch domain notes
prod prod NEW 48 GB / 12-core box release thermograph.org Caddy TLS; sized up (8/4/16g)
beta beta old box 75.119.132.91 main (none) testing/beta; no public domain yet

The dev branch is out of scope here — it deploys to the LAN dev server via deploy/deploy-dev.sh (a self-hosted GitHub Actions runner), not Terraform.

Per host, over SSH provisioners, Terraform:

  • installs Docker + the compose plugin if missing;
  • configures a ufw firewall (22/80/443 always; on a host with no domain it also opens the app port 8137);
  • ensures the git checkout at app_dir exists (clones on a fresh box) and resets it to the host's branch;
  • renders /etc/thermograph.env from Terraform variables (secrets injected from tfvars, pushed via provisioner content so they never touch local disk) and installs it root-owned 0640;
  • for a host with a domain, installs a rendered Caddyfile and reloads Caddy;
  • brings the stack up: docker compose <-f each compose file> up -d --build, running docker as root with /etc/thermograph.env sourced in the same shell;
  • health-checks http://127.0.0.1:8137/.

A change to the rendered env, the compose files, the branch, or the sizing flips the null_resource trigger and re-runs the provisioners on next apply.

Container sizing is env-driven

docker-compose.yml reads WORKERS, APP_CPUS, DB_CPUS, and DB_MEMORY from the environment (defaults 4 / 4 / 2 / 8g, identical to before). Terraform sets them per host through /etc/thermograph.env, so the big prod box can run larger caps without a compose edit. The Postgres internal memory budget (shared_buffers, effective_cache_size, work_mem, maintenance_work_mem) is derived from the same DB_MEMORY by deploy/db/init/20-tuning.sh — so raising db_memory scales the container cap and the tuning together (prod 16g → shared_buffers 4 GB). The tuning applies on a fresh DB volume; on an existing volume re-run it by hand (see the script header).

Prerequisites

  • Terraform >= 1.6 (v1.15 is installed).
  • SSH key access to both hosts as a sudo-capable user (default deploy). Point ssh_private_key_path at that key (~ is expanded).
  • The hosts are Debian/Ubuntu with apt and outbound internet (Docker/Caddy installs pull from the network). Docker may already be present — installs are conditional.
  • For the prod host: DNS for thermograph.org must point at the new box before apply, or Caddy's first-request cert issuance will fail.

Use

cd terraform
cp terraform.tfvars.example terraform.tfvars   # then edit: real IPs + secrets
terraform init
terraform plan
terraform apply

Target one host with -target='module.host["beta"]' if you want to apply to just one.

Local state + secrets caveat

The backend is local: terraform.tfstate is written next to the config and holds every secret in cleartext (the rendered env, VAPID keys, DB password, …). It is gitignored (terraform/.gitignore and the root .gitignore). Keep it off shared disks and back it up somewhere private. terraform.tfvars is likewise gitignored; only terraform.tfvars.example (dummy values) is committed. .terraform.lock.hcl is committed so provider versions are pinned across machines.

WARNING — applying against live prod

terraform apply runs remote-exec on the server: it resets the checkout to the branch, rewrites /etc/thermograph.env, and runs docker compose up -d --build (rebuilding images and recreating containers — a brief app restart). Against the live production host this is a real deploy. Review the plan, apply in a maintenance window, and prefer -target to touch one host at a time.

This is separate from the Postgres data cutover in deploy/POSTGRES-MIGRATION.md. Terraform provisions the host and starts the stack; it does not migrate the SQLite→Postgres accounts data. Sequence them deliberately: for a first cutover on a host, follow the migration doc's freeze/backup/copy steps around the point where Terraform brings the stack up — don't let Terraform recreate containers mid-migration.

Assumptions / notes

  • beta has no public domain by default. With compose_files = ["docker-compose.yml"] the app binds 127.0.0.1:8137 (loopback), so opening the port in ufw alone does not expose it. Reach beta via an SSH tunnel, or set domain = "beta.thermograph.org" (adds Caddy TLS) — or add the 0.0.0.0-publishing dev overlay to compose_files — to make it reachable. COOKIE_SECURE is auto-set to 0 when there's no domain (a Secure cookie is never sent over plain HTTP) and 1 behind Caddy TLS.
  • The rendered Caddyfile only reverse-proxies the app. The repo's deploy/Caddyfile additionally serves the emigriffith.dev portfolio and legacy redirects; those are host-specific and not templated here.
  • Provisioner-based by design: the hosts already exist, so this is not a create-from-scratch cloud config. Re-applying is idempotent (installs are guarded, git reset --hard, compose up reconciles).