Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
locals {
|
|
|
|
|
# Repo root (one level above this terraform/ dir). The module hashes the compose
|
|
|
|
|
# files here so a compose change re-triggers the remote deploy, and this is the
|
|
|
|
|
# tree the host's checkout mirrors over git.
|
|
|
|
|
repo_root = abspath("${path.root}/..")
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
# One module instance per host. The module is entirely SSH-provisioner driven — it
|
|
|
|
|
# configures an already-existing VPS and hands the app off to docker compose.
|
|
|
|
|
module "host" {
|
|
|
|
|
source = "./modules/thermograph-host"
|
|
|
|
|
for_each = var.hosts
|
|
|
|
|
|
|
|
|
|
# Per-host config
|
|
|
|
|
name = each.key
|
|
|
|
|
host = each.value.host
|
|
|
|
|
ssh_user = each.value.ssh_user
|
|
|
|
|
ssh_private_key_path = each.value.ssh_private_key_path
|
|
|
|
|
role = each.value.role
|
|
|
|
|
git_branch = each.value.git_branch
|
|
|
|
|
domain = each.value.domain
|
|
|
|
|
compose_files = each.value.compose_files
|
|
|
|
|
app_dir = each.value.app_dir
|
|
|
|
|
workers = each.value.workers
|
|
|
|
|
app_cpus = each.value.app_cpus
|
|
|
|
|
db_cpus = each.value.db_cpus
|
|
|
|
|
db_memory = each.value.db_memory
|
Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235)
* Split web/worker duties with THERMOGRAPH_ROLE
Background work (the subscription notifier) is welded to the same process
that serves requests, so scaling the web tier to N replicas would also scale
notifier instances unless something restricts it further than leader
election alone.
Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process
behavior). Every replica runs the same image; ROLE only gates whether a
process is allowed to own the notifier at all, layered on top of the
existing leader election: web replicas never start it even if they'd win
leader election, worker replicas start it if they win. The decision is
pulled into _should_run_notifier() so it's unit-testable without booting the
full app (DB init, places index, neighbor warmer).
Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE)
so a worker replica - which serves no real traffic - still has something
Swarm can health-check.
* Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate
Three changes toward the hop-1 interim cutover, all inert until Track B
stands up the platform:
docker-stack.yml: the Swarm stack file for the interim cutover, distinct
from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a
pre-built image (IMAGE_TAG) instead of building in place; app/worker publish
no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing
mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted),
reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard
needs a smaller MTU or large payloads silently stall); db is placement-
pinned to a labelled node; app/worker skip inline migrations
(RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing
that ever runs Alembic; secrets are real Swarm secrets mounted at
/run/secrets, read by the entrypoint shim rather than plain env vars.
TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default
latest-pg18, today's behavior unchanged), wired through Terraform
(timescaledb_tag, default "latest-pg18") so it can actually be pinned to an
exact minor without hand-editing the host - required before any host of the
stack could replicate with another (a floating tag risks mismatched
extension minors, which blocks a physical replica and risks compressed-
chunk corruption on restore).
Caddy active health-gate: both the Terraform-rendered Caddyfile and the live
deploy/Caddyfile now health-check the app on the same cheap /healthz route
its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage,
so it's cheap enough for a tight interval and works identically for a
worker replica, which serves no public traffic at all) - Caddy won't
forward into a container that's still booting or unhealthy.
Verified live: built and booted the real image via docker compose - both
containers report healthy via the new /healthz-based HEALTHCHECK, and GET /
still renders the full SSR homepage unchanged. Both Caddyfiles validated
with the real caddy binary. docker-stack.yml validated with docker compose
config (required-var guards fire with clear messages; secrets correctly
mount at /run/secrets/<name>, matching the entrypoint shim's mapping).
docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside
the existing openmeteo overlay. terraform validate + fmt clean.
2026-07-21 00:39:48 +00:00
|
|
|
timescaledb_tag = each.value.timescaledb_tag
|
2026-07-20 13:16:56 +00:00
|
|
|
openmeteo = each.value.openmeteo
|
|
|
|
|
om_data_dir = each.value.om_data_dir
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
|
|
|
|
|
# Shared infra config
|
|
|
|
|
repo_root = local.repo_root
|
|
|
|
|
repo_url = var.repo_url
|
|
|
|
|
app_port = var.app_port
|
|
|
|
|
|
2026-07-20 13:16:56 +00:00
|
|
|
# Shared object-storage config (only used where openmeteo = true)
|
|
|
|
|
om_bucket_remote = var.om_bucket_remote
|
|
|
|
|
om_rclone_conf = var.om_rclone_conf
|
|
|
|
|
om_vfs_cache_max = var.om_vfs_cache_max
|
|
|
|
|
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
# Shared secrets -> /etc/thermograph.env
|
|
|
|
|
postgres_password = var.postgres_password
|
|
|
|
|
auth_secret = var.auth_secret
|
|
|
|
|
vapid_private_key = var.vapid_private_key
|
|
|
|
|
vapid_public_key = var.vapid_public_key
|
|
|
|
|
vapid_contact = var.vapid_contact
|
|
|
|
|
google_verify = var.google_verify
|
|
|
|
|
bing_verify = var.bing_verify
|
|
|
|
|
mail_backend = var.mail_backend
|
|
|
|
|
smtp_host = var.smtp_host
|
|
|
|
|
smtp_port = var.smtp_port
|
|
|
|
|
smtp_user = var.smtp_user
|
|
|
|
|
smtp_password = var.smtp_password
|
|
|
|
|
smtp_starttls = var.smtp_starttls
|
|
|
|
|
mail_from = var.mail_from
|
|
|
|
|
mail_reply_to = var.mail_reply_to
|
|
|
|
|
discord_webhook = var.discord_webhook
|
|
|
|
|
discord_public_key = var.discord_public_key
|
|
|
|
|
discord_app_id = var.discord_app_id
|
|
|
|
|
discord_bot_token = var.discord_bot_token
|
|
|
|
|
discord_client_secret = var.discord_client_secret
|
|
|
|
|
}
|