Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
locals {
|
|
|
|
|
# THERMOGRAPH_DATABASE_URL is built from the same password compose uses for the db
|
|
|
|
|
# container, so the app always matches the database it initialized.
|
|
|
|
|
database_url = "postgresql+asyncpg://thermograph:${var.postgres_password}@db:5432/thermograph"
|
|
|
|
|
|
|
|
|
|
# A Secure cookie is only sent over HTTPS, so enable it only where Caddy terminates
|
|
|
|
|
# TLS (domain set). A plain-HTTP dev box (domain "") would otherwise drop its login
|
|
|
|
|
# cookie and no one could stay signed in.
|
|
|
|
|
cookie_secure = var.domain != "" ? "1" : "0"
|
|
|
|
|
|
|
|
|
|
# Public base URL: the domain over HTTPS, else the raw host:port for a Caddy-less box.
|
|
|
|
|
base_url = var.domain != "" ? "https://${var.domain}" : "http://${var.host}:${var.app_port}"
|
|
|
|
|
|
2026-07-20 13:16:56 +00:00
|
|
|
# Layer the self-hosted Open-Meteo overlay on hosts that self-host the archive.
|
|
|
|
|
effective_compose_files = var.openmeteo ? concat(var.compose_files, ["docker-compose.openmeteo.yml"]) : var.compose_files
|
|
|
|
|
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
# `-f a -f b` for the compose invocations (dev layers the dev overlay).
|
2026-07-20 13:16:56 +00:00
|
|
|
compose_flags = join(" ", [for f in local.effective_compose_files : "-f ${f}"])
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
|
|
|
|
|
# Hash the local compose files so a compose edit re-triggers the remote deploy.
|
2026-07-20 13:16:56 +00:00
|
|
|
compose_files_sha = join(",", [for f in local.effective_compose_files : filesha256("${var.repo_root}/${f}")])
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
|
|
|
|
|
# Rendered /etc/thermograph.env (sensitive — carries every secret).
|
|
|
|
|
env_content = templatefile("${path.module}/templates/thermograph.env.tftpl", {
|
|
|
|
|
app_port = var.app_port
|
|
|
|
|
postgres_password = var.postgres_password
|
|
|
|
|
database_url = local.database_url
|
|
|
|
|
auth_secret = var.auth_secret
|
Have Terraform generate its own internal secrets, with sizing tiers (#239)
Terraform generates the secrets that have no external meaning
(POSTGRES_PASSWORD, AUTH_SECRET, METRICS_TOKEN, INDEXNOW_KEY) via the random
provider instead of requiring the operator to hand-generate and paste each
into terraform.tfvars. Each is pinned with a static keepers value (secrets.tf)
so apply never regenerates a value already in use - the exact incident class
this guards against: every session invalidated, the app<->DB password
mismatched. Rotation is now a deliberate keepers edit, never a side effect.
postgres_password/auth_secret move from required inputs to optional (default
"") - explicit var wins when supplied (seeding an EXISTING live secret during
a migration onto Terraform, hop-1 cutover runbook Stage 0), else Terraform
generates and owns it. metrics_token/indexnow_key are new: neither existed in
Terraform before, both previously left for the app's own fallback generation.
VAPID deliberately stays a required, non-generated input - an EC keypair
where regeneration breaks every existing push subscription outright, unlike
an opaque token.
Sizing tiers: a locals.sizes t-shirt map (nano/small/medium/large ->
{workers, app_cpus, db_cpus, db_memory}), toward the target Proxmox
sizing-tier model (architecture doc SS6) ahead of actually provisioning VMs -
Proxmox itself stays deferred; today a tier just sizes container caps on the
existing SSH-managed hosts. A host can reference one by name (hosts.<name>.
size) or keep hand-picking the four fields, so existing tfvars are
unaffected; prod's example now uses size = "large" (identical numbers),
beta keeps explicit numbers, and a commented uat example demonstrates the
shortcut for a future ephemeral host.
Strengthened terraform/README.md's local-state caveat: more Terraform-
generated secrets landing in tfstate raises the stakes of the existing
never-commit-cleartext-state guidance, not just the sizing.
Verified: terraform validate + fmt clean. A real `terraform plan` against
fake hosts (prod/beta/uat, mixing size="large"/explicit-numbers/size="nano")
resolved every sizing correctly (prod 8/8/4/16g, beta 4/4/2/8g, uat
1/1/1/1g) and planned exactly one instance of each random_password/random_id
resource. Applied just those four resources (real generation, -target to
avoid touching the fake SSH-only host resources) and re-planned: "No
changes" - confirming the keepers pinning holds. Adding an explicit
postgres_password override afterward left the random_password resource
itself completely untouched (0 replace/destroy), confirming the override
path never disturbs the generated resource.
2026-07-21 01:08:57 +00:00
|
|
|
metrics_token = var.metrics_token
|
|
|
|
|
indexnow_key = var.indexnow_key
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
workers = var.workers
|
|
|
|
|
app_cpus = var.app_cpus
|
|
|
|
|
db_cpus = var.db_cpus
|
|
|
|
|
db_memory = var.db_memory
|
Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235)
* Split web/worker duties with THERMOGRAPH_ROLE
Background work (the subscription notifier) is welded to the same process
that serves requests, so scaling the web tier to N replicas would also scale
notifier instances unless something restricts it further than leader
election alone.
Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process
behavior). Every replica runs the same image; ROLE only gates whether a
process is allowed to own the notifier at all, layered on top of the
existing leader election: web replicas never start it even if they'd win
leader election, worker replicas start it if they win. The decision is
pulled into _should_run_notifier() so it's unit-testable without booting the
full app (DB init, places index, neighbor warmer).
Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE)
so a worker replica - which serves no real traffic - still has something
Swarm can health-check.
* Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate
Three changes toward the hop-1 interim cutover, all inert until Track B
stands up the platform:
docker-stack.yml: the Swarm stack file for the interim cutover, distinct
from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a
pre-built image (IMAGE_TAG) instead of building in place; app/worker publish
no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing
mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted),
reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard
needs a smaller MTU or large payloads silently stall); db is placement-
pinned to a labelled node; app/worker skip inline migrations
(RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing
that ever runs Alembic; secrets are real Swarm secrets mounted at
/run/secrets, read by the entrypoint shim rather than plain env vars.
TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default
latest-pg18, today's behavior unchanged), wired through Terraform
(timescaledb_tag, default "latest-pg18") so it can actually be pinned to an
exact minor without hand-editing the host - required before any host of the
stack could replicate with another (a floating tag risks mismatched
extension minors, which blocks a physical replica and risks compressed-
chunk corruption on restore).
Caddy active health-gate: both the Terraform-rendered Caddyfile and the live
deploy/Caddyfile now health-check the app on the same cheap /healthz route
its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage,
so it's cheap enough for a tight interval and works identically for a
worker replica, which serves no public traffic at all) - Caddy won't
forward into a container that's still booting or unhealthy.
Verified live: built and booted the real image via docker compose - both
containers report healthy via the new /healthz-based HEALTHCHECK, and GET /
still renders the full SSR homepage unchanged. Both Caddyfiles validated
with the real caddy binary. docker-stack.yml validated with docker compose
config (required-var guards fire with clear messages; secrets correctly
mount at /run/secrets/<name>, matching the entrypoint shim's mapping).
docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside
the existing openmeteo overlay. terraform validate + fmt clean.
2026-07-21 00:39:48 +00:00
|
|
|
timescaledb_tag = var.timescaledb_tag
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
base = "/"
|
|
|
|
|
base_url = local.base_url
|
|
|
|
|
cookie_secure = local.cookie_secure
|
2026-07-20 13:16:56 +00:00
|
|
|
openmeteo = var.openmeteo
|
|
|
|
|
om_data_dir = var.om_data_dir
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
vapid_private_key = var.vapid_private_key
|
|
|
|
|
vapid_public_key = var.vapid_public_key
|
|
|
|
|
vapid_contact = var.vapid_contact
|
|
|
|
|
google_verify = var.google_verify
|
|
|
|
|
bing_verify = var.bing_verify
|
|
|
|
|
mail_backend = var.mail_backend
|
|
|
|
|
smtp_host = var.smtp_host
|
|
|
|
|
smtp_port = var.smtp_port
|
|
|
|
|
smtp_user = var.smtp_user
|
|
|
|
|
smtp_password = var.smtp_password
|
|
|
|
|
smtp_starttls = var.smtp_starttls
|
|
|
|
|
mail_from = var.mail_from
|
|
|
|
|
mail_reply_to = var.mail_reply_to
|
|
|
|
|
discord_webhook = var.discord_webhook
|
|
|
|
|
discord_public_key = var.discord_public_key
|
|
|
|
|
discord_app_id = var.discord_app_id
|
|
|
|
|
discord_bot_token = var.discord_bot_token
|
|
|
|
|
discord_client_secret = var.discord_client_secret
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
# Caddyfile is only meaningful on a host with a public domain.
|
|
|
|
|
caddy_content = var.domain != "" ? templatefile("${path.module}/templates/Caddyfile.tftpl", {
|
2026-07-21 20:01:30 +00:00
|
|
|
domain = var.domain
|
|
|
|
|
port = var.app_port
|
|
|
|
|
frontend_port = var.frontend_port
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
}) : "# No public domain on this host; Caddy is not managed here.\n"
|
2026-07-20 13:16:56 +00:00
|
|
|
|
|
|
|
|
# systemd unit that keeps the object-storage bucket rclone-mounted at om_data_dir,
|
|
|
|
|
# so the Open-Meteo containers read the ERA5 .om archive from it. Only installed on
|
|
|
|
|
# openmeteo hosts; --allow-other lets the container (root) read the FUSE mount.
|
|
|
|
|
rclone_unit = <<-UNIT
|
|
|
|
|
[Unit]
|
|
|
|
|
Description=rclone mount ERA5 archive (Thermograph)
|
|
|
|
|
After=network-online.target
|
|
|
|
|
Wants=network-online.target
|
|
|
|
|
|
|
|
|
|
[Service]
|
|
|
|
|
Type=notify
|
|
|
|
|
ExecStartPre=/bin/mkdir -p ${var.om_data_dir}
|
|
|
|
|
ExecStart=/usr/bin/rclone mount ${var.om_bucket_remote} ${var.om_data_dir} --config /etc/rclone/rclone.conf --vfs-cache-mode full --vfs-cache-max-size ${var.om_vfs_cache_max} --dir-cache-time 12h --allow-other --umask 000
|
|
|
|
|
ExecStop=/bin/fusermount -u ${var.om_data_dir}
|
|
|
|
|
Restart=on-failure
|
|
|
|
|
RestartSec=5
|
|
|
|
|
|
|
|
|
|
[Install]
|
|
|
|
|
WantedBy=multi-user.target
|
|
|
|
|
UNIT
|
|
|
|
|
|
|
|
|
|
rclone_unit_content = var.openmeteo ? local.rclone_unit : "# openmeteo disabled on this host\n"
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
resource "null_resource" "host" {
|
|
|
|
|
# Re-provision when the rendered env, the compose files, the branch, sizing, or the
|
|
|
|
|
# Caddyfile change. sha256 of the secret-bearing env is unwrapped with nonsensitive()
|
|
|
|
|
# (a one-way hash leaks nothing) so plans stay readable.
|
|
|
|
|
triggers = {
|
|
|
|
|
env_sha = nonsensitive(sha256(local.env_content))
|
|
|
|
|
compose_sha = local.compose_files_sha
|
|
|
|
|
compose_flags = local.compose_flags
|
|
|
|
|
caddy_sha = sha256(local.caddy_content)
|
|
|
|
|
branch = var.git_branch
|
|
|
|
|
app_dir = var.app_dir
|
Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235)
* Split web/worker duties with THERMOGRAPH_ROLE
Background work (the subscription notifier) is welded to the same process
that serves requests, so scaling the web tier to N replicas would also scale
notifier instances unless something restricts it further than leader
election alone.
Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process
behavior). Every replica runs the same image; ROLE only gates whether a
process is allowed to own the notifier at all, layered on top of the
existing leader election: web replicas never start it even if they'd win
leader election, worker replicas start it if they win. The decision is
pulled into _should_run_notifier() so it's unit-testable without booting the
full app (DB init, places index, neighbor warmer).
Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE)
so a worker replica - which serves no real traffic - still has something
Swarm can health-check.
* Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate
Three changes toward the hop-1 interim cutover, all inert until Track B
stands up the platform:
docker-stack.yml: the Swarm stack file for the interim cutover, distinct
from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a
pre-built image (IMAGE_TAG) instead of building in place; app/worker publish
no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing
mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted),
reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard
needs a smaller MTU or large payloads silently stall); db is placement-
pinned to a labelled node; app/worker skip inline migrations
(RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing
that ever runs Alembic; secrets are real Swarm secrets mounted at
/run/secrets, read by the entrypoint shim rather than plain env vars.
TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default
latest-pg18, today's behavior unchanged), wired through Terraform
(timescaledb_tag, default "latest-pg18") so it can actually be pinned to an
exact minor without hand-editing the host - required before any host of the
stack could replicate with another (a floating tag risks mismatched
extension minors, which blocks a physical replica and risks compressed-
chunk corruption on restore).
Caddy active health-gate: both the Terraform-rendered Caddyfile and the live
deploy/Caddyfile now health-check the app on the same cheap /healthz route
its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage,
so it's cheap enough for a tight interval and works identically for a
worker replica, which serves no public traffic at all) - Caddy won't
forward into a container that's still booting or unhealthy.
Verified live: built and booted the real image via docker compose - both
containers report healthy via the new /healthz-based HEALTHCHECK, and GET /
still renders the full SSR homepage unchanged. Both Caddyfiles validated
with the real caddy binary. docker-stack.yml validated with docker compose
config (required-var guards fire with clear messages; secrets correctly
mount at /run/secrets/<name>, matching the entrypoint shim's mapping).
docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside
the existing openmeteo overlay. terraform validate + fmt clean.
2026-07-21 00:39:48 +00:00
|
|
|
sizing = "${var.workers}/${var.app_cpus}/${var.db_cpus}/${var.db_memory}/${var.timescaledb_tag}"
|
2026-07-20 13:16:56 +00:00
|
|
|
# Re-provision when the archive self-hosting config changes. The rclone.conf is
|
|
|
|
|
# hashed (nonsensitive on a one-way digest) so a credential rotation redeploys.
|
|
|
|
|
openmeteo = "${var.openmeteo}/${var.om_data_dir}/${var.om_bucket_remote}/${var.om_vfs_cache_max}"
|
|
|
|
|
om_conf = var.openmeteo ? nonsensitive(sha256(var.om_rclone_conf)) : "off"
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
connection {
|
|
|
|
|
type = "ssh"
|
|
|
|
|
host = var.host
|
|
|
|
|
user = var.ssh_user
|
|
|
|
|
private_key = file(pathexpand(var.ssh_private_key_path))
|
|
|
|
|
timeout = "5m"
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
# Push the rendered secrets via `content` so they are never written to local disk.
|
|
|
|
|
provisioner "file" {
|
|
|
|
|
content = local.env_content
|
|
|
|
|
destination = "/tmp/thermograph.env"
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
provisioner "file" {
|
|
|
|
|
content = local.caddy_content
|
|
|
|
|
destination = "/tmp/thermograph.Caddyfile"
|
|
|
|
|
}
|
|
|
|
|
|
2026-07-20 13:16:56 +00:00
|
|
|
# rclone config (bucket credentials) + the mount unit. Pushed via `content` so the
|
|
|
|
|
# secret never touches local disk; harmless placeholders on non-openmeteo hosts.
|
|
|
|
|
provisioner "file" {
|
|
|
|
|
content = var.openmeteo ? var.om_rclone_conf : "# openmeteo disabled on this host\n"
|
|
|
|
|
destination = "/tmp/thermograph.rclone.conf"
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
provisioner "file" {
|
|
|
|
|
content = local.rclone_unit_content
|
|
|
|
|
destination = "/tmp/rclone-om.service"
|
|
|
|
|
}
|
|
|
|
|
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
# 1. Host setup: Docker + compose plugin, ufw firewall, and (domain hosts) Caddy.
|
|
|
|
|
provisioner "remote-exec" {
|
|
|
|
|
inline = [
|
|
|
|
|
<<-EOT
|
|
|
|
|
set -eu
|
|
|
|
|
echo "[${var.name}] setup: docker, compose plugin, firewall"
|
|
|
|
|
if ! command -v docker >/dev/null 2>&1; then
|
|
|
|
|
curl -fsSL https://get.docker.com | sudo sh
|
|
|
|
|
fi
|
|
|
|
|
sudo usermod -aG docker "$(id -un)" || true
|
|
|
|
|
if ! sudo docker compose version >/dev/null 2>&1; then
|
|
|
|
|
sudo apt-get update -y
|
|
|
|
|
sudo apt-get install -y docker-compose-plugin
|
|
|
|
|
fi
|
|
|
|
|
if ! command -v ufw >/dev/null 2>&1; then
|
|
|
|
|
sudo apt-get update -y
|
|
|
|
|
sudo apt-get install -y ufw
|
|
|
|
|
fi
|
|
|
|
|
sudo ufw allow 22/tcp
|
|
|
|
|
sudo ufw allow 80/tcp
|
|
|
|
|
sudo ufw allow 443/tcp
|
|
|
|
|
DOMAIN='${var.domain}'
|
|
|
|
|
if [ -z "$DOMAIN" ]; then
|
|
|
|
|
sudo ufw allow ${var.app_port}/tcp
|
|
|
|
|
fi
|
|
|
|
|
sudo ufw --force enable
|
|
|
|
|
if [ -n "$DOMAIN" ]; then
|
|
|
|
|
if ! command -v caddy >/dev/null 2>&1; then
|
|
|
|
|
sudo apt-get install -y debian-keyring debian-archive-keyring apt-transport-https curl
|
|
|
|
|
curl -1sLf 'https://dl.cloudsmith.io/public/caddy/stable/gpg.key' | sudo gpg --batch --yes --dearmor -o /usr/share/keyrings/caddy-stable-archive-keyring.gpg
|
|
|
|
|
curl -1sLf 'https://dl.cloudsmith.io/public/caddy/stable/debian.deb.txt' | sudo tee /etc/apt/sources.list.d/caddy-stable.list >/dev/null
|
|
|
|
|
sudo apt-get update -y
|
|
|
|
|
sudo apt-get install -y caddy
|
|
|
|
|
fi
|
|
|
|
|
sudo install -m 0644 /tmp/thermograph.Caddyfile /etc/caddy/Caddyfile
|
|
|
|
|
sudo systemctl reload caddy || sudo systemctl restart caddy
|
|
|
|
|
fi
|
|
|
|
|
rm -f /tmp/thermograph.Caddyfile
|
|
|
|
|
EOT
|
|
|
|
|
]
|
|
|
|
|
}
|
|
|
|
|
|
2026-07-20 13:16:56 +00:00
|
|
|
# 1b. Self-hosted archive: rclone-mount the ERA5 bucket before compose up so
|
|
|
|
|
# open-meteo-api can read .om from object storage. Skipped on non-openmeteo hosts.
|
|
|
|
|
provisioner "remote-exec" {
|
|
|
|
|
inline = [
|
|
|
|
|
<<-EOT
|
|
|
|
|
set -eu
|
|
|
|
|
if [ "${var.openmeteo}" != "true" ]; then
|
|
|
|
|
rm -f /tmp/thermograph.rclone.conf /tmp/rclone-om.service
|
2026-07-20 14:33:09 +00:00
|
|
|
# Undo any prior openmeteo setup so a toggled-off host doesn't wait on a mount.
|
|
|
|
|
if [ -e /etc/systemd/system/docker.service.d/10-wait-rclone.conf ]; then
|
|
|
|
|
sudo rm -f /etc/systemd/system/docker.service.d/10-wait-rclone.conf
|
|
|
|
|
sudo systemctl disable --now rclone-om >/dev/null 2>&1 || true
|
|
|
|
|
sudo systemctl daemon-reload
|
|
|
|
|
fi
|
2026-07-20 13:16:56 +00:00
|
|
|
echo "[${var.name}] openmeteo: disabled"
|
|
|
|
|
exit 0
|
|
|
|
|
fi
|
|
|
|
|
echo "[${var.name}] openmeteo: rclone mount ${var.om_bucket_remote} -> ${var.om_data_dir}"
|
|
|
|
|
if ! command -v rclone >/dev/null 2>&1; then
|
|
|
|
|
curl -fsSL https://rclone.org/install.sh | sudo bash
|
|
|
|
|
fi
|
|
|
|
|
# FUSE allow_other so the container (root) can read a mount owned by this user.
|
|
|
|
|
if ! grep -q '^user_allow_other' /etc/fuse.conf 2>/dev/null; then
|
|
|
|
|
echo user_allow_other | sudo tee -a /etc/fuse.conf >/dev/null
|
|
|
|
|
fi
|
|
|
|
|
sudo install -d -m 0755 /etc/rclone
|
|
|
|
|
sudo install -m 0600 /tmp/thermograph.rclone.conf /etc/rclone/rclone.conf
|
|
|
|
|
sudo install -m 0644 /tmp/rclone-om.service /etc/systemd/system/rclone-om.service
|
|
|
|
|
rm -f /tmp/thermograph.rclone.conf /tmp/rclone-om.service
|
|
|
|
|
sudo install -d -m 0755 ${var.om_data_dir}
|
|
|
|
|
sudo systemctl daemon-reload
|
|
|
|
|
sudo systemctl enable rclone-om
|
|
|
|
|
sudo systemctl restart rclone-om
|
|
|
|
|
# Wait for the mount before compose bind-mounts it.
|
|
|
|
|
i=0
|
|
|
|
|
while [ "$i" -lt 30 ]; do
|
|
|
|
|
if mountpoint -q ${var.om_data_dir}; then break; fi
|
|
|
|
|
i=$((i + 1)); sleep 2
|
|
|
|
|
done
|
|
|
|
|
if ! mountpoint -q ${var.om_data_dir}; then
|
|
|
|
|
echo "[${var.name}] RCLONE MOUNT FAILED" >&2
|
|
|
|
|
sudo systemctl status rclone-om --no-pager || true
|
|
|
|
|
exit 1
|
|
|
|
|
fi
|
2026-07-20 14:33:09 +00:00
|
|
|
# Order Docker after the mount on every boot, so the restart-policy containers
|
|
|
|
|
# never bind an empty mount point. rclone-om is Type=notify, so `After` waits
|
|
|
|
|
# until the mount is actually ready — not merely that the unit was launched.
|
|
|
|
|
sudo install -d -m 0755 /etc/systemd/system/docker.service.d
|
|
|
|
|
printf '[Unit]\nWants=rclone-om.service\nAfter=rclone-om.service\n' \
|
|
|
|
|
| sudo tee /etc/systemd/system/docker.service.d/10-wait-rclone.conf >/dev/null
|
|
|
|
|
sudo systemctl daemon-reload
|
2026-07-20 13:16:56 +00:00
|
|
|
echo "[${var.name}] openmeteo: mounted at ${var.om_data_dir}"
|
|
|
|
|
EOT
|
|
|
|
|
]
|
|
|
|
|
}
|
|
|
|
|
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
# 2. Sync the checkout to the host's branch (cloning first on a fresh box).
|
|
|
|
|
provisioner "remote-exec" {
|
|
|
|
|
inline = [
|
|
|
|
|
<<-EOT
|
|
|
|
|
set -eu
|
|
|
|
|
echo "[${var.name}] code: ${var.app_dir} -> origin/${var.git_branch}"
|
|
|
|
|
sudo mkdir -p ${var.app_dir}
|
|
|
|
|
sudo chown -R "$(id -un)":"$(id -gn)" ${var.app_dir}
|
|
|
|
|
if [ ! -d ${var.app_dir}/.git ]; then
|
|
|
|
|
git clone ${var.repo_url} ${var.app_dir}
|
|
|
|
|
fi
|
|
|
|
|
cd ${var.app_dir}
|
|
|
|
|
git fetch --prune origin ${var.git_branch}
|
|
|
|
|
git reset --hard origin/${var.git_branch}
|
|
|
|
|
EOT
|
|
|
|
|
]
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
# 3. Install /etc/thermograph.env, then bring the stack up. docker runs as root
|
|
|
|
|
# (sources the env file in the same shell) so it never depends on the docker
|
|
|
|
|
# group membership taking effect in this session.
|
|
|
|
|
provisioner "remote-exec" {
|
|
|
|
|
inline = [
|
|
|
|
|
<<-EOT
|
|
|
|
|
set -eu
|
|
|
|
|
echo "[${var.name}] deploy: install env + docker compose up"
|
|
|
|
|
sudo install -m 0640 -o root -g root /tmp/thermograph.env /etc/thermograph.env
|
|
|
|
|
rm -f /tmp/thermograph.env
|
|
|
|
|
sudo bash -c 'set -a; . /etc/thermograph.env; set +a; cd ${var.app_dir} && docker compose ${local.compose_flags} up -d --build'
|
|
|
|
|
EOT
|
|
|
|
|
]
|
|
|
|
|
}
|
|
|
|
|
|
2026-07-21 20:01:30 +00:00
|
|
|
# 4. Health check backend on loopback -- a plain "/" here already exercises
|
|
|
|
|
# the whole chain end to end even without Caddy in front (backend's own
|
|
|
|
|
# reverse-proxy fallback forwards to frontend), so one curl covers both
|
|
|
|
|
# services in every topology this module supports (repo-split Stage 4).
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
provisioner "remote-exec" {
|
|
|
|
|
inline = [
|
|
|
|
|
<<-EOT
|
|
|
|
|
set -eu
|
|
|
|
|
echo "[${var.name}] health: http://127.0.0.1:${var.app_port}/"
|
|
|
|
|
ok=0
|
|
|
|
|
i=0
|
|
|
|
|
while [ "$i" -lt 30 ]; do
|
|
|
|
|
if curl -fsS -o /dev/null "http://127.0.0.1:${var.app_port}/"; then ok=1; break; fi
|
|
|
|
|
i=$((i + 1))
|
|
|
|
|
sleep 2
|
|
|
|
|
done
|
|
|
|
|
if [ "$ok" != 1 ]; then
|
|
|
|
|
echo "[${var.name}] HEALTH CHECK FAILED" >&2
|
2026-07-21 20:01:30 +00:00
|
|
|
sudo bash -c 'cd ${var.app_dir} && docker compose ${local.compose_flags} ps; docker compose ${local.compose_flags} logs --tail=50 backend; docker compose ${local.compose_flags} logs --tail=50 frontend' || true
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
exit 1
|
|
|
|
|
fi
|
|
|
|
|
echo "[${var.name}] OK: serving on 127.0.0.1:${var.app_port}"
|
2026-07-21 20:01:30 +00:00
|
|
|
if ! curl -fsS -o /dev/null "http://127.0.0.1:${var.frontend_port}/healthz"; then
|
|
|
|
|
echo "[${var.name}] frontend's own /healthz failed directly (backend's proxy to it may still work) -- check separately" >&2
|
|
|
|
|
sudo bash -c 'cd ${var.app_dir} && docker compose ${local.compose_flags} logs --tail=50 frontend' || true
|
|
|
|
|
fi
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
EOT
|
|
|
|
|
]
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
output "host" {
|
|
|
|
|
description = "Address this module manages."
|
|
|
|
|
value = var.host
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
output "role" {
|
|
|
|
|
description = "Role of this host."
|
|
|
|
|
value = var.role
|
|
|
|
|
}
|