thermograph/terraform/modules/thermograph-host/variables.tf

236 lines
4.8 KiB
Terraform
Raw Normal View History

# Per-host inputs (all supplied by the root module's for_each).
variable "name" {
description = "Short host key (e.g. \"prod\", \"dev\"), used in log lines."
type = string
}
variable "host" {
description = "IP or hostname to SSH to."
type = string
}
variable "ssh_user" {
description = "SSH login user (must be able to sudo)."
type = string
}
variable "ssh_private_key_path" {
description = "Path to the private key file for ssh_user."
type = string
}
variable "role" {
description = "\"prod\" | \"dev\" — informational."
type = string
}
variable "git_branch" {
description = "Branch the host checkout is reset to."
type = string
}
variable "domain" {
description = "Public domain. \"\" => no Caddy/TLS (open the app port instead)."
type = string
}
variable "compose_files" {
description = "Compose files to layer, in order (dev appends docker-compose.dev.yml)."
type = list(string)
}
variable "openmeteo" {
description = "Self-host the ERA5 archive: layer docker-compose.openmeteo.yml + provision the host rclone mount."
type = bool
default = false
}
variable "om_data_dir" {
description = "Host rclone mount point for the archive bucket (OM_DATA_DIR the overlay bind-mounts)."
type = string
default = "/mnt/om-archive"
}
variable "om_bucket_remote" {
description = "rclone remote:path for the archive bucket (mounted at om_data_dir)."
type = string
default = ""
}
variable "om_rclone_conf" {
description = "rclone.conf contents installed to /etc/rclone/rclone.conf. Sensitive."
type = string
default = ""
sensitive = true
}
variable "om_vfs_cache_max" {
description = "rclone --vfs-cache-max-size for the mount's on-disk hot cache."
type = string
default = "80G"
}
variable "app_dir" {
description = "Checkout path on the host."
type = string
}
variable "repo_root" {
description = "Local repo root, used to hash the compose files for the re-apply trigger."
type = string
}
variable "repo_url" {
description = "Git remote to clone from if the host has no checkout yet."
type = string
}
variable "app_port" {
description = "Port backend binds / is health-checked on."
type = number
}
variable "frontend_port" {
description = "Port the frontend SSR service binds / is health-checked on (repo-split Stage 4). Loopback-only, never opened in ufw -- reached via Caddy's path-split or backend's own reverse-proxy fallback, never directly."
type = number
default = 8080
}
# ---- Sizing -------------------------------------------------------------------
variable "workers" {
description = "uvicorn worker count (WORKERS)."
type = number
}
variable "app_cpus" {
description = "App container CPU cap (APP_CPUS)."
type = number
}
variable "db_cpus" {
description = "DB container CPU cap (DB_CPUS)."
type = number
}
variable "db_memory" {
description = "DB container memory cap (DB_MEMORY), e.g. \"8g\"."
type = string
}
Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235) * Split web/worker duties with THERMOGRAPH_ROLE Background work (the subscription notifier) is welded to the same process that serves requests, so scaling the web tier to N replicas would also scale notifier instances unless something restricts it further than leader election alone. Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process behavior). Every replica runs the same image; ROLE only gates whether a process is allowed to own the notifier at all, layered on top of the existing leader election: web replicas never start it even if they'd win leader election, worker replicas start it if they win. The decision is pulled into _should_run_notifier() so it's unit-testable without booting the full app (DB init, places index, neighbor warmer). Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE) so a worker replica - which serves no real traffic - still has something Swarm can health-check. * Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate Three changes toward the hop-1 interim cutover, all inert until Track B stands up the platform: docker-stack.yml: the Swarm stack file for the interim cutover, distinct from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a pre-built image (IMAGE_TAG) instead of building in place; app/worker publish no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted), reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard needs a smaller MTU or large payloads silently stall); db is placement- pinned to a labelled node; app/worker skip inline migrations (RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing that ever runs Alembic; secrets are real Swarm secrets mounted at /run/secrets, read by the entrypoint shim rather than plain env vars. TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default latest-pg18, today's behavior unchanged), wired through Terraform (timescaledb_tag, default "latest-pg18") so it can actually be pinned to an exact minor without hand-editing the host - required before any host of the stack could replicate with another (a floating tag risks mismatched extension minors, which blocks a physical replica and risks compressed- chunk corruption on restore). Caddy active health-gate: both the Terraform-rendered Caddyfile and the live deploy/Caddyfile now health-check the app on the same cheap /healthz route its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage, so it's cheap enough for a tight interval and works identically for a worker replica, which serves no public traffic at all) - Caddy won't forward into a container that's still booting or unhealthy. Verified live: built and booted the real image via docker compose - both containers report healthy via the new /healthz-based HEALTHCHECK, and GET / still renders the full SSR homepage unchanged. Both Caddyfiles validated with the real caddy binary. docker-stack.yml validated with docker compose config (required-var guards fire with clear messages; secrets correctly mount at /run/secrets/<name>, matching the entrypoint shim's mapping). docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside the existing openmeteo overlay. terraform validate + fmt clean.
2026-07-21 00:39:48 +00:00
variable "timescaledb_tag" {
description = "TimescaleDB image tag (TIMESCALEDB_TAG), e.g. \"2.17.2-pg18\". \"latest-pg18\" (the default) matches today's behavior; pin an exact minor before any host could ever replicate with another."
type = string
default = "latest-pg18"
}
# ---- Secrets rendered into /etc/thermograph.env -------------------------------
variable "postgres_password" {
type = string
sensitive = true
}
variable "auth_secret" {
type = string
sensitive = true
}
Have Terraform generate its own internal secrets, with sizing tiers (#239) Terraform generates the secrets that have no external meaning (POSTGRES_PASSWORD, AUTH_SECRET, METRICS_TOKEN, INDEXNOW_KEY) via the random provider instead of requiring the operator to hand-generate and paste each into terraform.tfvars. Each is pinned with a static keepers value (secrets.tf) so apply never regenerates a value already in use - the exact incident class this guards against: every session invalidated, the app<->DB password mismatched. Rotation is now a deliberate keepers edit, never a side effect. postgres_password/auth_secret move from required inputs to optional (default "") - explicit var wins when supplied (seeding an EXISTING live secret during a migration onto Terraform, hop-1 cutover runbook Stage 0), else Terraform generates and owns it. metrics_token/indexnow_key are new: neither existed in Terraform before, both previously left for the app's own fallback generation. VAPID deliberately stays a required, non-generated input - an EC keypair where regeneration breaks every existing push subscription outright, unlike an opaque token. Sizing tiers: a locals.sizes t-shirt map (nano/small/medium/large -> {workers, app_cpus, db_cpus, db_memory}), toward the target Proxmox sizing-tier model (architecture doc SS6) ahead of actually provisioning VMs - Proxmox itself stays deferred; today a tier just sizes container caps on the existing SSH-managed hosts. A host can reference one by name (hosts.<name>. size) or keep hand-picking the four fields, so existing tfvars are unaffected; prod's example now uses size = "large" (identical numbers), beta keeps explicit numbers, and a commented uat example demonstrates the shortcut for a future ephemeral host. Strengthened terraform/README.md's local-state caveat: more Terraform- generated secrets landing in tfstate raises the stakes of the existing never-commit-cleartext-state guidance, not just the sizing. Verified: terraform validate + fmt clean. A real `terraform plan` against fake hosts (prod/beta/uat, mixing size="large"/explicit-numbers/size="nano") resolved every sizing correctly (prod 8/8/4/16g, beta 4/4/2/8g, uat 1/1/1/1g) and planned exactly one instance of each random_password/random_id resource. Applied just those four resources (real generation, -target to avoid touching the fake SSH-only host resources) and re-planned: "No changes" - confirming the keepers pinning holds. Adding an explicit postgres_password override afterward left the random_password resource itself completely untouched (0 replace/destroy), confirming the override path never disturbs the generated resource.
2026-07-21 01:08:57 +00:00
variable "metrics_token" {
type = string
sensitive = true
default = ""
}
variable "indexnow_key" {
type = string
sensitive = true
default = ""
}
variable "vapid_private_key" {
type = string
sensitive = true
}
variable "vapid_public_key" {
type = string
sensitive = true
}
variable "vapid_contact" {
type = string
sensitive = true
}
variable "google_verify" {
type = string
sensitive = true
}
variable "bing_verify" {
type = string
sensitive = true
}
variable "mail_backend" {
type = string
}
variable "smtp_host" {
type = string
}
variable "smtp_port" {
type = string
}
variable "smtp_user" {
type = string
sensitive = true
}
variable "smtp_password" {
type = string
sensitive = true
}
variable "smtp_starttls" {
type = string
}
variable "mail_from" {
type = string
}
variable "mail_reply_to" {
type = string
}
variable "discord_webhook" {
type = string
sensitive = true
}
variable "discord_public_key" {
type = string
}
variable "discord_app_id" {
type = string
}
variable "discord_bot_token" {
type = string
sensitive = true
}
variable "discord_client_secret" {
type = string
sensitive = true
}
variable "registry_token" {
type = string
sensitive = true
}