thermograph/infra/terraform/modules/thermograph-host/variables.tf

141 lines
4.5 KiB
Terraform
Raw Normal View History

# Per-host inputs (all supplied by the root module's for_each).
variable "name" {
description = "Short host key (e.g. \"prod\", \"dev\"), used in log lines."
type = string
}
variable "host" {
description = "IP or hostname to SSH to."
type = string
}
variable "ssh_user" {
description = "SSH login user (must be able to sudo)."
type = string
}
variable "ssh_private_key_path" {
description = "Path to the private key file for ssh_user."
type = string
}
variable "role" {
description = "\"prod\" | \"beta\" | \"dev\" — informational. One module instance is one (host, environment) pair, so vps2 (which runs both prod and beta) gets two instances, each with its own role/app_dir/image tags -- see the root module's `hosts` variable."
type = string
}
variable "git_branch" {
description = "This INFRA repo's branch the host checkout is reset to (independent of which app images are deployed — see backend_image_tag / frontend_image_tag)."
Decouple Terraform from the app repo; add a GCP host scaffold Content-change pass following the extraction from the app monorepo (this repo now stands alone, sourced via git filter-repo to preserve history): - terraform/variables.tf, secrets.tf, modules/thermograph-host: remove every app-secret Terraform variable (postgres_password, auth_secret, VAPID keys, registry_token, Discord/SMTP creds, ...) and the random_password/random_id generators. The SOPS+age vault (deploy/secrets/*.yaml) is now the sole source of app secrets, rendered at deploy time by deploy/render-secrets.sh; Terraform renders only a non-secret /etc/thermograph-topology.env (sizing, routing) via the renamed thermograph-topology.env.tftpl template. - hosts gains a required app_image_tag field: the host's own checkout is now this infra repo, not the app repo, so there is no "current commit" to derive an image tag from — every host pins one explicitly. repo_url now points at this repo (private; typically needs an embedded read token). - deploy.sh: IMAGE_TAG is now required from the environment instead of derived via `git rev-parse HEAD` of the (now infra-repo) checkout, which would have silently resolved to the wrong or a nonexistent tag. - New terraform/modules/gcp-host: creates a GCE VM + minimal VPC/firewall only, then feeds its IP into the same thermograph-host module every SSH-managed host already uses — one provisioning path regardless of how a host came to exist. var.gcp_hosts defaults to {}, so no google_* resource is planned and the provider is never invoked without it (verified: plan and validate succeed with no GCP credentials configured). - terraform/README.md, ACCESS.md (renamed from INFRA.md), README.md: updated for the new secrets model, the GCP scaffold, and this repo's own identity. Verified: terraform fmt/validate/init clean; plan succeeds against realistic dummy hosts (prod+beta shape) and against a populated gcp_hosts entry (plans 6 resources with no live credentials, confirming the composition wires correctly end to end).
2026-07-22 04:46:05 +00:00
type = string
}
variable "backend_image_tag" {
description = "Backend image tag to pull (emi/thermograph-backend/app), e.g. \"sha-<12 hex>\" (build-push.yml's tag for the backend-repo commit) or a semver tag. The host has no app-repo checkout to derive this from, so it's always explicit."
type = string
}
variable "frontend_image_tag" {
description = "Frontend image tag to pull (emi/thermograph-frontend/app), e.g. \"sha-<12 hex>\" (build-push.yml's tag for the frontend-repo commit) or a semver tag. Always explicit, same as backend_image_tag."
type = string
}
variable "domain" {
description = "Public domain. \"\" => no Caddy/TLS (open the app port instead)."
type = string
}
variable "compose_files" {
description = "Compose files to layer, in order (dev appends docker-compose.dev.yml)."
type = list(string)
}
variable "openmeteo" {
description = "Self-host the ERA5 archive: layer docker-compose.openmeteo.yml + provision the host rclone mount."
type = bool
default = false
}
variable "om_data_dir" {
description = "Host rclone mount point for the archive bucket (OM_DATA_DIR the overlay bind-mounts)."
type = string
default = "/mnt/om-archive"
}
variable "om_bucket_remote" {
description = "rclone remote:path for the archive bucket (mounted at om_data_dir)."
type = string
default = ""
}
variable "om_rclone_conf" {
description = "rclone.conf contents installed to /etc/rclone/rclone.conf. Sensitive."
type = string
default = ""
sensitive = true
}
variable "om_vfs_cache_max" {
description = "rclone --vfs-cache-max-size for the mount's on-disk hot cache."
type = string
default = "80G"
}
variable "app_dir" {
description = "Checkout path on the host."
type = string
}
variable "repo_root" {
description = "Local repo root, used to hash the compose files for the re-apply trigger."
type = string
}
variable "repo_url" {
description = "Git remote to clone from if the host has no checkout yet."
type = string
}
variable "app_port" {
description = "Port backend binds / is health-checked on."
type = number
}
variable "frontend_port" {
description = "Port the frontend SSR service binds / is health-checked on (repo-split Stage 4). Loopback-only, never opened in ufw -- reached via Caddy's path-split or backend's own reverse-proxy fallback, never directly."
type = number
default = 8080
}
# ---- Sizing -------------------------------------------------------------------
variable "workers" {
description = "uvicorn worker count (WORKERS)."
type = number
}
variable "app_cpus" {
description = "App container CPU cap (APP_CPUS)."
type = number
}
variable "db_cpus" {
description = "DB container CPU cap (DB_CPUS)."
type = number
}
variable "db_memory" {
description = "DB container memory cap (DB_MEMORY), e.g. \"8g\"."
type = string
}
Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235) * Split web/worker duties with THERMOGRAPH_ROLE Background work (the subscription notifier) is welded to the same process that serves requests, so scaling the web tier to N replicas would also scale notifier instances unless something restricts it further than leader election alone. Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process behavior). Every replica runs the same image; ROLE only gates whether a process is allowed to own the notifier at all, layered on top of the existing leader election: web replicas never start it even if they'd win leader election, worker replicas start it if they win. The decision is pulled into _should_run_notifier() so it's unit-testable without booting the full app (DB init, places index, neighbor warmer). Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE) so a worker replica - which serves no real traffic - still has something Swarm can health-check. * Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate Three changes toward the hop-1 interim cutover, all inert until Track B stands up the platform: docker-stack.yml: the Swarm stack file for the interim cutover, distinct from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a pre-built image (IMAGE_TAG) instead of building in place; app/worker publish no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted), reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard needs a smaller MTU or large payloads silently stall); db is placement- pinned to a labelled node; app/worker skip inline migrations (RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing that ever runs Alembic; secrets are real Swarm secrets mounted at /run/secrets, read by the entrypoint shim rather than plain env vars. TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default latest-pg18, today's behavior unchanged), wired through Terraform (timescaledb_tag, default "latest-pg18") so it can actually be pinned to an exact minor without hand-editing the host - required before any host of the stack could replicate with another (a floating tag risks mismatched extension minors, which blocks a physical replica and risks compressed- chunk corruption on restore). Caddy active health-gate: both the Terraform-rendered Caddyfile and the live deploy/Caddyfile now health-check the app on the same cheap /healthz route its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage, so it's cheap enough for a tight interval and works identically for a worker replica, which serves no public traffic at all) - Caddy won't forward into a container that's still booting or unhealthy. Verified live: built and booted the real image via docker compose - both containers report healthy via the new /healthz-based HEALTHCHECK, and GET / still renders the full SSR homepage unchanged. Both Caddyfiles validated with the real caddy binary. docker-stack.yml validated with docker compose config (required-var guards fire with clear messages; secrets correctly mount at /run/secrets/<name>, matching the entrypoint shim's mapping). docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside the existing openmeteo overlay. terraform validate + fmt clean.
2026-07-21 00:39:48 +00:00
variable "timescaledb_tag" {
description = "TimescaleDB image tag (TIMESCALEDB_TAG), e.g. \"2.17.2-pg18\". \"latest-pg18\" (the default) matches today's behavior; pin an exact minor before any host could ever replicate with another."
type = string
default = "latest-pg18"
}
Decouple Terraform from the app repo; add a GCP host scaffold Content-change pass following the extraction from the app monorepo (this repo now stands alone, sourced via git filter-repo to preserve history): - terraform/variables.tf, secrets.tf, modules/thermograph-host: remove every app-secret Terraform variable (postgres_password, auth_secret, VAPID keys, registry_token, Discord/SMTP creds, ...) and the random_password/random_id generators. The SOPS+age vault (deploy/secrets/*.yaml) is now the sole source of app secrets, rendered at deploy time by deploy/render-secrets.sh; Terraform renders only a non-secret /etc/thermograph-topology.env (sizing, routing) via the renamed thermograph-topology.env.tftpl template. - hosts gains a required app_image_tag field: the host's own checkout is now this infra repo, not the app repo, so there is no "current commit" to derive an image tag from — every host pins one explicitly. repo_url now points at this repo (private; typically needs an embedded read token). - deploy.sh: IMAGE_TAG is now required from the environment instead of derived via `git rev-parse HEAD` of the (now infra-repo) checkout, which would have silently resolved to the wrong or a nonexistent tag. - New terraform/modules/gcp-host: creates a GCE VM + minimal VPC/firewall only, then feeds its IP into the same thermograph-host module every SSH-managed host already uses — one provisioning path regardless of how a host came to exist. var.gcp_hosts defaults to {}, so no google_* resource is planned and the provider is never invoked without it (verified: plan and validate succeed with no GCP credentials configured). - terraform/README.md, ACCESS.md (renamed from INFRA.md), README.md: updated for the new secrets model, the GCP scaffold, and this repo's own identity. Verified: terraform fmt/validate/init clean; plan succeeds against realistic dummy hosts (prod+beta shape) and against a populated gcp_hosts entry (plans 6 resources with no live credentials, confirming the composition wires correctly end to end).
2026-07-22 04:46:05 +00:00
# Secrets (POSTGRES_PASSWORD, THERMOGRAPH_AUTH_SECRET, VAPID keys, REGISTRY_TOKEN,
# Discord/SMTP credentials, ...) are no longer Terraform variables -- they're
# rendered at deploy time from the SOPS+age vault (deploy/secrets/*.yaml) by
# deploy/render-secrets.sh, which deploy.sh calls. See main.tf's remote-exec step 3.