thermograph/infra/terraform/main.tf

144 lines
6.1 KiB
Terraform
Raw Normal View History

Decouple Terraform from the app repo; add a GCP host scaffold Content-change pass following the extraction from the app monorepo (this repo now stands alone, sourced via git filter-repo to preserve history): - terraform/variables.tf, secrets.tf, modules/thermograph-host: remove every app-secret Terraform variable (postgres_password, auth_secret, VAPID keys, registry_token, Discord/SMTP creds, ...) and the random_password/random_id generators. The SOPS+age vault (deploy/secrets/*.yaml) is now the sole source of app secrets, rendered at deploy time by deploy/render-secrets.sh; Terraform renders only a non-secret /etc/thermograph-topology.env (sizing, routing) via the renamed thermograph-topology.env.tftpl template. - hosts gains a required app_image_tag field: the host's own checkout is now this infra repo, not the app repo, so there is no "current commit" to derive an image tag from — every host pins one explicitly. repo_url now points at this repo (private; typically needs an embedded read token). - deploy.sh: IMAGE_TAG is now required from the environment instead of derived via `git rev-parse HEAD` of the (now infra-repo) checkout, which would have silently resolved to the wrong or a nonexistent tag. - New terraform/modules/gcp-host: creates a GCE VM + minimal VPC/firewall only, then feeds its IP into the same thermograph-host module every SSH-managed host already uses — one provisioning path regardless of how a host came to exist. var.gcp_hosts defaults to {}, so no google_* resource is planned and the provider is never invoked without it (verified: plan and validate succeed with no GCP credentials configured). - terraform/README.md, ACCESS.md (renamed from INFRA.md), README.md: updated for the new secrets model, the GCP scaffold, and this repo's own identity. Verified: terraform fmt/validate/init clean; plan succeeds against realistic dummy hosts (prod+beta shape) and against a populated gcp_hosts entry (plans 6 resources with no live credentials, confirming the composition wires correctly end to end).
2026-07-22 04:46:05 +00:00
# No project/region default here on purpose: each var.gcp_hosts entry carries its
# own project + zone (a fleet could span projects), and with var.gcp_hosts empty
# this provider is never invoked, so there's nothing to default. When you do
# populate var.gcp_hosts, authenticate via Application Default Credentials
# (`gcloud auth application-default login`) or GOOGLE_APPLICATION_CREDENTIALS --
# no credentials are configured in this repo.
provider "google" {}
locals {
# Repo root (one level above this terraform/ dir). The module hashes the compose
# files here so a compose change re-triggers the remote deploy, and this is the
# tree the host's checkout mirrors over git.
repo_root = abspath("${path.root}/..")
Have Terraform generate its own internal secrets, with sizing tiers (#239) Terraform generates the secrets that have no external meaning (POSTGRES_PASSWORD, AUTH_SECRET, METRICS_TOKEN, INDEXNOW_KEY) via the random provider instead of requiring the operator to hand-generate and paste each into terraform.tfvars. Each is pinned with a static keepers value (secrets.tf) so apply never regenerates a value already in use - the exact incident class this guards against: every session invalidated, the app<->DB password mismatched. Rotation is now a deliberate keepers edit, never a side effect. postgres_password/auth_secret move from required inputs to optional (default "") - explicit var wins when supplied (seeding an EXISTING live secret during a migration onto Terraform, hop-1 cutover runbook Stage 0), else Terraform generates and owns it. metrics_token/indexnow_key are new: neither existed in Terraform before, both previously left for the app's own fallback generation. VAPID deliberately stays a required, non-generated input - an EC keypair where regeneration breaks every existing push subscription outright, unlike an opaque token. Sizing tiers: a locals.sizes t-shirt map (nano/small/medium/large -> {workers, app_cpus, db_cpus, db_memory}), toward the target Proxmox sizing-tier model (architecture doc SS6) ahead of actually provisioning VMs - Proxmox itself stays deferred; today a tier just sizes container caps on the existing SSH-managed hosts. A host can reference one by name (hosts.<name>. size) or keep hand-picking the four fields, so existing tfvars are unaffected; prod's example now uses size = "large" (identical numbers), beta keeps explicit numbers, and a commented uat example demonstrates the shortcut for a future ephemeral host. Strengthened terraform/README.md's local-state caveat: more Terraform- generated secrets landing in tfstate raises the stakes of the existing never-commit-cleartext-state guidance, not just the sizing. Verified: terraform validate + fmt clean. A real `terraform plan` against fake hosts (prod/beta/uat, mixing size="large"/explicit-numbers/size="nano") resolved every sizing correctly (prod 8/8/4/16g, beta 4/4/2/8g, uat 1/1/1/1g) and planned exactly one instance of each random_password/random_id resource. Applied just those four resources (real generation, -target to avoid touching the fake SSH-only host resources) and re-planned: "No changes" - confirming the keepers pinning holds. Adding an explicit postgres_password override afterward left the random_password resource itself completely untouched (0 replace/destroy), confirming the override path never disturbs the generated resource.
2026-07-21 01:08:57 +00:00
# Reusable size presets — a host can reference one by name (hosts.<name>.size)
# instead of hand-picking workers/app_cpus/db_cpus/db_memory separately. Mirrors
# the target Proxmox sizing-tier model (architecture doc §6: nano/small/medium/
# large -> {cores, mem}) ahead of actually provisioning VMs — Proxmox itself is
# deferred; today these tiers just size the container caps on the existing
# SSH-managed VPS hosts. "large" matches the current 48 GB prod box's numbers.
sizes = {
nano = { workers = 1, app_cpus = 1, db_cpus = 1, db_memory = "1g" }
small = { workers = 4, app_cpus = 4, db_cpus = 2, db_memory = "8g" }
medium = { workers = 6, app_cpus = 6, db_cpus = 3, db_memory = "12g" }
large = { workers = 8, app_cpus = 8, db_cpus = 4, db_memory = "16g" }
}
Decouple Terraform from the app repo; add a GCP host scaffold Content-change pass following the extraction from the app monorepo (this repo now stands alone, sourced via git filter-repo to preserve history): - terraform/variables.tf, secrets.tf, modules/thermograph-host: remove every app-secret Terraform variable (postgres_password, auth_secret, VAPID keys, registry_token, Discord/SMTP creds, ...) and the random_password/random_id generators. The SOPS+age vault (deploy/secrets/*.yaml) is now the sole source of app secrets, rendered at deploy time by deploy/render-secrets.sh; Terraform renders only a non-secret /etc/thermograph-topology.env (sizing, routing) via the renamed thermograph-topology.env.tftpl template. - hosts gains a required app_image_tag field: the host's own checkout is now this infra repo, not the app repo, so there is no "current commit" to derive an image tag from — every host pins one explicitly. repo_url now points at this repo (private; typically needs an embedded read token). - deploy.sh: IMAGE_TAG is now required from the environment instead of derived via `git rev-parse HEAD` of the (now infra-repo) checkout, which would have silently resolved to the wrong or a nonexistent tag. - New terraform/modules/gcp-host: creates a GCE VM + minimal VPC/firewall only, then feeds its IP into the same thermograph-host module every SSH-managed host already uses — one provisioning path regardless of how a host came to exist. var.gcp_hosts defaults to {}, so no google_* resource is planned and the provider is never invoked without it (verified: plan and validate succeed with no GCP credentials configured). - terraform/README.md, ACCESS.md (renamed from INFRA.md), README.md: updated for the new secrets model, the GCP scaffold, and this repo's own identity. Verified: terraform fmt/validate/init clean; plan succeeds against realistic dummy hosts (prod+beta shape) and against a populated gcp_hosts entry (plans 6 resources with no live credentials, confirming the composition wires correctly end to end).
2026-07-22 04:46:05 +00:00
}
# GCP-created hosts (empty by default — see variables.tf and modules/gcp-host).
# Creates only the VM + minimal networking; every provisioning behavior comes from
# the SAME thermograph-host module every SSH-managed host below uses, via
# local.all_hosts merging its output IP in below. No live resources exist until
# var.gcp_hosts is populated.
module "gcp_vm" {
source = "./modules/gcp-host"
for_each = var.gcp_hosts
name = each.key
project = each.value.project
zone = each.value.zone
machine_type = each.value.machine_type
ssh_user = each.value.ssh_user
ssh_public_key_path = each.value.ssh_public_key_path
}
Have Terraform generate its own internal secrets, with sizing tiers (#239) Terraform generates the secrets that have no external meaning (POSTGRES_PASSWORD, AUTH_SECRET, METRICS_TOKEN, INDEXNOW_KEY) via the random provider instead of requiring the operator to hand-generate and paste each into terraform.tfvars. Each is pinned with a static keepers value (secrets.tf) so apply never regenerates a value already in use - the exact incident class this guards against: every session invalidated, the app<->DB password mismatched. Rotation is now a deliberate keepers edit, never a side effect. postgres_password/auth_secret move from required inputs to optional (default "") - explicit var wins when supplied (seeding an EXISTING live secret during a migration onto Terraform, hop-1 cutover runbook Stage 0), else Terraform generates and owns it. metrics_token/indexnow_key are new: neither existed in Terraform before, both previously left for the app's own fallback generation. VAPID deliberately stays a required, non-generated input - an EC keypair where regeneration breaks every existing push subscription outright, unlike an opaque token. Sizing tiers: a locals.sizes t-shirt map (nano/small/medium/large -> {workers, app_cpus, db_cpus, db_memory}), toward the target Proxmox sizing-tier model (architecture doc SS6) ahead of actually provisioning VMs - Proxmox itself stays deferred; today a tier just sizes container caps on the existing SSH-managed hosts. A host can reference one by name (hosts.<name>. size) or keep hand-picking the four fields, so existing tfvars are unaffected; prod's example now uses size = "large" (identical numbers), beta keeps explicit numbers, and a commented uat example demonstrates the shortcut for a future ephemeral host. Strengthened terraform/README.md's local-state caveat: more Terraform- generated secrets landing in tfstate raises the stakes of the existing never-commit-cleartext-state guidance, not just the sizing. Verified: terraform validate + fmt clean. A real `terraform plan` against fake hosts (prod/beta/uat, mixing size="large"/explicit-numbers/size="nano") resolved every sizing correctly (prod 8/8/4/16g, beta 4/4/2/8g, uat 1/1/1/1g) and planned exactly one instance of each random_password/random_id resource. Applied just those four resources (real generation, -target to avoid touching the fake SSH-only host resources) and re-planned: "No changes" - confirming the keepers pinning holds. Adding an explicit postgres_password override afterward left the random_password resource itself completely untouched (0 replace/destroy), confirming the override path never disturbs the generated resource.
2026-07-21 01:08:57 +00:00
Decouple Terraform from the app repo; add a GCP host scaffold Content-change pass following the extraction from the app monorepo (this repo now stands alone, sourced via git filter-repo to preserve history): - terraform/variables.tf, secrets.tf, modules/thermograph-host: remove every app-secret Terraform variable (postgres_password, auth_secret, VAPID keys, registry_token, Discord/SMTP creds, ...) and the random_password/random_id generators. The SOPS+age vault (deploy/secrets/*.yaml) is now the sole source of app secrets, rendered at deploy time by deploy/render-secrets.sh; Terraform renders only a non-secret /etc/thermograph-topology.env (sizing, routing) via the renamed thermograph-topology.env.tftpl template. - hosts gains a required app_image_tag field: the host's own checkout is now this infra repo, not the app repo, so there is no "current commit" to derive an image tag from — every host pins one explicitly. repo_url now points at this repo (private; typically needs an embedded read token). - deploy.sh: IMAGE_TAG is now required from the environment instead of derived via `git rev-parse HEAD` of the (now infra-repo) checkout, which would have silently resolved to the wrong or a nonexistent tag. - New terraform/modules/gcp-host: creates a GCE VM + minimal VPC/firewall only, then feeds its IP into the same thermograph-host module every SSH-managed host already uses — one provisioning path regardless of how a host came to exist. var.gcp_hosts defaults to {}, so no google_* resource is planned and the provider is never invoked without it (verified: plan and validate succeed with no GCP credentials configured). - terraform/README.md, ACCESS.md (renamed from INFRA.md), README.md: updated for the new secrets model, the GCP scaffold, and this repo's own identity. Verified: terraform fmt/validate/init clean; plan succeeds against realistic dummy hosts (prod+beta shape) and against a populated gcp_hosts entry (plans 6 resources with no live credentials, confirming the composition wires correctly end to end).
2026-07-22 04:46:05 +00:00
locals {
# ONE MODULE INSTANCE PER (HOST, ENVIRONMENT) PAIR, not per host.
#
# var.hosts is keyed by machine (vps1, vps2) and each machine carries an
# `environments` map, because vps2 runs prod AND beta. The module below
# provisions an ENVIRONMENT — a checkout, a branch, image tags, a Caddy site,
# container sizing — so it needs one instance per environment, with the two
# instances on vps2 sharing that machine's SSH identity.
#
# Flattened to keys like "vps2-prod" and "vps2-beta". Keeping them distinct
# module instances is what makes `app_dir` (required, no default in
# variables.tf) do its job: two environments on one box get two explicitly
# different checkouts, so neither apply can `git reset --hard` the other's
# tree.
ssh_environments = merge([
for hname, h in var.hosts : {
for ename, e in h.environments :
"${hname}-${ename}" => merge(e, {
host = h.host
ssh_user = h.ssh_user
ssh_private_key_path = h.ssh_private_key_path
})
}
]...)
# Every environment this config manages, SSH-only (flattened above) plus
# GCP-created (whose `host` is filled in from the VM Terraform just created) —
# one unified map so a single `module.host` for_each below handles both without
# duplicating any provisioning logic. GCP hosts stay one-entry-per-machine:
# nothing creates two environments on a created VM today, and the variable
# keeps its flat shape. They always use a named size tier (simpler than
Decouple Terraform from the app repo; add a GCP host scaffold Content-change pass following the extraction from the app monorepo (this repo now stands alone, sourced via git filter-repo to preserve history): - terraform/variables.tf, secrets.tf, modules/thermograph-host: remove every app-secret Terraform variable (postgres_password, auth_secret, VAPID keys, registry_token, Discord/SMTP creds, ...) and the random_password/random_id generators. The SOPS+age vault (deploy/secrets/*.yaml) is now the sole source of app secrets, rendered at deploy time by deploy/render-secrets.sh; Terraform renders only a non-secret /etc/thermograph-topology.env (sizing, routing) via the renamed thermograph-topology.env.tftpl template. - hosts gains a required app_image_tag field: the host's own checkout is now this infra repo, not the app repo, so there is no "current commit" to derive an image tag from — every host pins one explicitly. repo_url now points at this repo (private; typically needs an embedded read token). - deploy.sh: IMAGE_TAG is now required from the environment instead of derived via `git rev-parse HEAD` of the (now infra-repo) checkout, which would have silently resolved to the wrong or a nonexistent tag. - New terraform/modules/gcp-host: creates a GCE VM + minimal VPC/firewall only, then feeds its IP into the same thermograph-host module every SSH-managed host already uses — one provisioning path regardless of how a host came to exist. var.gcp_hosts defaults to {}, so no google_* resource is planned and the provider is never invoked without it (verified: plan and validate succeed with no GCP credentials configured). - terraform/README.md, ACCESS.md (renamed from INFRA.md), README.md: updated for the new secrets model, the GCP scaffold, and this repo's own identity. Verified: terraform fmt/validate/init clean; plan succeeds against realistic dummy hosts (prod+beta shape) and against a populated gcp_hosts entry (plans 6 resources with no live credentials, confirming the composition wires correctly end to end).
2026-07-22 04:46:05 +00:00
# exposing the four raw sizing fields on that variable too).
all_hosts = merge(
local.ssh_environments,
Decouple Terraform from the app repo; add a GCP host scaffold Content-change pass following the extraction from the app monorepo (this repo now stands alone, sourced via git filter-repo to preserve history): - terraform/variables.tf, secrets.tf, modules/thermograph-host: remove every app-secret Terraform variable (postgres_password, auth_secret, VAPID keys, registry_token, Discord/SMTP creds, ...) and the random_password/random_id generators. The SOPS+age vault (deploy/secrets/*.yaml) is now the sole source of app secrets, rendered at deploy time by deploy/render-secrets.sh; Terraform renders only a non-secret /etc/thermograph-topology.env (sizing, routing) via the renamed thermograph-topology.env.tftpl template. - hosts gains a required app_image_tag field: the host's own checkout is now this infra repo, not the app repo, so there is no "current commit" to derive an image tag from — every host pins one explicitly. repo_url now points at this repo (private; typically needs an embedded read token). - deploy.sh: IMAGE_TAG is now required from the environment instead of derived via `git rev-parse HEAD` of the (now infra-repo) checkout, which would have silently resolved to the wrong or a nonexistent tag. - New terraform/modules/gcp-host: creates a GCE VM + minimal VPC/firewall only, then feeds its IP into the same thermograph-host module every SSH-managed host already uses — one provisioning path regardless of how a host came to exist. var.gcp_hosts defaults to {}, so no google_* resource is planned and the provider is never invoked without it (verified: plan and validate succeed with no GCP credentials configured). - terraform/README.md, ACCESS.md (renamed from INFRA.md), README.md: updated for the new secrets model, the GCP scaffold, and this repo's own identity. Verified: terraform fmt/validate/init clean; plan succeeds against realistic dummy hosts (prod+beta shape) and against a populated gcp_hosts entry (plans 6 resources with no live credentials, confirming the composition wires correctly end to end).
2026-07-22 04:46:05 +00:00
{
for name, h in var.gcp_hosts : name => merge(h, {
host = module.gcp_vm[name].external_ip
workers = null
app_cpus = null
db_cpus = null
db_memory = null
openmeteo = false
om_data_dir = "/mnt/om-archive"
})
}
)
# Per-host resolved sizing: a named tier (host.size) wins when set; otherwise the
# host's own workers/app_cpus/db_cpus/db_memory fields (each individually
Have Terraform generate its own internal secrets, with sizing tiers (#239) Terraform generates the secrets that have no external meaning (POSTGRES_PASSWORD, AUTH_SECRET, METRICS_TOKEN, INDEXNOW_KEY) via the random provider instead of requiring the operator to hand-generate and paste each into terraform.tfvars. Each is pinned with a static keepers value (secrets.tf) so apply never regenerates a value already in use - the exact incident class this guards against: every session invalidated, the app<->DB password mismatched. Rotation is now a deliberate keepers edit, never a side effect. postgres_password/auth_secret move from required inputs to optional (default "") - explicit var wins when supplied (seeding an EXISTING live secret during a migration onto Terraform, hop-1 cutover runbook Stage 0), else Terraform generates and owns it. metrics_token/indexnow_key are new: neither existed in Terraform before, both previously left for the app's own fallback generation. VAPID deliberately stays a required, non-generated input - an EC keypair where regeneration breaks every existing push subscription outright, unlike an opaque token. Sizing tiers: a locals.sizes t-shirt map (nano/small/medium/large -> {workers, app_cpus, db_cpus, db_memory}), toward the target Proxmox sizing-tier model (architecture doc SS6) ahead of actually provisioning VMs - Proxmox itself stays deferred; today a tier just sizes container caps on the existing SSH-managed hosts. A host can reference one by name (hosts.<name>. size) or keep hand-picking the four fields, so existing tfvars are unaffected; prod's example now uses size = "large" (identical numbers), beta keeps explicit numbers, and a commented uat example demonstrates the shortcut for a future ephemeral host. Strengthened terraform/README.md's local-state caveat: more Terraform- generated secrets landing in tfstate raises the stakes of the existing never-commit-cleartext-state guidance, not just the sizing. Verified: terraform validate + fmt clean. A real `terraform plan` against fake hosts (prod/beta/uat, mixing size="large"/explicit-numbers/size="nano") resolved every sizing correctly (prod 8/8/4/16g, beta 4/4/2/8g, uat 1/1/1/1g) and planned exactly one instance of each random_password/random_id resource. Applied just those four resources (real generation, -target to avoid touching the fake SSH-only host resources) and re-planned: "No changes" - confirming the keepers pinning holds. Adding an explicit postgres_password override afterward left the random_password resource itself completely untouched (0 replace/destroy), confirming the override path never disturbs the generated resource.
2026-07-21 01:08:57 +00:00
# defaulted in variables.tf) — so existing tfvars with explicit numbers are
# unaffected, and a tier is purely an opt-in shortcut.
host_sizing = {
Decouple Terraform from the app repo; add a GCP host scaffold Content-change pass following the extraction from the app monorepo (this repo now stands alone, sourced via git filter-repo to preserve history): - terraform/variables.tf, secrets.tf, modules/thermograph-host: remove every app-secret Terraform variable (postgres_password, auth_secret, VAPID keys, registry_token, Discord/SMTP creds, ...) and the random_password/random_id generators. The SOPS+age vault (deploy/secrets/*.yaml) is now the sole source of app secrets, rendered at deploy time by deploy/render-secrets.sh; Terraform renders only a non-secret /etc/thermograph-topology.env (sizing, routing) via the renamed thermograph-topology.env.tftpl template. - hosts gains a required app_image_tag field: the host's own checkout is now this infra repo, not the app repo, so there is no "current commit" to derive an image tag from — every host pins one explicitly. repo_url now points at this repo (private; typically needs an embedded read token). - deploy.sh: IMAGE_TAG is now required from the environment instead of derived via `git rev-parse HEAD` of the (now infra-repo) checkout, which would have silently resolved to the wrong or a nonexistent tag. - New terraform/modules/gcp-host: creates a GCE VM + minimal VPC/firewall only, then feeds its IP into the same thermograph-host module every SSH-managed host already uses — one provisioning path regardless of how a host came to exist. var.gcp_hosts defaults to {}, so no google_* resource is planned and the provider is never invoked without it (verified: plan and validate succeed with no GCP credentials configured). - terraform/README.md, ACCESS.md (renamed from INFRA.md), README.md: updated for the new secrets model, the GCP scaffold, and this repo's own identity. Verified: terraform fmt/validate/init clean; plan succeeds against realistic dummy hosts (prod+beta shape) and against a populated gcp_hosts entry (plans 6 resources with no live credentials, confirming the composition wires correctly end to end).
2026-07-22 04:46:05 +00:00
for name, h in local.all_hosts : name => h.size != null ? local.sizes[h.size] : {
Have Terraform generate its own internal secrets, with sizing tiers (#239) Terraform generates the secrets that have no external meaning (POSTGRES_PASSWORD, AUTH_SECRET, METRICS_TOKEN, INDEXNOW_KEY) via the random provider instead of requiring the operator to hand-generate and paste each into terraform.tfvars. Each is pinned with a static keepers value (secrets.tf) so apply never regenerates a value already in use - the exact incident class this guards against: every session invalidated, the app<->DB password mismatched. Rotation is now a deliberate keepers edit, never a side effect. postgres_password/auth_secret move from required inputs to optional (default "") - explicit var wins when supplied (seeding an EXISTING live secret during a migration onto Terraform, hop-1 cutover runbook Stage 0), else Terraform generates and owns it. metrics_token/indexnow_key are new: neither existed in Terraform before, both previously left for the app's own fallback generation. VAPID deliberately stays a required, non-generated input - an EC keypair where regeneration breaks every existing push subscription outright, unlike an opaque token. Sizing tiers: a locals.sizes t-shirt map (nano/small/medium/large -> {workers, app_cpus, db_cpus, db_memory}), toward the target Proxmox sizing-tier model (architecture doc SS6) ahead of actually provisioning VMs - Proxmox itself stays deferred; today a tier just sizes container caps on the existing SSH-managed hosts. A host can reference one by name (hosts.<name>. size) or keep hand-picking the four fields, so existing tfvars are unaffected; prod's example now uses size = "large" (identical numbers), beta keeps explicit numbers, and a commented uat example demonstrates the shortcut for a future ephemeral host. Strengthened terraform/README.md's local-state caveat: more Terraform- generated secrets landing in tfstate raises the stakes of the existing never-commit-cleartext-state guidance, not just the sizing. Verified: terraform validate + fmt clean. A real `terraform plan` against fake hosts (prod/beta/uat, mixing size="large"/explicit-numbers/size="nano") resolved every sizing correctly (prod 8/8/4/16g, beta 4/4/2/8g, uat 1/1/1/1g) and planned exactly one instance of each random_password/random_id resource. Applied just those four resources (real generation, -target to avoid touching the fake SSH-only host resources) and re-planned: "No changes" - confirming the keepers pinning holds. Adding an explicit postgres_password override afterward left the random_password resource itself completely untouched (0 replace/destroy), confirming the override path never disturbs the generated resource.
2026-07-21 01:08:57 +00:00
workers = h.workers
app_cpus = h.app_cpus
db_cpus = h.db_cpus
db_memory = h.db_memory
}
}
}
Decouple Terraform from the app repo; add a GCP host scaffold Content-change pass following the extraction from the app monorepo (this repo now stands alone, sourced via git filter-repo to preserve history): - terraform/variables.tf, secrets.tf, modules/thermograph-host: remove every app-secret Terraform variable (postgres_password, auth_secret, VAPID keys, registry_token, Discord/SMTP creds, ...) and the random_password/random_id generators. The SOPS+age vault (deploy/secrets/*.yaml) is now the sole source of app secrets, rendered at deploy time by deploy/render-secrets.sh; Terraform renders only a non-secret /etc/thermograph-topology.env (sizing, routing) via the renamed thermograph-topology.env.tftpl template. - hosts gains a required app_image_tag field: the host's own checkout is now this infra repo, not the app repo, so there is no "current commit" to derive an image tag from — every host pins one explicitly. repo_url now points at this repo (private; typically needs an embedded read token). - deploy.sh: IMAGE_TAG is now required from the environment instead of derived via `git rev-parse HEAD` of the (now infra-repo) checkout, which would have silently resolved to the wrong or a nonexistent tag. - New terraform/modules/gcp-host: creates a GCE VM + minimal VPC/firewall only, then feeds its IP into the same thermograph-host module every SSH-managed host already uses — one provisioning path regardless of how a host came to exist. var.gcp_hosts defaults to {}, so no google_* resource is planned and the provider is never invoked without it (verified: plan and validate succeed with no GCP credentials configured). - terraform/README.md, ACCESS.md (renamed from INFRA.md), README.md: updated for the new secrets model, the GCP scaffold, and this repo's own identity. Verified: terraform fmt/validate/init clean; plan succeeds against realistic dummy hosts (prod+beta shape) and against a populated gcp_hosts entry (plans 6 resources with no live credentials, confirming the composition wires correctly end to end).
2026-07-22 04:46:05 +00:00
# One module instance per host (SSH-managed or GCP-created — see local.all_hosts).
# The module is entirely SSH-provisioner driven: it configures an already-running
# VM (however it came to exist) and hands the app off to docker compose.
module "host" {
source = "./modules/thermograph-host"
Decouple Terraform from the app repo; add a GCP host scaffold Content-change pass following the extraction from the app monorepo (this repo now stands alone, sourced via git filter-repo to preserve history): - terraform/variables.tf, secrets.tf, modules/thermograph-host: remove every app-secret Terraform variable (postgres_password, auth_secret, VAPID keys, registry_token, Discord/SMTP creds, ...) and the random_password/random_id generators. The SOPS+age vault (deploy/secrets/*.yaml) is now the sole source of app secrets, rendered at deploy time by deploy/render-secrets.sh; Terraform renders only a non-secret /etc/thermograph-topology.env (sizing, routing) via the renamed thermograph-topology.env.tftpl template. - hosts gains a required app_image_tag field: the host's own checkout is now this infra repo, not the app repo, so there is no "current commit" to derive an image tag from — every host pins one explicitly. repo_url now points at this repo (private; typically needs an embedded read token). - deploy.sh: IMAGE_TAG is now required from the environment instead of derived via `git rev-parse HEAD` of the (now infra-repo) checkout, which would have silently resolved to the wrong or a nonexistent tag. - New terraform/modules/gcp-host: creates a GCE VM + minimal VPC/firewall only, then feeds its IP into the same thermograph-host module every SSH-managed host already uses — one provisioning path regardless of how a host came to exist. var.gcp_hosts defaults to {}, so no google_* resource is planned and the provider is never invoked without it (verified: plan and validate succeed with no GCP credentials configured). - terraform/README.md, ACCESS.md (renamed from INFRA.md), README.md: updated for the new secrets model, the GCP scaffold, and this repo's own identity. Verified: terraform fmt/validate/init clean; plan succeeds against realistic dummy hosts (prod+beta shape) and against a populated gcp_hosts entry (plans 6 resources with no live credentials, confirming the composition wires correctly end to end).
2026-07-22 04:46:05 +00:00
for_each = local.all_hosts
# Per-host config
name = each.key
host = each.value.host
ssh_user = each.value.ssh_user
ssh_private_key_path = each.value.ssh_private_key_path
role = each.value.role
git_branch = each.value.git_branch
backend_image_tag = each.value.backend_image_tag
frontend_image_tag = each.value.frontend_image_tag
domain = each.value.domain
compose_files = each.value.compose_files
app_dir = each.value.app_dir
Have Terraform generate its own internal secrets, with sizing tiers (#239) Terraform generates the secrets that have no external meaning (POSTGRES_PASSWORD, AUTH_SECRET, METRICS_TOKEN, INDEXNOW_KEY) via the random provider instead of requiring the operator to hand-generate and paste each into terraform.tfvars. Each is pinned with a static keepers value (secrets.tf) so apply never regenerates a value already in use - the exact incident class this guards against: every session invalidated, the app<->DB password mismatched. Rotation is now a deliberate keepers edit, never a side effect. postgres_password/auth_secret move from required inputs to optional (default "") - explicit var wins when supplied (seeding an EXISTING live secret during a migration onto Terraform, hop-1 cutover runbook Stage 0), else Terraform generates and owns it. metrics_token/indexnow_key are new: neither existed in Terraform before, both previously left for the app's own fallback generation. VAPID deliberately stays a required, non-generated input - an EC keypair where regeneration breaks every existing push subscription outright, unlike an opaque token. Sizing tiers: a locals.sizes t-shirt map (nano/small/medium/large -> {workers, app_cpus, db_cpus, db_memory}), toward the target Proxmox sizing-tier model (architecture doc SS6) ahead of actually provisioning VMs - Proxmox itself stays deferred; today a tier just sizes container caps on the existing SSH-managed hosts. A host can reference one by name (hosts.<name>. size) or keep hand-picking the four fields, so existing tfvars are unaffected; prod's example now uses size = "large" (identical numbers), beta keeps explicit numbers, and a commented uat example demonstrates the shortcut for a future ephemeral host. Strengthened terraform/README.md's local-state caveat: more Terraform- generated secrets landing in tfstate raises the stakes of the existing never-commit-cleartext-state guidance, not just the sizing. Verified: terraform validate + fmt clean. A real `terraform plan` against fake hosts (prod/beta/uat, mixing size="large"/explicit-numbers/size="nano") resolved every sizing correctly (prod 8/8/4/16g, beta 4/4/2/8g, uat 1/1/1/1g) and planned exactly one instance of each random_password/random_id resource. Applied just those four resources (real generation, -target to avoid touching the fake SSH-only host resources) and re-planned: "No changes" - confirming the keepers pinning holds. Adding an explicit postgres_password override afterward left the random_password resource itself completely untouched (0 replace/destroy), confirming the override path never disturbs the generated resource.
2026-07-21 01:08:57 +00:00
workers = local.host_sizing[each.key].workers
app_cpus = local.host_sizing[each.key].app_cpus
db_cpus = local.host_sizing[each.key].db_cpus
db_memory = local.host_sizing[each.key].db_memory
Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235) * Split web/worker duties with THERMOGRAPH_ROLE Background work (the subscription notifier) is welded to the same process that serves requests, so scaling the web tier to N replicas would also scale notifier instances unless something restricts it further than leader election alone. Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process behavior). Every replica runs the same image; ROLE only gates whether a process is allowed to own the notifier at all, layered on top of the existing leader election: web replicas never start it even if they'd win leader election, worker replicas start it if they win. The decision is pulled into _should_run_notifier() so it's unit-testable without booting the full app (DB init, places index, neighbor warmer). Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE) so a worker replica - which serves no real traffic - still has something Swarm can health-check. * Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate Three changes toward the hop-1 interim cutover, all inert until Track B stands up the platform: docker-stack.yml: the Swarm stack file for the interim cutover, distinct from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a pre-built image (IMAGE_TAG) instead of building in place; app/worker publish no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted), reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard needs a smaller MTU or large payloads silently stall); db is placement- pinned to a labelled node; app/worker skip inline migrations (RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing that ever runs Alembic; secrets are real Swarm secrets mounted at /run/secrets, read by the entrypoint shim rather than plain env vars. TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default latest-pg18, today's behavior unchanged), wired through Terraform (timescaledb_tag, default "latest-pg18") so it can actually be pinned to an exact minor without hand-editing the host - required before any host of the stack could replicate with another (a floating tag risks mismatched extension minors, which blocks a physical replica and risks compressed- chunk corruption on restore). Caddy active health-gate: both the Terraform-rendered Caddyfile and the live deploy/Caddyfile now health-check the app on the same cheap /healthz route its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage, so it's cheap enough for a tight interval and works identically for a worker replica, which serves no public traffic at all) - Caddy won't forward into a container that's still booting or unhealthy. Verified live: built and booted the real image via docker compose - both containers report healthy via the new /healthz-based HEALTHCHECK, and GET / still renders the full SSR homepage unchanged. Both Caddyfiles validated with the real caddy binary. docker-stack.yml validated with docker compose config (required-var guards fire with clear messages; secrets correctly mount at /run/secrets/<name>, matching the entrypoint shim's mapping). docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside the existing openmeteo overlay. terraform validate + fmt clean.
2026-07-21 00:39:48 +00:00
timescaledb_tag = each.value.timescaledb_tag
openmeteo = each.value.openmeteo
om_data_dir = each.value.om_data_dir
# Shared infra config
repo_root = local.repo_root
repo_url = var.repo_url
app_port = var.app_port
# Shared object-storage config (only used where openmeteo = true)
om_bucket_remote = var.om_bucket_remote
om_rclone_conf = var.om_rclone_conf
om_vfs_cache_max = var.om_vfs_cache_max
}