infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home #103
No reviewers
Labels
No labels
Compat/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference: Jinemi/thermograph#103
Loading…
Reference in a new issue
No description provided.
Delete branch "worktree-refactor+host-topology-vps1-vps2"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Re-architects the estate so machines are named by role and an environment stops implying a host. Nothing live has moved — the ordered cutover is
infra/deploy/RUNBOOK-vps1-vps2-cutover.md, and every step has a verification and a rollback.Mesh IPs don't move; what moves is which environment lives where. After the cutover, every "beta = 75.119.132.91" reference points at the wrong box — that address is vps1.
The core of it
infra/deploy/env-topology.shis the single source of truth: env → host, checkout, branch, deploy mode, stack name, env file, LB ports, DB role/database, service prefix.THERMOGRAPH_ENVis the input, and on vps2 it is the only thing distinguishing a beta deploy from a prod one — sodeploy.shrefuses to run when it disagrees with the checkout it was invoked from. The host-widesecrets-env/deploy-modemarkers survive only as a fallback for a single-environment box; they cannot describe vps2.Beta gets its own stack file rather than an overlay: a compose merge can add and override but cannot remove the
dbservice, and beta having no database of its own is the whole design. Its services are prefixedbeta-*because Swarm registers a service's short name as a DNS alias on every network it joins — two stacks both calling a servicewebon the shared data network would let prod's frontend resolve a beta task. Prod's stack file, env and LB config are untouched.One Postgres, two databases, two roles.
deploy/db/provision-env-db.shcreatesthermograph_beta(NOSUPERUSER, owns only its own database,CONNECTrevoked fromPUBLIC) plus a<role>_rofor ad-hoc queries — and refuses to run for the environment that owns the instance, since doing so would demote prod's bootstrap superuser.Things that would have broken silently
VPS1_SSH_*,VPS2_SSH_*).SSH_*meant "beta" and "the Forgejo box" because those shared a machine — the same conflation that once had the prod backup dumping beta.thermographwas dumped.hostfrom the node. On vps2 that would have filed every beta container, access log and JSON line asprod, feeding prod's alert rules with beta's traffic. It's now per-source, with a newnodelabel for the machine, and beta's log volume arrives via a vps2-only overlay (pointing it at dev's volume on vps1 would have manufactured fake beta data).ops/dbq.sh,ops/iceberg.shand the secrets seed scripts derive their SSH target and env-file path from the topology. The seed scripts in particular would have read prod's live env file when seeding beta.DEV_BIND_ADDR(loopback by default, the mesh address on vps1) instead of0.0.0.0— correct for a home LAN box, a public exposure of unreviewed branches on a VPS.Also
Terraform's
hostsis keyed by machine with a nestedenvironmentsmap, andmain.tfflattens(host, environment)pairs so two environments on one box can't share a checkout. Caddy's single reference config splits per host. Beta stops pinging IndexNow (it was asking search engines to index a rehearsal environment) and stops warming the city archive (shared upstream quota, cache only beta reads).Docs, onboarding and runbooks updated throughout — including the security rationales co-residency makes false: separate SSH credentials no longer buy a host boundary between beta and prod, and what remains is the database and the filesystem. The meaningful boundary is now vps1 vs vps2.
Verified
shellcheck -xclean over all 36 scripts (CI's pinned v0.11.0) · every workflow and stack file parses ·alloy validatewith CI's pinned v1.9.1 binary · the alerting structural checker reports 0 problems ·terraform validate+fmtclean, and the flattening evaluated against the example tfvars · vault re-encrypts and round-trips · the dev overlay renders127.0.0.1by default and10.10.0.2under the topology · the env/checkout mismatch guard refuses a crossed beta/prod deploy.Not verified (impossible without executing): anything about the live estate.
Companion docs PRs in
thermograph-docs: #10 (the decision record) and #11 (a pointer fromsystem-overview.md, which keeps describing the live shape until the cutover runs).Centralis' own estate description — the
onboardingtool, the server'sinstructionsstring, thethermograph-orientation/thermograph-opsskills and several tool descriptions — lives inemi/centralisand still states the old topology. The runbook lists exactly which strings need changing.Environments stop being machines. Until now each box WAS an environment -- "beta" named both a deploy target and a host, "the desktop" named both the operator's computer and the dev server -- so every path could assume one environment per host and hardcode /opt/thermograph, /etc/thermograph.env and ports 8137/8080. That assumption ends here: vps1 75.119.132.91 Forgejo, Grafana/Loki, the portfolio site, and DEV (own Postgres, mesh-only on 10.10.0.2:8137) vps2 169.58.46.181 PROD and BETA as two Swarm stacks sharing one TimescaleDB instance, plus Centralis, Postfix, backups desktop AI model hosting + flex Swarm capacity, no environment Nothing live has moved; the ordered cutover is in infra/deploy/RUNBOOK-vps1-vps2-cutover.md. deploy/env-topology.sh is the single source of truth: env -> host, checkout, branch, deploy mode, stack name, env file, LB ports, DB role/database, service prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing distinguishing a beta deploy from a prod one -- deploy.sh refuses to run when it disagrees with the checkout it was invoked from. The host-wide secrets-env and deploy-mode markers survive only as a fallback for a single-environment box. Beta gets its own stack file rather than an overlay (a merge cannot REMOVE the db service, and beta having no database of its own is the point). Its services are prefixed beta-*: Swarm registers a service's short name as a DNS alias on every network it joins, so two stacks both calling a service `web` on the shared data network would let prod's frontend resolve a beta task. Prod's stack, env and LB config are untouched. One Postgres, two databases with two roles: deploy/db/provision-env-db.sh creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT revoked from PUBLIC) plus a <role>_ro for ad-hoc queries, and refuses to run for the environment that owns the instance -- doing so would demote prod's bootstrap superuser. Fixes that co-residency would otherwise have broken silently: - CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and also "the Forgejo box" because those shared a machine; that conflation is what once had the prod backup dumping beta. - The nightly backup dumps BOTH application databases to separate off-box prefixes, and fails loudly on a missing one instead of skipping it. - Alloy stopped deriving the `host` label from the node -- on vps2 that would have filed every beta line as prod, feeding prod's alert rules with beta's traffic. It is now derived per source, with a new `node` label for the machine, and beta's log volume is mounted via a vps2-only overlay. - ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derived their SSH target and env-file path from the topology instead of hardcoding beta to 75.119.132.91 -- which is vps1's address now. - Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on vps1) instead of 0.0.0.0: correct for a home LAN box, a public exposure of unreviewed branches on a VPS. Terraform's hosts variable is keyed by machine with a nested environments map, and main.tf flattens (host, environment) pairs so two environments on one box cannot share a checkout. Caddy's single reference config is split per host. Docs, onboarding and the runbooks are updated throughout, including the security rationales that co-residency makes false: separate SSH credentials no longer put a host boundary between beta and prod, and the boundary that remains is the database and the filesystem.