infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home #103

Merged
admin_emi merged 2 commits from worktree-refactor+host-topology-vps1-vps2 into dev 2026-07-26 06:56:39 +00:00
Owner

Re-architects the estate so machines are named by role and an environment stops implying a host. Nothing live has moved — the ordered cutover is infra/deploy/RUNBOOK-vps1-vps2-cutover.md, and every step has a verification and a rollback.

vps1  75.119.132.91   Forgejo, Grafana/Loki, emigriffith.dev, and DEV
                      (its own Postgres, mesh-only on 10.10.0.2:8137)
vps2  169.58.46.181   PROD and BETA as two Swarm stacks sharing one
                      TimescaleDB instance, + Centralis, Postfix, backups
desktop               AI model hosting + flex Swarm capacity, no environment

Mesh IPs don't move; what moves is which environment lives where. After the cutover, every "beta = 75.119.132.91" reference points at the wrong box — that address is vps1.

The core of it

infra/deploy/env-topology.sh is the single source of truth: env → host, checkout, branch, deploy mode, stack name, env file, LB ports, DB role/database, service prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing distinguishing a beta deploy from a prod one — so deploy.sh refuses to run when it disagrees with the checkout it was invoked from. The host-wide secrets-env/deploy-mode markers survive only as a fallback for a single-environment box; they cannot describe vps2.

Beta gets its own stack file rather than an overlay: a compose merge can add and override but cannot remove the db service, and beta having no database of its own is the whole design. Its services are prefixed beta-* because Swarm registers a service's short name as a DNS alias on every network it joins — two stacks both calling a service web on the shared data network would let prod's frontend resolve a beta task. Prod's stack file, env and LB config are untouched.

One Postgres, two databases, two roles. deploy/db/provision-env-db.sh creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT revoked from PUBLIC) plus a <role>_ro for ad-hoc queries — and refuses to run for the environment that owns the instance, since doing so would demote prod's bootstrap superuser.

Things that would have broken silently

  • CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and "the Forgejo box" because those shared a machine — the same conflation that once had the prod backup dumping beta.
  • Both application databases are backed up, to separate off-box prefixes, failing loudly on a missing one rather than skipping it. Previously only thermograph was dumped.
  • Alloy no longer derives host from the node. On vps2 that would have filed every beta container, access log and JSON line as prod, feeding prod's alert rules with beta's traffic. It's now per-source, with a new node label for the machine, and beta's log volume arrives via a vps2-only overlay (pointing it at dev's volume on vps1 would have manufactured fake beta data).
  • ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derive their SSH target and env-file path from the topology. The seed scripts in particular would have read prod's live env file when seeding beta.
  • Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on vps1) instead of 0.0.0.0 — correct for a home LAN box, a public exposure of unreviewed branches on a VPS.

Also

Terraform's hosts is keyed by machine with a nested environments map, and main.tf flattens (host, environment) pairs so two environments on one box can't share a checkout. Caddy's single reference config splits per host. Beta stops pinging IndexNow (it was asking search engines to index a rehearsal environment) and stops warming the city archive (shared upstream quota, cache only beta reads).

Docs, onboarding and runbooks updated throughout — including the security rationales co-residency makes false: separate SSH credentials no longer buy a host boundary between beta and prod, and what remains is the database and the filesystem. The meaningful boundary is now vps1 vs vps2.

Verified

shellcheck -x clean over all 36 scripts (CI's pinned v0.11.0) · every workflow and stack file parses · alloy validate with CI's pinned v1.9.1 binary · the alerting structural checker reports 0 problems · terraform validate + fmt clean, and the flattening evaluated against the example tfvars · vault re-encrypts and round-trips · the dev overlay renders 127.0.0.1 by default and 10.10.0.2 under the topology · the env/checkout mismatch guard refuses a crossed beta/prod deploy.

Not verified (impossible without executing): anything about the live estate.

Companion docs PRs in thermograph-docs: #10 (the decision record) and #11 (a pointer from system-overview.md, which keeps describing the live shape until the cutover runs).

Centralis' own estate description — the onboarding tool, the server's instructions string, the thermograph-orientation/thermograph-ops skills and several tool descriptions — lives in emi/centralis and still states the old topology. The runbook lists exactly which strings need changing.

Re-architects the estate so machines are named by role and an environment stops implying a host. **Nothing live has moved** — the ordered cutover is `infra/deploy/RUNBOOK-vps1-vps2-cutover.md`, and every step has a verification and a rollback. ``` vps1 75.119.132.91 Forgejo, Grafana/Loki, emigriffith.dev, and DEV (its own Postgres, mesh-only on 10.10.0.2:8137) vps2 169.58.46.181 PROD and BETA as two Swarm stacks sharing one TimescaleDB instance, + Centralis, Postfix, backups desktop AI model hosting + flex Swarm capacity, no environment ``` Mesh IPs don't move; what moves is which environment lives where. **After the cutover, every "beta = 75.119.132.91" reference points at the wrong box** — that address is vps1. ## The core of it `infra/deploy/env-topology.sh` is the single source of truth: env → host, checkout, branch, deploy mode, stack name, env file, LB ports, DB role/database, service prefix. `THERMOGRAPH_ENV` is the input, and on vps2 it is the *only* thing distinguishing a beta deploy from a prod one — so `deploy.sh` refuses to run when it disagrees with the checkout it was invoked from. The host-wide `secrets-env`/`deploy-mode` markers survive only as a fallback for a single-environment box; they cannot describe vps2. Beta gets its own stack file rather than an overlay: a compose merge can add and override but cannot **remove** the `db` service, and beta having no database of its own is the whole design. Its services are prefixed `beta-*` because Swarm registers a service's short name as a DNS alias on every network it joins — two stacks both calling a service `web` on the shared data network would let prod's frontend resolve a beta task. **Prod's stack file, env and LB config are untouched.** One Postgres, two databases, two roles. `deploy/db/provision-env-db.sh` creates `thermograph_beta` (NOSUPERUSER, owns only its own database, `CONNECT` revoked from `PUBLIC`) plus a `<role>_ro` for ad-hoc queries — and refuses to run for the environment that *owns* the instance, since doing so would demote prod's bootstrap superuser. ## Things that would have broken silently - **CI secrets are keyed by host** (`VPS1_SSH_*`, `VPS2_SSH_*`). `SSH_*` meant "beta" *and* "the Forgejo box" because those shared a machine — the same conflation that once had the prod backup dumping beta. - **Both application databases are backed up**, to separate off-box prefixes, failing loudly on a missing one rather than skipping it. Previously only `thermograph` was dumped. - **Alloy no longer derives `host` from the node.** On vps2 that would have filed every beta container, access log and JSON line as `prod`, feeding prod's alert rules with beta's traffic. It's now per-source, with a new `node` label for the machine, and beta's log volume arrives via a vps2-only overlay (pointing it at dev's volume on vps1 would have manufactured fake beta data). - **`ops/dbq.sh`, `ops/iceberg.sh` and the secrets seed scripts** derive their SSH target and env-file path from the topology. The seed scripts in particular would have read *prod's* live env file when seeding beta. - **Dev's overlay binds `DEV_BIND_ADDR`** (loopback by default, the mesh address on vps1) instead of `0.0.0.0` — correct for a home LAN box, a public exposure of unreviewed branches on a VPS. ## Also Terraform's `hosts` is keyed by machine with a nested `environments` map, and `main.tf` flattens `(host, environment)` pairs so two environments on one box can't share a checkout. Caddy's single reference config splits per host. Beta stops pinging IndexNow (it was asking search engines to index a rehearsal environment) and stops warming the city archive (shared upstream quota, cache only beta reads). Docs, onboarding and runbooks updated throughout — including the security rationales co-residency makes false: separate SSH credentials no longer buy a host boundary between beta and prod, and what remains is the database and the filesystem. The meaningful boundary is now vps1 vs vps2. ## Verified `shellcheck -x` clean over all 36 scripts (CI's pinned v0.11.0) · every workflow and stack file parses · `alloy validate` with CI's pinned v1.9.1 binary · the alerting structural checker reports 0 problems · `terraform validate` + `fmt` clean, and the flattening evaluated against the example tfvars · vault re-encrypts and round-trips · the dev overlay renders `127.0.0.1` by default and `10.10.0.2` under the topology · the env/checkout mismatch guard refuses a crossed beta/prod deploy. Not verified (impossible without executing): anything about the live estate. Companion docs PRs in `thermograph-docs`: #10 (the decision record) and #11 (a pointer from `system-overview.md`, which keeps describing the live shape until the cutover runs). Centralis' own estate description — the `onboarding` tool, the server's `instructions` string, the `thermograph-orientation`/`thermograph-ops` skills and several tool descriptions — lives in `emi/centralis` and still states the old topology. The runbook lists exactly which strings need changing.
admin_emi added 1 commit 2026-07-25 22:02:16 +00:00
infra: split the estate into vps1/vps2, moving beta next to prod and dev onto vps1
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
4e97d8e5dc
Environments stop being machines. Until now each box WAS an environment --
"beta" named both a deploy target and a host, "the desktop" named both the
operator's computer and the dev server -- so every path could assume one
environment per host and hardcode /opt/thermograph, /etc/thermograph.env and
ports 8137/8080. That assumption ends here:

  vps1  75.119.132.91  Forgejo, Grafana/Loki, the portfolio site, and DEV
                       (own Postgres, mesh-only on 10.10.0.2:8137)
  vps2  169.58.46.181  PROD and BETA as two Swarm stacks sharing one
                       TimescaleDB instance, plus Centralis, Postfix, backups
  desktop              AI model hosting + flex Swarm capacity, no environment

Nothing live has moved; the ordered cutover is in
infra/deploy/RUNBOOK-vps1-vps2-cutover.md.

deploy/env-topology.sh is the single source of truth: env -> host, checkout,
branch, deploy mode, stack name, env file, LB ports, DB role/database, service
prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing
distinguishing a beta deploy from a prod one -- deploy.sh refuses to run when it
disagrees with the checkout it was invoked from. The host-wide secrets-env and
deploy-mode markers survive only as a fallback for a single-environment box.

Beta gets its own stack file rather than an overlay (a merge cannot REMOVE the
db service, and beta having no database of its own is the point). Its services
are prefixed beta-*: Swarm registers a service's short name as a DNS alias on
every network it joins, so two stacks both calling a service `web` on the shared
data network would let prod's frontend resolve a beta task. Prod's stack, env
and LB config are untouched.

One Postgres, two databases with two roles: deploy/db/provision-env-db.sh
creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT
revoked from PUBLIC) plus a <role>_ro for ad-hoc queries, and refuses to run for
the environment that owns the instance -- doing so would demote prod's bootstrap
superuser.

Fixes that co-residency would otherwise have broken silently:

- CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and
  also "the Forgejo box" because those shared a machine; that conflation is what
  once had the prod backup dumping beta.
- The nightly backup dumps BOTH application databases to separate off-box
  prefixes, and fails loudly on a missing one instead of skipping it.
- Alloy stopped deriving the `host` label from the node -- on vps2 that would
  have filed every beta line as prod, feeding prod's alert rules with beta's
  traffic. It is now derived per source, with a new `node` label for the machine,
  and beta's log volume is mounted via a vps2-only overlay.
- ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derived their SSH
  target and env-file path from the topology instead of hardcoding beta to
  75.119.132.91 -- which is vps1's address now.
- Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on
  vps1) instead of 0.0.0.0: correct for a home LAN box, a public exposure of
  unreviewed branches on a VPS.

Terraform's hosts variable is keyed by machine with a nested environments map,
and main.tf flattens (host, environment) pairs so two environments on one box
cannot share a checkout. Caddy's single reference config is split per host.

Docs, onboarding and the runbooks are updated throughout, including the security
rationales that co-residency makes false: separate SSH credentials no longer put
a host boundary between beta and prod, and the boundary that remains is the
database and the filesystem.
admin_emi added 1 commit 2026-07-25 22:03:34 +00:00
runbook: spell out that the CI secrets must exist before the merge
All checks were successful
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Successful in 22s
PR build (required check) / gate (pull_request) Successful in 2s
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
f1664bc0f5
infra-sync fires on any push to dev/main touching infra/**, which this change
does heavily. Without VPS1_SSH_*/VPS2_SSH_* those jobs SSH to an empty host and
fail; sync-beta also targets a checkout that does not exist until step 1.
Nothing is damaged (the jobs only fetch and render) but the red run reads like
an outage. Also records that the app Deploy workflow does NOT fire here, being
path-filtered to backend/** and frontend/**.
admin_emi merged commit d138f00a20 into dev 2026-07-26 06:56:39 +00:00
admin_emi deleted branch worktree-refactor+host-topology-vps1-vps2 2026-07-26 06:56:40 +00:00
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference: Jinemi/thermograph#103
No description provided.