# Runbook — moving beta to vps2 and dev to vps1 **Status: not executed.** Everything this runbook describes is landed in the repo and live nowhere. The estate still runs the old shape until someone works through the steps below. ## What changes | | before | after | |---|---|---| | `75.119.132.91` (**vps1**) | beta + Forgejo + Grafana/Loki | Forgejo + Grafana/Loki + **dev** | | `169.58.46.181` (**vps2**) | prod + Centralis + Postfix + backups | **prod + beta** + Centralis + Postfix + backups | | desktop | LAN dev server + CI runner | AI models + flex Swarm worker, **no environment** | | databases | one Postgres per environment | one instance on vps2, `thermograph` + `thermograph_beta` | Mesh addresses do **not** move: vps1 stays `10.10.0.2`, vps2 stays `10.10.0.1`. Every "beta = 75.119.132.91" reference anywhere is wrong after this — that address is vps1, and sending a beta-intended command there is the most likely way to do damage during this cutover. ## Order, and why it is this order Beta moves first and completely, while dev stays where it is. Then dev moves. The two halves are independent, so a problem in one never forces a rollback of the other — and beta is the half that carries a public hostname and a TLS certificate, so it gets done while you have the most attention. Prod is touched twice, and both touches are additive: a new database + role on its Postgres instance, and a new site block in its Caddyfile. Prod's stack file, its services, its volumes and its env file are not modified by this work at all. **Expected downtime:** beta, roughly the length of a dump/restore plus a stack bring-up (~10–20 min). Dev, as long as you like. **Prod: none** — unless step 3 or 8 is done wrong, which is what the verification lines are for. --- ## 0. Before you start - [ ] **Lower the DNS TTL** on `beta.thermograph.org` to 300s, at least an hour before step 7. Do this first; it is the only step with a lead time. - [ ] **Create the new Forgejo Actions secrets** (Settings → Secrets). The workflows on this branch reference them and will fail without them: VPS1_SSH_HOST=75.119.132.91 VPS1_SSH_USER=agent VPS1_SSH_KEY= VPS1_SSH_PORT=22 VPS2_SSH_HOST=169.58.46.181 VPS2_SSH_USER=agent VPS2_SSH_KEY= VPS2_SSH_PORT=22 Keep the old `SSH_*` / `PROD_SSH_*` secrets until step 11 — they are the rollback path, and deleting them early strands you. **Create these BEFORE merging.** `infra-sync.yml` fires on any push to `dev` or `main` touching `infra/**`, and this change touches a great deal of it. Without the new secrets those jobs SSH to an empty host and fail; `sync-beta` additionally targets `/opt/thermograph-beta`, which does not exist until step 1. Nothing is damaged either way — the jobs only fetch and render — but you will get a red run and a misleading alert. The app `Deploy` workflow does **not** fire on this merge: it is path-filtered to `backend/**` and `frontend/**`, and this change touches neither. - [ ] **Take a fresh prod backup and confirm it landed**, rather than trusting the schedule: ``` # from a machine with repo access # Actions -> "Ops cron (backup + IndexNow)" -> Run workflow ssh agent@169.58.46.181 'ls -lh ~/thermograph-backups | tail -3' ``` - [ ] **Confirm you can decrypt the vault** (`sops -d infra/deploy/secrets/beta.yaml` from `infra/`). Steps 3 and 5 need the rendered beta password. - [ ] Have the PR merged to `dev` and promoted to `main`. Hosts pull `main`, so nothing below works from an unmerged branch. --- ## 1. Give beta a checkout on vps2 ``` ssh agent@169.58.46.181 sudo git clone --branch main http://10.10.0.2:3080/Jinemi/thermograph.git /opt/thermograph-beta sudo chown -R agent:agent /opt/thermograph-beta ``` Prod's checkout at `/opt/thermograph` is untouched. Two checkouts side by side is the whole point: a `git reset --hard` for beta must never be able to move prod's tree. **Verify:** `git -C /opt/thermograph-beta log --oneline -1` matches `git -C /opt/thermograph log --oneline -1`. --- ## 2. Render beta's env file on vps2 ``` ssh agent@169.58.46.181 cd /opt/thermograph-beta/infra . deploy/render-secrets.sh render_thermograph_secrets /opt/thermograph-beta/infra beta /etc/thermograph-beta.env ``` The age key is already at `/etc/thermograph/age.key` on this box (prod uses it); beta needs no second copy. **Verify** — beta's file must name beta's role and database, and prod's must be untouched: ``` sudo grep -c . /etc/thermograph-beta.env # non-zero sudo grep -o 'thermograph_beta' /etc/thermograph-beta.env | head -1 sudo grep -o '@db:5432/thermograph$' /etc/thermograph.env # prod: still 'thermograph' ``` --- ## 3. Create beta's role and database on the shared instance This is the first step that touches prod's Postgres. It only ever adds; it drops nothing. ``` ssh agent@169.58.46.181 sudo bash /opt/thermograph-beta/infra/deploy/db/provision-env-db.sh beta ``` **Verify the isolation is real** — this is the whole justification for one instance serving two environments, so actually run it: ``` DBC=$(docker ps -q --filter label=com.docker.swarm.service.name=thermograph_db | head -1) # beta's role must NOT be able to reach prod's database: docker exec "$DBC" psql -U thermograph_beta -d thermograph -c 'select 1' # must FAIL # beta's read-only role exists and really is read-only: docker exec "$DBC" psql -U thermograph_beta_ro -d thermograph_beta -c 'select 1' # must SUCCEED docker exec "$DBC" psql -U thermograph_beta_ro -d thermograph_beta \ -c 'create table _x(i int)' # must FAIL # prod's own role is UNTOUCHED — still a superuser: docker exec "$DBC" psql -U thermograph -d postgres -tAc \ "select rolsuper from pg_roles where rolname='thermograph'" # must be t # and prod is still healthy: curl -fsS -o /dev/null -w '%{http_code}\n' https://thermograph.org/api/health ``` The script **refuses to run for prod** — prod's `thermograph` role is the instance's bootstrap superuser, not a guest, and provisioning it would demote it. That guard is why the superuser check above should never fail; run it anyway. If the first command SUCCEEDS, stop. `CONNECT` was not revoked and beta would be able to open a session against the production database; re-run the provisioning script and re-check before going further. This is not hypothetical — it happened on the live cutover. The first version of the script revoked `PUBLIC` connect on the *guest's* database only, and a fresh Postgres database carries `=Tc` (TEMP **and** CONNECT) for `PUBLIC`, so `thermograph_beta` could still reach `thermograph`. The script now hardens the owner's database too, granting the owner's read-only role an explicit `CONNECT` first so the revoke cannot strip `ops/dbq.sh` of its access. The check above is what caught it. **Rollback:** `DROP DATABASE thermograph_beta; DROP ROLE thermograph_beta;` — prod is unaffected either way. --- ## 4. Move beta's data (optional) Beta's existing database still lives on vps1. Decide honestly whether you want it: beta's value is a rehearsal of prod, and a fresh schema with a handful of test accounts is often *better* than carrying old beta state across. If you do want it: ``` # on vps1 — dump the old beta database ssh agent@75.119.132.91 \ 'docker exec thermograph-db-1 pg_dump -U thermograph -d thermograph --format=custom' \ > /tmp/beta-premove.dump # on vps2 — restore into beta's database as beta's role scp /tmp/beta-premove.dump agent@169.58.46.181:/tmp/ ssh agent@169.58.46.181 ' DBC=$(docker ps -q --filter label=com.docker.swarm.service.name=thermograph_db | head -1) docker cp /tmp/beta-premove.dump "$DBC":/tmp/ docker exec "$DBC" pg_restore -U thermograph -d thermograph_beta --no-owner --role=thermograph_beta /tmp/beta-premove.dump ' ``` `--no-owner --role=thermograph_beta` matters: the dump's objects are owned by the old `thermograph` role, and restoring them verbatim would leave beta's tables owned by a role beta does not connect as. **Do not** restore a *prod* dump into beta. Beta's daemon runs the notification timers, and beta's subscriber table would then be full of real people. Beta's vault sets `THERMOGRAPH_MAIL_BACKEND=console` so nothing would actually send, but that is one `sops edit` away from not being true. --- ## 5. Bring beta's stack up on vps2 ``` ssh agent@169.58.46.181 cd /opt/thermograph-beta THERMOGRAPH_ENV=beta SERVICE=all \ BACKEND_IMAGE_TAG= \ FRONTEND_IMAGE_TAG= \ infra/deploy/deploy.sh ``` Get the current tags from the old beta host first (`cat /opt/thermograph/infra/deploy/.image-tags.env` on vps1), so beta comes up on exactly the code it was already running. **Verify** — beta answers on its own loopback ports, and prod's are unmoved: ``` curl -fsS -o /dev/null -w 'beta %{http_code}\n' http://127.0.0.1:8237/healthz curl -fsS -o /dev/null -w 'prod %{http_code}\n' http://127.0.0.1:8137/healthz docker stack services thermograph-beta # 5 services, all 1/1 docker stack services thermograph # prod: unchanged, still 1/1 ``` **Verify the DNS-collision fix actually held** — this is the failure mode that would be subtle and awful in production: ``` # prod's frontend must resolve prod's web, not beta's docker exec $(docker ps -q -f label=com.docker.swarm.service.name=thermograph_frontend | head -1) \ getent hosts web docker exec $(docker ps -q -f label=com.docker.swarm.service.name=thermograph_frontend | head -1) \ getent hosts beta-web # should NOT resolve on prod's network ``` **Rollback:** `docker stack rm thermograph-beta` and remove the LB container (`docker rm -f thermograph-beta-lb`). Beta on vps1 is still running and still serving; nothing has moved yet. --- ## 6. Add beta's site block to vps2's Caddy `infra/deploy/Caddyfile.vps2` is the reference copy. Append its `beta.thermograph.org` block to the live `/etc/caddy/Caddyfile` on vps2. ``` ssh agent@169.58.46.181 sudo caddy validate --config /etc/caddy/Caddyfile # BEFORE reloading # `validate` runs as root and CREATES any log file the config names, owned by # root — which the caddy user then cannot open, and the reload fails. Fix the # ownership before reloading, or the first reload rejects the whole config: sudo chown caddy:caddy /var/log/caddy/beta.log sudo systemctl reload caddy ``` Validate before reload, always: a malformed Caddyfile takes `thermograph.org` down with it, and prod is on this box now. The `chown` is not optional and it bit this cutover. `caddy validate` as root left `/var/log/caddy/beta.log` as `root:root`, so the reload failed with `permission denied` on the log writer. (The root-owned `centralis.log` sitting in that directory is the same mistake, made earlier.) Note what did NOT happen: Caddy rejects a bad config atomically and keeps serving the old one, so `thermograph.org` stayed up throughout — verify it anyway. Caddy cannot issue the certificate until DNS moves (step 7), so expect the beta hostname to fail TLS until then. That is fine and expected. --- ## 7. Move DNS Point `beta.thermograph.org`'s A record at **169.58.46.181**. Then wait for propagation and the certificate: ``` dig +short beta.thermograph.org # 169.58.46.181 curl -fsS -o /dev/null -w '%{http_code}\n' https://beta.thermograph.org/api/health ssh agent@169.58.46.181 'sudo journalctl -u caddy -n 30 --no-pager | grep -i certificate' ``` **Rollback:** point the A record back at 75.119.132.91. Beta on vps1 is still up (you have not stopped it yet — that is step 9), so this is a clean revert for as long as you leave it running. --- ## 8. Stand dev up on vps1 Independent of everything above; do it whenever. ``` ssh agent@75.119.132.91 sudo bash /opt/thermograph-dev/infra/deploy/provision-dev.sh # clones if absent ``` The script writes `/etc/thermograph/secrets-env` = `dev` and removes any `deploy-mode` marker. **That marker flip is load-bearing, and it is why this step comes after step 9.** dev's env file is `/etc/thermograph.env` — the very path beta uses while beta is still on vps1. `infra-sync`'s `sync-dev` job refuses to render unless the marker already says `dev`, so until you run this script it syncs the checkout and skips the render with a loud message. Flip the marker before beta has vacated and the next dev push overwrites beta's live env file with dev's credentials; beta's running containers survive it and then come up wrong on their next deploy. (The checkout itself is safe to create at any time — only the render is gated.) Then deploy: ``` cd /opt/thermograph-dev SERVICE=all BACKEND_IMAGE_TAG= FRONTEND_IMAGE_TAG= \ infra/deploy/deploy-dev.sh ``` **Verify Caddy is the only way in.** This is the security-relevant check of the whole cutover — vps1 is a public box and dev runs unreviewed branches: ``` ss -ltnp | grep 8137 # MUST show 127.0.0.1:8137, never 0.0.0.0 curl --max-time 5 http://75.119.132.91:8137/healthz # MUST fail/refuse curl -o /dev/null -w '%{http_code}\n' https://dev.thermograph.org/ # 401 without creds curl -o /dev/null -w '%{http_code}\n' -u dev: https://dev.thermograph.org/healthz # 200 ``` Also confirm dev did **not** get the fleet's shared production credentials — `dev.yaml` is rendered alone, and vps1 is exactly the box that must not hold them: ``` sudo grep -c 'THERMOGRAPH_S3_SECRET_KEY\|REGISTRY_TOKEN\|VAPID_PRIVATE' /etc/thermograph.env # expect 0 ``` --- ## 9. Retire beta from vps1 Only after beta has been serving from vps2 through step 7 for long enough that you believe it. ``` ssh agent@75.119.132.91 cd /opt/thermograph/infra docker compose down # stops the old beta app stack # remove beta's site block from /etc/caddy/Caddyfile (see Caddyfile.vps1) sudo caddy validate --config /etc/caddy/Caddyfile && sudo systemctl reload caddy ``` **Keep the old `thermograph_pgdata` volume on vps1** until you are certain beta on vps2 is healthy and backed up. It is the only copy of pre-move beta data. Delete it deliberately, later, not as part of this runbook. Forgejo, Grafana, Loki and the portfolio site all keep running on this box untouched — do not stop anything else here. --- ## 10. Repoint the Alloy agents Both nodes need the new variables; the label model changed (see `observability/alloy/config.alloy`). ``` # vps2 — two environments, so two log volumes and the beta overlay ssh agent@169.58.46.181 cd /opt/thermograph/observability/alloy ALLOY_NODE=vps2 ALLOY_ENV=prod \ APPLOGS_VOLUME=thermograph_applogs BETA_APPLOGS_VOLUME=thermograph-beta_applogs \ LOKI_URL=http://10.10.0.2:3100/loki/api/v1/push \ docker compose -f docker-compose.agent.yml -f docker-compose.agent.beta.yml up -d # vps1 ssh agent@75.119.132.91 cd /opt/thermograph/observability/alloy ALLOY_NODE=vps1 ALLOY_ENV=dev APPLOGS_VOLUME=thermograph-dev_applogs \ LOKI_URL=http://10.10.0.2:3100/loki/api/v1/push \ docker compose -f docker-compose.agent.yml up -d ``` **Verify each environment is labelled as itself** — the point of the change is that beta's logs on vps2 are not filed as prod: ``` # in Grafana Explore, or via the Loki API: # {host="beta"} |= "" -> should show beta traffic, from node="vps2" # {host="prod"} |= "" -> prod only # {host="dev"} |= "" -> from node="vps1" # count by (host) (count_over_time({job="docker"}[5m])) -> three values, not one ``` If `{host="beta"}` is empty while beta is clearly serving, the container-name relabel did not match — check `docker ps --format '{{.Names}}'` on vps2 against the `/thermograph-beta_.*` regex. --- ## 11. Clean up - [ ] Delete the old `SSH_*` and `PROD_SSH_*` Forgejo secrets (only now — they were the rollback path). - [ ] Run the ops cron by hand and confirm **both** databases are dumped: `ssh agent@169.58.46.181 'ls ~/thermograph-backups | tail -4'` should show both a `thermograph-*.dump` and a `thermograph_beta-*.dump`. - [ ] Confirm the Forgejo backup job still works — it now uses `VPS1_SSH_*`, and it follows Forgejo (vps1), not beta. - [ ] Restore the DNS TTL on `beta.thermograph.org`. - [ ] Desktop: stop the old LAN dev stack and free the box for the AI models. Leave it joined to the Swarm as a worker; nothing schedules onto it today (every app service is pinned `node.role == manager`), which is what makes it safe to also run models there. - [ ] Update Centralis — see below. --- ## Centralis (outside this repo) Centralis' own text still describes the old estate, and it is not in this repository. These need a separate change in the Centralis repo: - The `onboarding` tool's estate paragraph ("Four machines on a WireGuard mesh... beta (beta.thermograph.org, Forgejo, Grafana + Loki), the operator's desktop (LAN dev, the main CI runner)"). - The MCP server's own protocol-level `instructions` string, which carries a second, separately-worded copy of the same paragraph. - The `thermograph-orientation` skill's estate table, and `thermograph-ops`' release-flow diagram ("LAN dev box"), its "ssh -L works for beta and is impossible for prod" claim (beta is a Swarm overlay now, so it behaves like prod), and its prod-only Swarm service-name list (beta has `beta-`-prefixed ones now). - Tool descriptions that hardcode the old shape: `fleet_status` ("Swarm services on prod, compose containers on beta/dev"), `estate_status` and `rollback_to` ("Centralis cannot SSH to the dev desktop"), `logs_query`'s per-environment service lists (Forgejo/Grafana/Loki are listed under `beta`; they belong under `dev`/vps1 now), and `secrets_render`'s assumption that a host's own `secrets-env` marker identifies one environment — vps2 has two. --- ## If it goes wrong | symptom | most likely cause | action | |---|---|---| | prod 502s after step 5 | beta's stack collided with prod's service DNS | `docker stack rm thermograph-beta`; prod recovers on its own. Check beta's services are `beta-`-prefixed | | beta 502s, prod fine | beta's LB or stack not up | `docker ps \| grep beta-lb`; `docker stack ps thermograph-beta --no-trunc` | | beta can't reach the database | role/database missing, or the wrong network | re-run `provision-env-db.sh beta`; confirm beta's tasks are on `thermograph_internal` | | beta serves prod's data | beta's env file has prod's URL | check `/etc/thermograph-beta.env` names `thermograph_beta` twice; re-render (step 2). **Stop beta until fixed** | | prod's nightly backup fails | `thermograph_beta` doesn't exist yet | either finish step 3 or revert the ops-cron change; the job fails loudly by design rather than skipping silently | | dev reachable from the internet | `DEV_BIND_ADDR` not applied | `ss -ltnp \| grep 8137`; redeploy via `deploy-dev.sh`, which sets it from `env-topology.sh` |