19 KiB
Runbook — moving beta to vps2 and dev to vps1
Status: not executed. Everything this runbook describes is landed in the repo and live nowhere. The estate still runs the old shape until someone works through the steps below.
What changes
| before | after | |
|---|---|---|
75.119.132.91 (vps1) |
beta + Forgejo + Grafana/Loki | Forgejo + Grafana/Loki + dev |
169.58.46.181 (vps2) |
prod + Centralis + Postfix + backups | prod + beta + Centralis + Postfix + backups |
| desktop | LAN dev server + CI runner | AI models + flex Swarm worker, no environment |
| databases | one Postgres per environment | one instance on vps2, thermograph + thermograph_beta |
Mesh addresses do not move: vps1 stays 10.10.0.2, vps2 stays 10.10.0.1.
Every "beta = 75.119.132.91" reference anywhere is wrong after this — that
address is vps1, and sending a beta-intended command there is the most likely
way to do damage during this cutover.
Order, and why it is this order
Beta moves first and completely, while dev stays where it is. Then dev moves. The two halves are independent, so a problem in one never forces a rollback of the other — and beta is the half that carries a public hostname and a TLS certificate, so it gets done while you have the most attention.
Prod is touched twice, and both touches are additive: a new database + role on its Postgres instance, and a new site block in its Caddyfile. Prod's stack file, its services, its volumes and its env file are not modified by this work at all.
Expected downtime: beta, roughly the length of a dump/restore plus a stack bring-up (~10–20 min). Dev, as long as you like. Prod: none — unless step 3 or 8 is done wrong, which is what the verification lines are for.
0. Before you start
-
Lower the DNS TTL on
beta.thermograph.orgto 300s, at least an hour before step 7. Do this first; it is the only step with a lead time. -
Create the new Forgejo Actions secrets (Settings → Secrets). The workflows on this branch reference them and will fail without them:
VPS1_SSH_HOST=75.119.132.91 VPS1_SSH_USER=agent VPS1_SSH_KEY=<the CI deploy key> VPS1_SSH_PORT=22 VPS2_SSH_HOST=169.58.46.181 VPS2_SSH_USER=agent VPS2_SSH_KEY=<the CI deploy key> VPS2_SSH_PORT=22 Keep the old `SSH_*` / `PROD_SSH_*` secrets until step 11 — they are the rollback path, and deleting them early strands you. **Create these BEFORE merging.** `infra-sync.yml` fires on any push to `dev` or `main` touching `infra/**`, and this change touches a great deal of it. Without the new secrets those jobs SSH to an empty host and fail; `sync-beta` additionally targets `/opt/thermograph-beta`, which does not exist until step 1. Nothing is damaged either way — the jobs only fetch and render — but you will get a red run and a misleading alert. The app `Deploy` workflow does **not** fire on this merge: it is path-filtered to `backend/**` and `frontend/**`, and this change touches neither. -
Take a fresh prod backup and confirm it landed, rather than trusting the schedule:
``` # from a machine with repo access # Actions -> "Ops cron (backup + IndexNow)" -> Run workflow ssh agent@169.58.46.181 'ls -lh ~/thermograph-backups | tail -3' ``` -
Confirm you can decrypt the vault (
sops -d infra/deploy/secrets/beta.yamlfrominfra/). Steps 3 and 5 need the rendered beta password. -
Have the PR merged to
devand promoted tomain. Hosts pullmain, so nothing below works from an unmerged branch.
1. Give beta a checkout on vps2
ssh agent@169.58.46.181
sudo git clone --branch main http://10.10.0.2:3080/emi/thermograph.git /opt/thermograph-beta
sudo chown -R agent:agent /opt/thermograph-beta
Prod's checkout at /opt/thermograph is untouched. Two checkouts side by side
is the whole point: a git reset --hard for beta must never be able to move
prod's tree.
Verify: git -C /opt/thermograph-beta log --oneline -1 matches
git -C /opt/thermograph log --oneline -1.
2. Render beta's env file on vps2
ssh agent@169.58.46.181
cd /opt/thermograph-beta/infra
. deploy/render-secrets.sh
render_thermograph_secrets /opt/thermograph-beta/infra beta /etc/thermograph-beta.env
The age key is already at /etc/thermograph/age.key on this box (prod uses it);
beta needs no second copy.
Verify — beta's file must name beta's role and database, and prod's must be untouched:
sudo grep -c . /etc/thermograph-beta.env # non-zero
sudo grep -o 'thermograph_beta' /etc/thermograph-beta.env | head -1
sudo grep -o '@db:5432/thermograph$' /etc/thermograph.env # prod: still 'thermograph'
3. Create beta's role and database on the shared instance
This is the first step that touches prod's Postgres. It only ever adds; it drops nothing.
ssh agent@169.58.46.181
sudo bash /opt/thermograph-beta/infra/deploy/db/provision-env-db.sh beta
Verify the isolation is real — this is the whole justification for one instance serving two environments, so actually run it:
DBC=$(docker ps -q --filter label=com.docker.swarm.service.name=thermograph_db | head -1)
# beta's role must NOT be able to reach prod's database:
docker exec "$DBC" psql -U thermograph_beta -d thermograph -c 'select 1' # must FAIL
# beta's read-only role exists and really is read-only:
docker exec "$DBC" psql -U thermograph_beta_ro -d thermograph_beta -c 'select 1' # must SUCCEED
docker exec "$DBC" psql -U thermograph_beta_ro -d thermograph_beta \
-c 'create table _x(i int)' # must FAIL
# prod's own role is UNTOUCHED — still a superuser:
docker exec "$DBC" psql -U thermograph -d postgres -tAc \
"select rolsuper from pg_roles where rolname='thermograph'" # must be t
# and prod is still healthy:
curl -fsS -o /dev/null -w '%{http_code}\n' https://thermograph.org/api/health
The script refuses to run for prod — prod's thermograph role is the
instance's bootstrap superuser, not a guest, and provisioning it would demote
it. That guard is why the superuser check above should never fail; run it
anyway.
If the first command SUCCEEDS, stop. CONNECT was not revoked and beta would be
able to open a session against the production database; re-run the provisioning
script and re-check before going further.
This is not hypothetical — it happened on the live cutover. The first version of
the script revoked PUBLIC connect on the guest's database only, and a fresh
Postgres database carries =Tc (TEMP and CONNECT) for PUBLIC, so
thermograph_beta could still reach thermograph. The script now hardens the
owner's database too, granting the owner's read-only role an explicit CONNECT
first so the revoke cannot strip ops/dbq.sh of its access. The check above is
what caught it.
Rollback: DROP DATABASE thermograph_beta; DROP ROLE thermograph_beta; —
prod is unaffected either way.
4. Move beta's data (optional)
Beta's existing database still lives on vps1. Decide honestly whether you want it: beta's value is a rehearsal of prod, and a fresh schema with a handful of test accounts is often better than carrying old beta state across. If you do want it:
# on vps1 — dump the old beta database
ssh agent@75.119.132.91 \
'docker exec thermograph-db-1 pg_dump -U thermograph -d thermograph --format=custom' \
> /tmp/beta-premove.dump
# on vps2 — restore into beta's database as beta's role
scp /tmp/beta-premove.dump agent@169.58.46.181:/tmp/
ssh agent@169.58.46.181 '
DBC=$(docker ps -q --filter label=com.docker.swarm.service.name=thermograph_db | head -1)
docker cp /tmp/beta-premove.dump "$DBC":/tmp/
docker exec "$DBC" pg_restore -U thermograph -d thermograph_beta --no-owner --role=thermograph_beta /tmp/beta-premove.dump
'
--no-owner --role=thermograph_beta matters: the dump's objects are owned by the
old thermograph role, and restoring them verbatim would leave beta's tables
owned by a role beta does not connect as.
Do not restore a prod dump into beta. Beta's daemon runs the notification
timers, and beta's subscriber table would then be full of real people. Beta's
vault sets THERMOGRAPH_MAIL_BACKEND=console so nothing would actually send,
but that is one sops edit away from not being true.
5. Bring beta's stack up on vps2
ssh agent@169.58.46.181
cd /opt/thermograph-beta
THERMOGRAPH_ENV=beta SERVICE=all \
BACKEND_IMAGE_TAG=<current beta backend tag> \
FRONTEND_IMAGE_TAG=<current beta frontend tag> \
infra/deploy/deploy.sh
Get the current tags from the old beta host first
(cat /opt/thermograph/infra/deploy/.image-tags.env on vps1), so beta comes up
on exactly the code it was already running.
Verify — beta answers on its own loopback ports, and prod's are unmoved:
curl -fsS -o /dev/null -w 'beta %{http_code}\n' http://127.0.0.1:8237/healthz
curl -fsS -o /dev/null -w 'prod %{http_code}\n' http://127.0.0.1:8137/healthz
docker stack services thermograph-beta # 5 services, all 1/1
docker stack services thermograph # prod: unchanged, still 1/1
Verify the DNS-collision fix actually held — this is the failure mode that would be subtle and awful in production:
# prod's frontend must resolve prod's web, not beta's
docker exec $(docker ps -q -f label=com.docker.swarm.service.name=thermograph_frontend | head -1) \
getent hosts web
docker exec $(docker ps -q -f label=com.docker.swarm.service.name=thermograph_frontend | head -1) \
getent hosts beta-web # should NOT resolve on prod's network
Rollback: docker stack rm thermograph-beta and remove the LB container
(docker rm -f thermograph-beta-lb). Beta on vps1 is still running and still
serving; nothing has moved yet.
6. Add beta's site block to vps2's Caddy
infra/deploy/Caddyfile.vps2 is the reference copy. Append its
beta.thermograph.org block to the live /etc/caddy/Caddyfile on vps2.
ssh agent@169.58.46.181
sudo caddy validate --config /etc/caddy/Caddyfile # BEFORE reloading
# `validate` runs as root and CREATES any log file the config names, owned by
# root — which the caddy user then cannot open, and the reload fails. Fix the
# ownership before reloading, or the first reload rejects the whole config:
sudo chown caddy:caddy /var/log/caddy/beta.log
sudo systemctl reload caddy
Validate before reload, always: a malformed Caddyfile takes thermograph.org
down with it, and prod is on this box now.
The chown is not optional and it bit this cutover. caddy validate as root
left /var/log/caddy/beta.log as root:root, so the reload failed with
permission denied on the log writer. (The root-owned centralis.log sitting
in that directory is the same mistake, made earlier.) Note what did NOT happen:
Caddy rejects a bad config atomically and keeps serving the old one, so
thermograph.org stayed up throughout — verify it anyway.
Caddy cannot issue the certificate until DNS moves (step 7), so expect the beta hostname to fail TLS until then. That is fine and expected.
7. Move DNS
Point beta.thermograph.org's A record at 169.58.46.181.
Then wait for propagation and the certificate:
dig +short beta.thermograph.org # 169.58.46.181
curl -fsS -o /dev/null -w '%{http_code}\n' https://beta.thermograph.org/api/health
ssh agent@169.58.46.181 'sudo journalctl -u caddy -n 30 --no-pager | grep -i certificate'
Rollback: point the A record back at 75.119.132.91. Beta on vps1 is still up (you have not stopped it yet — that is step 9), so this is a clean revert for as long as you leave it running.
8. Stand dev up on vps1
Independent of everything above; do it whenever.
ssh agent@75.119.132.91
sudo bash /opt/thermograph-dev/infra/deploy/provision-dev.sh # clones if absent
The script writes /etc/thermograph/secrets-env = dev and removes any
deploy-mode marker.
That marker flip is load-bearing, and it is why this step comes after step 9.
dev's env file is /etc/thermograph.env — the very path beta uses while beta is
still on vps1. infra-sync's sync-dev job refuses to render unless the marker
already says dev, so until you run this script it syncs the checkout and skips
the render with a loud message. Flip the marker before beta has vacated and the
next dev push overwrites beta's live env file with dev's credentials; beta's
running containers survive it and then come up wrong on their next deploy.
(The checkout itself is safe to create at any time — only the render is gated.)
Then deploy:
cd /opt/thermograph-dev
SERVICE=all BACKEND_IMAGE_TAG=<dev tag> FRONTEND_IMAGE_TAG=<dev tag> \
infra/deploy/deploy-dev.sh
Verify it is mesh-only. This is the security-relevant check of the whole cutover — vps1 is a public box and dev runs unreviewed branches:
ss -ltnp | grep 8137 # MUST show 10.10.0.2:8137, never 0.0.0.0:8137
curl -fsS -o /dev/null -w '%{http_code}\n' http://10.10.0.2:8137/healthz # from the mesh
curl --max-time 5 http://75.119.132.91:8137/healthz # MUST fail/refuse
Also confirm dev did not get the fleet's shared production credentials —
dev.yaml is rendered alone, and vps1 is exactly the box that must not hold
them:
sudo grep -c 'THERMOGRAPH_S3_SECRET_KEY\|REGISTRY_TOKEN\|VAPID_PRIVATE' /etc/thermograph.env # expect 0
9. Retire beta from vps1
Only after beta has been serving from vps2 through step 7 for long enough that you believe it.
ssh agent@75.119.132.91
cd /opt/thermograph/infra
docker compose down # stops the old beta app stack
# remove beta's site block from /etc/caddy/Caddyfile (see Caddyfile.vps1)
sudo caddy validate --config /etc/caddy/Caddyfile && sudo systemctl reload caddy
Keep the old thermograph_pgdata volume on vps1 until you are certain beta
on vps2 is healthy and backed up. It is the only copy of pre-move beta data.
Delete it deliberately, later, not as part of this runbook.
Forgejo, Grafana, Loki and the portfolio site all keep running on this box untouched — do not stop anything else here.
10. Repoint the Alloy agents
Both nodes need the new variables; the label model changed (see
observability/alloy/config.alloy).
# vps2 — two environments, so two log volumes and the beta overlay
ssh agent@169.58.46.181
cd /opt/thermograph/observability/alloy
ALLOY_NODE=vps2 ALLOY_ENV=prod \
APPLOGS_VOLUME=thermograph_applogs BETA_APPLOGS_VOLUME=thermograph-beta_applogs \
LOKI_URL=http://10.10.0.2:3100/loki/api/v1/push \
docker compose -f docker-compose.agent.yml -f docker-compose.agent.beta.yml up -d
# vps1
ssh agent@75.119.132.91
cd /opt/thermograph/observability/alloy
ALLOY_NODE=vps1 ALLOY_ENV=dev APPLOGS_VOLUME=thermograph-dev_applogs \
LOKI_URL=http://10.10.0.2:3100/loki/api/v1/push \
docker compose -f docker-compose.agent.yml up -d
Verify each environment is labelled as itself — the point of the change is that beta's logs on vps2 are not filed as prod:
# in Grafana Explore, or via the Loki API:
# {host="beta"} |= "" -> should show beta traffic, from node="vps2"
# {host="prod"} |= "" -> prod only
# {host="dev"} |= "" -> from node="vps1"
# count by (host) (count_over_time({job="docker"}[5m])) -> three values, not one
If {host="beta"} is empty while beta is clearly serving, the container-name
relabel did not match — check docker ps --format '{{.Names}}' on vps2 against
the /thermograph-beta_.* regex.
11. Clean up
- Delete the old
SSH_*andPROD_SSH_*Forgejo secrets (only now — they were the rollback path). - Run the ops cron by hand and confirm both databases are dumped:
ssh agent@169.58.46.181 'ls ~/thermograph-backups | tail -4'should show both athermograph-*.dumpand athermograph_beta-*.dump. - Confirm the Forgejo backup job still works — it now uses
VPS1_SSH_*, and it follows Forgejo (vps1), not beta. - Restore the DNS TTL on
beta.thermograph.org. - Desktop: stop the old LAN dev stack and free the box for the AI models.
Leave it joined to the Swarm as a worker; nothing schedules onto it today
(every app service is pinned
node.role == manager), which is what makes it safe to also run models there. - Update Centralis — see below.
Centralis (outside this repo)
Centralis' own text still describes the old estate, and it is not in this repository. These need a separate change in the Centralis repo:
- The
onboardingtool's estate paragraph ("Four machines on a WireGuard mesh... beta (beta.thermograph.org, Forgejo, Grafana + Loki), the operator's desktop (LAN dev, the main CI runner)"). - The MCP server's own protocol-level
instructionsstring, which carries a second, separately-worded copy of the same paragraph. - The
thermograph-orientationskill's estate table, andthermograph-ops' release-flow diagram ("LAN dev box"), its "ssh -L works for beta and is impossible for prod" claim (beta is a Swarm overlay now, so it behaves like prod), and its prod-only Swarm service-name list (beta hasbeta--prefixed ones now). - Tool descriptions that hardcode the old shape:
fleet_status("Swarm services on prod, compose containers on beta/dev"),estate_statusandrollback_to("Centralis cannot SSH to the dev desktop"),logs_query's per-environment service lists (Forgejo/Grafana/Loki are listed underbeta; they belong underdev/vps1 now), andsecrets_render's assumption that a host's ownsecrets-envmarker identifies one environment — vps2 has two.
If it goes wrong
| symptom | most likely cause | action |
|---|---|---|
| prod 502s after step 5 | beta's stack collided with prod's service DNS | docker stack rm thermograph-beta; prod recovers on its own. Check beta's services are beta--prefixed |
| beta 502s, prod fine | beta's LB or stack not up | docker ps | grep beta-lb; docker stack ps thermograph-beta --no-trunc |
| beta can't reach the database | role/database missing, or the wrong network | re-run provision-env-db.sh beta; confirm beta's tasks are on thermograph_internal |
| beta serves prod's data | beta's env file has prod's URL | check /etc/thermograph-beta.env names thermograph_beta twice; re-render (step 2). Stop beta until fixed |
| prod's nightly backup fails | thermograph_beta doesn't exist yet |
either finish step 3 or revert the ops-cron change; the job fails loudly by design rather than skipping silently |
| dev reachable from the internet | DEV_BIND_ADDR not applied |
ss -ltnp | grep 8137; redeploy via deploy-dev.sh, which sets it from env-topology.sh |