thermograph/.forgejo/workflows/deploy.yml
Emi Griffith 4e97d8e5dc
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
infra: split the estate into vps1/vps2, moving beta next to prod and dev onto vps1
Environments stop being machines. Until now each box WAS an environment --
"beta" named both a deploy target and a host, "the desktop" named both the
operator's computer and the dev server -- so every path could assume one
environment per host and hardcode /opt/thermograph, /etc/thermograph.env and
ports 8137/8080. That assumption ends here:

  vps1  75.119.132.91  Forgejo, Grafana/Loki, the portfolio site, and DEV
                       (own Postgres, mesh-only on 10.10.0.2:8137)
  vps2  169.58.46.181  PROD and BETA as two Swarm stacks sharing one
                       TimescaleDB instance, plus Centralis, Postfix, backups
  desktop              AI model hosting + flex Swarm capacity, no environment

Nothing live has moved; the ordered cutover is in
infra/deploy/RUNBOOK-vps1-vps2-cutover.md.

deploy/env-topology.sh is the single source of truth: env -> host, checkout,
branch, deploy mode, stack name, env file, LB ports, DB role/database, service
prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing
distinguishing a beta deploy from a prod one -- deploy.sh refuses to run when it
disagrees with the checkout it was invoked from. The host-wide secrets-env and
deploy-mode markers survive only as a fallback for a single-environment box.

Beta gets its own stack file rather than an overlay (a merge cannot REMOVE the
db service, and beta having no database of its own is the point). Its services
are prefixed beta-*: Swarm registers a service's short name as a DNS alias on
every network it joins, so two stacks both calling a service `web` on the shared
data network would let prod's frontend resolve a beta task. Prod's stack, env
and LB config are untouched.

One Postgres, two databases with two roles: deploy/db/provision-env-db.sh
creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT
revoked from PUBLIC) plus a <role>_ro for ad-hoc queries, and refuses to run for
the environment that owns the instance -- doing so would demote prod's bootstrap
superuser.

Fixes that co-residency would otherwise have broken silently:

- CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and
  also "the Forgejo box" because those shared a machine; that conflation is what
  once had the prod backup dumping beta.
- The nightly backup dumps BOTH application databases to separate off-box
  prefixes, and fails loudly on a missing one instead of skipping it.
- Alloy stopped deriving the `host` label from the node -- on vps2 that would
  have filed every beta line as prod, feeding prod's alert rules with beta's
  traffic. It is now derived per source, with a new `node` label for the machine,
  and beta's log volume is mounted via a vps2-only overlay.
- ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derived their SSH
  target and env-file path from the topology instead of hardcoding beta to
  75.119.132.91 -- which is vps1's address now.
- Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on
  vps1) instead of 0.0.0.0: correct for a home LAN box, a public exposure of
  unreviewed branches on a VPS.

Terraform's hosts variable is keyed by machine with a nested environments map,
and main.tf flattens (host, environment) pairs so two environments on one box
cannot share a checkout. Caddy's single reference config is split per host.

Docs, onboarding and the runbooks are updated throughout, including the security
rationales that co-residency makes false: separate SSH credentials no longer put
a host boundary between beta and prod, and the boundary that remains is the
database and the filesystem.
2026-07-25 15:01:29 -07:00

212 lines
10 KiB
YAML

name: Deploy
# The six per-domain-per-environment deploy workflows, collapsed into one.
#
# backend-deploy{,-dev,-prod}.yml and frontend-deploy{,-dev,-prod}.yml were the
# same file written six times: they differed only in a branch name, a paths
# filter, a concurrency group, the service name, its *_IMAGE_TAG variable and a
# secret prefix. The contract into infra/deploy/deploy.sh
# (SERVICE + BACKEND_IMAGE_TAG/FRONTEND_IMAGE_TAG) was already fully
# parameterised, so the duplication bought nothing and cost six files to keep in
# step.
#
# DEV IS A REAL DEPLOY TARGET AGAIN. It previously was not: dev meant a
# sudo-free compose stack on the operator's desktop, and the two
# *-deploy-dev.yml workflows that drove it were inert. Dev now lives on vps1 at
# /opt/thermograph-dev — a normal fleet host, reached over SSH exactly like beta
# and prod — so it gets a leg here.
#
# SECRETS ARE KEYED BY HOST, NOT BY ENVIRONMENT. `SSH_*` used to mean "beta" and
# `PROD_SSH_*` "prod", which worked only while each environment owned a box. It
# stopped being true in two directions at once: vps2 now runs beta AND prod, and
# the box `SSH_*` pointed at is now vps1, which runs dev, Forgejo and Grafana.
# The old names would have made "the beta secret" and "the Forgejo box secret"
# the same value by accident — the ops-cron file already records one incident
# caused by exactly that conflation. So: VPS1_SSH_* and VPS2_SSH_*, named for
# the machine, and the environment is passed separately as THERMOGRAPH_ENV.
#
# BETA AND PROD DEPLOYING TO ONE BOX DO NOT RACE. They have separate checkouts
# (/opt/thermograph-beta and /opt/thermograph), so the `git reset --hard` in one
# cannot pull the tree out from under the other, and deploy.sh's flock is per
# checkout. The concurrency groups below stay keyed by ref+service, which keeps
# two pushes to the SAME branch serialised — the property that actually matters.
#
# DELIBERATELY BORING EXPRESSIONS. No dynamic matrix (fromJSON), no
# `cond && secrets.A || secrets.B` ternary. Those are GitHub idioms that a
# Forgejo/act runner may evaluate differently, and the failure mode here is
# "production does not deploy" or, worse, "deploys with an empty SSH host". The
# three environments therefore get three explicit, mutually exclusive steps.
#
# What is preserved from the originals, all of it load-bearing:
# - fetch-depth: 0, because the image tag is keyed to the LAST COMMIT THAT
# TOUCHED THAT DOMAIN, not the branch tip. In a path-filtered monorepo the
# tip is often an unrelated domain's commit and a depth-1 clone cannot see
# past it.
# - The 12-hex truncation, matching build-push exactly.
# - Per-service, per-environment concurrency with cancel-in-progress: false --
# a half-finished deploy must never be cancelled by a newer one.
# - Separate credentials per host. Note what this does and does not buy now:
# it keeps vps1 (Forgejo, its CI, and whatever unreviewed branch dev is
# running) away from vps2 entirely. It no longer puts a host boundary
# between beta and prod, because they share vps2 by design — that boundary
# is now at the database (separate roles and databases) and the filesystem
# (separate checkouts and rendered env files).
# - appleboy/ssh-action by full URL; it is not mirrored in Forgejo's default
# action registry.
on:
push:
branches: [dev, main, release]
paths:
- 'backend/**'
- 'frontend/**'
workflow_dispatch: {}
jobs:
deploy:
strategy:
fail-fast: false
# One physical runner, and deploy.sh serialises host-side with flock
# anyway. Rolling one service at a time keeps the logs readable and
# matches what the six separate workflows effectively did.
max-parallel: 1
matrix:
service: [backend, frontend]
runs-on: docker
# Keyed by ref AND service, reproducing the old per-file groups
# (beta-deploy-backend, prod-deploy-frontend, ...).
concurrency:
group: deploy-${{ github.ref_name }}-${{ matrix.service }}
cancel-in-progress: false
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Plan this leg
id: plan
run: |
set -euo pipefail
# Branch selects the environment; the environment selects the host and
# the script path. dev -> dev (vps1), main -> beta (vps2),
# release -> prod (vps2). The paths differ because each environment
# has its own checkout — two environments on vps2 must never share
# one — and dev has its own entry point for its secrets policy.
case "${{ github.ref_name }}" in
dev) environment=dev ; script=/opt/thermograph-dev/infra/deploy/deploy-dev.sh ;;
main) environment=beta ; script=/opt/thermograph-beta/infra/deploy/deploy.sh ;;
release) environment=prod ; script=/opt/thermograph/infra/deploy/deploy.sh ;;
*) echo "::error::Deploy triggered on unexpected ref '${{ github.ref_name }}'"; exit 1 ;;
esac
# Did THIS push touch THIS domain? The workflow-level paths filter only
# tells us backend OR frontend moved; without this refinement a
# backend-only push would also roll the frontend, losing the
# independent-deploy property the FE/BE split exists for.
before="${{ github.event.before }}"
if [ -z "$before" ] || [ "$before" = "0000000000000000000000000000000000000000" ] \
|| ! git cat-file -e "$before^{commit}" 2>/dev/null; then
# No usable base (first push, force push, or workflow_dispatch).
# Deploy rather than skip: rolling onto the tag already running is a
# no-op for deploy.sh, whereas skipping silently strands a change.
changed=true
echo "no usable before-sha; defaulting to deploy"
elif git diff --name-only "$before" "${{ github.sha }}" | grep -q "^${{ matrix.service }}/"; then
changed=true
else
changed=false
fi
# The tag is the last commit that touched this domain, truncated to 12
# hex to match build-push.yml exactly. Actions expressions have no
# substring function, which is why this is computed in shell.
domain_sha="$(git log -1 --format=%H -- "${{ matrix.service }}/")"
tag="sha-${domain_sha:0:12}"
{
echo "environment=$environment"
echo "script=$script"
echo "changed=$changed"
echo "tag=$tag"
} >> "$GITHUB_OUTPUT"
# Export the deploy.sh contract as real environment variables, chosen
# in shell rather than with a `matrix.service == 'x' && a || b`
# expression. That idiom is a GitHub convention a Forgejo/act runner
# may evaluate differently, and the failure here would be silent: an
# empty *_IMAGE_TAG makes deploy.sh fall back to the tag already
# running, so the job goes green having deployed nothing.
{
echo "SERVICE=${{ matrix.service }}"
if [ "${{ matrix.service }}" = "backend" ]; then
echo "BACKEND_IMAGE_TAG=$tag"
else
echo "FRONTEND_IMAGE_TAG=$tag"
fi
} >> "$GITHUB_ENV"
echo "==> ${{ matrix.service }} -> $environment | changed=$changed | tag=$tag"
- name: Deploy to dev (vps1)
if: steps.plan.outputs.changed == 'true' && steps.plan.outputs.environment == 'dev'
uses: https://github.com/appleboy/ssh-action@v1.2.0
with:
host: ${{ secrets.VPS1_SSH_HOST }}
username: ${{ secrets.VPS1_SSH_USER }}
key: ${{ secrets.VPS1_SSH_KEY }}
port: ${{ secrets.VPS1_SSH_PORT }}
envs: SERVICE,BACKEND_IMAGE_TAG,FRONTEND_IMAGE_TAG,THERMOGRAPH_ENV
script: ${{ steps.plan.outputs.script }}
env:
SERVICE: ${{ matrix.service }}
THERMOGRAPH_ENV: dev
BACKEND_IMAGE_TAG: ${{ matrix.service == 'backend' && steps.plan.outputs.tag || '' }}
FRONTEND_IMAGE_TAG: ${{ matrix.service == 'frontend' && steps.plan.outputs.tag || '' }}
- name: Deploy to beta (vps2)
if: steps.plan.outputs.changed == 'true' && steps.plan.outputs.environment == 'beta'
uses: https://github.com/appleboy/ssh-action@v1.2.0
with:
host: ${{ secrets.VPS2_SSH_HOST }}
username: ${{ secrets.VPS2_SSH_USER }}
key: ${{ secrets.VPS2_SSH_KEY }}
port: ${{ secrets.VPS2_SSH_PORT }}
envs: SERVICE,BACKEND_IMAGE_TAG,FRONTEND_IMAGE_TAG,THERMOGRAPH_ENV
script: ${{ steps.plan.outputs.script }}
env:
SERVICE: ${{ matrix.service }}
# THERMOGRAPH_ENV is what makes this a BETA deploy rather than a prod
# one: same host, same credentials, same script — the environment is
# the only thing that differs, and deploy.sh refuses to run if it
# disagrees with the checkout it was invoked from.
THERMOGRAPH_ENV: beta
# Only the matching one is read by deploy.sh for a single-service roll;
# the other stays empty and the persisted .image-tags.env supplies the
# sibling's live tag, so this roll never disturbs it.
BACKEND_IMAGE_TAG: ${{ matrix.service == 'backend' && steps.plan.outputs.tag || '' }}
FRONTEND_IMAGE_TAG: ${{ matrix.service == 'frontend' && steps.plan.outputs.tag || '' }}
- name: Deploy to prod (vps2)
if: steps.plan.outputs.changed == 'true' && steps.plan.outputs.environment == 'prod'
uses: https://github.com/appleboy/ssh-action@v1.2.0
with:
host: ${{ secrets.VPS2_SSH_HOST }}
username: ${{ secrets.VPS2_SSH_USER }}
key: ${{ secrets.VPS2_SSH_KEY }}
port: ${{ secrets.VPS2_SSH_PORT }}
envs: SERVICE,BACKEND_IMAGE_TAG,FRONTEND_IMAGE_TAG,THERMOGRAPH_ENV
script: ${{ steps.plan.outputs.script }}
env:
SERVICE: ${{ matrix.service }}
THERMOGRAPH_ENV: prod
BACKEND_IMAGE_TAG: ${{ matrix.service == 'backend' && steps.plan.outputs.tag || '' }}
FRONTEND_IMAGE_TAG: ${{ matrix.service == 'frontend' && steps.plan.outputs.tag || '' }}
- name: Skipped
if: steps.plan.outputs.changed != 'true'
run: echo "${{ matrix.service }} unchanged in this push — nothing to roll."