2026-07-11 00:29:47 +00:00
|
|
|
#!/usr/bin/env bash
|
2026-07-22 18:39:16 +00:00
|
|
|
# Pull this INFRA repo's checkout up to date, then roll ONE (or all) of the
|
|
|
|
|
# docker-compose app services onto its separately-published image. Run on the
|
|
|
|
|
# VPS — each app repo's Forgejo Actions workflow invokes this over SSH (see
|
|
|
|
|
# .forgejo/workflows/deploy.yml in thermograph-backend / thermograph-frontend),
|
|
|
|
|
# and you can run it by hand too.
|
2026-07-11 00:29:47 +00:00
|
|
|
#
|
2026-07-22 18:39:16 +00:00
|
|
|
# # roll just the backend onto a specific image:
|
2026-07-23 05:11:33 +00:00
|
|
|
# ssh deploy@vps 'SERVICE=backend BACKEND_IMAGE_TAG=sha-<12hex> /opt/thermograph/infra/deploy/deploy.sh'
|
2026-07-22 18:39:16 +00:00
|
|
|
# # roll just the frontend:
|
2026-07-23 05:11:33 +00:00
|
|
|
# ssh deploy@vps 'SERVICE=frontend FRONTEND_IMAGE_TAG=sha-<12hex> /opt/thermograph/infra/deploy/deploy.sh'
|
2026-07-22 18:39:16 +00:00
|
|
|
# # bring the whole stack up (both tags required):
|
2026-07-23 05:11:33 +00:00
|
|
|
# ssh deploy@vps 'SERVICE=all BACKEND_IMAGE_TAG=sha-<a> FRONTEND_IMAGE_TAG=sha-<b> /opt/thermograph/infra/deploy/deploy.sh'
|
Decouple Terraform from the app repo; add a GCP host scaffold
Content-change pass following the extraction from the app monorepo (this repo
now stands alone, sourced via git filter-repo to preserve history):
- terraform/variables.tf, secrets.tf, modules/thermograph-host: remove every
app-secret Terraform variable (postgres_password, auth_secret, VAPID keys,
registry_token, Discord/SMTP creds, ...) and the random_password/random_id
generators. The SOPS+age vault (deploy/secrets/*.yaml) is now the sole
source of app secrets, rendered at deploy time by deploy/render-secrets.sh;
Terraform renders only a non-secret /etc/thermograph-topology.env (sizing,
routing) via the renamed thermograph-topology.env.tftpl template.
- hosts gains a required app_image_tag field: the host's own checkout is now
this infra repo, not the app repo, so there is no "current commit" to
derive an image tag from — every host pins one explicitly. repo_url now
points at this repo (private; typically needs an embedded read token).
- deploy.sh: IMAGE_TAG is now required from the environment instead of
derived via `git rev-parse HEAD` of the (now infra-repo) checkout, which
would have silently resolved to the wrong or a nonexistent tag.
- New terraform/modules/gcp-host: creates a GCE VM + minimal VPC/firewall
only, then feeds its IP into the same thermograph-host module every
SSH-managed host already uses — one provisioning path regardless of how a
host came to exist. var.gcp_hosts defaults to {}, so no google_* resource
is planned and the provider is never invoked without it (verified: plan
and validate succeed with no GCP credentials configured).
- terraform/README.md, ACCESS.md (renamed from INFRA.md), README.md: updated
for the new secrets model, the GCP scaffold, and this repo's own identity.
Verified: terraform fmt/validate/init clean; plan succeeds against realistic
dummy hosts (prod+beta shape) and against a populated gcp_hosts entry (plans
6 resources with no live credentials, confirming the composition wires
correctly end to end).
2026-07-22 04:46:05 +00:00
|
|
|
#
|
2026-07-22 18:39:16 +00:00
|
|
|
# FE/BE CI-CD split: backend and frontend are published from separate repos as
|
2026-07-23 05:11:33 +00:00
|
|
|
# separate images (emi/thermograph/backend, emi/thermograph/frontend),
|
2026-07-22 18:39:16 +00:00
|
|
|
# so a deploy targets ONE service and leaves the other's running container +
|
|
|
|
|
# tag untouched. Each service's live tag is persisted host-side in
|
|
|
|
|
# deploy/.image-tags.env (untracked -- survives the git reset below) so a
|
|
|
|
|
# single-service roll re-renders compose with BOTH services' real tags and
|
|
|
|
|
# never accidentally recreates or downgrades the sibling.
|
|
|
|
|
#
|
2026-07-23 05:11:33 +00:00
|
|
|
# This checkout is the monorepo (compose + deploy live under infra/), not an app-only checkout: BRANCH is this repo's
|
Decouple Terraform from the app repo; add a GCP host scaffold
Content-change pass following the extraction from the app monorepo (this repo
now stands alone, sourced via git filter-repo to preserve history):
- terraform/variables.tf, secrets.tf, modules/thermograph-host: remove every
app-secret Terraform variable (postgres_password, auth_secret, VAPID keys,
registry_token, Discord/SMTP creds, ...) and the random_password/random_id
generators. The SOPS+age vault (deploy/secrets/*.yaml) is now the sole
source of app secrets, rendered at deploy time by deploy/render-secrets.sh;
Terraform renders only a non-secret /etc/thermograph-topology.env (sizing,
routing) via the renamed thermograph-topology.env.tftpl template.
- hosts gains a required app_image_tag field: the host's own checkout is now
this infra repo, not the app repo, so there is no "current commit" to
derive an image tag from — every host pins one explicitly. repo_url now
points at this repo (private; typically needs an embedded read token).
- deploy.sh: IMAGE_TAG is now required from the environment instead of
derived via `git rev-parse HEAD` of the (now infra-repo) checkout, which
would have silently resolved to the wrong or a nonexistent tag.
- New terraform/modules/gcp-host: creates a GCE VM + minimal VPC/firewall
only, then feeds its IP into the same thermograph-host module every
SSH-managed host already uses — one provisioning path regardless of how a
host came to exist. var.gcp_hosts defaults to {}, so no google_* resource
is planned and the provider is never invoked without it (verified: plan
and validate succeed with no GCP credentials configured).
- terraform/README.md, ACCESS.md (renamed from INFRA.md), README.md: updated
for the new secrets model, the GCP scaffold, and this repo's own identity.
Verified: terraform fmt/validate/init clean; plan succeeds against realistic
dummy hosts (prod+beta shape) and against a populated gcp_hosts entry (plans
6 resources with no live credentials, confirming the composition wires
correctly end to end).
2026-07-22 04:46:05 +00:00
|
|
|
# branch (compose files, db init, the secrets vault, this script itself);
|
2026-07-22 18:39:16 +00:00
|
|
|
# the *_IMAGE_TAG values are the separately-published app images to run -- the
|
|
|
|
|
# two axes are independent and rarely change together.
|
2026-07-11 00:29:47 +00:00
|
|
|
set -euo pipefail
|
|
|
|
|
|
infra: split the estate into vps1/vps2, moving beta next to prod and dev onto vps1
Environments stop being machines. Until now each box WAS an environment --
"beta" named both a deploy target and a host, "the desktop" named both the
operator's computer and the dev server -- so every path could assume one
environment per host and hardcode /opt/thermograph, /etc/thermograph.env and
ports 8137/8080. That assumption ends here:
vps1 75.119.132.91 Forgejo, Grafana/Loki, the portfolio site, and DEV
(own Postgres, mesh-only on 10.10.0.2:8137)
vps2 169.58.46.181 PROD and BETA as two Swarm stacks sharing one
TimescaleDB instance, plus Centralis, Postfix, backups
desktop AI model hosting + flex Swarm capacity, no environment
Nothing live has moved; the ordered cutover is in
infra/deploy/RUNBOOK-vps1-vps2-cutover.md.
deploy/env-topology.sh is the single source of truth: env -> host, checkout,
branch, deploy mode, stack name, env file, LB ports, DB role/database, service
prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing
distinguishing a beta deploy from a prod one -- deploy.sh refuses to run when it
disagrees with the checkout it was invoked from. The host-wide secrets-env and
deploy-mode markers survive only as a fallback for a single-environment box.
Beta gets its own stack file rather than an overlay (a merge cannot REMOVE the
db service, and beta having no database of its own is the point). Its services
are prefixed beta-*: Swarm registers a service's short name as a DNS alias on
every network it joins, so two stacks both calling a service `web` on the shared
data network would let prod's frontend resolve a beta task. Prod's stack, env
and LB config are untouched.
One Postgres, two databases with two roles: deploy/db/provision-env-db.sh
creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT
revoked from PUBLIC) plus a <role>_ro for ad-hoc queries, and refuses to run for
the environment that owns the instance -- doing so would demote prod's bootstrap
superuser.
Fixes that co-residency would otherwise have broken silently:
- CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and
also "the Forgejo box" because those shared a machine; that conflation is what
once had the prod backup dumping beta.
- The nightly backup dumps BOTH application databases to separate off-box
prefixes, and fails loudly on a missing one instead of skipping it.
- Alloy stopped deriving the `host` label from the node -- on vps2 that would
have filed every beta line as prod, feeding prod's alert rules with beta's
traffic. It is now derived per source, with a new `node` label for the machine,
and beta's log volume is mounted via a vps2-only overlay.
- ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derived their SSH
target and env-file path from the topology instead of hardcoding beta to
75.119.132.91 -- which is vps1's address now.
- Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on
vps1) instead of 0.0.0.0: correct for a home LAN box, a public exposure of
unreviewed branches on a VPS.
Terraform's hosts variable is keyed by machine with a nested environments map,
and main.tf flattens (host, environment) pairs so two environments on one box
cannot share a checkout. Caddy's single reference config is split per host.
Docs, onboarding and the runbooks are updated throughout, including the security
rationales that co-residency makes false: separate SSH credentials no longer put
a host boundary between beta and prod, and the boundary that remains is the
database and the filesystem.
2026-07-25 22:01:29 +00:00
|
|
|
# --- which environment is this? ------------------------------------------------
|
|
|
|
|
# vps2 runs beta AND prod, so "the host" no longer answers this — the caller does,
|
|
|
|
|
# via THERMOGRAPH_ENV (the deploy workflow always passes it). A by-hand run on a
|
|
|
|
|
# single-environment box still falls back to the host marker, and a box with
|
|
|
|
|
# neither still defaults to prod's historical paths, so nothing about an existing
|
|
|
|
|
# single-env host changes.
|
|
|
|
|
#
|
|
|
|
|
# Resolved from the SCRIPT'S OWN LOCATION, not a hardcoded /opt/thermograph:
|
|
|
|
|
# invoking /opt/thermograph-beta/infra/deploy/deploy.sh must act on the beta
|
|
|
|
|
# checkout even if something in the environment says otherwise. That is also the
|
|
|
|
|
# check that catches the one genuinely dangerous mistake on a two-environment
|
|
|
|
|
# host — running prod's checkout with THERMOGRAPH_ENV=beta, or the reverse.
|
|
|
|
|
SELF_DIR=$(cd "$(dirname "$0")" && pwd)
|
|
|
|
|
SELF_APP_DIR=$(cd "$SELF_DIR/../.." && pwd)
|
|
|
|
|
|
|
|
|
|
# Guarded exactly like render-secrets.sh below: the deploy that INTRODUCES this
|
|
|
|
|
# file runs with a checkout that predates it (it arrives with the git reset
|
|
|
|
|
# further down, after which deploy.sh re-execs). Missing => keep the pre-split
|
|
|
|
|
# behaviour rather than fail.
|
|
|
|
|
if [ -f "$SELF_DIR/env-topology.sh" ]; then
|
|
|
|
|
# shellcheck source=infra/deploy/env-topology.sh
|
|
|
|
|
. "$SELF_DIR/env-topology.sh"
|
|
|
|
|
ENV_NAME=$(thermograph_env_name)
|
|
|
|
|
# Nothing to go on anywhere: this is a pre-split host whose paths are prod's.
|
|
|
|
|
[ -n "$ENV_NAME" ] || ENV_NAME=prod
|
|
|
|
|
thermograph_topology "$ENV_NAME"
|
|
|
|
|
if [ "$SELF_APP_DIR" != "$TG_APP_DIR" ] && [ -z "${APP_DIR:-}" ]; then
|
|
|
|
|
echo "!! environment/checkout mismatch: THERMOGRAPH_ENV=$ENV_NAME expects" >&2
|
|
|
|
|
echo "!! $TG_APP_DIR but this script lives in $SELF_APP_DIR." >&2
|
|
|
|
|
echo "!! On vps2 that means beta and prod have been crossed. Refusing to deploy." >&2
|
|
|
|
|
echo "!! If this checkout really is $ENV_NAME (a rehearsal copy, a relocated" >&2
|
|
|
|
|
echo "!! checkout), say so explicitly: APP_DIR=$SELF_APP_DIR ..." >&2
|
|
|
|
|
exit 2
|
|
|
|
|
fi
|
|
|
|
|
APP_DIR="${APP_DIR:-$TG_APP_DIR}"
|
|
|
|
|
BRANCH="${BRANCH:-$TG_BRANCH}"
|
|
|
|
|
ENV_FILE="$TG_ENV_FILE"
|
|
|
|
|
DEPLOY_MODE="$TG_DEPLOY_MODE"
|
|
|
|
|
[ "$TG_SKIP_COMMON" = 1 ] && export THERMOGRAPH_SECRETS_SKIP_COMMON=1
|
|
|
|
|
# The second pass (after the re-exec) must resolve the SAME environment, and
|
|
|
|
|
# the stack path needs it too.
|
|
|
|
|
export THERMOGRAPH_ENV="$ENV_NAME"
|
|
|
|
|
else
|
|
|
|
|
ENV_NAME=""
|
|
|
|
|
APP_DIR="${APP_DIR:-/opt/thermograph}"
|
|
|
|
|
BRANCH="${BRANCH:-main}"
|
|
|
|
|
ENV_FILE=/etc/thermograph.env
|
|
|
|
|
DEPLOY_MODE=""
|
|
|
|
|
fi
|
|
|
|
|
|
2026-07-23 05:11:33 +00:00
|
|
|
# Monorepo layout: git operations act on the checkout root ($APP_DIR); all
|
|
|
|
|
# compose files and deploy assets live under infra/.
|
|
|
|
|
INFRA_DIR="$APP_DIR/infra"
|
2026-07-22 18:39:16 +00:00
|
|
|
# Which service this deploy rolls: backend | frontend | all. Defaults to `all`
|
|
|
|
|
# (a full-stack bring-up) so a by-hand run with both tags still works; the
|
|
|
|
|
# per-repo deploy.yml workflows always pass an explicit single service.
|
|
|
|
|
SERVICE="${SERVICE:-all}"
|
Containerize the app and move the databases to PostgreSQL 18 (#220)
Run Thermograph as a docker-compose stack (app + Postgres 18) and standardize the
data layer on Postgres, while keeping the test suite on SQLite.
- accounts/db.py: DSN-driven engines. On Postgres, a per-worker read-write +
read-only asyncpg pair (the RO engine pins read-only transactions, used by the
pure-GET endpoints) plus a sync psycopg engine for the notifier thread; the
SQLite path is preserved for tests/local (selected when THERMOGRAPH_DATABASE_URL
is unset). models.py: boolean server_default -> sa.false().
- store.py / metrics.py: dialect-flexible — Postgres UNLOGGED tables via psycopg
when configured, else the existing raw-sqlite3 paths byte-for-byte; sync
interfaces and every fail-soft contract preserved.
- Alembic (backend/alembic/) manages the accounts schema; the container entrypoint
runs `alembic upgrade head` before uvicorn (4 workers). migrate_accounts_to_pg.py
copies the accounts data SQLite->PG through the ORM (UUID/bool/JSON coerced),
skips access_token, and resets identity sequences.
- Dockerfile + docker-compose.yml: app image (uvicorn, 4 workers, loopback 8137)
and a Postgres 18 db (2 CPUs) running pg_duckdb (deploy/db/) so the parquet
climate cache is queryable in-DB via read_parquet('/parquet/cache/*.parquet').
- deploy.sh/thermograph.service rewired to manage the compose stack; env example,
Makefile targets (up/down/db-up), and deploy/POSTGRES-MIGRATION.md cutover runbook.
Tests stay on SQLite (dialect fallback) — 323 pass. The full Postgres stack was
verified via docker compose: alembic migrations, register/login, the RO endpoint,
store/metrics round-trips, and the accounts data migration.
2026-07-20 06:28:23 +00:00
|
|
|
HEALTH_PORT="${HEALTH_PORT:-8137}"
|
2026-07-11 00:29:47 +00:00
|
|
|
cd "$APP_DIR"
|
|
|
|
|
|
2026-07-22 18:39:16 +00:00
|
|
|
case "$SERVICE" in
|
|
|
|
|
backend|frontend|all) ;;
|
|
|
|
|
*) echo "!! SERVICE must be backend|frontend|all, got '$SERVICE'" >&2; exit 2 ;;
|
|
|
|
|
esac
|
|
|
|
|
|
2026-07-22 23:39:29 +00:00
|
|
|
# Serialize deploys on this host. Backend and frontend deploy from SEPARATE
|
|
|
|
|
# repos whose workflows can fire for the same push within seconds of each
|
|
|
|
|
# other, and both SSH into this one checkout: concurrent runs race on the
|
|
|
|
|
# `git reset` below, the shared compose project, and the .image-tags.env
|
|
|
|
|
# read/modify/write (a lost update there re-rolls the sibling onto a stale
|
|
|
|
|
# tag). flock makes the second deploy wait its turn instead. The lock fd is
|
|
|
|
|
# inherited across the self re-exec below, so the lock spans the whole run;
|
|
|
|
|
# -w 600 bounds the wait (a deploy holding the lock >10 min is already
|
|
|
|
|
# broken), after which this exits non-zero and the CI job fails loudly.
|
2026-07-23 05:11:33 +00:00
|
|
|
DEPLOY_LOCK="$INFRA_DIR/deploy/.deploy.lock"
|
2026-07-22 23:39:29 +00:00
|
|
|
if [ -z "${DEPLOY_SH_FLOCKED:-}" ]; then
|
|
|
|
|
export DEPLOY_SH_FLOCKED=1
|
|
|
|
|
exec flock -w 600 "$DEPLOY_LOCK" "$0" "$@"
|
|
|
|
|
fi
|
|
|
|
|
|
Containerize the app and move the databases to PostgreSQL 18 (#220)
Run Thermograph as a docker-compose stack (app + Postgres 18) and standardize the
data layer on Postgres, while keeping the test suite on SQLite.
- accounts/db.py: DSN-driven engines. On Postgres, a per-worker read-write +
read-only asyncpg pair (the RO engine pins read-only transactions, used by the
pure-GET endpoints) plus a sync psycopg engine for the notifier thread; the
SQLite path is preserved for tests/local (selected when THERMOGRAPH_DATABASE_URL
is unset). models.py: boolean server_default -> sa.false().
- store.py / metrics.py: dialect-flexible — Postgres UNLOGGED tables via psycopg
when configured, else the existing raw-sqlite3 paths byte-for-byte; sync
interfaces and every fail-soft contract preserved.
- Alembic (backend/alembic/) manages the accounts schema; the container entrypoint
runs `alembic upgrade head` before uvicorn (4 workers). migrate_accounts_to_pg.py
copies the accounts data SQLite->PG through the ORM (UUID/bool/JSON coerced),
skips access_token, and resets identity sequences.
- Dockerfile + docker-compose.yml: app image (uvicorn, 4 workers, loopback 8137)
and a Postgres 18 db (2 CPUs) running pg_duckdb (deploy/db/) so the parquet
climate cache is queryable in-DB via read_parquet('/parquet/cache/*.parquet').
- deploy.sh/thermograph.service rewired to manage the compose stack; env example,
Makefile targets (up/down/db-up), and deploy/POSTGRES-MIGRATION.md cutover runbook.
Tests stay on SQLite (dialect fallback) — 323 pass. The full Postgres stack was
verified via docker compose: alembic migrations, register/login, the RO endpoint,
store/metrics round-trips, and the accounts data migration.
2026-07-20 06:28:23 +00:00
|
|
|
# Secrets (POSTGRES_PASSWORD, VAPID keys, AUTH_SECRET, ...) drive compose
|
2026-07-21 20:01:30 +00:00
|
|
|
# interpolation and are also loaded into the backend container via env_file.
|
2026-07-22 03:21:46 +00:00
|
|
|
#
|
|
|
|
|
# When this host is configured for SOPS (an age key + /etc/thermograph/secrets-env),
|
|
|
|
|
# first render /etc/thermograph.env from the committed encrypted source of truth
|
|
|
|
|
# (deploy/secrets/*.yaml) so a key rotation is just an edit+commit+deploy. The guard
|
|
|
|
|
# on the helper's existence keeps the very deploy that INTRODUCES this file safe: on
|
|
|
|
|
# the first pass the checkout may predate it (it arrives with the git reset below,
|
|
|
|
|
# after which deploy.sh re-execs), so a missing helper simply falls back to the
|
|
|
|
|
# existing /etc/thermograph.env. Then source it so a by-hand run interpolates the
|
|
|
|
|
# same as the systemd unit does. See deploy/render-secrets.sh + deploy/secrets/.
|
infra: split the estate into vps1/vps2, moving beta next to prod and dev onto vps1
Environments stop being machines. Until now each box WAS an environment --
"beta" named both a deploy target and a host, "the desktop" named both the
operator's computer and the dev server -- so every path could assume one
environment per host and hardcode /opt/thermograph, /etc/thermograph.env and
ports 8137/8080. That assumption ends here:
vps1 75.119.132.91 Forgejo, Grafana/Loki, the portfolio site, and DEV
(own Postgres, mesh-only on 10.10.0.2:8137)
vps2 169.58.46.181 PROD and BETA as two Swarm stacks sharing one
TimescaleDB instance, plus Centralis, Postfix, backups
desktop AI model hosting + flex Swarm capacity, no environment
Nothing live has moved; the ordered cutover is in
infra/deploy/RUNBOOK-vps1-vps2-cutover.md.
deploy/env-topology.sh is the single source of truth: env -> host, checkout,
branch, deploy mode, stack name, env file, LB ports, DB role/database, service
prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing
distinguishing a beta deploy from a prod one -- deploy.sh refuses to run when it
disagrees with the checkout it was invoked from. The host-wide secrets-env and
deploy-mode markers survive only as a fallback for a single-environment box.
Beta gets its own stack file rather than an overlay (a merge cannot REMOVE the
db service, and beta having no database of its own is the point). Its services
are prefixed beta-*: Swarm registers a service's short name as a DNS alias on
every network it joins, so two stacks both calling a service `web` on the shared
data network would let prod's frontend resolve a beta task. Prod's stack, env
and LB config are untouched.
One Postgres, two databases with two roles: deploy/db/provision-env-db.sh
creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT
revoked from PUBLIC) plus a <role>_ro for ad-hoc queries, and refuses to run for
the environment that owns the instance -- doing so would demote prod's bootstrap
superuser.
Fixes that co-residency would otherwise have broken silently:
- CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and
also "the Forgejo box" because those shared a machine; that conflation is what
once had the prod backup dumping beta.
- The nightly backup dumps BOTH application databases to separate off-box
prefixes, and fails loudly on a missing one instead of skipping it.
- Alloy stopped deriving the `host` label from the node -- on vps2 that would
have filed every beta line as prod, feeding prod's alert rules with beta's
traffic. It is now derived per source, with a new `node` label for the machine,
and beta's log volume is mounted via a vps2-only overlay.
- ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derived their SSH
target and env-file path from the topology instead of hardcoding beta to
75.119.132.91 -- which is vps1's address now.
- Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on
vps1) instead of 0.0.0.0: correct for a home LAN box, a public exposure of
unreviewed branches on a VPS.
Terraform's hosts variable is keyed by machine with a nested environments map,
and main.tf flattens (host, environment) pairs so two environments on one box
cannot share a checkout. Caddy's single reference config is split per host.
Docs, onboarding and the runbooks are updated throughout, including the security
rationales that co-residency makes false: separate SSH credentials no longer put
a host boundary between beta and prod, and the boundary that remains is the
database and the filesystem.
2026-07-25 22:01:29 +00:00
|
|
|
#
|
|
|
|
|
# $ENV_FILE, not a hardcoded /etc/thermograph.env: on vps2 prod renders
|
|
|
|
|
# prod.yaml there while beta renders beta.yaml to /etc/thermograph-beta.env.
|
|
|
|
|
# One file per environment, never shared.
|
2026-07-23 05:11:33 +00:00
|
|
|
if [ -f "$INFRA_DIR/deploy/render-secrets.sh" ]; then
|
2026-07-23 22:26:05 +00:00
|
|
|
# shellcheck source=infra/deploy/render-secrets.sh
|
2026-07-23 05:11:33 +00:00
|
|
|
. "$INFRA_DIR/deploy/render-secrets.sh"
|
infra: split the estate into vps1/vps2, moving beta next to prod and dev onto vps1
Environments stop being machines. Until now each box WAS an environment --
"beta" named both a deploy target and a host, "the desktop" named both the
operator's computer and the dev server -- so every path could assume one
environment per host and hardcode /opt/thermograph, /etc/thermograph.env and
ports 8137/8080. That assumption ends here:
vps1 75.119.132.91 Forgejo, Grafana/Loki, the portfolio site, and DEV
(own Postgres, mesh-only on 10.10.0.2:8137)
vps2 169.58.46.181 PROD and BETA as two Swarm stacks sharing one
TimescaleDB instance, plus Centralis, Postfix, backups
desktop AI model hosting + flex Swarm capacity, no environment
Nothing live has moved; the ordered cutover is in
infra/deploy/RUNBOOK-vps1-vps2-cutover.md.
deploy/env-topology.sh is the single source of truth: env -> host, checkout,
branch, deploy mode, stack name, env file, LB ports, DB role/database, service
prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing
distinguishing a beta deploy from a prod one -- deploy.sh refuses to run when it
disagrees with the checkout it was invoked from. The host-wide secrets-env and
deploy-mode markers survive only as a fallback for a single-environment box.
Beta gets its own stack file rather than an overlay (a merge cannot REMOVE the
db service, and beta having no database of its own is the point). Its services
are prefixed beta-*: Swarm registers a service's short name as a DNS alias on
every network it joins, so two stacks both calling a service `web` on the shared
data network would let prod's frontend resolve a beta task. Prod's stack, env
and LB config are untouched.
One Postgres, two databases with two roles: deploy/db/provision-env-db.sh
creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT
revoked from PUBLIC) plus a <role>_ro for ad-hoc queries, and refuses to run for
the environment that owns the instance -- doing so would demote prod's bootstrap
superuser.
Fixes that co-residency would otherwise have broken silently:
- CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and
also "the Forgejo box" because those shared a machine; that conflation is what
once had the prod backup dumping beta.
- The nightly backup dumps BOTH application databases to separate off-box
prefixes, and fails loudly on a missing one instead of skipping it.
- Alloy stopped deriving the `host` label from the node -- on vps2 that would
have filed every beta line as prod, feeding prod's alert rules with beta's
traffic. It is now derived per source, with a new `node` label for the machine,
and beta's log volume is mounted via a vps2-only overlay.
- ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derived their SSH
target and env-file path from the topology instead of hardcoding beta to
75.119.132.91 -- which is vps1's address now.
- Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on
vps1) instead of 0.0.0.0: correct for a home LAN box, a public exposure of
unreviewed branches on a VPS.
Terraform's hosts variable is keyed by machine with a nested environments map,
and main.tf flattens (host, environment) pairs so two environments on one box
cannot share a checkout. Caddy's single reference config is split per host.
Docs, onboarding and the runbooks are updated throughout, including the security
rationales that co-residency makes false: separate SSH credentials no longer put
a host boundary between beta and prod, and the boundary that remains is the
database and the filesystem.
2026-07-25 22:01:29 +00:00
|
|
|
render_thermograph_secrets "$INFRA_DIR" "$ENV_NAME" "$ENV_FILE"
|
2026-07-22 03:21:46 +00:00
|
|
|
fi
|
infra: split the estate into vps1/vps2, moving beta next to prod and dev onto vps1
Environments stop being machines. Until now each box WAS an environment --
"beta" named both a deploy target and a host, "the desktop" named both the
operator's computer and the dev server -- so every path could assume one
environment per host and hardcode /opt/thermograph, /etc/thermograph.env and
ports 8137/8080. That assumption ends here:
vps1 75.119.132.91 Forgejo, Grafana/Loki, the portfolio site, and DEV
(own Postgres, mesh-only on 10.10.0.2:8137)
vps2 169.58.46.181 PROD and BETA as two Swarm stacks sharing one
TimescaleDB instance, plus Centralis, Postfix, backups
desktop AI model hosting + flex Swarm capacity, no environment
Nothing live has moved; the ordered cutover is in
infra/deploy/RUNBOOK-vps1-vps2-cutover.md.
deploy/env-topology.sh is the single source of truth: env -> host, checkout,
branch, deploy mode, stack name, env file, LB ports, DB role/database, service
prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing
distinguishing a beta deploy from a prod one -- deploy.sh refuses to run when it
disagrees with the checkout it was invoked from. The host-wide secrets-env and
deploy-mode markers survive only as a fallback for a single-environment box.
Beta gets its own stack file rather than an overlay (a merge cannot REMOVE the
db service, and beta having no database of its own is the point). Its services
are prefixed beta-*: Swarm registers a service's short name as a DNS alias on
every network it joins, so two stacks both calling a service `web` on the shared
data network would let prod's frontend resolve a beta task. Prod's stack, env
and LB config are untouched.
One Postgres, two databases with two roles: deploy/db/provision-env-db.sh
creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT
revoked from PUBLIC) plus a <role>_ro for ad-hoc queries, and refuses to run for
the environment that owns the instance -- doing so would demote prod's bootstrap
superuser.
Fixes that co-residency would otherwise have broken silently:
- CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and
also "the Forgejo box" because those shared a machine; that conflation is what
once had the prod backup dumping beta.
- The nightly backup dumps BOTH application databases to separate off-box
prefixes, and fails loudly on a missing one instead of skipping it.
- Alloy stopped deriving the `host` label from the node -- on vps2 that would
have filed every beta line as prod, feeding prod's alert rules with beta's
traffic. It is now derived per source, with a new `node` label for the machine,
and beta's log volume is mounted via a vps2-only overlay.
- ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derived their SSH
target and env-file path from the topology instead of hardcoding beta to
75.119.132.91 -- which is vps1's address now.
- Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on
vps1) instead of 0.0.0.0: correct for a home LAN box, a public exposure of
unreviewed branches on a VPS.
Terraform's hosts variable is keyed by machine with a nested environments map,
and main.tf flattens (host, environment) pairs so two environments on one box
cannot share a checkout. Caddy's single reference config is split per host.
Docs, onboarding and the runbooks are updated throughout, including the security
rationales that co-residency makes false: separate SSH credentials no longer put
a host boundary between beta and prod, and the boundary that remains is the
database and the filesystem.
2026-07-25 22:01:29 +00:00
|
|
|
# $ENV_FILE is rendered at deploy time from the SOPS vault — it cannot exist at
|
|
|
|
|
# lint time, so don't ask shellcheck to follow it.
|
2026-07-23 22:26:05 +00:00
|
|
|
set -a
|
|
|
|
|
# shellcheck source=/dev/null
|
infra: split the estate into vps1/vps2, moving beta next to prod and dev onto vps1
Environments stop being machines. Until now each box WAS an environment --
"beta" named both a deploy target and a host, "the desktop" named both the
operator's computer and the dev server -- so every path could assume one
environment per host and hardcode /opt/thermograph, /etc/thermograph.env and
ports 8137/8080. That assumption ends here:
vps1 75.119.132.91 Forgejo, Grafana/Loki, the portfolio site, and DEV
(own Postgres, mesh-only on 10.10.0.2:8137)
vps2 169.58.46.181 PROD and BETA as two Swarm stacks sharing one
TimescaleDB instance, plus Centralis, Postfix, backups
desktop AI model hosting + flex Swarm capacity, no environment
Nothing live has moved; the ordered cutover is in
infra/deploy/RUNBOOK-vps1-vps2-cutover.md.
deploy/env-topology.sh is the single source of truth: env -> host, checkout,
branch, deploy mode, stack name, env file, LB ports, DB role/database, service
prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing
distinguishing a beta deploy from a prod one -- deploy.sh refuses to run when it
disagrees with the checkout it was invoked from. The host-wide secrets-env and
deploy-mode markers survive only as a fallback for a single-environment box.
Beta gets its own stack file rather than an overlay (a merge cannot REMOVE the
db service, and beta having no database of its own is the point). Its services
are prefixed beta-*: Swarm registers a service's short name as a DNS alias on
every network it joins, so two stacks both calling a service `web` on the shared
data network would let prod's frontend resolve a beta task. Prod's stack, env
and LB config are untouched.
One Postgres, two databases with two roles: deploy/db/provision-env-db.sh
creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT
revoked from PUBLIC) plus a <role>_ro for ad-hoc queries, and refuses to run for
the environment that owns the instance -- doing so would demote prod's bootstrap
superuser.
Fixes that co-residency would otherwise have broken silently:
- CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and
also "the Forgejo box" because those shared a machine; that conflation is what
once had the prod backup dumping beta.
- The nightly backup dumps BOTH application databases to separate off-box
prefixes, and fails loudly on a missing one instead of skipping it.
- Alloy stopped deriving the `host` label from the node -- on vps2 that would
have filed every beta line as prod, feeding prod's alert rules with beta's
traffic. It is now derived per source, with a new `node` label for the machine,
and beta's log volume is mounted via a vps2-only overlay.
- ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derived their SSH
target and env-file path from the topology instead of hardcoding beta to
75.119.132.91 -- which is vps1's address now.
- Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on
vps1) instead of 0.0.0.0: correct for a home LAN box, a public exposure of
unreviewed branches on a VPS.
Terraform's hosts variable is keyed by machine with a nested environments map,
and main.tf flattens (host, environment) pairs so two environments on one box
cannot share a checkout. Caddy's single reference config is split per host.
Docs, onboarding and the runbooks are updated throughout, including the security
rationales that co-residency makes false: separate SSH credentials no longer put
a host boundary between beta and prod, and the boundary that remains is the
database and the filesystem.
2026-07-25 22:01:29 +00:00
|
|
|
. "$ENV_FILE" 2>/dev/null || true
|
2026-07-23 22:26:05 +00:00
|
|
|
set +a
|
Containerize the app and move the databases to PostgreSQL 18 (#220)
Run Thermograph as a docker-compose stack (app + Postgres 18) and standardize the
data layer on Postgres, while keeping the test suite on SQLite.
- accounts/db.py: DSN-driven engines. On Postgres, a per-worker read-write +
read-only asyncpg pair (the RO engine pins read-only transactions, used by the
pure-GET endpoints) plus a sync psycopg engine for the notifier thread; the
SQLite path is preserved for tests/local (selected when THERMOGRAPH_DATABASE_URL
is unset). models.py: boolean server_default -> sa.false().
- store.py / metrics.py: dialect-flexible — Postgres UNLOGGED tables via psycopg
when configured, else the existing raw-sqlite3 paths byte-for-byte; sync
interfaces and every fail-soft contract preserved.
- Alembic (backend/alembic/) manages the accounts schema; the container entrypoint
runs `alembic upgrade head` before uvicorn (4 workers). migrate_accounts_to_pg.py
copies the accounts data SQLite->PG through the ORM (UUID/bool/JSON coerced),
skips access_token, and resets identity sequences.
- Dockerfile + docker-compose.yml: app image (uvicorn, 4 workers, loopback 8137)
and a Postgres 18 db (2 CPUs) running pg_duckdb (deploy/db/) so the parquet
climate cache is queryable in-DB via read_parquet('/parquet/cache/*.parquet').
- deploy.sh/thermograph.service rewired to manage the compose stack; env example,
Makefile targets (up/down/db-up), and deploy/POSTGRES-MIGRATION.md cutover runbook.
Tests stay on SQLite (dialect fallback) — 323 pass. The full Postgres stack was
verified via docker compose: alembic migrations, register/login, the RO endpoint,
store/metrics round-trips, and the accounts data migration.
2026-07-20 06:28:23 +00:00
|
|
|
|
SEO: add 250 English-market city pages; auto-warm archives on deploy (#97)
The population-ranked global top-500 skewed to Asian megacities and missed
high-English-search-demand cities. gen_cities.py now tops up with the top ~250
cities from English-speaking countries (US/GB/CA/AU/NZ/IE/ZA) not already in the
global set, so US coverage goes 13->146, GB 2->42, CA 3->29, etc. (Seattle, Boston,
Manchester, Melbourne, Auckland, Dublin, ...). cities.json regenerated to 750.
Both deploy scripts now launch warm_cities.py automatically after the health check,
detached (dev: a systemd --user transient unit; prod: setsid/nohup), so the city
pages serve from cache without a manual step; idempotent, so only the first deploy
does the full warm. DEPLOY.md updated.
2026-07-16 00:11:14 +00:00
|
|
|
# Pre-warm the ~750 city-page archives so /climate pages serve from cache and a
|
2026-07-21 20:01:30 +00:00
|
|
|
# search-engine crawl never bursts the archive API quota. Detached inside the
|
|
|
|
|
# backend container (compose exec -d), idempotent (skips already-cached cells),
|
|
|
|
|
# so it never blocks the deploy or health check and is cheap on every deploy
|
|
|
|
|
# after the first full warm.
|
SEO: add 250 English-market city pages; auto-warm archives on deploy (#97)
The population-ranked global top-500 skewed to Asian megacities and missed
high-English-search-demand cities. gen_cities.py now tops up with the top ~250
cities from English-speaking countries (US/GB/CA/AU/NZ/IE/ZA) not already in the
global set, so US coverage goes 13->146, GB 2->42, CA 3->29, etc. (Seattle, Boston,
Manchester, Melbourne, Auckland, Dublin, ...). cities.json regenerated to 750.
Both deploy scripts now launch warm_cities.py automatically after the health check,
detached (dev: a systemd --user transient unit; prod: setsid/nohup), so the city
pages serve from cache without a manual step; idempotent, so only the first deploy
does the full warm. DEPLOY.md updated.
2026-07-16 00:11:14 +00:00
|
|
|
warm_city_archives() {
|
2026-07-21 20:01:30 +00:00
|
|
|
echo "==> Warming city-page archives in the background (backend:/app/logs/warm-cities.log)"
|
|
|
|
|
docker compose exec -d backend sh -c \
|
Containerize the app and move the databases to PostgreSQL 18 (#220)
Run Thermograph as a docker-compose stack (app + Postgres 18) and standardize the
data layer on Postgres, while keeping the test suite on SQLite.
- accounts/db.py: DSN-driven engines. On Postgres, a per-worker read-write +
read-only asyncpg pair (the RO engine pins read-only transactions, used by the
pure-GET endpoints) plus a sync psycopg engine for the notifier thread; the
SQLite path is preserved for tests/local (selected when THERMOGRAPH_DATABASE_URL
is unset). models.py: boolean server_default -> sa.false().
- store.py / metrics.py: dialect-flexible — Postgres UNLOGGED tables via psycopg
when configured, else the existing raw-sqlite3 paths byte-for-byte; sync
interfaces and every fail-soft contract preserved.
- Alembic (backend/alembic/) manages the accounts schema; the container entrypoint
runs `alembic upgrade head` before uvicorn (4 workers). migrate_accounts_to_pg.py
copies the accounts data SQLite->PG through the ORM (UUID/bool/JSON coerced),
skips access_token, and resets identity sequences.
- Dockerfile + docker-compose.yml: app image (uvicorn, 4 workers, loopback 8137)
and a Postgres 18 db (2 CPUs) running pg_duckdb (deploy/db/) so the parquet
climate cache is queryable in-DB via read_parquet('/parquet/cache/*.parquet').
- deploy.sh/thermograph.service rewired to manage the compose stack; env example,
Makefile targets (up/down/db-up), and deploy/POSTGRES-MIGRATION.md cutover runbook.
Tests stay on SQLite (dialect fallback) — 323 pass. The full Postgres stack was
verified via docker compose: alembic migrations, register/login, the RO endpoint,
store/metrics round-trips, and the accounts data migration.
2026-07-20 06:28:23 +00:00
|
|
|
'python warm_cities.py --pace 2 >> /app/logs/warm-cities.log 2>&1' || true
|
SEO: add 250 English-market city pages; auto-warm archives on deploy (#97)
The population-ranked global top-500 skewed to Asian megacities and missed
high-English-search-demand cities. gen_cities.py now tops up with the top ~250
cities from English-speaking countries (US/GB/CA/AU/NZ/IE/ZA) not already in the
global set, so US coverage goes 13->146, GB 2->42, CA 3->29, etc. (Seattle, Boston,
Manchester, Melbourne, Auckland, Dublin, ...). cities.json regenerated to 750.
Both deploy scripts now launch warm_cities.py automatically after the health check,
detached (dev: a systemd --user transient unit; prod: setsid/nohup), so the city
pages serve from cache without a manual step; idempotent, so only the first deploy
does the full warm. DEPLOY.md updated.
2026-07-16 00:11:14 +00:00
|
|
|
}
|
|
|
|
|
|
2026-07-16 20:38:59 +00:00
|
|
|
# Notify IndexNow (Bing / DuckDuckGo / Yandex) of the site's URLs, but only when
|
Containerize the app and move the databases to PostgreSQL 18 (#220)
Run Thermograph as a docker-compose stack (app + Postgres 18) and standardize the
data layer on Postgres, while keeping the test suite on SQLite.
- accounts/db.py: DSN-driven engines. On Postgres, a per-worker read-write +
read-only asyncpg pair (the RO engine pins read-only transactions, used by the
pure-GET endpoints) plus a sync psycopg engine for the notifier thread; the
SQLite path is preserved for tests/local (selected when THERMOGRAPH_DATABASE_URL
is unset). models.py: boolean server_default -> sa.false().
- store.py / metrics.py: dialect-flexible — Postgres UNLOGGED tables via psycopg
when configured, else the existing raw-sqlite3 paths byte-for-byte; sync
interfaces and every fail-soft contract preserved.
- Alembic (backend/alembic/) manages the accounts schema; the container entrypoint
runs `alembic upgrade head` before uvicorn (4 workers). migrate_accounts_to_pg.py
copies the accounts data SQLite->PG through the ORM (UUID/bool/JSON coerced),
skips access_token, and resets identity sequences.
- Dockerfile + docker-compose.yml: app image (uvicorn, 4 workers, loopback 8137)
and a Postgres 18 db (2 CPUs) running pg_duckdb (deploy/db/) so the parquet
climate cache is queryable in-DB via read_parquet('/parquet/cache/*.parquet').
- deploy.sh/thermograph.service rewired to manage the compose stack; env example,
Makefile targets (up/down/db-up), and deploy/POSTGRES-MIGRATION.md cutover runbook.
Tests stay on SQLite (dialect fallback) — 323 pass. The full Postgres stack was
verified via docker compose: alembic migrations, register/login, the RO endpoint,
store/metrics round-trips, and the accounts data migration.
2026-07-20 06:28:23 +00:00
|
|
|
# the set of pages actually changed (a new/removed city) — code-only deploys skip.
|
|
|
|
|
# Best-effort: never fails the deploy.
|
2026-07-16 20:38:59 +00:00
|
|
|
ping_indexnow() {
|
|
|
|
|
echo "==> Pinging IndexNow (only if the URL set changed)"
|
Containerize the app and move the databases to PostgreSQL 18 (#220)
Run Thermograph as a docker-compose stack (app + Postgres 18) and standardize the
data layer on Postgres, while keeping the test suite on SQLite.
- accounts/db.py: DSN-driven engines. On Postgres, a per-worker read-write +
read-only asyncpg pair (the RO engine pins read-only transactions, used by the
pure-GET endpoints) plus a sync psycopg engine for the notifier thread; the
SQLite path is preserved for tests/local (selected when THERMOGRAPH_DATABASE_URL
is unset). models.py: boolean server_default -> sa.false().
- store.py / metrics.py: dialect-flexible — Postgres UNLOGGED tables via psycopg
when configured, else the existing raw-sqlite3 paths byte-for-byte; sync
interfaces and every fail-soft contract preserved.
- Alembic (backend/alembic/) manages the accounts schema; the container entrypoint
runs `alembic upgrade head` before uvicorn (4 workers). migrate_accounts_to_pg.py
copies the accounts data SQLite->PG through the ORM (UUID/bool/JSON coerced),
skips access_token, and resets identity sequences.
- Dockerfile + docker-compose.yml: app image (uvicorn, 4 workers, loopback 8137)
and a Postgres 18 db (2 CPUs) running pg_duckdb (deploy/db/) so the parquet
climate cache is queryable in-DB via read_parquet('/parquet/cache/*.parquet').
- deploy.sh/thermograph.service rewired to manage the compose stack; env example,
Makefile targets (up/down/db-up), and deploy/POSTGRES-MIGRATION.md cutover runbook.
Tests stay on SQLite (dialect fallback) — 323 pass. The full Postgres stack was
verified via docker compose: alembic migrations, register/login, the RO endpoint,
store/metrics round-trips, and the accounts data migration.
2026-07-20 06:28:23 +00:00
|
|
|
local base="${THERMOGRAPH_BASE_URL:-https://thermograph.org}"
|
2026-07-21 20:01:30 +00:00
|
|
|
docker compose exec -T backend python indexnow.py --if-changed "$base" \
|
Containerize the app and move the databases to PostgreSQL 18 (#220)
Run Thermograph as a docker-compose stack (app + Postgres 18) and standardize the
data layer on Postgres, while keeping the test suite on SQLite.
- accounts/db.py: DSN-driven engines. On Postgres, a per-worker read-write +
read-only asyncpg pair (the RO engine pins read-only transactions, used by the
pure-GET endpoints) plus a sync psycopg engine for the notifier thread; the
SQLite path is preserved for tests/local (selected when THERMOGRAPH_DATABASE_URL
is unset). models.py: boolean server_default -> sa.false().
- store.py / metrics.py: dialect-flexible — Postgres UNLOGGED tables via psycopg
when configured, else the existing raw-sqlite3 paths byte-for-byte; sync
interfaces and every fail-soft contract preserved.
- Alembic (backend/alembic/) manages the accounts schema; the container entrypoint
runs `alembic upgrade head` before uvicorn (4 workers). migrate_accounts_to_pg.py
copies the accounts data SQLite->PG through the ORM (UUID/bool/JSON coerced),
skips access_token, and resets identity sequences.
- Dockerfile + docker-compose.yml: app image (uvicorn, 4 workers, loopback 8137)
and a Postgres 18 db (2 CPUs) running pg_duckdb (deploy/db/) so the parquet
climate cache is queryable in-DB via read_parquet('/parquet/cache/*.parquet').
- deploy.sh/thermograph.service rewired to manage the compose stack; env example,
Makefile targets (up/down/db-up), and deploy/POSTGRES-MIGRATION.md cutover runbook.
Tests stay on SQLite (dialect fallback) — 323 pass. The full Postgres stack was
verified via docker compose: alembic migrations, register/login, the RO endpoint,
store/metrics round-trips, and the accounts data migration.
2026-07-20 06:28:23 +00:00
|
|
|
|| echo "!! IndexNow ping failed (non-fatal)" >&2
|
2026-07-16 20:38:59 +00:00
|
|
|
}
|
|
|
|
|
|
2026-07-11 00:29:47 +00:00
|
|
|
echo "==> Fetching $BRANCH"
|
|
|
|
|
git fetch --prune origin "$BRANCH"
|
|
|
|
|
git reset --hard "origin/$BRANCH"
|
|
|
|
|
|
2026-07-22 00:15:33 +00:00
|
|
|
# Re-exec: git reset --hard just rewrote this very file's bytes on disk while
|
|
|
|
|
# it's still running. bash reads a script via buffered, byte-offset I/O, so
|
|
|
|
|
# anything AFTER this point in the OLD execution can read from the wrong
|
|
|
|
|
# offset once the file's size/content changed underneath it -- a classic
|
|
|
|
|
# self-modifying-script footgun. Confirmed live: after this PR added ~15
|
|
|
|
|
# lines above, one deploy ran with the OLD "Building images" log lines even
|
|
|
|
|
# though `git status` showed the checkout correctly at the NEW commit --
|
|
|
|
|
# the file changed under a running interpreter, not the checkout. Restart
|
|
|
|
|
# fresh from the now-updated file so everything after this line is
|
|
|
|
|
# guaranteed self-consistent. Guarded so the second invocation doesn't
|
|
|
|
|
# fetch+reset+re-exec forever.
|
|
|
|
|
if [ -z "${DEPLOY_SH_REEXECED:-}" ]; then
|
|
|
|
|
export DEPLOY_SH_REEXECED=1
|
|
|
|
|
exec "$0" "$@"
|
|
|
|
|
fi
|
|
|
|
|
|
infra: split the estate into vps1/vps2, moving beta next to prod and dev onto vps1
Environments stop being machines. Until now each box WAS an environment --
"beta" named both a deploy target and a host, "the desktop" named both the
operator's computer and the dev server -- so every path could assume one
environment per host and hardcode /opt/thermograph, /etc/thermograph.env and
ports 8137/8080. That assumption ends here:
vps1 75.119.132.91 Forgejo, Grafana/Loki, the portfolio site, and DEV
(own Postgres, mesh-only on 10.10.0.2:8137)
vps2 169.58.46.181 PROD and BETA as two Swarm stacks sharing one
TimescaleDB instance, plus Centralis, Postfix, backups
desktop AI model hosting + flex Swarm capacity, no environment
Nothing live has moved; the ordered cutover is in
infra/deploy/RUNBOOK-vps1-vps2-cutover.md.
deploy/env-topology.sh is the single source of truth: env -> host, checkout,
branch, deploy mode, stack name, env file, LB ports, DB role/database, service
prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing
distinguishing a beta deploy from a prod one -- deploy.sh refuses to run when it
disagrees with the checkout it was invoked from. The host-wide secrets-env and
deploy-mode markers survive only as a fallback for a single-environment box.
Beta gets its own stack file rather than an overlay (a merge cannot REMOVE the
db service, and beta having no database of its own is the point). Its services
are prefixed beta-*: Swarm registers a service's short name as a DNS alias on
every network it joins, so two stacks both calling a service `web` on the shared
data network would let prod's frontend resolve a beta task. Prod's stack, env
and LB config are untouched.
One Postgres, two databases with two roles: deploy/db/provision-env-db.sh
creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT
revoked from PUBLIC) plus a <role>_ro for ad-hoc queries, and refuses to run for
the environment that owns the instance -- doing so would demote prod's bootstrap
superuser.
Fixes that co-residency would otherwise have broken silently:
- CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and
also "the Forgejo box" because those shared a machine; that conflation is what
once had the prod backup dumping beta.
- The nightly backup dumps BOTH application databases to separate off-box
prefixes, and fails loudly on a missing one instead of skipping it.
- Alloy stopped deriving the `host` label from the node -- on vps2 that would
have filed every beta line as prod, feeding prod's alert rules with beta's
traffic. It is now derived per source, with a new `node` label for the machine,
and beta's log volume is mounted via a vps2-only overlay.
- ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derived their SSH
target and env-file path from the topology instead of hardcoding beta to
75.119.132.91 -- which is vps1's address now.
- Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on
vps1) instead of 0.0.0.0: correct for a home LAN box, a public exposure of
unreviewed branches on a VPS.
Terraform's hosts variable is keyed by machine with a nested environments map,
and main.tf flattens (host, environment) pairs so two environments on one box
cannot share a checkout. Caddy's single reference config is split per host.
Docs, onboarding and the runbooks are updated throughout, including the security
rationales that co-residency makes false: separate SSH credentials no longer put
a host boundary between beta and prod, and the boundary that remains is the
database and the filesystem.
2026-07-25 22:01:29 +00:00
|
|
|
# Stack-mode routing: prod and beta are both Swarm stacks, dev is compose. The
|
|
|
|
|
# mode is a property of the ENVIRONMENT (env-topology.sh), not of the host --
|
|
|
|
|
# vps2 runs two stacks, and a host-wide /etc/thermograph/deploy-mode marker
|
|
|
|
|
# cannot describe a host that runs more than one environment. The marker is
|
|
|
|
|
# still honoured when the topology file is absent (a checkout that predates it).
|
|
|
|
|
# Checked AFTER the reset+re-exec so the stack script is always the freshly
|
|
|
|
|
# pulled one, and the SERVICE/tag contract passes through unchanged -- the
|
|
|
|
|
# deploy workflow never needs to know which mode an environment runs.
|
|
|
|
|
if [ -z "$DEPLOY_MODE" ]; then
|
|
|
|
|
DEPLOY_MODE=$(cat /etc/thermograph/deploy-mode 2>/dev/null || true)
|
|
|
|
|
fi
|
|
|
|
|
if [ "$DEPLOY_MODE" = "stack" ]; then
|
2026-07-23 05:11:33 +00:00
|
|
|
exec bash "$INFRA_DIR/deploy/stack/deploy-stack.sh"
|
2026-07-23 04:13:12 +00:00
|
|
|
fi
|
|
|
|
|
|
2026-07-23 05:11:33 +00:00
|
|
|
# Monorepo: run every docker compose command from infra/ (the git operations
|
|
|
|
|
# above ran at the checkout root). The compose project name is pinned in the
|
|
|
|
|
# file itself (`name: thermograph`), so moving the working directory does NOT
|
|
|
|
|
# rename the project and recreate the stack under a second name.
|
|
|
|
|
cd "$INFRA_DIR"
|
|
|
|
|
|
infra: split the estate into vps1/vps2, moving beta next to prod and dev onto vps1
Environments stop being machines. Until now each box WAS an environment --
"beta" named both a deploy target and a host, "the desktop" named both the
operator's computer and the dev server -- so every path could assume one
environment per host and hardcode /opt/thermograph, /etc/thermograph.env and
ports 8137/8080. That assumption ends here:
vps1 75.119.132.91 Forgejo, Grafana/Loki, the portfolio site, and DEV
(own Postgres, mesh-only on 10.10.0.2:8137)
vps2 169.58.46.181 PROD and BETA as two Swarm stacks sharing one
TimescaleDB instance, plus Centralis, Postfix, backups
desktop AI model hosting + flex Swarm capacity, no environment
Nothing live has moved; the ordered cutover is in
infra/deploy/RUNBOOK-vps1-vps2-cutover.md.
deploy/env-topology.sh is the single source of truth: env -> host, checkout,
branch, deploy mode, stack name, env file, LB ports, DB role/database, service
prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing
distinguishing a beta deploy from a prod one -- deploy.sh refuses to run when it
disagrees with the checkout it was invoked from. The host-wide secrets-env and
deploy-mode markers survive only as a fallback for a single-environment box.
Beta gets its own stack file rather than an overlay (a merge cannot REMOVE the
db service, and beta having no database of its own is the point). Its services
are prefixed beta-*: Swarm registers a service's short name as a DNS alias on
every network it joins, so two stacks both calling a service `web` on the shared
data network would let prod's frontend resolve a beta task. Prod's stack, env
and LB config are untouched.
One Postgres, two databases with two roles: deploy/db/provision-env-db.sh
creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT
revoked from PUBLIC) plus a <role>_ro for ad-hoc queries, and refuses to run for
the environment that owns the instance -- doing so would demote prod's bootstrap
superuser.
Fixes that co-residency would otherwise have broken silently:
- CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and
also "the Forgejo box" because those shared a machine; that conflation is what
once had the prod backup dumping beta.
- The nightly backup dumps BOTH application databases to separate off-box
prefixes, and fails loudly on a missing one instead of skipping it.
- Alloy stopped deriving the `host` label from the node -- on vps2 that would
have filed every beta line as prod, feeding prod's alert rules with beta's
traffic. It is now derived per source, with a new `node` label for the machine,
and beta's log volume is mounted via a vps2-only overlay.
- ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derived their SSH
target and env-file path from the topology instead of hardcoding beta to
75.119.132.91 -- which is vps1's address now.
- Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on
vps1) instead of 0.0.0.0: correct for a home LAN box, a public exposure of
unreviewed branches on a VPS.
Terraform's hosts variable is keyed by machine with a nested environments map,
and main.tf flattens (host, environment) pairs so two environments on one box
cannot share a checkout. Caddy's single reference config is split per host.
Docs, onboarding and the runbooks are updated throughout, including the security
rationales that co-residency makes false: separate SSH credentials no longer put
a host boundary between beta and prod, and the boundary that remains is the
database and the filesystem.
2026-07-25 22:01:29 +00:00
|
|
|
# Compose-mode environments (dev) carry their project name, file list and bind
|
|
|
|
|
# address in the topology table. Set only when not already in the environment,
|
|
|
|
|
# so deploy-dev.sh's own exports and a by-hand override both still win.
|
|
|
|
|
if [ -n "${TG_COMPOSE_PROJECT:-}" ]; then
|
|
|
|
|
export COMPOSE_PROJECT_NAME="${COMPOSE_PROJECT_NAME:-$TG_COMPOSE_PROJECT}"
|
|
|
|
|
export COMPOSE_FILE="${COMPOSE_FILE:-$TG_COMPOSE_FILE}"
|
|
|
|
|
fi
|
|
|
|
|
# Which address the dev overlay publishes on. Defaults to loopback in the
|
|
|
|
|
# compose file; dev on vps1 sets the mesh address here, because a public VPS
|
|
|
|
|
# must never publish an unreviewed branch's stack on 0.0.0.0.
|
|
|
|
|
if [ -n "${TG_BIND_ADDR:-}" ]; then
|
|
|
|
|
export DEV_BIND_ADDR="${DEV_BIND_ADDR:-$TG_BIND_ADDR}"
|
|
|
|
|
fi
|
|
|
|
|
|
2026-07-22 18:39:16 +00:00
|
|
|
# Registry-pull cutover: pull the image each app repo's build-push.yml already
|
|
|
|
|
# built and pushed, instead of building in place. This checkout is
|
|
|
|
|
# thermograph-infra, not an app repo, so there's no "current commit" to derive
|
|
|
|
|
# a tag from -- the caller (the deploying repo's deploy.yml) exports its own
|
|
|
|
|
# service's tag (BACKEND_IMAGE_TAG or FRONTEND_IMAGE_TAG = sha-<12 hex> of the
|
|
|
|
|
# app commit, or a semver tag).
|
2026-07-21 21:36:03 +00:00
|
|
|
REGISTRY_HOST="${REGISTRY_HOST:-git.thermograph.org}"
|
2026-07-22 18:39:16 +00:00
|
|
|
export REGISTRY_HOST BACKEND_IMAGE_PATH FRONTEND_IMAGE_PATH
|
|
|
|
|
|
|
|
|
|
# Load the last-deployed tag for BOTH services first, so a single-service roll
|
|
|
|
|
# still renders compose with the sibling's real, currently-running tag (never a
|
|
|
|
|
# bare `local` that would recreate/downgrade it). The incoming env for the
|
|
|
|
|
# service being deployed then overrides its line below. This file is untracked
|
|
|
|
|
# (see .gitignore), so `git reset --hard` above leaves it in place.
|
2026-07-23 05:11:33 +00:00
|
|
|
TAGS_FILE="$INFRA_DIR/deploy/.image-tags.env"
|
2026-07-22 22:48:23 +00:00
|
|
|
# Capture the tags the caller explicitly passed BEFORE sourcing -- they must
|
|
|
|
|
# WIN. The file only supplies the *sibling's* last-known tag; sourcing it
|
|
|
|
|
# unconditionally would clobber an incoming tag (e.g. a frontend deploy whose
|
|
|
|
|
# FRONTEND_IMAGE_TAG got overwritten by the stale `local` the first backend-only
|
|
|
|
|
# deploy persisted for the not-yet-known sibling -> pull `:local` -> "manifest
|
|
|
|
|
# unknown"). So source for the sibling, then re-apply the caller's own value.
|
|
|
|
|
_incoming_backend="${BACKEND_IMAGE_TAG:-}"
|
|
|
|
|
_incoming_frontend="${FRONTEND_IMAGE_TAG:-}"
|
2026-07-22 18:39:16 +00:00
|
|
|
if [ -f "$TAGS_FILE" ]; then
|
2026-07-23 22:26:05 +00:00
|
|
|
set -a
|
|
|
|
|
# shellcheck source=/dev/null # untracked runtime artifact this script writes below
|
|
|
|
|
. "$TAGS_FILE"
|
|
|
|
|
set +a
|
2026-07-22 18:39:16 +00:00
|
|
|
fi
|
2026-07-22 22:48:23 +00:00
|
|
|
[ -n "$_incoming_backend" ] && BACKEND_IMAGE_TAG="$_incoming_backend"
|
|
|
|
|
[ -n "$_incoming_frontend" ] && FRONTEND_IMAGE_TAG="$_incoming_frontend"
|
2026-07-22 18:39:16 +00:00
|
|
|
|
|
|
|
|
# Guard: the service(s) being rolled MUST have a concrete tag supplied now (the
|
|
|
|
|
# sibling's may come from the persisted file). `all` needs both.
|
|
|
|
|
case "$SERVICE" in
|
|
|
|
|
backend) : "${BACKEND_IMAGE_TAG:?set BACKEND_IMAGE_TAG=sha-<12-hex> for a backend deploy}" ;;
|
|
|
|
|
frontend) : "${FRONTEND_IMAGE_TAG:?set FRONTEND_IMAGE_TAG=sha-<12-hex> for a frontend deploy}" ;;
|
|
|
|
|
all)
|
|
|
|
|
: "${BACKEND_IMAGE_TAG:?set BACKEND_IMAGE_TAG=sha-<12-hex> (SERVICE=all needs both)}"
|
|
|
|
|
: "${FRONTEND_IMAGE_TAG:?set FRONTEND_IMAGE_TAG=sha-<12-hex> (SERVICE=all needs both)}"
|
|
|
|
|
;;
|
|
|
|
|
esac
|
|
|
|
|
# Compose interpolates both vars for the whole file even when we act on one
|
|
|
|
|
# service; default the not-yet-known sibling (first-ever deploy) to `local` so
|
|
|
|
|
# interpolation doesn't warn -- harmless since --no-deps never touches it.
|
|
|
|
|
export BACKEND_IMAGE_TAG="${BACKEND_IMAGE_TAG:-local}"
|
|
|
|
|
export FRONTEND_IMAGE_TAG="${FRONTEND_IMAGE_TAG:-local}"
|
|
|
|
|
|
|
|
|
|
# Which compose services this run pulls/rolls.
|
daemon: move the Discord gateway and scheduler out of the web process into Go (#21)
The gateway bot and APScheduler were long-lived stateful I/O loops running
inside the async web app under a leader election. They move into a single Go
binary that owns ONLY that I/O -- websocket, RESUME, heartbeat, backoff, timers.
It owns no grading logic. Anything needing data calls back over a new
internal-only surface (/internal/discord/grade, /internal/jobs/*). Grading
depends on polars and the parquet cache; reimplementing it in Go would let the
bot's grades drift from the API's. The grade route returns gateway-ready JSON
and Go relays the bytes verbatim.
The binary ships in the backend image and runs as a second compose service off
the same tag, so the two ends of the /internal/* contract can never skew.
deploy.sh rolls daemon alongside backend -- without that the service would never
be created, since a single-service deploy uses --no-deps. It also probes the
image first and skips the daemon when rolling a tag that predates the binary:
infra tracks main while image tags are env-staged, so a host can legitimately be
asked to roll an older backend image, and creating the service anyway would
leave a container crash-looping on a missing binary.
replicas: 1 with order: stop-first replaces the leader election -- Discord
permits one gateway connection per bot token.
THERMOGRAPH_INTERNAL_TOKEN is optional: both ends derive it from
THERMOGRAPH_AUTH_SECRET via HMAC under a domain-separation label, so this needs
no new vault entry. The derivation is pinned to a shared cross-language test
vector asserted on both sides, so drift fails CI instead of 401ing every call.
Fail closed when neither secret is set.
Improvements over the Python: a close intended for RESUME uses 4000 rather than
1000 (Discord invalidates a session closed 1000, so the old default defeated its
own resume); MESSAGE_CREATE runs on a bounded worker pool; and a malformed HELLO
returns an error rather than a clean reconnect, which would otherwise reset
backoff and hot-loop against the gateway.
365 Python tests pass; Go build/vet/test -race clean; shellcheck 0 findings.
2026-07-23 22:49:54 +00:00
|
|
|
#
|
2026-07-24 00:06:52 +00:00
|
|
|
# `daemon` and `lake` ride with `backend` and are never targets on their own.
|
|
|
|
|
# Both run the SAME image at the SAME tag (a second binary and a second role
|
|
|
|
|
# inside the backend image): the daemon talks to the web process over the
|
|
|
|
|
# /internal/* contract, the lake serves it history — rolling one without the
|
|
|
|
|
# other is exactly the version skew those contracts have no negotiation for.
|
|
|
|
|
# Pairing them here is also what makes the services exist at all: a
|
|
|
|
|
# single-service deploy runs `up -d --no-deps <targets>`, so a service listed
|
daemon: move the Discord gateway and scheduler out of the web process into Go (#21)
The gateway bot and APScheduler were long-lived stateful I/O loops running
inside the async web app under a leader election. They move into a single Go
binary that owns ONLY that I/O -- websocket, RESUME, heartbeat, backoff, timers.
It owns no grading logic. Anything needing data calls back over a new
internal-only surface (/internal/discord/grade, /internal/jobs/*). Grading
depends on polars and the parquet cache; reimplementing it in Go would let the
bot's grades drift from the API's. The grade route returns gateway-ready JSON
and Go relays the bytes verbatim.
The binary ships in the backend image and runs as a second compose service off
the same tag, so the two ends of the /internal/* contract can never skew.
deploy.sh rolls daemon alongside backend -- without that the service would never
be created, since a single-service deploy uses --no-deps. It also probes the
image first and skips the daemon when rolling a tag that predates the binary:
infra tracks main while image tags are env-staged, so a host can legitimately be
asked to roll an older backend image, and creating the service anyway would
leave a container crash-looping on a missing binary.
replicas: 1 with order: stop-first replaces the leader election -- Discord
permits one gateway connection per bot token.
THERMOGRAPH_INTERNAL_TOKEN is optional: both ends derive it from
THERMOGRAPH_AUTH_SECRET via HMAC under a domain-separation label, so this needs
no new vault entry. The derivation is pinned to a shared cross-language test
vector asserted on both sides, so drift fails CI instead of 401ing every call.
Fail closed when neither secret is set.
Improvements over the Python: a close intended for RESUME uses 4000 rather than
1000 (Discord invalidates a session closed 1000, so the old default defeated its
own resume); MESSAGE_CREATE runs on a bounded worker pool; and a malformed HELLO
returns an error rather than a clean reconnect, which would otherwise reset
backoff and hot-loop against the gateway.
365 Python tests pass; Go build/vet/test -race clean; shellcheck 0 findings.
2026-07-23 22:49:54 +00:00
|
|
|
# only in compose but absent from TARGETS would never be created on a
|
2026-07-24 00:06:52 +00:00
|
|
|
# backend-only deploy.
|
2026-07-22 18:39:16 +00:00
|
|
|
case "$SERVICE" in
|
2026-07-24 00:06:52 +00:00
|
|
|
backend) TARGETS=(backend lake daemon) ;;
|
2026-07-22 18:39:16 +00:00
|
|
|
frontend) TARGETS=(frontend) ;;
|
2026-07-24 00:06:52 +00:00
|
|
|
all) TARGETS=(backend lake frontend daemon) ;;
|
2026-07-22 18:39:16 +00:00
|
|
|
esac
|
2026-07-21 21:36:03 +00:00
|
|
|
|
2026-07-22 22:27:30 +00:00
|
|
|
# Login only when a token is supplied. The SSH/CI deploy paths (deploy.yml,
|
|
|
|
|
# deploy-prod.yml, deploy-dev on the LAN runner) don't pass REGISTRY_TOKEN --
|
|
|
|
|
# the host is already `docker login`ed to the registry (persistent cred in
|
|
|
|
|
# ~/.docker/config.json), so an unconditional login with an empty token would
|
|
|
|
|
# abort the deploy under `set -e`. Use the token if present, else trust the
|
|
|
|
|
# host's existing cred; a genuine auth problem then fails loudly at `pull`.
|
|
|
|
|
if [ -n "${REGISTRY_TOKEN:-}" ]; then
|
|
|
|
|
echo "==> Logging in to the registry ($REGISTRY_HOST)"
|
|
|
|
|
echo "$REGISTRY_TOKEN" | docker login "$REGISTRY_HOST" --username emi --password-stdin
|
|
|
|
|
else
|
|
|
|
|
echo "==> No REGISTRY_TOKEN in env; relying on the host's existing docker login to $REGISTRY_HOST"
|
|
|
|
|
fi
|
2026-07-21 21:36:03 +00:00
|
|
|
|
2026-07-22 18:39:16 +00:00
|
|
|
echo "==> Pulling images (backend=$BACKEND_IMAGE_TAG frontend=$FRONTEND_IMAGE_TAG; rolling: ${TARGETS[*]})"
|
|
|
|
|
# Retry: build-push.yml (triggered by the same push) has no ordering guarantee
|
|
|
|
|
# against this deploy -- Forgejo Actions `needs:` only works between jobs in ONE
|
|
|
|
|
# workflow file, not across the separate build-push.yml triggered by the same
|
|
|
|
|
# event. Confirmed live: a deploy raced ahead of the push and failed with "not
|
|
|
|
|
# found". A bounded retry (~5 min) covers a normal build; a genuine problem
|
|
|
|
|
# (bad tag, registry down) still fails loudly after that.
|
2026-07-22 02:12:44 +00:00
|
|
|
pull_ok=0
|
|
|
|
|
for i in $(seq 1 30); do
|
2026-07-22 18:39:16 +00:00
|
|
|
if docker compose pull "${TARGETS[@]}"; then
|
2026-07-22 02:12:44 +00:00
|
|
|
pull_ok=1
|
|
|
|
|
break
|
|
|
|
|
fi
|
|
|
|
|
echo " pull attempt $i/30 failed (image may not be pushed yet); retrying in 10s..." >&2
|
|
|
|
|
sleep 10
|
|
|
|
|
done
|
|
|
|
|
if [ "$pull_ok" != 1 ]; then
|
|
|
|
|
echo "!! docker compose pull failed after 30 attempts" >&2
|
|
|
|
|
exit 1
|
|
|
|
|
fi
|
2026-07-20 06:09:15 +00:00
|
|
|
|
daemon: move the Discord gateway and scheduler out of the web process into Go (#21)
The gateway bot and APScheduler were long-lived stateful I/O loops running
inside the async web app under a leader election. They move into a single Go
binary that owns ONLY that I/O -- websocket, RESUME, heartbeat, backoff, timers.
It owns no grading logic. Anything needing data calls back over a new
internal-only surface (/internal/discord/grade, /internal/jobs/*). Grading
depends on polars and the parquet cache; reimplementing it in Go would let the
bot's grades drift from the API's. The grade route returns gateway-ready JSON
and Go relays the bytes verbatim.
The binary ships in the backend image and runs as a second compose service off
the same tag, so the two ends of the /internal/* contract can never skew.
deploy.sh rolls daemon alongside backend -- without that the service would never
be created, since a single-service deploy uses --no-deps. It also probes the
image first and skips the daemon when rolling a tag that predates the binary:
infra tracks main while image tags are env-staged, so a host can legitimately be
asked to roll an older backend image, and creating the service anyway would
leave a container crash-looping on a missing binary.
replicas: 1 with order: stop-first replaces the leader election -- Discord
permits one gateway connection per bot token.
THERMOGRAPH_INTERNAL_TOKEN is optional: both ends derive it from
THERMOGRAPH_AUTH_SECRET via HMAC under a domain-separation label, so this needs
no new vault entry. The derivation is pinned to a shared cross-language test
vector asserted on both sides, so drift fails CI instead of 401ing every call.
Fail closed when neither secret is set.
Improvements over the Python: a close intended for RESUME uses 4000 rather than
1000 (Discord invalidates a session closed 1000, so the old default defeated its
own resume); MESSAGE_CREATE runs on a bounded worker pool; and a malformed HELLO
returns an error rather than a clean reconnect, which would otherwise reset
backoff and hot-loop against the gateway.
365 Python tests pass; Go build/vet/test -race clean; shellcheck 0 findings.
2026-07-23 22:49:54 +00:00
|
|
|
# The daemon binary ships INSIDE the backend image, but infra and app images
|
|
|
|
|
# advance on different axes by design: this checkout tracks `main`, while image
|
|
|
|
|
# tags are env-staged (prod's come from `release`). So a host can legitimately
|
|
|
|
|
# be asked to roll a backend image OLDER than this compose file -- one built
|
|
|
|
|
# before the daemon existed and with no /usr/local/bin/thermograph-daemon in it.
|
|
|
|
|
# Creating the service anyway would leave a container crash-looping on a missing
|
|
|
|
|
# binary, on a deploy that otherwise succeeded. Probe the image we are actually
|
|
|
|
|
# about to roll and drop the daemon from this run if it can't support it; the
|
|
|
|
|
# next deploy of a tag that has it picks it up with no further action.
|
|
|
|
|
case " ${TARGETS[*]} " in
|
|
|
|
|
*" daemon "*)
|
2026-07-25 04:13:47 +00:00
|
|
|
# Built directly from the same vars docker-compose.yml's `daemon.image:`
|
|
|
|
|
# interpolates (REGISTRY_HOST/BACKEND_IMAGE_PATH/BACKEND_IMAGE_TAG),
|
|
|
|
|
# NOT via `docker compose config --images daemon`: that command does not
|
|
|
|
|
# actually filter to the named service (confirmed live on Compose
|
|
|
|
|
# v5.3.1 -- it prints every service's image, one per line, in file
|
|
|
|
|
# order) so `| head -1` silently grabbed db's image instead. The probe
|
|
|
|
|
# then always found no daemon binary in a Postgres image and dropped
|
|
|
|
|
# daemon from EVERY backend deploy, regardless of what the real backend
|
|
|
|
|
# image contained -- reproduced and confirmed against beta directly.
|
|
|
|
|
daemon_img="${REGISTRY_HOST:-git.thermograph.org}/${BACKEND_IMAGE_PATH:-emi/thermograph/backend}:${BACKEND_IMAGE_TAG}"
|
|
|
|
|
if ! docker run --rm --entrypoint sh "$daemon_img" -c 'test -x /usr/local/bin/thermograph-daemon' 2>/dev/null; then
|
daemon: move the Discord gateway and scheduler out of the web process into Go (#21)
The gateway bot and APScheduler were long-lived stateful I/O loops running
inside the async web app under a leader election. They move into a single Go
binary that owns ONLY that I/O -- websocket, RESUME, heartbeat, backoff, timers.
It owns no grading logic. Anything needing data calls back over a new
internal-only surface (/internal/discord/grade, /internal/jobs/*). Grading
depends on polars and the parquet cache; reimplementing it in Go would let the
bot's grades drift from the API's. The grade route returns gateway-ready JSON
and Go relays the bytes verbatim.
The binary ships in the backend image and runs as a second compose service off
the same tag, so the two ends of the /internal/* contract can never skew.
deploy.sh rolls daemon alongside backend -- without that the service would never
be created, since a single-service deploy uses --no-deps. It also probes the
image first and skips the daemon when rolling a tag that predates the binary:
infra tracks main while image tags are env-staged, so a host can legitimately be
asked to roll an older backend image, and creating the service anyway would
leave a container crash-looping on a missing binary.
replicas: 1 with order: stop-first replaces the leader election -- Discord
permits one gateway connection per bot token.
THERMOGRAPH_INTERNAL_TOKEN is optional: both ends derive it from
THERMOGRAPH_AUTH_SECRET via HMAC under a domain-separation label, so this needs
no new vault entry. The derivation is pinned to a shared cross-language test
vector asserted on both sides, so drift fails CI instead of 401ing every call.
Fail closed when neither secret is set.
Improvements over the Python: a close intended for RESUME uses 4000 rather than
1000 (Discord invalidates a session closed 1000, so the old default defeated its
own resume); MESSAGE_CREATE runs on a bounded worker pool; and a malformed HELLO
returns an error rather than a clean reconnect, which would otherwise reset
backoff and hot-loop against the gateway.
365 Python tests pass; Go build/vet/test -race clean; shellcheck 0 findings.
2026-07-23 22:49:54 +00:00
|
|
|
echo "==> $daemon_img predates the daemon binary; rolling without the daemon service this run"
|
|
|
|
|
kept=()
|
|
|
|
|
for t in "${TARGETS[@]}"; do
|
|
|
|
|
[ "$t" = daemon ] || kept+=("$t")
|
|
|
|
|
done
|
|
|
|
|
TARGETS=("${kept[@]}")
|
|
|
|
|
fi
|
|
|
|
|
;;
|
|
|
|
|
esac
|
|
|
|
|
|
2026-07-22 18:39:16 +00:00
|
|
|
# Roll only the target service(s). Backend schema migrations run inside its own
|
|
|
|
|
# entrypoint (alembic upgrade head) before uvicorn, so there's no separate
|
|
|
|
|
# migrate step; frontend is stateless.
|
|
|
|
|
#
|
|
|
|
|
# Single-service rolls use --no-deps so recreating backend doesn't also bounce
|
|
|
|
|
# db, and recreating frontend doesn't touch backend -- that independence is the
|
|
|
|
|
# whole point of the split. A full `all` deploy instead uses --remove-orphans,
|
|
|
|
|
# which matters when the service topology itself changes (a renamed-away
|
|
|
|
|
# service's old container would otherwise keep running and squat its port --
|
|
|
|
|
# confirmed live, this is what blocked beta's first dual-service deploy with
|
|
|
|
|
# "port is already allocated").
|
|
|
|
|
echo "==> Rolling ${TARGETS[*]}"
|
|
|
|
|
if [ "$SERVICE" = all ]; then
|
|
|
|
|
docker compose up -d --remove-orphans
|
|
|
|
|
else
|
|
|
|
|
docker compose up -d --no-deps "${TARGETS[@]}"
|
|
|
|
|
fi
|
|
|
|
|
|
|
|
|
|
# Persist the now-live tags so the next single-service deploy knows the
|
|
|
|
|
# sibling's real tag. Written after `up` so a failed pull never records a tag
|
|
|
|
|
# that isn't actually running.
|
|
|
|
|
mkdir -p "$(dirname "$TAGS_FILE")"
|
|
|
|
|
cat > "$TAGS_FILE" <<EOF
|
|
|
|
|
# Written by deploy.sh -- the image tag each service is currently running.
|
|
|
|
|
# Untracked (see .gitignore); lets a single-service deploy leave the other alone.
|
|
|
|
|
BACKEND_IMAGE_TAG=$BACKEND_IMAGE_TAG
|
|
|
|
|
FRONTEND_IMAGE_TAG=$FRONTEND_IMAGE_TAG
|
|
|
|
|
EOF
|
|
|
|
|
|
|
|
|
|
# Health check the service(s) we rolled: backend on 8137, frontend on 8080.
|
|
|
|
|
# (For `all`, backend's `/` serving is the readiness signal the old script used
|
|
|
|
|
# and the frontend depends_on backend anyway.)
|
2026-07-22 22:48:23 +00:00
|
|
|
# Health via each container's own HEALTHCHECK (docker inspect), NOT a host-port
|
|
|
|
|
# curl: the dev overlay leaves the frontend port UNpublished (reached through the
|
|
|
|
|
# backend's _proxy_to_frontend), so a localhost:8080 curl spuriously fails there.
|
|
|
|
|
# Both images HEALTHCHECK-curl /healthz internally, so this works published or not.
|
2026-07-22 18:39:16 +00:00
|
|
|
health_ok=1
|
|
|
|
|
for svc in "${TARGETS[@]}"; do
|
2026-07-22 22:48:23 +00:00
|
|
|
cid=$(docker compose ps -q "$svc" 2>/dev/null)
|
|
|
|
|
echo "==> Health check: $svc (container health)"
|
|
|
|
|
ok=0; st=unknown
|
|
|
|
|
for i in $(seq 1 40); do
|
|
|
|
|
st=$(docker inspect --format '{{if .State.Health}}{{.State.Health.Status}}{{else}}none{{end}}' "$cid" 2>/dev/null || echo gone)
|
|
|
|
|
[ "$st" = healthy ] && { ok=1; break; }
|
|
|
|
|
if [ "$st" = none ] && [ "$(docker inspect --format '{{.State.Status}}' "$cid" 2>/dev/null)" = running ]; then ok=1; break; fi
|
|
|
|
|
sleep 2
|
2026-07-22 18:39:16 +00:00
|
|
|
done
|
|
|
|
|
if [ "$ok" = 1 ]; then
|
2026-07-22 22:48:23 +00:00
|
|
|
echo "==> OK: $svc is healthy"
|
2026-07-22 18:39:16 +00:00
|
|
|
else
|
2026-07-22 22:48:23 +00:00
|
|
|
echo "!! Health check failed for $svc (status=$st)" >&2
|
2026-07-22 18:39:16 +00:00
|
|
|
health_ok=0
|
2026-07-11 00:29:47 +00:00
|
|
|
fi
|
|
|
|
|
done
|
2026-07-22 18:39:16 +00:00
|
|
|
|
|
|
|
|
if [ "$health_ok" != 1 ]; then
|
|
|
|
|
docker compose ps || true
|
|
|
|
|
for svc in "${TARGETS[@]}"; do docker compose logs --tail=50 "$svc" || true; done
|
|
|
|
|
exit 1
|
|
|
|
|
fi
|
|
|
|
|
|
2026-07-22 23:39:29 +00:00
|
|
|
# Every deploy pulls a new sha-tagged image and nothing ever removed the old
|
|
|
|
|
# ones -- hosts accumulate gigabytes of dead tags at one per app commit. Keep
|
|
|
|
|
# only the tags recorded as now-live (both services) and delete other tags of
|
|
|
|
|
# the two app-image repos. Best-effort and after the health gate, so a failed
|
|
|
|
|
# roll never garbage-collects the image a rollback would need; docker also
|
|
|
|
|
# refuses to remove an image any container still uses.
|
|
|
|
|
echo "==> Pruning old app-image tags"
|
2026-07-23 05:11:33 +00:00
|
|
|
_be_repo="${REGISTRY_HOST}/${BACKEND_IMAGE_PATH:-emi/thermograph/backend}"
|
|
|
|
|
_fe_repo="${REGISTRY_HOST}/${FRONTEND_IMAGE_PATH:-emi/thermograph/frontend}"
|
2026-07-22 23:39:29 +00:00
|
|
|
docker images --format '{{.Repository}}:{{.Tag}}' \
|
|
|
|
|
| grep -E "^(${_be_repo}|${_fe_repo}):" \
|
|
|
|
|
| grep -v -e "^${_be_repo}:${BACKEND_IMAGE_TAG}$" -e "^${_fe_repo}:${FRONTEND_IMAGE_TAG}$" \
|
|
|
|
|
| xargs -r docker rmi 2>/dev/null || true
|
|
|
|
|
|
2026-07-22 18:39:16 +00:00
|
|
|
# Post-deploy warm/IndexNow only make sense once the backend is (re)deployed --
|
|
|
|
|
# they exec inside the backend container. Skip them on a frontend-only roll.
|
|
|
|
|
if [ "$SERVICE" = backend ] || [ "$SERVICE" = all ]; then
|
|
|
|
|
warm_city_archives
|
|
|
|
|
ping_indexnow
|
|
|
|
|
fi
|
|
|
|
|
exit 0
|