All checks were successful
Sync infra to hosts / sync-beta (push) Successful in 13s
Sync infra to hosts / sync-prod (push) Successful in 12s
secrets-guard / encrypted (push) Successful in 7s
shell-lint / shellcheck (push) Successful in 8s
Build + push frontend image (Forgejo registry) / build-push (push) Successful in 53s
Deploy frontend to beta VPS / deploy (push) Successful in 1m16s
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 6s
Ports frontend/ (Jinja2/FastAPI, ~1180 LOC) to Go with html/template. No climate math, no DB, no auth -- every route fetches from the backend's /content/* API. Verified with a golden-HTML diff, not just unit tests: both the Python original and the Go rewrite were run against the same committed fixtures and every route compared byte-for-byte, confirmed programmatically. That process caught defects unit tests alone missed, since map[string]any has no compile-time field check: - Render-context keys were snake_case throughout while the templates read PascalCase fields. A missing map key doesn't error, it silently renders empty -- title, meta description, canonical URL, OpenGraph tags, and the homepage's entire ranked list were blank on every page despite every route returning 200. Fixed by renaming every key to match each template's own documented field contract, and passing API structs straight through wherever their fields already matched (removes a whole layer of future drift risk). - Three pages 500'd: ToolHref needed a composed href, not a bare "lat,lon" fragment; the records table needed the raw API struct. - JSON-LD was double-encoded: <script type="application/ld+json"> is JAVASCRIPT context to html/template's escaper regardless of the script's type attribute, so template.HTML gets re-escaped as a quoted JS string. Needed template.JS. The glossary term page's JSON-LD was never built at all -- added. - html/template silently strips literal HTML and JS comments from parsed output (verified in isolation) -- both need a FuncMap function returning template.HTML/template.JS to survive. Packaging: 187MB -> 22.6MB. Two defects caught before reaching a host: the Swarm stack's entrypoint override with no explicit command drops the image's CMD entirely (every deploy would have exited 127), and COPY --chown by name fails under the classic Docker builder on Alpine. Both fixed. go build/vet/test -race clean; docker build passes its embedded test step under both BuildKit and the classic builder; shellcheck 0 findings.
362 lines
15 KiB
YAML
362 lines
15 KiB
YAML
# Docker Swarm stack for prod — the autoscaling successor to docker-compose.yml
|
|
# on that host (beta and LAN dev stay on plain compose; deploy.sh routes by the
|
|
# /etc/thermograph/deploy-mode marker). Two-image world: web/worker run the
|
|
# backend image, frontend its own — per-service tags, same contract as compose.
|
|
#
|
|
# Deployed by deploy/stack/deploy-stack.sh, which:
|
|
# - sources /etc/thermograph.env (SOPS-rendered) so ${VARS} here interpolate,
|
|
# - installs a uid-10001-readable copy at /etc/thermograph/stack.env that
|
|
# deploy/stack/env-entrypoint.sh sources inside each app task (set-if-unset,
|
|
# so `environment:` blocks below always win) — this replaces compose's
|
|
# env_file:, which `docker stack deploy` does not support,
|
|
# - runs migrations as a ONE-SHOT task before rolling (RUN_MIGRATIONS=0 in
|
|
# every replica — N replicas must never race Alembic),
|
|
# - manages the loopback LB bridge (see lb/README note below): Swarm's mesh
|
|
# can only publish on 0.0.0.0 (would expose the plaintext app un-fronted),
|
|
# so nothing here has `ports:`. A plain container on this attachable
|
|
# overlay binds 127.0.0.1:8137/8080 for the host Caddy and proxies to the
|
|
# service VIPs — Swarm's VIP does the actual load balancing across
|
|
# replicas.
|
|
#
|
|
# Scaling model: `web` is stateless (ROLE=web never runs the notifier) and
|
|
# scales 1..N — the autoscaler service adjusts replicas between
|
|
# WEB_MIN_REPLICAS/WEB_MAX_REPLICAS on task CPU. `worker` owns the notifier +
|
|
# scheduler: exactly 1 replica, with the cluster-wide Postgres advisory lock
|
|
# (THERMOGRAPH_SINGLETON_PG) as belt-and-suspenders. `db` is exactly 1 — a
|
|
# database does not scale by container count on one host; its levers are
|
|
# DB_CPUS/DB_MEMORY (and, multi-host later, Patroni replicas per the topology
|
|
# doc). Everything is pinned to the manager node: all volumes are local to
|
|
# prod today. That constraint is the ONLY thing to relax when a second app
|
|
# node joins.
|
|
#
|
|
# TIMESCALEDB_IMAGE must be the exact image (digest-pinned) the compose stack
|
|
# was running — see hazard #7 in the hop-1 runbook: a floating tag can change
|
|
# the extension minor under an existing volume. deploy-stack.sh resolves it
|
|
# from the running/last-known container automatically.
|
|
|
|
services:
|
|
db:
|
|
image: ${TIMESCALEDB_IMAGE:?required}
|
|
environment:
|
|
POSTGRES_USER: thermograph
|
|
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?required}
|
|
POSTGRES_DB: thermograph
|
|
DB_MEMORY: ${DB_MEMORY:-16g}
|
|
volumes:
|
|
- pgdata:/var/lib/postgresql
|
|
- /opt/thermograph/infra/deploy/db/init:/docker-entrypoint-initdb.d:ro
|
|
networks:
|
|
- internal
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "pg_isready -U thermograph -d thermograph"]
|
|
interval: 5s
|
|
timeout: 5s
|
|
retries: 10
|
|
deploy:
|
|
replicas: 1
|
|
placement:
|
|
constraints: ["node.role == manager"]
|
|
resources:
|
|
limits:
|
|
cpus: "${DB_CPUS:-4}"
|
|
memory: ${DB_MEMORY:-16g}
|
|
restart_policy:
|
|
condition: on-failure
|
|
|
|
web:
|
|
image: ${REGISTRY_HOST:-git.thermograph.org}/${BACKEND_IMAGE_PATH:-emi/thermograph/backend}:${BACKEND_IMAGE_TAG:?required}
|
|
entrypoint: ["/host/env-entrypoint.sh"]
|
|
environment:
|
|
THERMOGRAPH_DATABASE_URL: postgresql+asyncpg://thermograph:${POSTGRES_PASSWORD}@db:5432/thermograph
|
|
THERMOGRAPH_BASE: /
|
|
PORT: 8137
|
|
THERMOGRAPH_SERVICE_ROLE: backend
|
|
THERMOGRAPH_FRONTEND_BASE_INTERNAL: http://frontend:8080
|
|
WORKERS: ${WEB_WORKERS:-4}
|
|
THERMOGRAPH_DATA_DIR: /state
|
|
# web NEVER runs the notifier/scheduler, even if it would win election —
|
|
# that's the worker service's job. This is what makes web replicas safe.
|
|
THERMOGRAPH_ROLE: web
|
|
RUN_MIGRATIONS: "0"
|
|
# Overlay tasks reach the HOST's Postfix via the docker_gwbridge gateway,
|
|
# not the compose bridge's 172.19.0.1 (an overlay has no host gateway).
|
|
# provision-mail.sh's DOCKER_MAIL_GATEWAY/SUBNET cover this listener.
|
|
THERMOGRAPH_SMTP_HOST: ${STACK_SMTP_HOST:-172.18.0.1}
|
|
# History reads ask the lake service (Swarm VIP over 1..2 replicas)
|
|
# before any third-party source; unset on beta/LAN, where compose has no
|
|
# lake service and climate.py falls straight through to NASA.
|
|
THERMOGRAPH_LAKE_URL: http://lake:8141
|
|
volumes:
|
|
- appdata:/state
|
|
- applogs:/app/logs
|
|
- /opt/thermograph/infra/deploy/stack/env-entrypoint.sh:/host/env-entrypoint.sh:ro
|
|
- /etc/thermograph/stack.env:/host/thermograph.env:ro
|
|
networks:
|
|
- internal
|
|
deploy:
|
|
replicas: ${WEB_MIN_REPLICAS:-1}
|
|
placement:
|
|
constraints: ["node.role == manager"]
|
|
resources:
|
|
limits:
|
|
cpus: "${WEB_CPUS:-4}"
|
|
restart_policy:
|
|
condition: on-failure
|
|
update_config:
|
|
# New task must pass the image's own HEALTHCHECK before the old one
|
|
# stops — zero-downtime single-service rolls.
|
|
order: start-first
|
|
failure_action: rollback
|
|
|
|
worker:
|
|
image: ${REGISTRY_HOST:-git.thermograph.org}/${BACKEND_IMAGE_PATH:-emi/thermograph/backend}:${BACKEND_IMAGE_TAG:?required}
|
|
entrypoint: ["/host/env-entrypoint.sh"]
|
|
environment:
|
|
THERMOGRAPH_DATABASE_URL: postgresql+asyncpg://thermograph:${POSTGRES_PASSWORD}@db:5432/thermograph
|
|
THERMOGRAPH_BASE: /
|
|
PORT: 8137
|
|
THERMOGRAPH_SERVICE_ROLE: backend
|
|
THERMOGRAPH_FRONTEND_BASE_INTERNAL: http://frontend:8080
|
|
WORKERS: "1"
|
|
THERMOGRAPH_DATA_DIR: /state
|
|
THERMOGRAPH_ROLE: worker
|
|
# Cluster-wide Postgres advisory lock, not the host flock: correct at
|
|
# replicas=1 today and stays correct if a second worker ever appears.
|
|
THERMOGRAPH_SINGLETON_PG: "1"
|
|
RUN_MIGRATIONS: "0"
|
|
THERMOGRAPH_SMTP_HOST: ${STACK_SMTP_HOST:-172.18.0.1}
|
|
THERMOGRAPH_LAKE_URL: http://lake:8141
|
|
volumes:
|
|
- appdata:/state
|
|
- applogs:/app/logs
|
|
- /opt/thermograph/infra/deploy/stack/env-entrypoint.sh:/host/env-entrypoint.sh:ro
|
|
- /etc/thermograph/stack.env:/host/thermograph.env:ro
|
|
networks:
|
|
- internal
|
|
deploy:
|
|
replicas: 1
|
|
placement:
|
|
constraints: ["node.role == manager"]
|
|
resources:
|
|
limits:
|
|
cpus: "${WORKER_CPUS:-2}"
|
|
restart_policy:
|
|
condition: on-failure
|
|
|
|
# The ERA5 lake indexer: SQL-style, partition-pruned reads over the
|
|
# era5-thermograph bucket with a node-local parquet cache, so a cell's
|
|
# history is one LAN hop for web instead of an S3 fetch (or a NASA call).
|
|
# Same backend image; THERMOGRAPH_ROLE=lake makes the entrypoint start
|
|
# lake_app (no database, no migrations). Its S3 credentials
|
|
# (THERMOGRAPH_LAKE_S3_ACCESS_KEY/_SECRET_KEY) arrive via the same rendered
|
|
# /host/thermograph.env as every other secret. Without them it stays
|
|
# healthy and answers 503/404, and web falls through to NASA — the lake is
|
|
# an accelerator, never a point of failure.
|
|
lake:
|
|
image: ${REGISTRY_HOST:-git.thermograph.org}/${BACKEND_IMAGE_PATH:-emi/thermograph/backend}:${BACKEND_IMAGE_TAG:?required}
|
|
entrypoint: ["/host/env-entrypoint.sh"]
|
|
environment:
|
|
THERMOGRAPH_ROLE: lake
|
|
PORT: 8141
|
|
THERMOGRAPH_SERVICE_ROLE: backend
|
|
WORKERS: "1"
|
|
THERMOGRAPH_LAKE_CACHE: /state/lake-cache
|
|
volumes:
|
|
# Replicas share the cache; lake_app's tmp+rename writes make that safe.
|
|
- lakecache:/state
|
|
- /opt/thermograph/infra/deploy/stack/env-entrypoint.sh:/host/env-entrypoint.sh:ro
|
|
- /etc/thermograph/stack.env:/host/thermograph.env:ro
|
|
networks:
|
|
- internal
|
|
deploy:
|
|
replicas: ${LAKE_MIN_REPLICAS:-1}
|
|
placement:
|
|
constraints: ["node.role == manager"]
|
|
resources:
|
|
limits:
|
|
cpus: "${LAKE_CPUS:-2}"
|
|
restart_policy:
|
|
condition: on-failure
|
|
update_config:
|
|
order: start-first
|
|
failure_action: rollback
|
|
|
|
daemon:
|
|
# SAME image and tag as web/worker on purpose: the daemon binary
|
|
# (/usr/local/bin/thermograph-daemon, built into the backend image) and the
|
|
# app share the /internal/* API contract, so one BACKEND_IMAGE_TAG rolls
|
|
# them together and they can never skew.
|
|
image: ${REGISTRY_HOST:-git.thermograph.org}/${BACKEND_IMAGE_PATH:-emi/thermograph/backend}:${BACKEND_IMAGE_TAG:?required}
|
|
# Not env-entrypoint.sh: that shim's handoff execs the image's
|
|
# deploy/entrypoint.sh (Alembic + uvicorn) whenever it exists, which is
|
|
# exactly what this service must never run — migrations belong to the
|
|
# one-shot task, and a second migrator racing it is a real hazard. Instead
|
|
# source the same host-rendered env directly (it is shell-sourceable by
|
|
# contract — deploy.sh sources it too) and exec the binary. stack.env holds
|
|
# vault secrets, never service-topology URLs, so it cannot collide with the
|
|
# `environment:` block below.
|
|
entrypoint:
|
|
- /bin/bash
|
|
- -c
|
|
- 'set -a; [ -f /host/thermograph.env ] && . /host/thermograph.env; set +a; exec /usr/local/bin/thermograph-daemon'
|
|
environment:
|
|
# The service VIP: the daemon's /internal/* calls load-balance across web
|
|
# replicas like any other client of the app.
|
|
THERMOGRAPH_API_BASE_INTERNAL: http://web:8137
|
|
# Only the env mount — no appdata/applogs: warm-cities/IndexNow work
|
|
# executes inside the web process (the daemon just fires the internal
|
|
# endpoints on a timer), so the daemon holds no state and logs to stdout.
|
|
volumes:
|
|
- /etc/thermograph/stack.env:/host/thermograph.env:ro
|
|
networks:
|
|
- internal
|
|
deploy:
|
|
# EXACTLY 1, load-bearing — do not scale this service, ever. Discord
|
|
# permits ONE gateway connection per bot token; the Python bot enforced
|
|
# that with a leader election (core/singleton.claim_leader), and this
|
|
# pinned replica count is what REPLACES that election. A second replica
|
|
# means two gateway sessions fighting over the token (Discord kicks
|
|
# both into a reconnect storm) and every timer job firing twice. The
|
|
# autoscaler cannot touch this — it scales ${STACK_NAME}_web only.
|
|
replicas: 1
|
|
placement:
|
|
constraints: ["node.role == manager"]
|
|
resources:
|
|
limits:
|
|
cpus: "0.5"
|
|
memory: 128m
|
|
restart_policy:
|
|
condition: on-failure
|
|
update_config:
|
|
# stop-first (the default), NOT start-first: a start-first roll would
|
|
# briefly run two daemons — two gateway sessions on one token, the
|
|
# exact thing replicas:1 exists to prevent. A few seconds of gateway
|
|
# downtime per deploy is the correct trade.
|
|
order: stop-first
|
|
failure_action: rollback
|
|
|
|
frontend:
|
|
image: ${REGISTRY_HOST:-git.thermograph.org}/${FRONTEND_IMAGE_PATH:-emi/thermograph/frontend}:${FRONTEND_IMAGE_TAG:?required}
|
|
entrypoint: ["/host/env-entrypoint.sh"]
|
|
# REQUIRED, not cosmetic: overriding `entrypoint:` with no `command:` drops
|
|
# the image's own CMD entirely (Docker/Swarm semantics, not merged) --
|
|
# env-entrypoint.sh then sees zero args and falls through to its hardcoded
|
|
# `exec uvicorn app:app` fallback, which no longer exists in this (Go)
|
|
# image. Without this line the frontend task exits 127 on every deploy.
|
|
command: ["/usr/local/bin/thermograph-frontend"]
|
|
environment:
|
|
THERMOGRAPH_BASE: /
|
|
PORT: 8080
|
|
THERMOGRAPH_SERVICE_ROLE: frontend
|
|
THERMOGRAPH_API_BASE_INTERNAL: http://web:8137
|
|
volumes:
|
|
- /opt/thermograph/infra/deploy/stack/env-entrypoint.sh:/host/env-entrypoint.sh:ro
|
|
- /etc/thermograph/stack.env:/host/thermograph.env:ro
|
|
networks:
|
|
- internal
|
|
deploy:
|
|
replicas: 1
|
|
placement:
|
|
constraints: ["node.role == manager"]
|
|
resources:
|
|
limits:
|
|
cpus: "${FRONTEND_CPUS:-2}"
|
|
restart_policy:
|
|
condition: on-failure
|
|
update_config:
|
|
order: start-first
|
|
failure_action: rollback
|
|
|
|
# Scales `web` between WEB_MIN_REPLICAS and WEB_MAX_REPLICAS on sustained
|
|
# task CPU (docker stats on this node — every task is pinned here today).
|
|
# Deliberately a dumb shell loop with hysteresis + cooldown, not an
|
|
# autoscaling framework: Swarm has no native autoscaler and this workload
|
|
# doesn't justify one (topology doc §6 — declarative scaling as a feature).
|
|
autoscaler:
|
|
image: docker:27-cli
|
|
entrypoint: ["/bin/sh", "/host/autoscale.sh"]
|
|
environment:
|
|
STACK_NAME: ${STACK_NAME:-thermograph}
|
|
MIN_REPLICAS: ${WEB_MIN_REPLICAS:-1}
|
|
MAX_REPLICAS: ${WEB_MAX_REPLICAS:-3}
|
|
# Thresholds are avg docker-stats CPU% PER TASK (host-core-relative: 100
|
|
# = one full core). Up fast, down slow.
|
|
SCALE_UP_CPU: ${SCALE_UP_CPU:-220}
|
|
SCALE_DOWN_CPU: ${SCALE_DOWN_CPU:-60}
|
|
POLL_SECONDS: ${POLL_SECONDS:-15}
|
|
UP_SAMPLES: ${UP_SAMPLES:-3}
|
|
DOWN_SAMPLES: ${DOWN_SAMPLES:-20}
|
|
COOLDOWN_SECONDS: ${COOLDOWN_SECONDS:-180}
|
|
volumes:
|
|
- /var/run/docker.sock:/var/run/docker.sock
|
|
- /opt/thermograph/infra/deploy/stack/autoscale.sh:/host/autoscale.sh:ro
|
|
networks:
|
|
- internal
|
|
deploy:
|
|
replicas: 1
|
|
placement:
|
|
constraints: ["node.role == manager"]
|
|
resources:
|
|
limits:
|
|
cpus: "0.2"
|
|
memory: 64m
|
|
restart_policy:
|
|
condition: any
|
|
|
|
# Second autoscaler instance, watching `lake` (1..2 replicas). The lake is
|
|
# read-only and cache-heavy, so it saturates on decode CPU well below web's
|
|
# thresholds — scale up sooner, still down slowly.
|
|
autoscaler-lake:
|
|
image: docker:27-cli
|
|
entrypoint: ["/bin/sh", "/host/autoscale.sh"]
|
|
environment:
|
|
STACK_NAME: ${STACK_NAME:-thermograph}
|
|
TARGET_SERVICE: lake
|
|
MIN_REPLICAS: ${LAKE_MIN_REPLICAS:-1}
|
|
MAX_REPLICAS: ${LAKE_MAX_REPLICAS:-2}
|
|
SCALE_UP_CPU: ${LAKE_SCALE_UP_CPU:-120}
|
|
SCALE_DOWN_CPU: ${LAKE_SCALE_DOWN_CPU:-30}
|
|
POLL_SECONDS: ${POLL_SECONDS:-15}
|
|
UP_SAMPLES: ${UP_SAMPLES:-3}
|
|
DOWN_SAMPLES: ${DOWN_SAMPLES:-20}
|
|
COOLDOWN_SECONDS: ${COOLDOWN_SECONDS:-180}
|
|
volumes:
|
|
- /var/run/docker.sock:/var/run/docker.sock
|
|
- /opt/thermograph/infra/deploy/stack/autoscale.sh:/host/autoscale.sh:ro
|
|
networks:
|
|
- internal
|
|
deploy:
|
|
replicas: 1
|
|
placement:
|
|
constraints: ["node.role == manager"]
|
|
resources:
|
|
limits:
|
|
cpus: "0.2"
|
|
memory: 64m
|
|
restart_policy:
|
|
condition: any
|
|
|
|
networks:
|
|
internal:
|
|
driver: overlay
|
|
# Attachable so the loopback LB bridge (a PLAIN container — only plain
|
|
# containers can bind 127.0.0.1; Swarm port configs cannot) can join and
|
|
# reach the service VIPs.
|
|
attachable: true
|
|
|
|
volumes:
|
|
# External and explicitly named: the real stack REUSES the volumes the
|
|
# compose stack created (same data, zero migration). deploy-stack.sh's test
|
|
# mode points these at throwaway names instead.
|
|
pgdata:
|
|
external: true
|
|
name: ${PGDATA_VOLUME:-thermograph_pgdata}
|
|
appdata:
|
|
external: true
|
|
name: ${APPDATA_VOLUME:-thermograph_appdata}
|
|
applogs:
|
|
external: true
|
|
name: ${APPLOGS_VOLUME:-thermograph_applogs}
|
|
# Lake parquet cache: node-local and disposable (a cold cache just refills
|
|
# from the bucket), so NOT external — the stack owns its lifecycle.
|
|
lakecache: {}
|