thermograph/docker-compose.yml
Emi Griffith c41ade74b1 Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235)
* Split web/worker duties with THERMOGRAPH_ROLE

Background work (the subscription notifier) is welded to the same process
that serves requests, so scaling the web tier to N replicas would also scale
notifier instances unless something restricts it further than leader
election alone.

Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process
behavior). Every replica runs the same image; ROLE only gates whether a
process is allowed to own the notifier at all, layered on top of the
existing leader election: web replicas never start it even if they'd win
leader election, worker replicas start it if they win. The decision is
pulled into _should_run_notifier() so it's unit-testable without booting the
full app (DB init, places index, neighbor warmer).

Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE)
so a worker replica - which serves no real traffic - still has something
Swarm can health-check.

* Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate

Three changes toward the hop-1 interim cutover, all inert until Track B
stands up the platform:

docker-stack.yml: the Swarm stack file for the interim cutover, distinct
from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a
pre-built image (IMAGE_TAG) instead of building in place; app/worker publish
no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing
mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted),
reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard
needs a smaller MTU or large payloads silently stall); db is placement-
pinned to a labelled node; app/worker skip inline migrations
(RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing
that ever runs Alembic; secrets are real Swarm secrets mounted at
/run/secrets, read by the entrypoint shim rather than plain env vars.

TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default
latest-pg18, today's behavior unchanged), wired through Terraform
(timescaledb_tag, default "latest-pg18") so it can actually be pinned to an
exact minor without hand-editing the host - required before any host of the
stack could replicate with another (a floating tag risks mismatched
extension minors, which blocks a physical replica and risks compressed-
chunk corruption on restore).

Caddy active health-gate: both the Terraform-rendered Caddyfile and the live
deploy/Caddyfile now health-check the app on the same cheap /healthz route
its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage,
so it's cheap enough for a tight interval and works identically for a
worker replica, which serves no public traffic at all) - Caddy won't
forward into a container that's still booting or unhealthy.

Verified live: built and booted the real image via docker compose - both
containers report healthy via the new /healthz-based HEALTHCHECK, and GET /
still renders the full SSR homepage unchanged. Both Caddyfiles validated
with the real caddy binary. docker-stack.yml validated with docker compose
config (required-var guards fire with clear messages; secrets correctly
mount at /run/secrets/<name>, matching the entrypoint shim's mapping).
docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside
the existing openmeteo overlay. terraform validate + fmt clean.
2026-07-21 00:39:48 +00:00

117 lines
5.7 KiB
YAML

# Thermograph production stack: the FastAPI app plus its TimescaleDB (PostgreSQL 18)
# database.
#
# docker compose up -d --build # or: make up
#
# POSTGRES_PASSWORD must be set at `docker compose` time — compose reads it from
# the repo-root .env for local runs (copy .env.example -> .env), and in prod the
# systemd unit's EnvironmentFile=/etc/thermograph.env puts it in the environment
# so `docker compose up` can interpolate it. It is used BOTH to initialize the db
# container and to build the app's THERMOGRAPH_DATABASE_URL below.
services:
db:
# TimescaleDB on PostgreSQL 18 (the stock image already sets
# shared_preload_libraries=timescaledb). The app's climate record — the full
# daily archive and the hourly recent+forecast bundle — lives in hypertables
# here (see backend/data/climate_store.py), so the DB, not the filesystem, is the
# source of truth. The init script CREATE EXTENSIONs timescaledb on a fresh volume
# (Alembic also does, idempotently, at app boot).
#
# TIMESCALEDB_TAG defaults to the floating latest-pg18 tag (today's behavior,
# unchanged) so a plain `docker compose up` keeps working with no setup. Pin it
# to an exact minor (e.g. 2.17.2-pg18) before any host of this stack could ever
# replicate with another — a floating tag risks two hosts landing on different
# extension minors, which blocks a physical replica and risks compressed-chunk
# corruption on restore. docker-stack.yml (the Swarm interim stack) REQUIRES an
# exact pin for exactly this reason; use the SAME tag on both once you set one.
image: timescale/timescaledb:${TIMESCALEDB_TAG:-latest-pg18}
environment:
POSTGRES_USER: thermograph
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD}
POSTGRES_DB: thermograph
# The init tuning script (deploy/db/init/20-tuning.sh) scales the Postgres +
# DuckDB memory GUCs from this budget, matching the mem_limit below. One knob
# per host: Terraform sets DB_MEMORY (prod 16g), local/beta default 8g.
DB_MEMORY: ${DB_MEMORY:-8g}
volumes:
# Mount the volume at the PARENT of the data dir and let the image pick its own
# PGDATA subdir under it (the timescaledb image defaults to
# /var/lib/postgresql/data). Pinning PGDATA directly at the mountpoint trips an
# initdb chmod on some Docker setups; the whole tree still persists this way.
- pgdata:/var/lib/postgresql
- ./deploy/db/init:/docker-entrypoint-initdb.d
healthcheck:
test: ["CMD-SHELL", "pg_isready -U thermograph -d thermograph"]
interval: 5s
timeout: 5s
retries: 10
# Give Postgres room to cache + process. mem_limit is the hard ceiling; the actual
# budget is tuned in deploy/db/init/20-tuning.sh, which scales shared_buffers (25%),
# effective_cache_size (75%), work_mem and maintenance_work_mem from DB_MEMORY — so
# raising DB_MEMORY raises both the cap and the tuning together. shm_size backs
# parallel-query shared memory (the 64MB docker default is
# too small once shared_buffers/parallelism grow).
# Sized via env (Terraform sets DB_CPUS/DB_MEMORY per host); defaults match the
# historical 2 CPU / 8 GB budget so a plain `docker compose up` is unchanged.
mem_limit: ${DB_MEMORY:-8g}
shm_size: 1gb
# Cap the DB at DB_CPUS CPUs. Compose v2 honors the top-level `cpus:`; the
# deploy.resources block is the Swarm-style equivalent, kept for parity.
cpus: ${DB_CPUS:-2}
deploy:
resources:
limits:
cpus: "${DB_CPUS:-2}"
# Must match mem_limit above (compose rejects distinct values).
memory: ${DB_MEMORY:-8g}
restart: unless-stopped
# No host port on purpose: the app reaches Postgres as db:5432 on the
# compose network. Nothing outside the stack should touch the database.
app:
build: .
depends_on:
db:
condition: service_healthy
environment:
# Built from POSTGRES_PASSWORD; this `environment` value wins over anything
# in env_file, so the URL always matches the db container's password.
THERMOGRAPH_DATABASE_URL: postgresql+asyncpg://thermograph:${POSTGRES_PASSWORD}@db:5432/thermograph
THERMOGRAPH_BASE: /
PORT: 8137
# Worker count is env-driven so Terraform can raise it on a bigger host;
# defaults to 4 to keep a plain `docker compose up` identical to before.
WORKERS: ${WORKERS:-4}
# One worker wins this lock and runs the subscription notifier / homepage
# sweep; it lives on the appdata volume so it's shared across workers.
THERMOGRAPH_SINGLETON_LOCK: /app/data/notifier.lock
# Prod secrets live in /etc/thermograph.env: POSTGRES_PASSWORD,
# THERMOGRAPH_AUTH_SECRET, THERMOGRAPH_VAPID_PRIVATE_KEY/_PUBLIC_KEY,
# THERMOGRAPH_COOKIE_SECURE=1, mail/Discord keys, ... (see
# deploy/thermograph.env.example). `required: false` so local `docker compose
# up` works without that file — it reads POSTGRES_PASSWORD from repo-root .env.
env_file:
- path: /etc/thermograph.env
required: false
volumes:
# Parquet cache, notifier.lock, homepage.json, vapid.json persist here.
- appdata:/app/data
- applogs:/app/logs
# Sized via env (Terraform sets APP_CPUS per host); defaults to 4 CPUs. The
# deploy.resources block mirrors the top-level `cpus:` for Swarm parity; the
# dev overlay drops both so dev runs uncapped.
cpus: ${APP_CPUS:-4}
deploy:
resources:
limits:
cpus: "${APP_CPUS:-4}"
ports:
# Loopback only — host Caddy terminates TLS and reverse-proxies to this.
- "127.0.0.1:8137:8137"
restart: unless-stopped
volumes:
pgdata: {}
appdata: {}
applogs: {}