thermograph/infra/docker-compose.dev.yml

86 lines
4.3 KiB
YAML
Raw Normal View History

# Dev overlay for the LAN dev server (deploy/deploy-dev.sh):
#
# docker compose -f docker-compose.yml -f docker-compose.dev.yml up -d
#
# Split-repo adaptation: the monorepo's docker-compose.dev.yml (see
# thermograph/docker-compose.dev.yml) overlaid a base file where backend and
# frontend had `build: .` -- dev's whole point there was `--build` in place from
# the working checkout. This infra repo holds no Dockerfile at all (each service's
# lives in its own app repo, thermograph-backend / thermograph-frontend), so the
# base docker-compose.yml already has NO `build:` for either service, only
# `image: .../${BACKEND_IMAGE_PATH}:${BACKEND_IMAGE_TAG}` (and the frontend
# equivalent) -- same registry-pull model as prod/beta, just pointed at a dev tag
# by deploy-dev.sh. This overlay must NOT reintroduce `build:`; it only relaxes
# resource caps and LAN-exposes a port, same as the monorepo overlay did.
#
# Differences from the prod stack:
# 1. backend is published on ${DEV_BIND_ADDR}, defaulting to LOOPBACK.
#
# This used to be a flat 0.0.0.0:8137, from when dev ran on the operator's
# LAN box and the point was for phones on the Wi-Fi to reach it. Dev now
# runs on vps1, a public VPS, where 0.0.0.0 would publish whatever
# unreviewed branch is in flight to the entire internet — with no Caddy, no
# TLS and no auth in front of it.
#
# So the default is 127.0.0.1 — safe anywhere, including a laptop running
# `make dev-up` — and on vps1 deploy/env-topology.sh keeps it there. Caddy
# fronts the stack at dev.thermograph.org with TLS, HTTP basic auth and
# X-Robots-Tag: noindex (deploy/Caddyfile.vps1). Loopback is what makes that
# guard total: the site block is the only route in, so nothing reaches the
# app without the password — not even from the WireGuard mesh.
# 2. frontend's port publish is dropped entirely -- dev has no Caddy to reach it
# directly, so it stays compose-internal-only, reached solely through
# backend's own reverse-proxy fallback (THERMOGRAPH_FRONTEND_BASE_INTERNAL,
# see backend/web/app.py's _proxy_to_frontend in the backend repo). This
# keeps the dev stack serving the one URL it always has, at :8137.
# 3. The CPU caps are removed -- dev runs UNCAPPED (no thread/CPU limits), unlike
# prod's backend=4 / frontend=2 / db=2 allocation. `!reset` drops the base
# value (both the top-level `cpus:` and the Swarm-style `deploy.resources`
# block the base file carries for parity).
# 4. backend/daemon/lake get a second env_file entry pointing at an OPTIONAL
# copy of the render under $APP_DIR (`required: false`, so it is a no-op
# when absent). It exists for snap-packaged Docker, which is confined to
# $HOME and cannot see /etc/thermograph.env at all — not a permissions
# error; the file simply does not exist as far as snap-confined Docker is
# concerned, so the base file's env_file entry silently loads nothing.
# deploy-dev.sh now produces the mirror ONLY when it detects snap Docker,
# which vps1 (ordinary apt Docker) is not — there, dev reads
# /etc/thermograph.env like beta and prod do, and no plaintext copy of the
# render is written into the checkout. compose appends env_file lists across
# overlays (last-wins on duplicate keys), so this stays additive.
services:
backend:
ports: !override
- "${DEV_BIND_ADDR:-127.0.0.1}:8137:8137"
cpus: !reset null
deploy: !reset null
env_file:
- path: ./deploy/dev-secrets.env
required: false
frontend:
ports: !reset null
cpus: !reset null
deploy: !reset null
daemon: move the Discord gateway and scheduler out of the web process into Go (#21) The gateway bot and APScheduler were long-lived stateful I/O loops running inside the async web app under a leader election. They move into a single Go binary that owns ONLY that I/O -- websocket, RESUME, heartbeat, backoff, timers. It owns no grading logic. Anything needing data calls back over a new internal-only surface (/internal/discord/grade, /internal/jobs/*). Grading depends on polars and the parquet cache; reimplementing it in Go would let the bot's grades drift from the API's. The grade route returns gateway-ready JSON and Go relays the bytes verbatim. The binary ships in the backend image and runs as a second compose service off the same tag, so the two ends of the /internal/* contract can never skew. deploy.sh rolls daemon alongside backend -- without that the service would never be created, since a single-service deploy uses --no-deps. It also probes the image first and skips the daemon when rolling a tag that predates the binary: infra tracks main while image tags are env-staged, so a host can legitimately be asked to roll an older backend image, and creating the service anyway would leave a container crash-looping on a missing binary. replicas: 1 with order: stop-first replaces the leader election -- Discord permits one gateway connection per bot token. THERMOGRAPH_INTERNAL_TOKEN is optional: both ends derive it from THERMOGRAPH_AUTH_SECRET via HMAC under a domain-separation label, so this needs no new vault entry. The derivation is pinned to a shared cross-language test vector asserted on both sides, so drift fails CI instead of 401ing every call. Fail closed when neither secret is set. Improvements over the Python: a close intended for RESUME uses 4000 rather than 1000 (Discord invalidates a session closed 1000, so the old default defeated its own resume); MESSAGE_CREATE runs on a bounded worker pool; and a malformed HELLO returns an error rather than a clean reconnect, which would otherwise reset backoff and hot-loop against the gateway. 365 Python tests pass; Go build/vet/test -race clean; shellcheck 0 findings.
2026-07-23 22:49:54 +00:00
# Same uncapped-on-dev treatment for the daemon; it has no ports to touch
# (outbound-only in every environment).
daemon:
cpus: !reset null
deploy: !reset null
env_file:
- path: ./deploy/dev-secrets.env
required: false
db:
cpus: !reset null
deploy: !reset null
lake:
cpus: !reset null
deploy: !reset null
# Uncapped memory on dev too (prod ceilings it at 8g). The Postgres/DuckDB
# memory *budget* is still the ~8 GB derived by deploy/db/init/20-tuning.sh from
# the default DB_MEMORY (raise DB_MEMORY to give dev more); shm_size stays (it's
# required shared memory for parallel query, not a limit).
mem_limit: !reset null
env_file:
- path: ./deploy/dev-secrets.env
required: false