thermograph/infra/docker-compose.dev.yml

53 lines
2.6 KiB
YAML
Raw Normal View History

# Dev overlay for the LAN dev server (deploy/deploy-dev.sh):
#
# docker compose -f docker-compose.yml -f docker-compose.dev.yml up -d
#
# Split-repo adaptation: the monorepo's docker-compose.dev.yml (see
# thermograph/docker-compose.dev.yml) overlaid a base file where backend and
# frontend had `build: .` -- dev's whole point there was `--build` in place from
# the working checkout. This infra repo holds no Dockerfile at all (each service's
# lives in its own app repo, thermograph-backend / thermograph-frontend), so the
# base docker-compose.yml already has NO `build:` for either service, only
# `image: .../${BACKEND_IMAGE_PATH}:${BACKEND_IMAGE_TAG}` (and the frontend
# equivalent) -- same registry-pull model as prod/beta, just pointed at a dev tag
# by deploy-dev.sh. This overlay must NOT reintroduce `build:`; it only relaxes
# resource caps and LAN-exposes a port, same as the monorepo overlay did.
#
# Differences from the prod stack (unchanged intent from the monorepo overlay):
# 1. backend is published on ALL interfaces (0.0.0.0:8137), not loopback, so
# phones and other devices on the Wi-Fi can reach the dev server directly --
# dev has no Caddy in front (prod does, which is why the base file binds
# 127.0.0.1 only).
# 2. frontend's port publish is dropped entirely -- dev has no Caddy to reach it
# directly, so it stays compose-internal-only, reached solely through
# backend's own reverse-proxy fallback (THERMOGRAPH_FRONTEND_BASE_INTERNAL,
# see backend/web/app.py's _proxy_to_frontend in the backend repo). This
# keeps the dev stack serving the one URL it always has, at :8137.
# 3. The CPU caps are removed -- dev runs UNCAPPED (no thread/CPU limits), unlike
# prod's backend=4 / frontend=2 / db=2 allocation. `!reset` drops the base
# value (both the top-level `cpus:` and the Swarm-style `deploy.resources`
# block the base file carries for parity).
services:
backend:
ports: !override
- "8137:8137"
cpus: !reset null
deploy: !reset null
frontend:
ports: !reset null
cpus: !reset null
deploy: !reset null
daemon: move the Discord gateway and scheduler out of the web process into Go (#21) The gateway bot and APScheduler were long-lived stateful I/O loops running inside the async web app under a leader election. They move into a single Go binary that owns ONLY that I/O -- websocket, RESUME, heartbeat, backoff, timers. It owns no grading logic. Anything needing data calls back over a new internal-only surface (/internal/discord/grade, /internal/jobs/*). Grading depends on polars and the parquet cache; reimplementing it in Go would let the bot's grades drift from the API's. The grade route returns gateway-ready JSON and Go relays the bytes verbatim. The binary ships in the backend image and runs as a second compose service off the same tag, so the two ends of the /internal/* contract can never skew. deploy.sh rolls daemon alongside backend -- without that the service would never be created, since a single-service deploy uses --no-deps. It also probes the image first and skips the daemon when rolling a tag that predates the binary: infra tracks main while image tags are env-staged, so a host can legitimately be asked to roll an older backend image, and creating the service anyway would leave a container crash-looping on a missing binary. replicas: 1 with order: stop-first replaces the leader election -- Discord permits one gateway connection per bot token. THERMOGRAPH_INTERNAL_TOKEN is optional: both ends derive it from THERMOGRAPH_AUTH_SECRET via HMAC under a domain-separation label, so this needs no new vault entry. The derivation is pinned to a shared cross-language test vector asserted on both sides, so drift fails CI instead of 401ing every call. Fail closed when neither secret is set. Improvements over the Python: a close intended for RESUME uses 4000 rather than 1000 (Discord invalidates a session closed 1000, so the old default defeated its own resume); MESSAGE_CREATE runs on a bounded worker pool; and a malformed HELLO returns an error rather than a clean reconnect, which would otherwise reset backoff and hot-loop against the gateway. 365 Python tests pass; Go build/vet/test -race clean; shellcheck 0 findings.
2026-07-23 22:49:54 +00:00
# Same uncapped-on-dev treatment for the daemon; it has no ports to touch
# (outbound-only in every environment).
daemon:
cpus: !reset null
deploy: !reset null
db:
cpus: !reset null
deploy: !reset null
# Uncapped memory on dev too (prod ceilings it at 8g). The Postgres/DuckDB
# memory *budget* is still the ~8 GB derived by deploy/db/init/20-tuning.sh from
# the default DB_MEMORY (raise DB_MEMORY to give dev more); shm_size stays (it's
# required shared memory for parallel query, not a limit).
mem_limit: !reset null