thermograph/infra/docker-compose.dev.yml
emi 2d3f37c474
All checks were successful
Sync infra to hosts / sync-beta (push) Successful in 8s
Sync infra to hosts / sync-prod (push) Successful in 7s
secrets-guard / encrypted (push) Successful in 8s
shell-lint / shellcheck (push) Successful in 10s
Build + push backend image (Forgejo registry) / build-push (push) Successful in 1m14s
Deploy backend to beta VPS / deploy (push) Successful in 2m4s
daemon: move the Discord gateway and scheduler out of the web process into Go (#21)
The gateway bot and APScheduler were long-lived stateful I/O loops running
inside the async web app under a leader election. They move into a single Go
binary that owns ONLY that I/O -- websocket, RESUME, heartbeat, backoff, timers.

It owns no grading logic. Anything needing data calls back over a new
internal-only surface (/internal/discord/grade, /internal/jobs/*). Grading
depends on polars and the parquet cache; reimplementing it in Go would let the
bot's grades drift from the API's. The grade route returns gateway-ready JSON
and Go relays the bytes verbatim.

The binary ships in the backend image and runs as a second compose service off
the same tag, so the two ends of the /internal/* contract can never skew.
deploy.sh rolls daemon alongside backend -- without that the service would never
be created, since a single-service deploy uses --no-deps. It also probes the
image first and skips the daemon when rolling a tag that predates the binary:
infra tracks main while image tags are env-staged, so a host can legitimately be
asked to roll an older backend image, and creating the service anyway would
leave a container crash-looping on a missing binary.

replicas: 1 with order: stop-first replaces the leader election -- Discord
permits one gateway connection per bot token.

THERMOGRAPH_INTERNAL_TOKEN is optional: both ends derive it from
THERMOGRAPH_AUTH_SECRET via HMAC under a domain-separation label, so this needs
no new vault entry. The derivation is pinned to a shared cross-language test
vector asserted on both sides, so drift fails CI instead of 401ing every call.
Fail closed when neither secret is set.

Improvements over the Python: a close intended for RESUME uses 4000 rather than
1000 (Discord invalidates a session closed 1000, so the old default defeated its
own resume); MESSAGE_CREATE runs on a bounded worker pool; and a malformed HELLO
returns an error rather than a clean reconnect, which would otherwise reset
backoff and hot-loop against the gateway.

365 Python tests pass; Go build/vet/test -race clean; shellcheck 0 findings.
2026-07-23 22:49:54 +00:00

52 lines
2.6 KiB
YAML

# Dev overlay for the LAN dev server (deploy/deploy-dev.sh):
#
# docker compose -f docker-compose.yml -f docker-compose.dev.yml up -d
#
# Split-repo adaptation: the monorepo's docker-compose.dev.yml (see
# thermograph/docker-compose.dev.yml) overlaid a base file where backend and
# frontend had `build: .` -- dev's whole point there was `--build` in place from
# the working checkout. This infra repo holds no Dockerfile at all (each service's
# lives in its own app repo, thermograph-backend / thermograph-frontend), so the
# base docker-compose.yml already has NO `build:` for either service, only
# `image: .../${BACKEND_IMAGE_PATH}:${BACKEND_IMAGE_TAG}` (and the frontend
# equivalent) -- same registry-pull model as prod/beta, just pointed at a dev tag
# by deploy-dev.sh. This overlay must NOT reintroduce `build:`; it only relaxes
# resource caps and LAN-exposes a port, same as the monorepo overlay did.
#
# Differences from the prod stack (unchanged intent from the monorepo overlay):
# 1. backend is published on ALL interfaces (0.0.0.0:8137), not loopback, so
# phones and other devices on the Wi-Fi can reach the dev server directly --
# dev has no Caddy in front (prod does, which is why the base file binds
# 127.0.0.1 only).
# 2. frontend's port publish is dropped entirely -- dev has no Caddy to reach it
# directly, so it stays compose-internal-only, reached solely through
# backend's own reverse-proxy fallback (THERMOGRAPH_FRONTEND_BASE_INTERNAL,
# see backend/web/app.py's _proxy_to_frontend in the backend repo). This
# keeps the dev stack serving the one URL it always has, at :8137.
# 3. The CPU caps are removed -- dev runs UNCAPPED (no thread/CPU limits), unlike
# prod's backend=4 / frontend=2 / db=2 allocation. `!reset` drops the base
# value (both the top-level `cpus:` and the Swarm-style `deploy.resources`
# block the base file carries for parity).
services:
backend:
ports: !override
- "8137:8137"
cpus: !reset null
deploy: !reset null
frontend:
ports: !reset null
cpus: !reset null
deploy: !reset null
# Same uncapped-on-dev treatment for the daemon; it has no ports to touch
# (outbound-only in every environment).
daemon:
cpus: !reset null
deploy: !reset null
db:
cpus: !reset null
deploy: !reset null
# Uncapped memory on dev too (prod ceilings it at 8g). The Postgres/DuckDB
# memory *budget* is still the ~8 GB derived by deploy/db/init/20-tuning.sh from
# the default DB_MEMORY (raise DB_MEMORY to give dev more); shm_size stays (it's
# required shared memory for parallel query, not a limit).
mem_limit: !reset null