2026-07-22 19:20:05 +00:00
|
|
|
# Dev overlay for the LAN dev server (deploy/deploy-dev.sh):
|
|
|
|
|
#
|
|
|
|
|
# docker compose -f docker-compose.yml -f docker-compose.dev.yml up -d
|
|
|
|
|
#
|
|
|
|
|
# Split-repo adaptation: the monorepo's docker-compose.dev.yml (see
|
|
|
|
|
# thermograph/docker-compose.dev.yml) overlaid a base file where backend and
|
|
|
|
|
# frontend had `build: .` -- dev's whole point there was `--build` in place from
|
|
|
|
|
# the working checkout. This infra repo holds no Dockerfile at all (each service's
|
|
|
|
|
# lives in its own app repo, thermograph-backend / thermograph-frontend), so the
|
|
|
|
|
# base docker-compose.yml already has NO `build:` for either service, only
|
|
|
|
|
# `image: .../${BACKEND_IMAGE_PATH}:${BACKEND_IMAGE_TAG}` (and the frontend
|
|
|
|
|
# equivalent) -- same registry-pull model as prod/beta, just pointed at a dev tag
|
|
|
|
|
# by deploy-dev.sh. This overlay must NOT reintroduce `build:`; it only relaxes
|
|
|
|
|
# resource caps and LAN-exposes a port, same as the monorepo overlay did.
|
|
|
|
|
#
|
|
|
|
|
# Differences from the prod stack (unchanged intent from the monorepo overlay):
|
|
|
|
|
# 1. backend is published on ALL interfaces (0.0.0.0:8137), not loopback, so
|
|
|
|
|
# phones and other devices on the Wi-Fi can reach the dev server directly --
|
|
|
|
|
# dev has no Caddy in front (prod does, which is why the base file binds
|
|
|
|
|
# 127.0.0.1 only).
|
|
|
|
|
# 2. frontend's port publish is dropped entirely -- dev has no Caddy to reach it
|
|
|
|
|
# directly, so it stays compose-internal-only, reached solely through
|
|
|
|
|
# backend's own reverse-proxy fallback (THERMOGRAPH_FRONTEND_BASE_INTERNAL,
|
|
|
|
|
# see backend/web/app.py's _proxy_to_frontend in the backend repo). This
|
|
|
|
|
# keeps the dev stack serving the one URL it always has, at :8137.
|
|
|
|
|
# 3. The CPU caps are removed -- dev runs UNCAPPED (no thread/CPU limits), unlike
|
|
|
|
|
# prod's backend=4 / frontend=2 / db=2 allocation. `!reset` drops the base
|
|
|
|
|
# value (both the top-level `cpus:` and the Swarm-style `deploy.resources`
|
|
|
|
|
# block the base file carries for parity).
|
|
|
|
|
services:
|
|
|
|
|
backend:
|
|
|
|
|
ports: !override
|
|
|
|
|
- "8137:8137"
|
|
|
|
|
cpus: !reset null
|
|
|
|
|
deploy: !reset null
|
|
|
|
|
frontend:
|
|
|
|
|
ports: !reset null
|
|
|
|
|
cpus: !reset null
|
|
|
|
|
deploy: !reset null
|
daemon: move the Discord gateway and scheduler out of the web process into Go (#21)
The gateway bot and APScheduler were long-lived stateful I/O loops running
inside the async web app under a leader election. They move into a single Go
binary that owns ONLY that I/O -- websocket, RESUME, heartbeat, backoff, timers.
It owns no grading logic. Anything needing data calls back over a new
internal-only surface (/internal/discord/grade, /internal/jobs/*). Grading
depends on polars and the parquet cache; reimplementing it in Go would let the
bot's grades drift from the API's. The grade route returns gateway-ready JSON
and Go relays the bytes verbatim.
The binary ships in the backend image and runs as a second compose service off
the same tag, so the two ends of the /internal/* contract can never skew.
deploy.sh rolls daemon alongside backend -- without that the service would never
be created, since a single-service deploy uses --no-deps. It also probes the
image first and skips the daemon when rolling a tag that predates the binary:
infra tracks main while image tags are env-staged, so a host can legitimately be
asked to roll an older backend image, and creating the service anyway would
leave a container crash-looping on a missing binary.
replicas: 1 with order: stop-first replaces the leader election -- Discord
permits one gateway connection per bot token.
THERMOGRAPH_INTERNAL_TOKEN is optional: both ends derive it from
THERMOGRAPH_AUTH_SECRET via HMAC under a domain-separation label, so this needs
no new vault entry. The derivation is pinned to a shared cross-language test
vector asserted on both sides, so drift fails CI instead of 401ing every call.
Fail closed when neither secret is set.
Improvements over the Python: a close intended for RESUME uses 4000 rather than
1000 (Discord invalidates a session closed 1000, so the old default defeated its
own resume); MESSAGE_CREATE runs on a bounded worker pool; and a malformed HELLO
returns an error rather than a clean reconnect, which would otherwise reset
backoff and hot-loop against the gateway.
365 Python tests pass; Go build/vet/test -race clean; shellcheck 0 findings.
2026-07-23 22:49:54 +00:00
|
|
|
# Same uncapped-on-dev treatment for the daemon; it has no ports to touch
|
|
|
|
|
# (outbound-only in every environment).
|
|
|
|
|
daemon:
|
|
|
|
|
cpus: !reset null
|
|
|
|
|
deploy: !reset null
|
2026-07-22 19:20:05 +00:00
|
|
|
db:
|
|
|
|
|
cpus: !reset null
|
|
|
|
|
deploy: !reset null
|
|
|
|
|
# Uncapped memory on dev too (prod ceilings it at 8g). The Postgres/DuckDB
|
|
|
|
|
# memory *budget* is still the ~8 GB derived by deploy/db/init/20-tuning.sh from
|
|
|
|
|
# the default DB_MEMORY (raise DB_MEMORY to give dev more); shm_size stays (it's
|
|
|
|
|
# required shared memory for parallel query, not a limit).
|
|
|
|
|
mem_limit: !reset null
|