All checks were successful
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / build-backend (pull_request) Successful in 1m18s
PR build (required check) / changes (pull_request) Successful in 6s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 3s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / validate-observability (pull_request) Has been skipped
web/app.py started two long-lived background jobs under a leader election: the
Discord gateway bot and an APScheduler. Both are stateful I/O loops -- reconnect,
RESUME, heartbeat, backoff, interval timers -- living inside an async web app
that also has to serve requests. This moves them into a single Go binary.
Go owns ONLY the stateful I/O. It owns no climate or grading logic: anything
needing data calls back into Python over a new internal-only HTTP surface
(/internal/discord/grade, /internal/jobs/warm-cities, /internal/jobs/indexnow).
Grading depends on polars and the parquet cache; reimplementing it in Go would
make the bot's grades drift from the API's, and the slash-command path
deliberately shares one grade builder so the two can never disagree. The grade
route returns gateway-ready JSON -- including the ephemeral-flag drop that
discord_bot.py used to do -- and Go relays those bytes verbatim without parsing
the embed.
Packaging: the binary is built by a golang:1.26 stage in the backend Dockerfile
and shipped in the SAME image, run as a second compose service off the SAME tag.
The daemon and backend share the /internal/* contract, so they must never skew
versions; one image makes that structural rather than a convention. Its
entrypoint bypasses entrypoint.sh -- the backend owns alembic, and two racing
migrators is a real hazard.
replicas: 1 in the Swarm stack is load-bearing. Discord permits exactly one
gateway connection per bot token; the pin replaces core/singleton.claim_leader
for this workload. update_config uses order: stop-first, since start-first would
briefly run two gateways. autoscale.sh targets ${STACK_NAME}_web only, so it
cannot scale this.
Security: the internal routes compare the token with hmac.compare_digest and the
whole router 404s when THERMOGRAPH_INTERNAL_TOKEN is unset -- fail closed, never
default open. Caddy only routes /api/*, /digest and /discord/interactions to the
backend, so /internal/* was never publicly reachable; the token is defence in
depth. The router mounts before the catch-all frontend proxy so /internal/*
cannot fall through to it. The daemon refuses to start without the token.
Behaviour preserved from the Python, with the reasoning carried into the Go
comments: non-privileged intents (no MESSAGE_CONTENT, so no portal review);
fatal close codes 4004/4010-4014 stop rather than loop; the bot-author and
self-author mention-loop guard; allowed_mentions locked to {"parse":[],
"replied_user":true} so a crafted query cannot turn a reply into an @everyone
ping; the first cron tick deferred one full interval rather than firing at boot,
since warm-cities already runs at deploy time; and no overlapping warm-cities
run, which would double-spend the archive-fetch quota.
Two deliberate improvements over the Python. A close intended for RESUME now
uses 4000 rather than 1000 -- Discord invalidates a session closed 1000/1001, so
the Python's default close silently defeated its own resume. And MESSAGE_CREATE
is handled on a bounded worker pool rather than an unbounded thread hand-off, so
a flood of mentions cannot spawn unbounded work against the backend.
A .dockerignore is added because a disposable backend/.venv was being swallowed
by COPY . /app/ and duplicated again by the chown layer, inflating the image to
1.8 GB; it builds at 578 MB.
Tests: 29 Go gateway tests covering every behaviour the deleted
test_discord_bot.py asserted, plus cron/config/apiclient suites; 10 new Python
tests for the internal routes (fail-closed, auth, flag drop, per-job 409 guard).
Full suite 359 passed / 7 skipped; go build, vet and test -race clean.
52 lines
2.6 KiB
YAML
52 lines
2.6 KiB
YAML
# Dev overlay for the LAN dev server (deploy/deploy-dev.sh):
|
|
#
|
|
# docker compose -f docker-compose.yml -f docker-compose.dev.yml up -d
|
|
#
|
|
# Split-repo adaptation: the monorepo's docker-compose.dev.yml (see
|
|
# thermograph/docker-compose.dev.yml) overlaid a base file where backend and
|
|
# frontend had `build: .` -- dev's whole point there was `--build` in place from
|
|
# the working checkout. This infra repo holds no Dockerfile at all (each service's
|
|
# lives in its own app repo, thermograph-backend / thermograph-frontend), so the
|
|
# base docker-compose.yml already has NO `build:` for either service, only
|
|
# `image: .../${BACKEND_IMAGE_PATH}:${BACKEND_IMAGE_TAG}` (and the frontend
|
|
# equivalent) -- same registry-pull model as prod/beta, just pointed at a dev tag
|
|
# by deploy-dev.sh. This overlay must NOT reintroduce `build:`; it only relaxes
|
|
# resource caps and LAN-exposes a port, same as the monorepo overlay did.
|
|
#
|
|
# Differences from the prod stack (unchanged intent from the monorepo overlay):
|
|
# 1. backend is published on ALL interfaces (0.0.0.0:8137), not loopback, so
|
|
# phones and other devices on the Wi-Fi can reach the dev server directly --
|
|
# dev has no Caddy in front (prod does, which is why the base file binds
|
|
# 127.0.0.1 only).
|
|
# 2. frontend's port publish is dropped entirely -- dev has no Caddy to reach it
|
|
# directly, so it stays compose-internal-only, reached solely through
|
|
# backend's own reverse-proxy fallback (THERMOGRAPH_FRONTEND_BASE_INTERNAL,
|
|
# see backend/web/app.py's _proxy_to_frontend in the backend repo). This
|
|
# keeps the dev stack serving the one URL it always has, at :8137.
|
|
# 3. The CPU caps are removed -- dev runs UNCAPPED (no thread/CPU limits), unlike
|
|
# prod's backend=4 / frontend=2 / db=2 allocation. `!reset` drops the base
|
|
# value (both the top-level `cpus:` and the Swarm-style `deploy.resources`
|
|
# block the base file carries for parity).
|
|
services:
|
|
backend:
|
|
ports: !override
|
|
- "8137:8137"
|
|
cpus: !reset null
|
|
deploy: !reset null
|
|
frontend:
|
|
ports: !reset null
|
|
cpus: !reset null
|
|
deploy: !reset null
|
|
# Same uncapped-on-dev treatment for the daemon; it has no ports to touch
|
|
# (outbound-only in every environment).
|
|
daemon:
|
|
cpus: !reset null
|
|
deploy: !reset null
|
|
db:
|
|
cpus: !reset null
|
|
deploy: !reset null
|
|
# Uncapped memory on dev too (prod ceilings it at 8g). The Postgres/DuckDB
|
|
# memory *budget* is still the ~8 GB derived by deploy/db/init/20-tuning.sh from
|
|
# the default DB_MEMORY (raise DB_MEMORY to give dev more); shm_size stays (it's
|
|
# required shared memory for parallel query, not a limit).
|
|
mem_limit: !reset null
|