|
All checks were successful
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / build-backend (pull_request) Successful in 1m18s
PR build (required check) / changes (pull_request) Successful in 6s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 3s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / validate-observability (pull_request) Has been skipped
web/app.py started two long-lived background jobs under a leader election: the
Discord gateway bot and an APScheduler. Both are stateful I/O loops -- reconnect,
RESUME, heartbeat, backoff, interval timers -- living inside an async web app
that also has to serve requests. This moves them into a single Go binary.
Go owns ONLY the stateful I/O. It owns no climate or grading logic: anything
needing data calls back into Python over a new internal-only HTTP surface
(/internal/discord/grade, /internal/jobs/warm-cities, /internal/jobs/indexnow).
Grading depends on polars and the parquet cache; reimplementing it in Go would
make the bot's grades drift from the API's, and the slash-command path
deliberately shares one grade builder so the two can never disagree. The grade
route returns gateway-ready JSON -- including the ephemeral-flag drop that
discord_bot.py used to do -- and Go relays those bytes verbatim without parsing
the embed.
Packaging: the binary is built by a golang:1.26 stage in the backend Dockerfile
and shipped in the SAME image, run as a second compose service off the SAME tag.
The daemon and backend share the /internal/* contract, so they must never skew
versions; one image makes that structural rather than a convention. Its
entrypoint bypasses entrypoint.sh -- the backend owns alembic, and two racing
migrators is a real hazard.
replicas: 1 in the Swarm stack is load-bearing. Discord permits exactly one
gateway connection per bot token; the pin replaces core/singleton.claim_leader
for this workload. update_config uses order: stop-first, since start-first would
briefly run two gateways. autoscale.sh targets ${STACK_NAME}_web only, so it
cannot scale this.
Security: the internal routes compare the token with hmac.compare_digest and the
whole router 404s when THERMOGRAPH_INTERNAL_TOKEN is unset -- fail closed, never
default open. Caddy only routes /api/*, /digest and /discord/interactions to the
backend, so /internal/* was never publicly reachable; the token is defence in
depth. The router mounts before the catch-all frontend proxy so /internal/*
cannot fall through to it. The daemon refuses to start without the token.
Behaviour preserved from the Python, with the reasoning carried into the Go
comments: non-privileged intents (no MESSAGE_CONTENT, so no portal review);
fatal close codes 4004/4010-4014 stop rather than loop; the bot-author and
self-author mention-loop guard; allowed_mentions locked to {"parse":[],
"replied_user":true} so a crafted query cannot turn a reply into an @everyone
ping; the first cron tick deferred one full interval rather than firing at boot,
since warm-cities already runs at deploy time; and no overlapping warm-cities
run, which would double-spend the archive-fetch quota.
Two deliberate improvements over the Python. A close intended for RESUME now
uses 4000 rather than 1000 -- Discord invalidates a session closed 1000/1001, so
the Python's default close silently defeated its own resume. And MESSAGE_CREATE
is handled on a bounded worker pool rather than an unbounded thread hand-off, so
a flood of mentions cannot spawn unbounded work against the backend.
A .dockerignore is added because a disposable backend/.venv was being swallowed
by COPY . /app/ and duplicated again by the chown layer, inflating the image to
1.8 GB; it builds at 578 MB.
Tests: 29 Go gateway tests covering every behaviour the deleted
test_discord_bot.py asserted, plus cron/config/apiclient suites; 10 new Python
tests for the internal routes (fail-closed, auth, flag drop, per-job 409 guard).
Full suite 359 passed / 7 skipped; go build, vet and test -race clean.
|
||
|---|---|---|
| .. | ||
| .claude/skills/key-gaps | ||
| deploy | ||
| terraform | ||
| .env.example | ||
| .gitignore | ||
| .sops.yaml | ||
| ACCESS.md | ||
| CLAUDE.md | ||
| DEPLOY-DEV.md | ||
| DEPLOY.md | ||
| docker-compose.dev.yml | ||
| docker-compose.openmeteo.yml | ||
| docker-compose.yml | ||
| docker-stack.yml | ||
| Makefile | ||
| README.md | ||
thermograph-infra
Infrastructure for Thermograph: Terraform host
provisioning, the SOPS+age secrets vault, Docker Swarm/WireGuard networking,
Forgejo, Caddy, and the deploy scripts that run the already-built app image on
each host. Extracted from the app monorepo (emi/thermograph) — the app repo
owns building and testing the app; this repo owns running it.
terraform/— provisions/configures hosts (SSH-driven by default; an optional GCP-creating module is scaffolded, no live resources yet) and triggers each deploy. Seeterraform/README.md.deploy/secrets/— the git-native SOPS+age secrets vault (every app secret, encrypted at rest, rendered at deploy time). Seedeploy/secrets/README.md.deploy/swarm/,deploy/forgejo/— the 3-node WireGuard/Swarm cluster that hosts Forgejo (git + CI + registry); does not run the app itself. SeeACCESS.mdand the READMEs under each directory.deploy/deploy.sh— pulls the pinned app image (IMAGE_TAG) and rolls the compose stack; invoked by Terraform and by the app repo's.forgejo/workflows/deploy.ymlover SSH.docker-compose*.yml,docker-stack.yml— how the app image runs (compose in production today;docker-stack.ymlis a design record for a possible future Swarm-based app deploy, not currently live).
The app's own source, Dockerfile, and build/test CI stay in the app repo —
this repo never checks out app source; hosts only pull tagged images from the
registry. See ACCESS.md for host access and the Swarm/Forgejo topology, and
terraform/README.md for the day-to-day plan/apply workflow.
Branches & how changes reach each environment
main— what prod and beta run: their/opt/thermographcheckoutsgit reset --hard origin/mainat the start of every deploy (deploy/deploy.sh). A merge tomainreaches those hosts on the next app deploy (or a by-handdeploy.shrun); there is no separate infra deploy trigger.dev— what LAN dev runs:~/thermograph-devresets to it viadeploy/deploy-dev.sh. Keep it fast-forwarded tomain(infra changes are not environment-staged today; the branches exist so LAN dev can trail or lead when needed).release— currently consumed by nothing (prod tracksmain, notrelease). It exists to mirror the app repos' dev→main→release promotion shape if per-environment infra staging is ever wanted; until then, treatmainas live-everywhere.
Note the asymmetry with the app repos: app code IS environment-staged (dev→main→release maps to LAN→beta→prod via image tags), infra is not.