thermograph/infra
emi 2d3f37c474
All checks were successful
Sync infra to hosts / sync-beta (push) Successful in 8s
Sync infra to hosts / sync-prod (push) Successful in 7s
secrets-guard / encrypted (push) Successful in 8s
shell-lint / shellcheck (push) Successful in 10s
Build + push backend image (Forgejo registry) / build-push (push) Successful in 1m14s
Deploy backend to beta VPS / deploy (push) Successful in 2m4s
daemon: move the Discord gateway and scheduler out of the web process into Go (#21)
The gateway bot and APScheduler were long-lived stateful I/O loops running
inside the async web app under a leader election. They move into a single Go
binary that owns ONLY that I/O -- websocket, RESUME, heartbeat, backoff, timers.

It owns no grading logic. Anything needing data calls back over a new
internal-only surface (/internal/discord/grade, /internal/jobs/*). Grading
depends on polars and the parquet cache; reimplementing it in Go would let the
bot's grades drift from the API's. The grade route returns gateway-ready JSON
and Go relays the bytes verbatim.

The binary ships in the backend image and runs as a second compose service off
the same tag, so the two ends of the /internal/* contract can never skew.
deploy.sh rolls daemon alongside backend -- without that the service would never
be created, since a single-service deploy uses --no-deps. It also probes the
image first and skips the daemon when rolling a tag that predates the binary:
infra tracks main while image tags are env-staged, so a host can legitimately be
asked to roll an older backend image, and creating the service anyway would
leave a container crash-looping on a missing binary.

replicas: 1 with order: stop-first replaces the leader election -- Discord
permits one gateway connection per bot token.

THERMOGRAPH_INTERNAL_TOKEN is optional: both ends derive it from
THERMOGRAPH_AUTH_SECRET via HMAC under a domain-separation label, so this needs
no new vault entry. The derivation is pinned to a shared cross-language test
vector asserted on both sides, so drift fails CI instead of 401ing every call.
Fail closed when neither secret is set.

Improvements over the Python: a close intended for RESUME uses 4000 rather than
1000 (Discord invalidates a session closed 1000, so the old default defeated its
own resume); MESSAGE_CREATE runs on a bounded worker pool; and a malformed HELLO
returns an error rather than a clean reconnect, which would otherwise reset
backoff and hot-loop against the gateway.

365 Python tests pass; Go build/vet/test -race clean; shellcheck 0 findings.
2026-07-23 22:49:54 +00:00
..
.claude/skills/key-gaps Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
deploy daemon: move the Discord gateway and scheduler out of the web process into Go (#21) 2026-07-23 22:49:54 +00:00
terraform Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
.env.example daemon: move the Discord gateway and scheduler out of the web process into Go (#21) 2026-07-23 22:49:54 +00:00
.gitignore Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
.sops.yaml Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
ACCESS.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
CLAUDE.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
DEPLOY-DEV.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
DEPLOY.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
docker-compose.dev.yml daemon: move the Discord gateway and scheduler out of the web process into Go (#21) 2026-07-23 22:49:54 +00:00
docker-compose.openmeteo.yml Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
docker-compose.yml daemon: move the Discord gateway and scheduler out of the web process into Go (#21) 2026-07-23 22:49:54 +00:00
docker-stack.yml Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
Makefile Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
README.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00

thermograph-infra

Infrastructure for Thermograph: Terraform host provisioning, the SOPS+age secrets vault, Docker Swarm/WireGuard networking, Forgejo, Caddy, and the deploy scripts that run the already-built app image on each host. Extracted from the app monorepo (emi/thermograph) — the app repo owns building and testing the app; this repo owns running it.

  • terraform/ — provisions/configures hosts (SSH-driven by default; an optional GCP-creating module is scaffolded, no live resources yet) and triggers each deploy. See terraform/README.md.
  • deploy/secrets/ — the git-native SOPS+age secrets vault (every app secret, encrypted at rest, rendered at deploy time). See deploy/secrets/README.md.
  • deploy/swarm/, deploy/forgejo/ — the 3-node WireGuard/Swarm cluster that hosts Forgejo (git + CI + registry); does not run the app itself. See ACCESS.md and the READMEs under each directory.
  • deploy/deploy.sh — pulls the pinned app image (IMAGE_TAG) and rolls the compose stack; invoked by Terraform and by the app repo's .forgejo/workflows/deploy.yml over SSH.
  • docker-compose*.yml, docker-stack.yml — how the app image runs (compose in production today; docker-stack.yml is a design record for a possible future Swarm-based app deploy, not currently live).

The app's own source, Dockerfile, and build/test CI stay in the app repo — this repo never checks out app source; hosts only pull tagged images from the registry. See ACCESS.md for host access and the Swarm/Forgejo topology, and terraform/README.md for the day-to-day plan/apply workflow.

Branches & how changes reach each environment

  • main — what prod and beta run: their /opt/thermograph checkouts git reset --hard origin/main at the start of every deploy (deploy/deploy.sh). A merge to main reaches those hosts on the next app deploy (or a by-hand deploy.sh run); there is no separate infra deploy trigger.
  • dev — what LAN dev runs: ~/thermograph-dev resets to it via deploy/deploy-dev.sh. Keep it fast-forwarded to main (infra changes are not environment-staged today; the branches exist so LAN dev can trail or lead when needed).
  • release — currently consumed by nothing (prod tracks main, not release). It exists to mirror the app repos' dev→main→release promotion shape if per-environment infra staging is ever wanted; until then, treat main as live-everywhere.

Note the asymmetry with the app repos: app code IS environment-staged (dev→main→release maps to LAN→beta→prod via image tags), infra is not.