|
All checks were successful
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-backend (pull_request) Successful in 1m33s
PR build (required check) / gate (pull_request) Successful in 3s
web/app.py started two long-lived background jobs under a leader election: the
Discord gateway bot and an APScheduler. Both are stateful I/O loops -- reconnect,
RESUME, heartbeat, backoff, interval timers -- living inside an async web app
that also has to serve requests. This moves them into a single Go binary.
Go owns ONLY the stateful I/O. It owns no climate or grading logic: anything
needing data calls back into Python over a new internal-only HTTP surface
(/internal/discord/grade, /internal/jobs/warm-cities, /internal/jobs/indexnow).
Grading depends on polars and the parquet cache; reimplementing it in Go would
make the bot's grades drift from the API's, and the slash-command path
deliberately shares one grade builder so the two can never disagree. The grade
route returns gateway-ready JSON -- including the ephemeral-flag drop that
discord_bot.py used to do -- and Go relays those bytes verbatim without parsing
the embed.
Packaging: the binary is built by a golang:1.26 stage in the backend Dockerfile
and shipped in the SAME image, run as a second compose service off the SAME tag.
The daemon and backend share the /internal/* contract, so they must never skew
versions; one image makes that structural rather than a convention. Its
entrypoint bypasses entrypoint.sh -- the backend owns alembic, and two racing
migrators is a real hazard.
replicas: 1 in the Swarm stack is load-bearing. Discord permits exactly one
gateway connection per bot token; the pin replaces core/singleton.claim_leader
for this workload. update_config uses order: stop-first, since start-first would
briefly run two gateways. autoscale.sh targets ${STACK_NAME}_web only, so it
cannot scale this.
Security: the internal routes compare the token with hmac.compare_digest and the
whole router 404s when THERMOGRAPH_INTERNAL_TOKEN is unset -- fail closed, never
default open. Caddy only routes /api/*, /digest and /discord/interactions to the
backend, so /internal/* was never publicly reachable; the token is defence in
depth. The router mounts before the catch-all frontend proxy so /internal/*
cannot fall through to it. The daemon refuses to start without the token.
Behaviour preserved from the Python, with the reasoning carried into the Go
comments: non-privileged intents (no MESSAGE_CONTENT, so no portal review);
fatal close codes 4004/4010-4014 stop rather than loop; the bot-author and
self-author mention-loop guard; allowed_mentions locked to {"parse":[],
"replied_user":true} so a crafted query cannot turn a reply into an @everyone
ping; the first cron tick deferred one full interval rather than firing at boot,
since warm-cities already runs at deploy time; and no overlapping warm-cities
run, which would double-spend the archive-fetch quota.
Two deliberate improvements over the Python. A close intended for RESUME now
uses 4000 rather than 1000 -- Discord invalidates a session closed 1000/1001, so
the Python's default close silently defeated its own resume. And MESSAGE_CREATE
is handled on a bounded worker pool rather than an unbounded thread hand-off, so
a flood of mentions cannot spawn unbounded work against the backend.
A .dockerignore is added because a disposable backend/.venv was being swallowed
by COPY . /app/ and duplicated again by the chown layer, inflating the image to
1.8 GB; it builds at 578 MB.
Tests: 29 Go gateway tests covering every behaviour the deleted
test_discord_bot.py asserted, plus cron/config/apiclient suites; 10 new Python
tests for the internal routes (fail-closed, auth, flag drop, per-job 409 guard).
Full suite 359 passed / 7 skipped; go build, vet and test -race clean.
|
||
|---|---|---|
| .. | ||
| accounts | ||
| alembic | ||
| api | ||
| core | ||
| daemon | ||
| data | ||
| deploy | ||
| notifications | ||
| scripts | ||
| tests | ||
| web | ||
| .dockerignore | ||
| .gitignore | ||
| alembic.ini | ||
| app.py | ||
| cities.json | ||
| cities_flavor.json | ||
| CLAUDE.md | ||
| docker-compose.test.yml | ||
| Dockerfile | ||
| drift_check.py | ||
| gen_cities.py | ||
| gen_flavor.py | ||
| indexnow.py | ||
| Makefile | ||
| migrate.py | ||
| migrate_accounts_to_pg.py | ||
| migrate_cache_to_pg.py | ||
| paths.py | ||
| README.md | ||
| requirements-dev.txt | ||
| requirements-seed.txt | ||
| requirements.txt | ||
| seed_era5.py | ||
| warm_cities.py | ||
thermograph-backend
The Thermograph API service: grades recent local weather against ~45 years of
climate history, and hosts the accounts, notification, and SSR-content
back-end that the rest of the split Thermograph stack (thermograph-frontend)
talks to over HTTP. Split from the emi/thermograph monorepo — this repo owns
the API/DB/accounts/notifications layer only; it renders no HTML/CSS/JS of its
own.
See CLAUDE.md for the full split topology, deploy flow, and
API version contract, and thermograph-docs (a sibling
repo) for cross-cutting architecture decisions and operator runbooks.
How it works
- Grid — a lat/lon is snapped to a stable ~4 sq mi cell
(
data/grid.py); longitude spacing is scaled bycos(latitude)so cells stay roughly square at any latitude. The cell id is the cache key, so the same spot always resolves to the same data. - Data (on-demand, cached to parquet) — the first request for a cell
fetches the full 1980–present daily record (max/min temp, precip) from
the free Open-Meteo archive (ERA5) and writes it
to
data/cache/<cell_id>.parquet(zstd, ~200 KB for 45 years). Later requests read the parquet directly. Recent days come from Open-Meteo's forecast API (past_days), so history and grading share one source. - Percentiles & grading (
data/grading.py) — each day is graded against every historical day within ±7 days of it (a 15-day seasonal window, wrapping year-end), as an empirical mid-rank percentile. Temperature uses a symmetric tier ladder (TEMP_BANDS: Near Record / High / Above Normal / Normal / Below Normal / Low / Near Record); precipitation is graded separately (RAIN_BANDS) since most days are dry — a rainy day is ranked only among rain days in its window, dry days are colored by dry-streak length instead. - Caching & ETags — every derived payload (grade/calendar/day/SSR
content) is cached in SQLite (
data/store.py) under(kind, cell_id, key), validated by a token that only advances when the cell's history actually changes. That token doubles as a weak ETag, so an unchanged request costs a 304 with no payload rebuild (api/payloads.py,web/app.py).
Layout
accounts/ fastapi-users models/schemas/db + api_accounts routes
alembic/ Postgres schema migrations (alembic upgrade head on boot)
api/ versioned payload builders (payloads.py, content_payloads.py)
+ route wiring (content_routes.py, sitemap.py, homepage.py)
core/ metrics, audit/access logging, a singleton helper
data/ grid snapping, climate fetch/cache, grading/scoring,
places/cities, the derived-payload store
notifications/ push (VAPID), email, Discord bot (interactions + linking),
monthly digest, the in-process scheduler (city warming, IndexNow)
web/app.py the FastAPI app (routes, CORS, ETag/versioning, middleware)
deploy/ container entrypoint.sh (alembic migrate, then serve)
app.py shim re-exporting web.app:app — keeps the launch target
`app:app` stable regardless of internal package layout
scripts/ one-off admin scripts (Discord slash-command registration)
tests/ pytest suite, hermetic (see tests/conftest.py)
cities.json,
cities_flavor.json bundled reference data (generated by gen_cities.py /
gen_flavor.py), not runtime state
How it fits the split
thermograph-frontendcalls this service'sGET /api/v2/...endpoints (grade, geocode, calendar) and the SSR content endpoints under/content/...; it negotiates compatibility viaGET /api/version.thermograph-infraowns the deploy/Compose/Terraform layer — this repo only builds and publishes its own container image (git.thermograph.org/emi/thermograph-backend/app) and hands infra a tag to roll out (SERVICE=backend+BACKEND_IMAGE_TAGinto infra'sdeploy/deploy.sh).thermograph-docsholds the cross-repo architecture/runbook docs; this README only covers what's local to this service.
Build & run
docker build -t thermograph-backend .
docker run -p 8137:8137 --env-file .env thermograph-backend
The image runs deploy/entrypoint.sh: alembic upgrade head against
THERMOGRAPH_DATABASE_URL (retried, since a fresh Postgres volume can still be
starting up), then uvicorn app:app on $PORT (default 8137) with
$WORKERS workers (default 4). /healthz is an I/O-free liveness probe.
For local development without a container, see the "Run / test locally"
section of CLAUDE.md — there is no Makefile in this repo yet,
so it's a plain venv + pytest/uvicorn invocation.
Notifications: Discord slash commands
notifications/discord_interactions.py answers Discord's HTTP Interactions
endpoint (/discord/interactions) for the /grade <city> slash command.
Registering (or updating) the command definition with Discord's REST API is a
one-off admin action, not part of the running app:
THERMOGRAPH_DISCORD_APP_ID=... THERMOGRAPH_DISCORD_BOT_TOKEN=... \
python3 scripts/register_discord_commands.py
Global command changes can take up to an hour to propagate.