|
All checks were successful
PR build (required check) / changes (pull_request) Successful in 7s
secrets-guard / encrypted (pull_request) Successful in 6s
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-backend (pull_request) Successful in 1m9s
PR build (required check) / gate (pull_request) Successful in 2s
geocode_nominatim referenced _REVGEO_LOCK, which was removed when reverse geocoding moved from a bare lock to the dedicated worker/queue design -- the forward-geocode function was never updated to match, so every call raised NameError, degrading to a 502 on the frontend. Live on prod, beta, dev, main, and release since the lock was removed (~30h): every comma-qualified name, postcode, and non-cities1000 place 502'd, while the local GeoNames index kept bare city names working, so the failure was invisible to simple smoke checks. The three existing tests covering this path all monkeypatch geocode_nominatim itself, so none of them ever executed the broken body. Fixed by routing forward-geocode jobs through the same worker/queue that already serializes reverse-geocode jobs, instead of reintroducing a second lock -- both job kinds are now drained by the one worker thread, so _revgeo_last still has exactly one writer and the shared ~1/sec Nominatim pacing the docstring always claimed is now actually enforced, not just asserted in a comment. Adds regression coverage that exercises the real queue/worker plumbing (not a monkeypatch of geocode_nominatim) plus a stubbed-HTTP test of _fetch_geocode_forward's own body, the function whose earlier version never once executed successfully. |
||
|---|---|---|
| .. | ||
| accounts | ||
| alembic | ||
| api | ||
| core | ||
| daemon | ||
| data | ||
| deploy | ||
| notifications | ||
| scripts | ||
| tests | ||
| web | ||
| .dockerignore | ||
| .gitignore | ||
| alembic.ini | ||
| app.py | ||
| cities.json | ||
| cities_flavor.json | ||
| CLAUDE.md | ||
| docker-compose.test.yml | ||
| Dockerfile | ||
| drift_check.py | ||
| gen_cities.py | ||
| gen_era5_lake.py | ||
| gen_flavor.py | ||
| indexnow.py | ||
| lake_app.py | ||
| Makefile | ||
| migrate.py | ||
| migrate_accounts_to_pg.py | ||
| migrate_cache_to_pg.py | ||
| paths.py | ||
| README.md | ||
| requirements-dev.txt | ||
| requirements-seed.txt | ||
| requirements.txt | ||
| seed_era5.py | ||
| warm_cities.py | ||
thermograph-backend
The Thermograph API service: grades recent local weather against ~45 years of
climate history, and hosts the accounts, notification, and SSR-content
back-end that the rest of the split Thermograph stack (thermograph-frontend)
talks to over HTTP. Split from the emi/thermograph monorepo — this repo owns
the API/DB/accounts/notifications layer only; it renders no HTML/CSS/JS of its
own.
See CLAUDE.md for the full split topology, deploy flow, and
API version contract, and thermograph-docs (a sibling
repo) for cross-cutting architecture decisions and operator runbooks.
How it works
- Grid — a lat/lon is snapped to a stable ~4 sq mi cell
(
data/grid.py); longitude spacing is scaled bycos(latitude)so cells stay roughly square at any latitude. The cell id is the cache key, so the same spot always resolves to the same data. - Data (on-demand, cached to parquet) — the first request for a cell
fetches the full 1980–present daily record (max/min temp, precip) from
the free Open-Meteo archive (ERA5) and writes it
to
data/cache/<cell_id>.parquet(zstd, ~200 KB for 45 years). Later requests read the parquet directly. Recent days come from Open-Meteo's forecast API (past_days), so history and grading share one source. - Percentiles & grading (
data/grading.py) — each day is graded against every historical day within ±7 days of it (a 15-day seasonal window, wrapping year-end), as an empirical mid-rank percentile. Temperature uses a symmetric tier ladder (TEMP_BANDS: Near Record / High / Above Normal / Normal / Below Normal / Low / Near Record); precipitation is graded separately (RAIN_BANDS) since most days are dry — a rainy day is ranked only among rain days in its window, dry days are colored by dry-streak length instead. - Caching & ETags — every derived payload (grade/calendar/day/SSR
content) is cached in SQLite (
data/store.py) under(kind, cell_id, key), validated by a token that only advances when the cell's history actually changes. That token doubles as a weak ETag, so an unchanged request costs a 304 with no payload rebuild (api/payloads.py,web/app.py).
Layout
accounts/ fastapi-users models/schemas/db + api_accounts routes
alembic/ Postgres schema migrations (alembic upgrade head on boot)
api/ versioned payload builders (payloads.py, content_payloads.py)
+ route wiring (content_routes.py, sitemap.py, homepage.py)
core/ metrics, audit/access logging, a singleton helper
data/ grid snapping, climate fetch/cache, grading/scoring,
places/cities, the derived-payload store
notifications/ push (VAPID), email, Discord bot (interactions + linking),
monthly digest, the in-process scheduler (city warming, IndexNow)
web/app.py the FastAPI app (routes, CORS, ETag/versioning, middleware)
deploy/ container entrypoint.sh (alembic migrate, then serve)
app.py shim re-exporting web.app:app — keeps the launch target
`app:app` stable regardless of internal package layout
scripts/ one-off admin scripts (Discord slash-command registration)
tests/ pytest suite, hermetic (see tests/conftest.py)
cities.json,
cities_flavor.json bundled reference data (generated by gen_cities.py /
gen_flavor.py), not runtime state
How it fits the split
thermograph-frontendcalls this service'sGET /api/v2/...endpoints (grade, geocode, calendar) and the SSR content endpoints under/content/...; it negotiates compatibility viaGET /api/version.thermograph-infraowns the deploy/Compose/Terraform layer — this repo only builds and publishes its own container image (git.thermograph.org/emi/thermograph-backend/app) and hands infra a tag to roll out (SERVICE=backend+BACKEND_IMAGE_TAGinto infra'sdeploy/deploy.sh).thermograph-docsholds the cross-repo architecture/runbook docs; this README only covers what's local to this service.
Build & run
docker build -t thermograph-backend .
docker run -p 8137:8137 --env-file .env thermograph-backend
The image runs deploy/entrypoint.sh: alembic upgrade head against
THERMOGRAPH_DATABASE_URL (retried, since a fresh Postgres volume can still be
starting up), then uvicorn app:app on $PORT (default 8137) with
$WORKERS workers (default 4). /healthz is an I/O-free liveness probe.
For local development without a container, see the "Run / test locally"
section of CLAUDE.md — there is no Makefile in this repo yet,
so it's a plain venv + pytest/uvicorn invocation.
Notifications: Discord slash commands
notifications/discord_interactions.py answers Discord's HTTP Interactions
endpoint (/discord/interactions) for the /grade <city> slash command.
Registering (or updating) the command definition with Discord's REST API is a
one-off admin action, not part of the running app:
THERMOGRAPH_DISCORD_APP_ID=... THERMOGRAPH_DISCORD_BOT_TOKEN=... \
python3 scripts/register_discord_commands.py
Global command changes can take up to an hour to propagate.