thermograph/backend/requirements.txt
Emi Griffith 0f34594ce0
All checks were successful
PR build (required check) / changes (pull_request) Successful in 12s
secrets-guard / encrypted (pull_request) Successful in 11s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-backend (pull_request) Successful in 1m40s
PR build (required check) / gate (pull_request) Successful in 4s
ERA5 lake: bucket-hosted history primary + SQL indexer service
Move climate history to ERA5 served from our own object storage, with a
prod-only query service in front:

- data/era5lake.py: lake layout (per-point whole-record 1940+ serving files,
  a tile/year/month hive table, an Iceberg-style manifest) plus grid math and
  the read client (local dir -> lake service -> bucket).
- gen_era5_lake.py: tile-aligned extractor from the Earthmover Icechunk ERA5
  archive (one 86-year pull covers all 144 points of a chunk tile; resumable
  from the manifest; --cities / --tiles / --land).
- lake_app.py + THERMOGRAPH_ROLE=lake: the lake service. /history serves a
  point's parquet off a disk cache (11ms cold / 3ms warm in rehearsal);
  /query runs SELECT-only SQL on DuckDB over the hive table (79ms pruned
  aggregate, 88ms full scan of a 4.5M-row tile). httpfs is baked at image
  build.
- climate.py: history chain is now era5-lake -> NASA POWER -> Open-Meteo; the
  lake slice starts at START_DATE so grading windows are unchanged, and an
  unconfigured lake costs nothing (beta/LAN unchanged).
- stack: lake service (1..2 replicas behind the VIP, own cache volume) plus a
  second autoscaler instance (autoscale.sh gains TARGET_SERVICE); web/worker
  get THERMOGRAPH_LAKE_URL.
- seed_era5.py: constants corrected against the live store (icechunkV2,
  single/temporal, valid_time, ECMWF short names, pcodec).

Bucket creds (THERMOGRAPH_LAKE_S3_ACCESS_KEY/_SECRET_KEY) go in the prod
vault; until they land the lake stays healthy and everything falls through to
NASA exactly as before.
2026-07-23 14:17:50 -07:00

30 lines
1.4 KiB
Text

fastapi==0.115.6
uvicorn[standard]==0.34.0
httpx==0.28.1
polars==1.42.1
numpy==2.2.1
# Accounts + notification subscriptions (see accounts/db.py, users.py, notify.py).
# Postgres in prod/containers: asyncpg (async web/store) + psycopg (sync notifier);
# aiosqlite is kept for the SQLite fallback the test suite runs on. alembic manages
# the accounts schema. The `pool` extra pulls in psycopg_pool, used by
# data/climate_store.py + data/store.py to bound the raw connections those modules
# open on Postgres (see their module docstrings) instead of one-per-thread forever.
fastapi-users[sqlalchemy]==15.0.5
aiosqlite==0.22.1
asyncpg==0.30.0
psycopg[binary,pool]==3.2.3
alembic==1.14.0
# Web Push (VAPID) delivery of notifications (see notifications/push.py). Pulls in
# py-vapid, cryptography, and http-ece.
pywebpush==2.0.0
# Ed25519 verification of Discord interaction webhooks (see notifications/discord_interactions.py).
PyNaCl==1.5.0
# Discord gateway bot's websocket client (notifications/discord_bot.py). Already
# pulled in transitively by uvicorn[standard]; pinned here since we import it directly.
websockets==16.0
# Recurring worker-tier jobs — city warming, IndexNow pings (see
# notifications/scheduler.py). In-process; no broker/queue needed at this scale.
apscheduler==3.11.3
# The lake role's SQL engine (lake_app.py /query): embedded, parallel
# partition-pruned scans over the ERA5 hive table. Only lake_app imports it.
duckdb==1.4.2