All checks were successful
PR build (required check) / changes (pull_request) Successful in 12s
secrets-guard / encrypted (pull_request) Successful in 11s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-backend (pull_request) Successful in 1m40s
PR build (required check) / gate (pull_request) Successful in 4s
Move climate history to ERA5 served from our own object storage, with a prod-only query service in front: - data/era5lake.py: lake layout (per-point whole-record 1940+ serving files, a tile/year/month hive table, an Iceberg-style manifest) plus grid math and the read client (local dir -> lake service -> bucket). - gen_era5_lake.py: tile-aligned extractor from the Earthmover Icechunk ERA5 archive (one 86-year pull covers all 144 points of a chunk tile; resumable from the manifest; --cities / --tiles / --land). - lake_app.py + THERMOGRAPH_ROLE=lake: the lake service. /history serves a point's parquet off a disk cache (11ms cold / 3ms warm in rehearsal); /query runs SELECT-only SQL on DuckDB over the hive table (79ms pruned aggregate, 88ms full scan of a 4.5M-row tile). httpfs is baked at image build. - climate.py: history chain is now era5-lake -> NASA POWER -> Open-Meteo; the lake slice starts at START_DATE so grading windows are unchanged, and an unconfigured lake costs nothing (beta/LAN unchanged). - stack: lake service (1..2 replicas behind the VIP, own cache volume) plus a second autoscaler instance (autoscale.sh gains TARGET_SERVICE); web/worker get THERMOGRAPH_LAKE_URL. - seed_era5.py: constants corrected against the live store (icechunkV2, single/temporal, valid_time, ECMWF short names, pcodec). Bucket creds (THERMOGRAPH_LAKE_S3_ACCESS_KEY/_SECRET_KEY) go in the prod vault; until they land the lake stays healthy and everything falls through to NASA exactly as before.
16 lines
676 B
Text
16 lines
676 B
Text
# Seed-only dependencies for seed_era5.py (the one-time ERA5 backfill). NOT installed
|
|
# in the app image or CI — kept out of requirements.txt so the runtime stays lean.
|
|
# Install on the machine that runs the seed: pip install -r requirements-seed.txt
|
|
# Versions intentionally unpinned: the icechunk/xarray/zarr stack moves fast; install
|
|
# the current compatible set on the seed box and verify with `python seed_era5.py --dry-run`.
|
|
-r requirements.txt
|
|
icechunk
|
|
xarray
|
|
zarr
|
|
# The live store compresses with PCodec; without this extra every read fails
|
|
# with "codec not available: 'pcodec'".
|
|
numcodecs[pcodec]
|
|
# gen_era5_lake.py's bucket writer.
|
|
boto3
|
|
numcodecs[pcodec]
|
|
boto3
|