All checks were successful
PR build (required check) / changes (pull_request) Successful in 9s
secrets-guard / encrypted (pull_request) Successful in 6s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-backend (pull_request) Successful in 44s
PR build (required check) / gate (pull_request) Successful in 2s
Registers the hive part files under era5/daily into an Apache Iceberg v2 table at iceberg/era5_daily via pyiceberg add_files -- the metadata points at the existing parquet in place, no data rewrite. Incremental: each run diffs era5/manifest.parquet (the completeness signal while extraction is running) against the thermograph.synced-tiles table property, updated in the same commit as the files, so runs are idempotent and resume at batch boundaries. Partitioned by truncate[12](lat_idx), truncate[12](lon_idx), year(date), month(date) -- order-preserving transforms add_files derives from footer stats; queries prune on plain column predicates. Local sqlite catalog that re-registers from the latest metadata JSON if lost; version-hint.text refreshed per run so DuckDB can iceberg_scan the bare table root. Per-batch retries with backoff for Contabo SLOW_DOWN and slow-transfer timeouts. Slim pinned one-shot image; local-filesystem pytest suite (no network).
5 lines
211 B
Text
5 lines
211 B
Text
# Pinned to the versions the sync was validated with (add_files partition
|
|
# inference from footer stats, sqlite catalog, path-style S3).
|
|
pyiceberg[sql-sqlite,pyarrow,s3fs]==0.11.1
|
|
pyarrow==25.0.0
|
|
s3fs==2026.6.0
|