thermograph/infra/deploy/db/README.md
Emi Griffith ae1d9bb534 Subtree-merge thermograph-infra (origin/main) into infra/
git-subtree-dir: infra
git-subtree-mainline: d6df04eab2
git-subtree-split: 99b4b3f78d
2026-07-22 22:01:11 -07:00

87 lines
3.7 KiB
Markdown

# Thermograph DB — TimescaleDB (PostgreSQL 18)
The `db` service runs the stock **`timescale/timescaledb:latest-pg18`** image
(TimescaleDB 2.24+, genuine PostgreSQL 18). The app's climate record lives here in
**hypertables** — the DB, not the filesystem, is the source of truth:
- **`climate_history`** — a hypertable of the full daily archive per grid cell
(`cell_id, date, tmax, tmin, precip, wind, gust, humid, fmax, fmin, feels`), back
to 1980. Durable/LOGGED (a 45-year, rate-limited refetch is expensive),
range-partitioned on `date` (5-year chunks), compressed for chunks older than a
year.
- **`climate_recent`** — the recent-observations + forward-forecast bundle (a plain
table: it holds future dates and is rewritten hourly).
- **`climate_sync`** — per-cell freshness (epoch seconds) that replaces the old
parquet file mtimes: it drives the hourly history top-up, the 1-hour forecast
TTL, and the `recent_stamp` token embedded in derived-payload validity.
The schema is created by Alembic (`backend/alembic/versions/0002_climate_hypertables.py`,
run at app boot via `deploy/entrypoint.sh`). The app reads/writes it through
`backend/data/climate_store.py` (psycopg + polars). See
`deploy/POSTGRES-MIGRATION.md` for the parquet→hypertable cutover.
## Why the stock image (no custom Dockerfile)
The previous DB image was a custom `pgduckdb/pgduckdb:18` build whose only purpose
was ad-hoc `read_parquet()` over the parquet cache. Now the climate record is in
real tables, so that capability is gone and the DB is the **stock TimescaleDB
image** — no build step. The image already sets
`shared_preload_libraries=timescaledb`; never `ALTER SYSTEM SET
shared_preload_libraries` (it would land in `postgresql.auto.conf` and override the
image's preload).
## Files here
- **`init/10-timescaledb.sql`** — `CREATE EXTENSION IF NOT EXISTS timescaledb;`
(runs from `/docker-entrypoint-initdb.d` on first cluster init; Alembic also does
this idempotently at boot).
- **`init/20-tuning.sh`** — scales `shared_buffers` (25%), `effective_cache_size`
(75%), `work_mem`, and `maintenance_work_mem` from `DB_MEMORY` via `ALTER SYSTEM`.
## Compose `db` service
```yaml
db:
image: timescale/timescaledb:latest-pg18
environment:
POSTGRES_USER: thermograph
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD}
POSTGRES_DB: thermograph
DB_MEMORY: ${DB_MEMORY:-8g}
volumes:
- pgdata:/var/lib/postgresql
- ./deploy/db/init:/docker-entrypoint-initdb.d
# healthcheck / cpus / mem_limit / shm_size: unchanged
```
Notes:
- The volume is mounted at the **parent** of the data dir; the image picks its own
PGDATA subdir (`/var/lib/postgresql/data`) under it. The whole tree persists on
the named volume.
- No host port on purpose: the app reaches Postgres as `db:5432` on the compose
network. Nothing outside the stack should touch the database.
- Init scripts only run when PGDATA is empty. On an **existing** database, enable
the extension once by hand:
```
docker compose exec db psql -U thermograph -d thermograph \
-c 'CREATE EXTENSION IF NOT EXISTS timescaledb;'
```
## Inspecting the hypertable
```sql
-- Chunk / compression overview
SELECT hypertable_name, num_chunks, compression_enabled
FROM timescaledb_information.hypertables;
-- One cell, most recent archived days
SELECT date, tmax, tmin, precip
FROM climate_history WHERE cell_id = '1026_-2857'
ORDER BY date DESC LIMIT 5;
-- Per-year highs for one cell
SELECT EXTRACT(YEAR FROM date) AS yr,
ROUND(AVG(tmax)::numeric, 1) AS avg_tmax, MAX(tmax) AS record_high
FROM climate_history WHERE cell_id = '1026_-2857' AND date >= '2020-01-01'
GROUP BY yr ORDER BY yr;
```