thermograph/infra/deploy/db/README.md
Emi Griffith ae1d9bb534 Subtree-merge thermograph-infra (origin/main) into infra/
git-subtree-dir: infra
git-subtree-mainline: d6df04eab2
git-subtree-split: 99b4b3f78d
2026-07-22 22:01:11 -07:00

3.7 KiB

Thermograph DB — TimescaleDB (PostgreSQL 18)

The db service runs the stock timescale/timescaledb:latest-pg18 image (TimescaleDB 2.24+, genuine PostgreSQL 18). The app's climate record lives here in hypertables — the DB, not the filesystem, is the source of truth:

  • climate_history — a hypertable of the full daily archive per grid cell (cell_id, date, tmax, tmin, precip, wind, gust, humid, fmax, fmin, feels), back to 1980. Durable/LOGGED (a 45-year, rate-limited refetch is expensive), range-partitioned on date (5-year chunks), compressed for chunks older than a year.
  • climate_recent — the recent-observations + forward-forecast bundle (a plain table: it holds future dates and is rewritten hourly).
  • climate_sync — per-cell freshness (epoch seconds) that replaces the old parquet file mtimes: it drives the hourly history top-up, the 1-hour forecast TTL, and the recent_stamp token embedded in derived-payload validity.

The schema is created by Alembic (backend/alembic/versions/0002_climate_hypertables.py, run at app boot via deploy/entrypoint.sh). The app reads/writes it through backend/data/climate_store.py (psycopg + polars). See deploy/POSTGRES-MIGRATION.md for the parquet→hypertable cutover.

Why the stock image (no custom Dockerfile)

The previous DB image was a custom pgduckdb/pgduckdb:18 build whose only purpose was ad-hoc read_parquet() over the parquet cache. Now the climate record is in real tables, so that capability is gone and the DB is the stock TimescaleDB image — no build step. The image already sets shared_preload_libraries=timescaledb; never ALTER SYSTEM SET shared_preload_libraries (it would land in postgresql.auto.conf and override the image's preload).

Files here

  • init/10-timescaledb.sqlCREATE EXTENSION IF NOT EXISTS timescaledb; (runs from /docker-entrypoint-initdb.d on first cluster init; Alembic also does this idempotently at boot).
  • init/20-tuning.sh — scales shared_buffers (25%), effective_cache_size (75%), work_mem, and maintenance_work_mem from DB_MEMORY via ALTER SYSTEM.

Compose db service

  db:
    image: timescale/timescaledb:latest-pg18
    environment:
      POSTGRES_USER: thermograph
      POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD}
      POSTGRES_DB: thermograph
      DB_MEMORY: ${DB_MEMORY:-8g}
    volumes:
      - pgdata:/var/lib/postgresql
      - ./deploy/db/init:/docker-entrypoint-initdb.d
    # healthcheck / cpus / mem_limit / shm_size: unchanged

Notes:

  • The volume is mounted at the parent of the data dir; the image picks its own PGDATA subdir (/var/lib/postgresql/data) under it. The whole tree persists on the named volume.
  • No host port on purpose: the app reaches Postgres as db:5432 on the compose network. Nothing outside the stack should touch the database.
  • Init scripts only run when PGDATA is empty. On an existing database, enable the extension once by hand:
    docker compose exec db psql -U thermograph -d thermograph \
      -c 'CREATE EXTENSION IF NOT EXISTS timescaledb;'
    

Inspecting the hypertable

-- Chunk / compression overview
SELECT hypertable_name, num_chunks, compression_enabled
FROM timescaledb_information.hypertables;

-- One cell, most recent archived days
SELECT date, tmax, tmin, precip
FROM climate_history WHERE cell_id = '1026_-2857'
ORDER BY date DESC LIMIT 5;

-- Per-year highs for one cell
SELECT EXTRACT(YEAR FROM date) AS yr,
       ROUND(AVG(tmax)::numeric, 1) AS avg_tmax, MAX(tmax) AS record_high
FROM climate_history WHERE cell_id = '1026_-2857' AND date >= '2020-01-01'
GROUP BY yr ORDER BY yr;