thermograph/deploy/db
Emi Griffith dc4ee9a8db Scale DB tuning from DB_MEMORY; order Docker after the rclone mount (#226)
Two prod-readiness hardening changes:

DB tuning scales with the container budget. Replace the fixed 8 GB
20-tuning.sql with 20-tuning.sh, which derives shared_buffers (25%),
effective_cache_size (75%), work_mem, maintenance_work_mem and
duckdb.max_memory (50%) from the DB_MEMORY the compose db service now passes in.
The ratios reproduce the historical 8 GB tuning exactly and scale linearly, so
the 48 GB prod box (db_memory 16g) gets shared_buffers 4 GB / duckdb 8 GB with
no separate edit. Beta/local (8g default) are unchanged. Docs that told
operators to raise the tuning by hand are updated.

Boot ordering for the self-hosted archive. On an openmeteo host, install a
docker.service drop-in (Wants/After rclone-om.service) so Docker starts after
the object-storage mount is ready on every boot — the restart-policy containers
never bind an empty mount point. rclone-om is Type=notify, so After waits for
the mount to actually be ready. Cleaned up when openmeteo is toggled off.
2026-07-20 14:33:09 +00:00
..
init Scale DB tuning from DB_MEMORY; order Docker after the rclone mount (#226) 2026-07-20 14:33:09 +00:00
Dockerfile.db Containerize the app and move the databases to PostgreSQL 18 (#220) 2026-07-20 06:28:23 +00:00
README.md Containerize the app and move the databases to PostgreSQL 18 (#220) 2026-07-20 06:28:23 +00:00

Thermograph DB image — Postgres 18 + parquet reads

The db service can read the app's Parquet climate cache (data/cache/*.parquet) directly from SQL, for ad-hoc analytics, via pg_duckdb — DuckDB embedded inside Postgres.

SELECT r['date'] AS date, r['tmax'] AS tmax, r['tmin'] AS tmin
FROM read_parquet('/parquet/1026_-2857.parquet') r
ORDER BY r['date'] LIMIT 5;

Chosen extension: pg_duckdb — and why

Option PG18? Fit Verdict
pg_duckdb (duckdb / MotherDuck) Yes — official image pgduckdb/pgduckdb:18-v1.1.1 is genuine PG18.1 read_parquet('…') in plain SQL; globs, union_by_name, full DuckDB analytics engine Chosen
pg_parquet (Crunchy Data) Yes (1418) COPY … TO/FROM '…' (format 'parquet') — import/export, not a query engine Viable, but COPY-oriented; no standalone official image (ships via Crunchy Bridge/CPK), so it'd need a Rust/pgrx source build
parquet_fdw No prebuilt PG18 support; low activity Foreign tables over parquet Rejected — oldest, weakest PG18 story

pg_duckdb wins for the stated goal (ad-hoc analytics): it exposes DuckDB's read_parquet directly in SQL, so you query cache files like tables — no import step, no foreign-table DDL, and you get aggregation/joins/globs across many cells at once.

PG-version reality for PG18 (empirically verified 2026-07-19)

No version delta. PG18 support is real, not a fallback. The pulled image reports:

PostgreSQL 18.1 (Debian 18.1-1.pgdg12+2) on x86_64-pc-linux-gnu
pg_extension: pg_duckdb 1.1.0
shared_preload_libraries: pg_duckdb

The image is built on Debian 12 bookworm — the same base as the official postgres:18 image — and uses the standard docker-entrypoint.sh. So POSTGRES_USER / POSTGRES_PASSWORD / POSTGRES_DB / PGDATA / /docker-entrypoint-initdb.d / pg_isready all behave exactly as with postgres:18. It is a drop-in replacement for the db service; nothing else in the stack changes.

We FROM the official pg_duckdb image (pinned to 18-v1.1.1, not 18-main) rather than FROM postgres:18 + compile, because pg_duckdb links a full DuckDB build — compiling from source in the Dockerfile means the DuckDB toolchain and a long, fragile build for no benefit over the maintainers' official PG18 image.

Files here

  • Dockerfile.dbFROM pgduckdb/pgduckdb:18-v1.1.1, plus COPY of the init script so the image enables the extension on first init even without the compose bind mount.
  • init/10-parquet.sqlCREATE EXTENSION IF NOT EXISTS pg_duckdb; (runs from /docker-entrypoint-initdb.d on first cluster init).

Compose snippet to merge into the db service

Replace image: postgres:18 with the build: block; add the read-only parquet bind and the init mount. Everything else in the db service stays as-is.

  db:
    # image: postgres:18            # <- remove; build the parquet-capable image
    build:
      context: .
      dockerfile: deploy/db/Dockerfile.db
    environment:
      POSTGRES_USER: thermograph
      POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD}
      POSTGRES_DB: thermograph
      PGDATA: /var/lib/postgresql/data
    volumes:
      - pgdata:/var/lib/postgresql/data
      - ./data/cache:/parquet:ro                 # read-only parquet cache
      - ./deploy/db/init:/docker-entrypoint-initdb.d
    # healthcheck / cpus / deploy / restart: unchanged

Notes:

  • The bind mount at /docker-entrypoint-initdb.d replaces the base image's own init scripts (its 0001-install-pg_duckdb.sql and the MotherDuck-only 0002-enable-md-pg_duckdb.sql). That's intended: our 10-parquet.sql still runs CREATE EXTENSION, and we don't use MotherDuck. If you prefer to keep the image's baked scripts, drop the :/docker-entrypoint-initdb.d line — the Dockerfile already bakes 10-parquet.sql in.
  • :ro keeps the DB from ever mutating the app's cache. In prod the app writes the cache to the appdata volume; point this bind at wherever that lives on the host if you want the DB to see the live cache rather than the repo copy.
  • Init scripts only run when PGDATA is empty. On an existing database, enable it once by hand:
    docker compose exec db psql -U thermograph -d thermograph \
      -c 'CREATE EXTENSION IF NOT EXISTS pg_duckdb;'
    

Usage — real cache file at /parquet/…

Cache columns: date, tmax, tmin, precip, wind, gust, humid, fmax, fmin, feels. pg_duckdb ≥ 0.3 uses the r['colname'] subscript syntax (not AS t(col type, …)).

-- one cell, first rows
SELECT r['date'] AS date, r['tmax'] AS tmax, r['tmin'] AS tmin, r['precip'] AS precip
FROM read_parquet('/parquet/1026_-2857.parquet') r
ORDER BY r['date'] LIMIT 5;

-- per-year analytics over one cell
SELECT EXTRACT(YEAR FROM r['date']::timestamp) AS yr,
       ROUND(AVG(r['tmax'])::numeric, 1) AS avg_tmax,
       MAX(r['tmax']) AS record_high
FROM read_parquet('/parquet/1026_-2857.parquet') r
WHERE r['date'] >= '2020-01-01'
GROUP BY yr ORDER BY yr;

-- glob across every cached cell (union_by_name handles the _rf / _forecast
-- files whose column sets differ)
SELECT COUNT(*) FROM read_parquet('/parquet/*.parquet', union_by_name := true) r;

Proof of work (verified 2026-07-19)

Built deploy/db/Dockerfile.db, ran a throwaway container with the real data/cache bind-mounted read-only at /parquet, then:

$ psql -c "SELECT version();"
 PostgreSQL 18.1 (Debian 18.1-1.pgdg12+2) on x86_64-pc-linux-gnu ...

$ psql -c "SELECT extname, extversion FROM pg_extension WHERE extname='pg_duckdb';"
  extname  | extversion
-----------+------------
 pg_duckdb | 1.1.0

$ psql -c "SELECT COUNT(*) FROM read_parquet('/parquet/1026_-2857.parquet') r;"
 row_count
-----------
     16982

$ psql -c "SELECT r['date'] AS date, r['tmax'] AS tmax, r['tmin'] AS tmin,
           r['precip'] AS precip, r['humid'] AS humid
           FROM read_parquet('/parquet/1026_-2857.parquet') r
           ORDER BY r['date'] LIMIT 5;"
        date         | tmax | tmin | precip | humid
---------------------+------+------+--------+-------
 1980-01-01 00:00:00 | 56.2 | 34.7 |      0 |    73
 1980-01-02 00:00:00 | 63.8 |   38 |      0 |    79
 1980-01-03 00:00:00 | 60.1 | 46.1 |  0.315 |    83
 1980-01-04 00:00:00 | 51.7 |   40 |      0 |    68
 1980-01-05 00:00:00 | 56.5 | 33.7 |      0 |    78

$ psql -c "SELECT EXTRACT(YEAR FROM r['date']::timestamp) AS yr,
           ROUND(AVG(r['tmax'])::numeric,1) AS avg_tmax, MAX(r['tmax']) AS record_high
           FROM read_parquet('/parquet/1026_-2857.parquet') r
           WHERE r['date'] >= '2020-01-01' GROUP BY yr ORDER BY yr;"
  yr  | avg_tmax | record_high
------+----------+-------------
 2020 |     78.8 |          99
 2021 |     77.0 |          93
 2022 |     79.2 |       101.4
 2023 |     80.8 |       107.1
 2024 |     79.8 |        97.5
 2025 |     80.0 |        99.4
 2026 |     77.8 |        95.6

$ psql -c "SELECT COUNT(*) FROM read_parquet('/parquet/*.parquet', union_by_name := true) r;"
 count
--------
 525421   -- rows across all 45 cached cells