thermograph/deploy/db/README.md
Emi Griffith 6fd2d7c981 Containerize the app and move the databases to PostgreSQL 18 (#220)
Run Thermograph as a docker-compose stack (app + Postgres 18) and standardize the
data layer on Postgres, while keeping the test suite on SQLite.

- accounts/db.py: DSN-driven engines. On Postgres, a per-worker read-write +
  read-only asyncpg pair (the RO engine pins read-only transactions, used by the
  pure-GET endpoints) plus a sync psycopg engine for the notifier thread; the
  SQLite path is preserved for tests/local (selected when THERMOGRAPH_DATABASE_URL
  is unset). models.py: boolean server_default -> sa.false().
- store.py / metrics.py: dialect-flexible — Postgres UNLOGGED tables via psycopg
  when configured, else the existing raw-sqlite3 paths byte-for-byte; sync
  interfaces and every fail-soft contract preserved.
- Alembic (backend/alembic/) manages the accounts schema; the container entrypoint
  runs `alembic upgrade head` before uvicorn (4 workers). migrate_accounts_to_pg.py
  copies the accounts data SQLite->PG through the ORM (UUID/bool/JSON coerced),
  skips access_token, and resets identity sequences.
- Dockerfile + docker-compose.yml: app image (uvicorn, 4 workers, loopback 8137)
  and a Postgres 18 db (2 CPUs) running pg_duckdb (deploy/db/) so the parquet
  climate cache is queryable in-DB via read_parquet('/parquet/cache/*.parquet').
- deploy.sh/thermograph.service rewired to manage the compose stack; env example,
  Makefile targets (up/down/db-up), and deploy/POSTGRES-MIGRATION.md cutover runbook.

Tests stay on SQLite (dialect fallback) — 323 pass. The full Postgres stack was
verified via docker compose: alembic migrations, register/login, the RO endpoint,
store/metrics round-trips, and the accounts data migration.
2026-07-20 06:28:23 +00:00

171 lines
7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Thermograph DB image — Postgres 18 + parquet reads
The `db` service can read the app's Parquet climate cache
(`data/cache/*.parquet`) directly from SQL, for ad-hoc analytics, via
**pg_duckdb** — DuckDB embedded inside Postgres.
```sql
SELECT r['date'] AS date, r['tmax'] AS tmax, r['tmin'] AS tmin
FROM read_parquet('/parquet/1026_-2857.parquet') r
ORDER BY r['date'] LIMIT 5;
```
## Chosen extension: pg_duckdb — and why
| Option | PG18? | Fit | Verdict |
| --- | --- | --- | --- |
| **pg_duckdb** (duckdb / MotherDuck) | **Yes** — official image `pgduckdb/pgduckdb:18-v1.1.1` is genuine PG18.1 | `read_parquet('…')` in plain SQL; globs, `union_by_name`, full DuckDB analytics engine | **Chosen** |
| pg_parquet (Crunchy Data) | Yes (1418) | `COPY … TO/FROM '…' (format 'parquet')` — import/export, not a query engine | Viable, but COPY-oriented; no standalone official image (ships via Crunchy Bridge/CPK), so it'd need a Rust/pgrx source build |
| parquet_fdw | No prebuilt PG18 support; low activity | Foreign tables over parquet | Rejected — oldest, weakest PG18 story |
pg_duckdb wins for the stated goal (ad-hoc analytics): it exposes DuckDB's
`read_parquet` directly in SQL, so you query cache files like tables — no
import step, no foreign-table DDL, and you get aggregation/joins/globs across
many cells at once.
## PG-version reality for PG18 (empirically verified 2026-07-19)
**No version delta.** PG18 support is real, not a fallback. The pulled image
reports:
```
PostgreSQL 18.1 (Debian 18.1-1.pgdg12+2) on x86_64-pc-linux-gnu
pg_extension: pg_duckdb 1.1.0
shared_preload_libraries: pg_duckdb
```
The image is built on Debian 12 bookworm — the **same base as the official
`postgres:18` image** — and uses the standard `docker-entrypoint.sh`. So
`POSTGRES_USER` / `POSTGRES_PASSWORD` / `POSTGRES_DB` / `PGDATA` /
`/docker-entrypoint-initdb.d` / `pg_isready` all behave exactly as with
`postgres:18`. It is a drop-in replacement for the `db` service; nothing else
in the stack changes.
We `FROM` the official pg_duckdb image (pinned to `18-v1.1.1`, not `18-main`)
rather than `FROM postgres:18` + compile, because pg_duckdb links a full DuckDB
build — compiling from source in the Dockerfile means the DuckDB toolchain and
a long, fragile build for no benefit over the maintainers' official PG18 image.
## Files here
- **`Dockerfile.db`** — `FROM pgduckdb/pgduckdb:18-v1.1.1`, plus `COPY` of the
init script so the image enables the extension on first init even without the
compose bind mount.
- **`init/10-parquet.sql`** — `CREATE EXTENSION IF NOT EXISTS pg_duckdb;` (runs
from `/docker-entrypoint-initdb.d` on first cluster init).
## Compose snippet to merge into the `db` service
Replace `image: postgres:18` with the `build:` block; add the read-only parquet
bind and the init mount. Everything else in the `db` service stays as-is.
```yaml
db:
# image: postgres:18 # <- remove; build the parquet-capable image
build:
context: .
dockerfile: deploy/db/Dockerfile.db
environment:
POSTGRES_USER: thermograph
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD}
POSTGRES_DB: thermograph
PGDATA: /var/lib/postgresql/data
volumes:
- pgdata:/var/lib/postgresql/data
- ./data/cache:/parquet:ro # read-only parquet cache
- ./deploy/db/init:/docker-entrypoint-initdb.d
# healthcheck / cpus / deploy / restart: unchanged
```
Notes:
- The bind mount at `/docker-entrypoint-initdb.d` **replaces** the base image's
own init scripts (its `0001-install-pg_duckdb.sql` and the MotherDuck-only
`0002-enable-md-pg_duckdb.sql`). That's intended: our `10-parquet.sql` still
runs `CREATE EXTENSION`, and we don't use MotherDuck. If you prefer to keep
the image's baked scripts, drop the `:/docker-entrypoint-initdb.d` line — the
`Dockerfile` already bakes `10-parquet.sql` in.
- `:ro` keeps the DB from ever mutating the app's cache. In prod the app writes
the cache to the `appdata` volume; point this bind at wherever that lives on
the host if you want the DB to see the live cache rather than the repo copy.
- Init scripts only run when `PGDATA` is empty. On an **existing** database,
enable it once by hand:
```
docker compose exec db psql -U thermograph -d thermograph \
-c 'CREATE EXTENSION IF NOT EXISTS pg_duckdb;'
```
## Usage — real cache file at `/parquet/…`
Cache columns: `date, tmax, tmin, precip, wind, gust, humid, fmax, fmin, feels`.
pg_duckdb ≥ 0.3 uses the `r['colname']` subscript syntax (not
`AS t(col type, …)`).
```sql
-- one cell, first rows
SELECT r['date'] AS date, r['tmax'] AS tmax, r['tmin'] AS tmin, r['precip'] AS precip
FROM read_parquet('/parquet/1026_-2857.parquet') r
ORDER BY r['date'] LIMIT 5;
-- per-year analytics over one cell
SELECT EXTRACT(YEAR FROM r['date']::timestamp) AS yr,
ROUND(AVG(r['tmax'])::numeric, 1) AS avg_tmax,
MAX(r['tmax']) AS record_high
FROM read_parquet('/parquet/1026_-2857.parquet') r
WHERE r['date'] >= '2020-01-01'
GROUP BY yr ORDER BY yr;
-- glob across every cached cell (union_by_name handles the _rf / _forecast
-- files whose column sets differ)
SELECT COUNT(*) FROM read_parquet('/parquet/*.parquet', union_by_name := true) r;
```
## Proof of work (verified 2026-07-19)
Built `deploy/db/Dockerfile.db`, ran a throwaway container with the real
`data/cache` bind-mounted read-only at `/parquet`, then:
```
$ psql -c "SELECT version();"
PostgreSQL 18.1 (Debian 18.1-1.pgdg12+2) on x86_64-pc-linux-gnu ...
$ psql -c "SELECT extname, extversion FROM pg_extension WHERE extname='pg_duckdb';"
extname | extversion
-----------+------------
pg_duckdb | 1.1.0
$ psql -c "SELECT COUNT(*) FROM read_parquet('/parquet/1026_-2857.parquet') r;"
row_count
-----------
16982
$ psql -c "SELECT r['date'] AS date, r['tmax'] AS tmax, r['tmin'] AS tmin,
r['precip'] AS precip, r['humid'] AS humid
FROM read_parquet('/parquet/1026_-2857.parquet') r
ORDER BY r['date'] LIMIT 5;"
date | tmax | tmin | precip | humid
---------------------+------+------+--------+-------
1980-01-01 00:00:00 | 56.2 | 34.7 | 0 | 73
1980-01-02 00:00:00 | 63.8 | 38 | 0 | 79
1980-01-03 00:00:00 | 60.1 | 46.1 | 0.315 | 83
1980-01-04 00:00:00 | 51.7 | 40 | 0 | 68
1980-01-05 00:00:00 | 56.5 | 33.7 | 0 | 78
$ psql -c "SELECT EXTRACT(YEAR FROM r['date']::timestamp) AS yr,
ROUND(AVG(r['tmax'])::numeric,1) AS avg_tmax, MAX(r['tmax']) AS record_high
FROM read_parquet('/parquet/1026_-2857.parquet') r
WHERE r['date'] >= '2020-01-01' GROUP BY yr ORDER BY yr;"
yr | avg_tmax | record_high
------+----------+-------------
2020 | 78.8 | 99
2021 | 77.0 | 93
2022 | 79.2 | 101.4
2023 | 80.8 | 107.1
2024 | 79.8 | 97.5
2025 | 80.0 | 99.4
2026 | 77.8 | 95.6
$ psql -c "SELECT COUNT(*) FROM read_parquet('/parquet/*.parquet', union_by_name := true) r;"
count
--------
525421 -- rows across all 45 cached cells
```