thermograph/deploy/db/README.md

172 lines
7 KiB
Markdown
Raw Normal View History

Containerize the app and move the databases to PostgreSQL 18 (#220) Run Thermograph as a docker-compose stack (app + Postgres 18) and standardize the data layer on Postgres, while keeping the test suite on SQLite. - accounts/db.py: DSN-driven engines. On Postgres, a per-worker read-write + read-only asyncpg pair (the RO engine pins read-only transactions, used by the pure-GET endpoints) plus a sync psycopg engine for the notifier thread; the SQLite path is preserved for tests/local (selected when THERMOGRAPH_DATABASE_URL is unset). models.py: boolean server_default -> sa.false(). - store.py / metrics.py: dialect-flexible — Postgres UNLOGGED tables via psycopg when configured, else the existing raw-sqlite3 paths byte-for-byte; sync interfaces and every fail-soft contract preserved. - Alembic (backend/alembic/) manages the accounts schema; the container entrypoint runs `alembic upgrade head` before uvicorn (4 workers). migrate_accounts_to_pg.py copies the accounts data SQLite->PG through the ORM (UUID/bool/JSON coerced), skips access_token, and resets identity sequences. - Dockerfile + docker-compose.yml: app image (uvicorn, 4 workers, loopback 8137) and a Postgres 18 db (2 CPUs) running pg_duckdb (deploy/db/) so the parquet climate cache is queryable in-DB via read_parquet('/parquet/cache/*.parquet'). - deploy.sh/thermograph.service rewired to manage the compose stack; env example, Makefile targets (up/down/db-up), and deploy/POSTGRES-MIGRATION.md cutover runbook. Tests stay on SQLite (dialect fallback) — 323 pass. The full Postgres stack was verified via docker compose: alembic migrations, register/login, the RO endpoint, store/metrics round-trips, and the accounts data migration.
2026-07-20 06:28:23 +00:00
# Thermograph DB image — Postgres 18 + parquet reads
The `db` service can read the app's Parquet climate cache
(`data/cache/*.parquet`) directly from SQL, for ad-hoc analytics, via
**pg_duckdb** — DuckDB embedded inside Postgres.
```sql
SELECT r['date'] AS date, r['tmax'] AS tmax, r['tmin'] AS tmin
FROM read_parquet('/parquet/1026_-2857.parquet') r
ORDER BY r['date'] LIMIT 5;
```
## Chosen extension: pg_duckdb — and why
| Option | PG18? | Fit | Verdict |
| --- | --- | --- | --- |
| **pg_duckdb** (duckdb / MotherDuck) | **Yes** — official image `pgduckdb/pgduckdb:18-v1.1.1` is genuine PG18.1 | `read_parquet('…')` in plain SQL; globs, `union_by_name`, full DuckDB analytics engine | **Chosen** |
| pg_parquet (Crunchy Data) | Yes (1418) | `COPY … TO/FROM '…' (format 'parquet')` — import/export, not a query engine | Viable, but COPY-oriented; no standalone official image (ships via Crunchy Bridge/CPK), so it'd need a Rust/pgrx source build |
| parquet_fdw | No prebuilt PG18 support; low activity | Foreign tables over parquet | Rejected — oldest, weakest PG18 story |
pg_duckdb wins for the stated goal (ad-hoc analytics): it exposes DuckDB's
`read_parquet` directly in SQL, so you query cache files like tables — no
import step, no foreign-table DDL, and you get aggregation/joins/globs across
many cells at once.
## PG-version reality for PG18 (empirically verified 2026-07-19)
**No version delta.** PG18 support is real, not a fallback. The pulled image
reports:
```
PostgreSQL 18.1 (Debian 18.1-1.pgdg12+2) on x86_64-pc-linux-gnu
pg_extension: pg_duckdb 1.1.0
shared_preload_libraries: pg_duckdb
```
The image is built on Debian 12 bookworm — the **same base as the official
`postgres:18` image** — and uses the standard `docker-entrypoint.sh`. So
`POSTGRES_USER` / `POSTGRES_PASSWORD` / `POSTGRES_DB` / `PGDATA` /
`/docker-entrypoint-initdb.d` / `pg_isready` all behave exactly as with
`postgres:18`. It is a drop-in replacement for the `db` service; nothing else
in the stack changes.
We `FROM` the official pg_duckdb image (pinned to `18-v1.1.1`, not `18-main`)
rather than `FROM postgres:18` + compile, because pg_duckdb links a full DuckDB
build — compiling from source in the Dockerfile means the DuckDB toolchain and
a long, fragile build for no benefit over the maintainers' official PG18 image.
## Files here
- **`Dockerfile.db`** — `FROM pgduckdb/pgduckdb:18-v1.1.1`, plus `COPY` of the
init script so the image enables the extension on first init even without the
compose bind mount.
- **`init/10-parquet.sql`** — `CREATE EXTENSION IF NOT EXISTS pg_duckdb;` (runs
from `/docker-entrypoint-initdb.d` on first cluster init).
## Compose snippet to merge into the `db` service
Replace `image: postgres:18` with the `build:` block; add the read-only parquet
bind and the init mount. Everything else in the `db` service stays as-is.
```yaml
db:
# image: postgres:18 # <- remove; build the parquet-capable image
build:
context: .
dockerfile: deploy/db/Dockerfile.db
environment:
POSTGRES_USER: thermograph
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD}
POSTGRES_DB: thermograph
PGDATA: /var/lib/postgresql/data
volumes:
- pgdata:/var/lib/postgresql/data
- ./data/cache:/parquet:ro # read-only parquet cache
- ./deploy/db/init:/docker-entrypoint-initdb.d
# healthcheck / cpus / deploy / restart: unchanged
```
Notes:
- The bind mount at `/docker-entrypoint-initdb.d` **replaces** the base image's
own init scripts (its `0001-install-pg_duckdb.sql` and the MotherDuck-only
`0002-enable-md-pg_duckdb.sql`). That's intended: our `10-parquet.sql` still
runs `CREATE EXTENSION`, and we don't use MotherDuck. If you prefer to keep
the image's baked scripts, drop the `:/docker-entrypoint-initdb.d` line — the
`Dockerfile` already bakes `10-parquet.sql` in.
- `:ro` keeps the DB from ever mutating the app's cache. In prod the app writes
the cache to the `appdata` volume; point this bind at wherever that lives on
the host if you want the DB to see the live cache rather than the repo copy.
- Init scripts only run when `PGDATA` is empty. On an **existing** database,
enable it once by hand:
```
docker compose exec db psql -U thermograph -d thermograph \
-c 'CREATE EXTENSION IF NOT EXISTS pg_duckdb;'
```
## Usage — real cache file at `/parquet/…`
Cache columns: `date, tmax, tmin, precip, wind, gust, humid, fmax, fmin, feels`.
pg_duckdb ≥ 0.3 uses the `r['colname']` subscript syntax (not
`AS t(col type, …)`).
```sql
-- one cell, first rows
SELECT r['date'] AS date, r['tmax'] AS tmax, r['tmin'] AS tmin, r['precip'] AS precip
FROM read_parquet('/parquet/1026_-2857.parquet') r
ORDER BY r['date'] LIMIT 5;
-- per-year analytics over one cell
SELECT EXTRACT(YEAR FROM r['date']::timestamp) AS yr,
ROUND(AVG(r['tmax'])::numeric, 1) AS avg_tmax,
MAX(r['tmax']) AS record_high
FROM read_parquet('/parquet/1026_-2857.parquet') r
WHERE r['date'] >= '2020-01-01'
GROUP BY yr ORDER BY yr;
-- glob across every cached cell (union_by_name handles the _rf / _forecast
-- files whose column sets differ)
SELECT COUNT(*) FROM read_parquet('/parquet/*.parquet', union_by_name := true) r;
```
## Proof of work (verified 2026-07-19)
Built `deploy/db/Dockerfile.db`, ran a throwaway container with the real
`data/cache` bind-mounted read-only at `/parquet`, then:
```
$ psql -c "SELECT version();"
PostgreSQL 18.1 (Debian 18.1-1.pgdg12+2) on x86_64-pc-linux-gnu ...
$ psql -c "SELECT extname, extversion FROM pg_extension WHERE extname='pg_duckdb';"
extname | extversion
-----------+------------
pg_duckdb | 1.1.0
$ psql -c "SELECT COUNT(*) FROM read_parquet('/parquet/1026_-2857.parquet') r;"
row_count
-----------
16982
$ psql -c "SELECT r['date'] AS date, r['tmax'] AS tmax, r['tmin'] AS tmin,
r['precip'] AS precip, r['humid'] AS humid
FROM read_parquet('/parquet/1026_-2857.parquet') r
ORDER BY r['date'] LIMIT 5;"
date | tmax | tmin | precip | humid
---------------------+------+------+--------+-------
1980-01-01 00:00:00 | 56.2 | 34.7 | 0 | 73
1980-01-02 00:00:00 | 63.8 | 38 | 0 | 79
1980-01-03 00:00:00 | 60.1 | 46.1 | 0.315 | 83
1980-01-04 00:00:00 | 51.7 | 40 | 0 | 68
1980-01-05 00:00:00 | 56.5 | 33.7 | 0 | 78
$ psql -c "SELECT EXTRACT(YEAR FROM r['date']::timestamp) AS yr,
ROUND(AVG(r['tmax'])::numeric,1) AS avg_tmax, MAX(r['tmax']) AS record_high
FROM read_parquet('/parquet/1026_-2857.parquet') r
WHERE r['date'] >= '2020-01-01' GROUP BY yr ORDER BY yr;"
yr | avg_tmax | record_high
------+----------+-------------
2020 | 78.8 | 99
2021 | 77.0 | 93
2022 | 79.2 | 101.4
2023 | 80.8 | 107.1
2024 | 79.8 | 97.5
2025 | 80.0 | 99.4
2026 | 77.8 | 95.6
$ psql -c "SELECT COUNT(*) FROM read_parquet('/parquet/*.parquet', union_by_name := true) r;"
count
--------
525421 -- rows across all 45 cached cells
```