All checks were successful
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 10s
PR build (required check) / changes (pull_request) Successful in 16s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Successful in 20s
PR build (required check) / build-backend (pull_request) Successful in 45s
PR build (required check) / gate (pull_request) Successful in 5s
Audited the five CLAUDE.md files and all twenty-one README.md files against the
tree, machine-checking every in-repo path they name and verifying the testable
claims against the live hosts.
The one that matters is in the root file: dev was documented as reachable on
the mesh at 10.10.0.2:8137. It is not, and never was from anywhere but vps1 —
infra/docker-compose.yml binds the port to 127.0.0.1, and the address answers
from neither vps2 nor vps1 itself. Anyone following it gets a connection
refused with nothing to explain it.
The rest are stale paths, several from the reunification:
* assetlinks.json moved under frontend/static/ in the subtree merge; the TWA
README kept the pre-merge path in both places it names it. Following it
would put the file where nothing serves it and Android app-link
verification would fail silently.
* push.py and notify.py now live in backend/notifications/.
* INFRA.md and deploy/stack/README have never existed in this repo, in any
branch.
* the Caddyfile is at deploy/stack/lb/Caddyfile.
* three bare relative paths that resolve for a reader but not from the
directory the file sits in: units.js is the frontend's, deploy.sh is
infra's, entrypoint.sh is the backend's.
Also records why mesh clients must pin the ROOT_URL host and not only the image
host: the registry's bearer-token realm follows ROOT_URL, so pinning
git.thermograph.org alone still sends the token request out the public route,
where the /v2/* matcher returns 403 and docker falls back to anonymous. That
surfaces as `unauthorized: reqPackageAccess`, indistinguishable from a bad
credential.
Verified true and left alone: the four-domain layout, both .claude runbooks,
the absence of any domain-level .forgejo directory, the pinned compose project
name, the deploy contract, prod's eight stack services, beta's five prefixed
ones with no db of its own, dev's five, and every documented make target.
87 lines
3.7 KiB
Markdown
87 lines
3.7 KiB
Markdown
# Thermograph DB — TimescaleDB (PostgreSQL 18)
|
|
|
|
The `db` service runs the stock **`timescale/timescaledb:latest-pg18`** image
|
|
(TimescaleDB 2.24+, genuine PostgreSQL 18). The app's climate record lives here in
|
|
**hypertables** — the DB, not the filesystem, is the source of truth:
|
|
|
|
- **`climate_history`** — a hypertable of the full daily archive per grid cell
|
|
(`cell_id, date, tmax, tmin, precip, wind, gust, humid, fmax, fmin, feels`), back
|
|
to 1980. Durable/LOGGED (a 45-year, rate-limited refetch is expensive),
|
|
range-partitioned on `date` (5-year chunks), compressed for chunks older than a
|
|
year.
|
|
- **`climate_recent`** — the recent-observations + forward-forecast bundle (a plain
|
|
table: it holds future dates and is rewritten hourly).
|
|
- **`climate_sync`** — per-cell freshness (epoch seconds) that replaces the old
|
|
parquet file mtimes: it drives the hourly history top-up, the 1-hour forecast
|
|
TTL, and the `recent_stamp` token embedded in derived-payload validity.
|
|
|
|
The schema is created by Alembic (`backend/alembic/versions/0002_climate_hypertables.py`,
|
|
run at app boot via `backend/deploy/entrypoint.sh`). The app reads/writes it through
|
|
`backend/data/climate_store.py` (psycopg + polars). See
|
|
`deploy/POSTGRES-MIGRATION.md` for the parquet→hypertable cutover.
|
|
|
|
## Why the stock image (no custom Dockerfile)
|
|
|
|
The previous DB image was a custom `pgduckdb/pgduckdb:18` build whose only purpose
|
|
was ad-hoc `read_parquet()` over the parquet cache. Now the climate record is in
|
|
real tables, so that capability is gone and the DB is the **stock TimescaleDB
|
|
image** — no build step. The image already sets
|
|
`shared_preload_libraries=timescaledb`; never `ALTER SYSTEM SET
|
|
shared_preload_libraries` (it would land in `postgresql.auto.conf` and override the
|
|
image's preload).
|
|
|
|
## Files here
|
|
|
|
- **`init/10-timescaledb.sql`** — `CREATE EXTENSION IF NOT EXISTS timescaledb;`
|
|
(runs from `/docker-entrypoint-initdb.d` on first cluster init; Alembic also does
|
|
this idempotently at boot).
|
|
- **`init/20-tuning.sh`** — scales `shared_buffers` (25%), `effective_cache_size`
|
|
(75%), `work_mem`, and `maintenance_work_mem` from `DB_MEMORY` via `ALTER SYSTEM`.
|
|
|
|
## Compose `db` service
|
|
|
|
```yaml
|
|
db:
|
|
image: timescale/timescaledb:latest-pg18
|
|
environment:
|
|
POSTGRES_USER: thermograph
|
|
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD}
|
|
POSTGRES_DB: thermograph
|
|
DB_MEMORY: ${DB_MEMORY:-8g}
|
|
volumes:
|
|
- pgdata:/var/lib/postgresql
|
|
- ./deploy/db/init:/docker-entrypoint-initdb.d
|
|
# healthcheck / cpus / mem_limit / shm_size: unchanged
|
|
```
|
|
|
|
Notes:
|
|
- The volume is mounted at the **parent** of the data dir; the image picks its own
|
|
PGDATA subdir (`/var/lib/postgresql/data`) under it. The whole tree persists on
|
|
the named volume.
|
|
- No host port on purpose: the app reaches Postgres as `db:5432` on the compose
|
|
network. Nothing outside the stack should touch the database.
|
|
- Init scripts only run when PGDATA is empty. On an **existing** database, enable
|
|
the extension once by hand:
|
|
```
|
|
docker compose exec db psql -U thermograph -d thermograph \
|
|
-c 'CREATE EXTENSION IF NOT EXISTS timescaledb;'
|
|
```
|
|
|
|
## Inspecting the hypertable
|
|
|
|
```sql
|
|
-- Chunk / compression overview
|
|
SELECT hypertable_name, num_chunks, compression_enabled
|
|
FROM timescaledb_information.hypertables;
|
|
|
|
-- One cell, most recent archived days
|
|
SELECT date, tmax, tmin, precip
|
|
FROM climate_history WHERE cell_id = '1026_-2857'
|
|
ORDER BY date DESC LIMIT 5;
|
|
|
|
-- Per-year highs for one cell
|
|
SELECT EXTRACT(YEAR FROM date) AS yr,
|
|
ROUND(AVG(tmax)::numeric, 1) AS avg_tmax, MAX(tmax) AS record_high
|
|
FROM climate_history WHERE cell_id = '1026_-2857' AND date >= '2020-01-01'
|
|
GROUP BY yr ORDER BY yr;
|
|
```
|