diff --git a/CLAUDE.md b/CLAUDE.md
index d6cdf31..2395be1 100644
--- a/CLAUDE.md
+++ b/CLAUDE.md
@@ -1,27 +1,81 @@
# thermograph monorepo — agent instructions
-Reunified monorepo (2026-07-22): `backend/`, `frontend/`, `infra/`,
-`observability/` — each a former split repo subtree-merged with history.
-`thermograph-docs` remains a separate repo; cross-cutting decision docs and
-operator runbooks belong there, not here.
+One repo, four domains: `backend/`, `frontend/`, `infra/`, `observability/`.
+Reunified 2026-07-22 from four split repos, with history, via subtree merges.
+`thermograph-docs` is deliberately a separate repo — cross-cutting decision
+records and operator runbooks go there, not here.
-- **Each domain keeps its own `CLAUDE.md`** — read the one for the domain you
- touch. They still use repo-era wording ("this repo", sibling-repo names);
- mentally map `thermograph-backend` → `backend/` etc., and fix wording as you
- touch files near it.
-- **CI lives only in root `.forgejo/workflows/`**, path-filtered per domain
- (see README). Never re-add workflows under a domain's own `.forgejo/` — they
- are inert there and become a trap.
-- Images are `emi/thermograph/backend` and `emi/thermograph/frontend`
- (`sha-<12hex>`). Deploy = `infra/deploy/deploy.sh` over SSH, per-service.
+**Rule for this file and every domain `CLAUDE.md`: only statements that would
+break CI if they became false, or that name a file that exists.** Background,
+history and rationale belong in `thermograph-docs`. These files are read before
+every change, so a stale one is a correctness bug, not a documentation bug.
+
+## Branches and environments
+
+| Branch | Deploys to | Workflow |
+|---|---|---|
+| feature branch | nothing | PR into `dev` |
+| `dev` | LAN dev | `*-deploy-dev.yml` — **currently inert**, see below |
+| `main` | beta (beta.thermograph.org) | `*-deploy.yml` |
+| `release` | prod (thermograph.org) | `*-deploy-prod.yml` |
+
+`dev`, `main` and `release` are protected: **everything is a PR**, for humans and
+agents alike. Promotion is one PR per hop, `dev` → `main` → `release`.
+
+The two `*-deploy-dev.yml` workflows are inert — the LAN box's
+`~/thermograph-dev` is still a split-era `thermograph-infra` checkout, so the
+monorepo path they call does not exist there. Use `infra/`'s `make dev-up`
+locally instead.
+
+## The deploy contract
+
+One entry point, two modes, one contract:
+
+```
+SERVICE=backend|frontend|all BACKEND_IMAGE_TAG=sha-<12hex> FRONTEND_IMAGE_TAG=sha-<12hex> \
+ /opt/thermograph/infra/deploy/deploy.sh
+```
+
+`deploy.sh` resets the host checkout, renders secrets from the SOPS vault, then
+either rolls compose services or — if `/etc/thermograph/deploy-mode` contains
+`stack` — execs `infra/deploy/stack/deploy-stack.sh`.
+
+- **prod runs Swarm.** Its stack is `infra/deploy/stack/thermograph-stack.yml`
+ (db, web, worker, lake, daemon, frontend, autoscaler, autoscaler-lake).
+ `deploy-stack.sh` also offers `STACK_TEST=1`: a full parallel rehearsal on
+ throwaway volumes and ports that cannot touch live data.
+- **beta and LAN dev run compose**, from `infra/docker-compose.yml`
+ (db, backend, lake, daemon, frontend).
+- Each service's live tag is persisted host-side, so a single-service roll never
+ disturbs the sibling's running tag.
+
+Images are `emi/thermograph/backend` and `emi/thermograph/frontend`, tagged
+`sha-<12hex>` on every push and by semver on `v*.*.*` tags. They build and deploy
+**independently** — a backend change ships without a frontend deploy and vice
+versa. That independence is the point of the FE/BE split and survived
+reunification.
+
+## Rules that bind across domains
+
+- **CI lives only in root `.forgejo/workflows/`**, path-filtered per domain.
+ Never add workflows under a domain's own `.forgejo/` — they are inert there and
+ become a trap.
+- **Secrets only via the SOPS vault** (`infra/deploy/secrets/`). Never hand-edit
+ `/etc/thermograph.env` on a host — it is a rendered artifact. `secrets-guard`
+ CI rejects plaintext. `seed-from-live.sh` reads production secrets and is
+ explicitly **not** for an agent to run.
+- FE and BE ship out of lockstep, so the `/api/version` contract and
+ `PAYLOAD_VER` discipline are load-bearing — see `backend/CLAUDE.md`.
+- Anything named "prefetch" must never spend the Open-Meteo quota.
+ Nominatim ≤ 1 req/s.
- The compose project name is pinned (`name: thermograph` in
- `infra/docker-compose.yml`); LAN dev overrides via
- `COMPOSE_PROJECT_NAME=thermograph-dev`. Don't remove either half.
-- Cross-domain rules that still bind: FE/BE ship independently → keep the
- `/api/version` contract and `PAYLOAD_VER` discipline (see
- `backend/CLAUDE.md`); anything named "prefetch" must never spend the
- Open-Meteo quota; Nominatim ≤ 1 req/s; secrets only via the SOPS vault
- (`infra/deploy/secrets/`).
+ `infra/docker-compose.yml`); LAN dev overrides with
+ `COMPOSE_PROJECT_NAME=thermograph-dev`. Don't remove either half — the pinned
+ name is what makes the Swarm stack's external volume names line up.
- `CUTOVER-NOTES.md` is the source of truth for what is and isn't live yet.
-- Commits/PRs: concise and technical; never mention AI/assistants/automated
- authorship.
+
+## Commits & PRs
+
+Describe only the substance of the change; concise and technical. Never mention
+AI, Claude, assistants or automated authorship anywhere — no trailers,
+co-authors, emoji, or "as requested" narration.
diff --git a/README.md b/README.md
index 3b3ff55..e268e31 100644
--- a/README.md
+++ b/README.md
@@ -10,7 +10,7 @@ for: **per-domain images, per-domain deploys, and an async FE/BE contract**.
|---|---|---|
| `backend/` | FastAPI graded-climate API, accounts, notifications (Discord bot, push, mail), data pipeline | `backend-build-push` → image `emi/thermograph/backend`; `backend-deploy[-prod\|-dev]` |
| `frontend/` | Public client: static JS/CSS + SSR pages | `frontend-*` mirrors of the above; image `emi/thermograph/frontend` |
-| `infra/` | Compose, deploy scripts, terraform, SOPS secrets vault, ops cron | `infra-sync` (host checkout + secrets render), `secrets-guard`, `ops-cron` |
+| `infra/` | Compose (beta, LAN dev) + the Swarm stack (prod), deploy scripts, terraform, SOPS secrets vault, ops cron | `infra-sync` (host checkout + secrets render), `secrets-guard`, `ops-cron` |
| `observability/` | Loki + Grafana + Alloy stack | `observability-validate` |
`thermograph-docs` deliberately **stays its own repo** (ADRs + runbooks, no
diff --git a/backend/CLAUDE.md b/backend/CLAUDE.md
index 0bd2165..ccc1b68 100644
--- a/backend/CLAUDE.md
+++ b/backend/CLAUDE.md
@@ -1,161 +1,66 @@
-# thermograph-backend — agent instructions
+# backend/ — agent instructions
-## What this repo is
+The Thermograph **API service**: a FastAPI app serving the graded-climate API
+(`/api/v2/...`), user accounts, notifications, and the SSR-content JSON API the
+frontend renders pages from. It owns the climate data pipeline and the
+derived-payload cache. It serves **no HTML/CSS/JS** — that is `frontend/`.
-The Thermograph **API service**: FastAPI app serving the graded-climate API
-(`/api/v2/...`), user accounts (`accounts/`, fastapi-users + Postgres via
-alembic), push/email/Discord notifications (`notifications/`), and the
-SSR-content JSON API (`api/content_routes.py` + `api/content_payloads.py`) that
-the separate `thermograph-frontend`'s server-rendered pages consume. It owns
-the climate data pipeline (`data/` — grid snapping, Open-Meteo fetch + parquet
-cache, percentile grading/scoring, places/cities) and the derived-payload
-cache/ETag store (`data/store.py`, `api/payloads.py`). It does **not** serve
-any HTML/CSS/JS itself — that moved to `thermograph-frontend` (repo-split
-Stage 7a); this process is a pure JSON API + DB-backed service, launched as
-`app:app` (a shim re-exporting `web.app:app`, kept stable for systemd/CI/`run.sh`
-callers that still expect that target).
+Read the root `CLAUDE.md` first for the branch model, deploy contract and image
+names.
-## Split topology
+## Build & verify
-Split out of the `emi/thermograph` monorepo (see
-`MONOREPO-CUTOVER-PLAN.md` at the workspace root during the transition). Sibling
-repos:
-- **`thermograph-frontend`** — the client the public sees: static
- JS/CSS/HTML + `frontend_ssr`'s server-rendered climate hub/city/month pages,
- calling this repo's `/api/v2/...` and `/content/...` JSON endpoints.
-- **`thermograph-infra`** — deploy/terraform/Swarm-Compose plumbing:
- `docker-compose.yml`, `deploy/deploy.sh`, secrets (SOPS-encrypted
- `deploy/secrets/*.yaml`), Terraform for prod/beta provisioning. This repo has
- no deploy scripts or terraform of its own — it only builds an image and hands
- infra a tag.
-- **`thermograph-docs`** — non-code planning layer: architecture decision
- records and operator runbooks. Nothing here should duplicate that; if you're
- writing a cross-cutting decision doc or a step-by-step operator runbook,
- it belongs there, not in this repo's README/CLAUDE.md.
-- **`thermograph-observability`** — monitoring/metrics stack, separate from
- this repo's own lightweight in-process `core/metrics.py`.
+- `make test` — builds a py3.12 venv and runs the hermetic pytest suite
+ (`scripts/test.sh`; pass `ARGS=...`). Tests are hermetic per `tests/conftest.py`:
+ no real Open-Meteo/Nominatim/GeoNames calls, throwaway SQLite, notifier thread
+ disabled.
+- `make smoke` — boots the built image plus a throwaway db and asserts
+ `/healthz` + `/api/version`.
+- CI runs this suite **inside the built image** (`build.yml`), so the exact
+ interpreter and deps that ship are what gets tested.
-## Branch, PR & deploy flow
+`app.py` at the root is a one-line re-export shim for `web/app.py`, kept stable
+for systemd/CI callers. Entry target is `app:app`.
-- `build-push.yml` builds **this repo's own image**
- (`git.thermograph.org/emi/thermograph-backend/app`, tagged `sha-<12hex>` on
- every push to `dev`/`main`/`release` and additionally by semver on `v*.*.*`
- tags) — no more shared monorepo image. `build.yml` (push/PR to `main`) is a
- build-only sanity check (`docker build`), not a boot/health check — repeated
- attempts at a live boot+healthz check hit this runner's own quirks (bridge-IP
- reachability, non-unique `$$`, squatted host ports) without ever going
- reliably green; the image's actual boot is verified by hand (`docker run` +
- alembic + `/healthz` 200 + a real `/api/v2/place` 200) instead.
-- **`main` → beta**: `deploy.yml` SSHes to beta and runs
- `thermograph-infra`'s `/opt/thermograph/deploy/deploy.sh` with
- `SERVICE=backend` + `BACKEND_IMAGE_TAG=sha-<12hex>` (must match
- build-push.yml's 12-hex SHA tag exactly — Actions expressions have no
- substring function for the full 40-char SHA). `deploy.sh` rolls only the
- `backend` compose service (`--no-deps`) and persists the tag in
- `deploy/.image-tags.env` host-side, so a backend-only deploy never touches
- the frontend's running container or tag.
-- **`release` → prod**: `deploy-prod.yml` is the same shape against a
- completely separate secret set (`PROD_SSH_HOST/USER/KEY/PORT`) so a beta
- credential leak can't reach prod. Also `SERVICE=backend` +
- `BACKEND_IMAGE_TAG`.
-- The frontend deploys itself the same way, independently — that
- independence (a backend change ships without a frontend deploy and vice
- versa) is the entire point of the FE/BE CI-CD split.
-- No `dev`-branch LAN deploy path exists in this repo (the monorepo's
- `deploy-dev.yml`/`docker-compose.dev.yml` never made it across the split —
- known gap, see the cutover plan if resurrecting it).
+## Contracts that break silently if you change them
-## Run / test locally
-
-There is **no Makefile in this repo** — the monorepo's app-level `make`
-targets (`lan-run`, `test`, `venv`, ...) haven't been given a new home yet
-(tracked as a gap in the cutover plan). Until that lands, drive it directly:
-
-```bash
-python3 -m venv .venv && .venv/bin/pip install -r requirements.txt -r requirements-dev.txt
-.venv/bin/uvicorn app:app --host 0.0.0.0 --port 8137 # serve
-.venv/bin/python -m pytest tests # test
-```
-
-The monorepo's Makefile preferred `uv venv --python 3.12 .venv` (falling back
-to plain `python3 -m venv` only if `uv` isn't installed) — prefer `uv` here too
-if it's on the box. There is no committed `.venv` in this repo; if one exists
-locally and looks broken (import errors, wrong Python), just `rm -rf .venv` and
-recreate rather than debugging it — it's gitignored, disposable state, not
-something the repo depends on being a specific way.
-
-Tests are hermetic (`tests/conftest.py`): no real Open-Meteo/Nominatim/GeoNames
-calls, a throwaway SQLite for accounts + the derived-payload store, and the
-notifier thread disabled. `THERMOGRAPH_FRONTEND_BASE_INTERNAL` is set to an
-unreachable placeholder in conftest — `web/app.py` fails loud at import if it's
-unset in a real run (it's the internal URL back to `frontend_ssr` for the
-non-API HTML routes this process no longer serves itself).
-
-## API version contract
-
-- Everything real lives under **`/api/v2/...`**; `/api/` and `/api/v1/` stay
- mounted as aliases of the same handlers (they always were — not new).
-- **`GET /api/version`** (`web/app.py`) is the capability/negotiation endpoint
- a split-out frontend can call at boot: `{backend_version, min_frontend,
- payload_ver}`, driven by two module constants next to it —
- `API_CONTRACT_VERSION` (bump only on a breaking change to any `/api/v{N}`
- route or payload shape) and `MIN_SUPPORTED_FRONTEND` (oldest frontend
- contract version this backend still serves correctly). A genuine breaking
- change ships as a new `v3` router mounted alongside `v2` re-registering only
- the changed handlers, `v2` kept working, `API_CONTRACT_VERSION` bumped in the
- same PR.
-- **`PAYLOAD_VER`** (`api/payloads.py`, currently `"p2"`) is a *separate*
- concern from URL versioning — it's the cache/ETag invalidation token baked
- into every derived-store validity token (`history_token()`). Bump it
- whenever a response payload's shape changes (new metric, renamed/removed
- field) so one bump atomically orphans every pre-upgrade cached row instead
- of a stale row's `history_end` happening to still validate under the old
- shape.
-- Both endpoints (`/healthz`, `/api/version`) are deliberately I/O-free so they
- stay cheap under tight healthcheck intervals.
-
-## Live cross-repo contracts (don't break silently)
-
-- **`/cell` bundle + ETag/If-None-Match**: every graded payload (`grade`,
- `calendar`, `day`) is cached in SQLite by `(kind, cell_id, key)` keyed to a
- validity token that only advances when the underlying history actually
- changes (`_etag_for` in `web/app.py`, `history_token`/`PAYLOAD_VER` in
- `api/payloads.py`). The same token doubles as a weak ETag; a client sending
- `If-None-Match` gets an empty 304 without the payload being rebuilt or even
- loaded. `expose_headers=["ETag"]` in the CORS middleware matters as much as
- `allow_origins` — without it the frontend's `cache.js` reading
- `res.headers.get("ETag")` cross-origin silently gets `null` and never
- revalidates correctly. Don't change the ETag derivation or the CORS
- `expose_headers` list without checking the frontend's cache layer.
-- **Percentile/unit parity with `thermograph-frontend`'s `static/shared.js`**:
- the frontend's percentile-to-ordinal display logic explicitly mirrors this
- repo's `data/grading.py::pct_ordinal()` (floors an empirical percentile into
- 1..99 — a rank against the sample can round to 0 or 100, and "100th
- percentile" is never shown). `TEMP_BANDS`/`RAIN_BANDS` tier boundaries in
- `data/grading.py` are the source of truth for the tier names/thresholds
- (Near Record / High / Above Normal / Normal / Below Normal / Low, and the
- five rain-intensity tiers) — changing them without a corresponding frontend
- update produces a client that draws differently-colored tiers than the
- labels the API returns.
-- **SSR content API** (`api/content_routes.py`): a separate ETag'd JSON surface
- (`/content/hub`, `/content/city/{slug}`, `/content/city/{slug}/month/{month}`,
- `/content/city/{slug}/records`, `/content/home`, `/content/sitemap`,
- `/content/indexnow-key`) that `frontend_ssr` calls in-process-free (no shared
- Python import) to render pages. Same caching pattern as the main API; keep
- payload shape changes coordinated with whatever consumes it on the frontend
- side.
+- **`GET /api/version`** → `{backend_version, min_frontend, payload_ver}`, driven
+ by `API_CONTRACT_VERSION` and `MIN_SUPPORTED_FRONTEND` in `web/app.py`. A
+ breaking change ships as a new `v3` router mounted **alongside** `v2`,
+ re-registering only changed handlers, with the constant bumped in the same PR.
+ `/api` and `/api/v1` stay mounted as aliases.
+- **`PAYLOAD_VER`** (`api/payloads.py`) is separate from URL versioning — it is
+ the cache/ETag invalidation token. Bump it whenever a response payload's shape
+ changes, so one bump atomically orphans every pre-upgrade cached row.
+- **ETag / `If-None-Match`** — every graded payload is cached against a validity
+ token that doubles as a weak ETag. `expose_headers=["ETag"]` in the CORS
+ middleware matters as much as `allow_origins`: without it the frontend's
+ `cache.js` reads `null` cross-origin and never revalidates. Don't change the
+ ETag derivation or that list without checking `frontend/static/cache.js`.
+- **`data/grading.py::pct_ordinal()`** is mirrored by `frontend`'s
+ `shared.js::pctOrd()` — floor a percentile into `1..99`, never 0 or 100.
+ `TEMP_BANDS`/`RAIN_BANDS` are the source of truth for tier names and
+ thresholds; changing them without the frontend produces tiers drawn in colours
+ that disagree with the labels the API returns.
+- **Fahrenheit country set** — `api/content_payloads.py`'s `F_COUNTRIES` must stay
+ identical to `frontend`'s `format.py::F_COUNTRIES` and `static/units.js`'s
+ `F_REGIONS`. There is a test asserting this.
+- **`/healthz` and `/api/version` are deliberately I/O-free** so they stay cheap
+ under tight healthcheck intervals.
## Layout
-`accounts/` (fastapi-users models/schemas/db, `alembic/` migrations),
-`api/` (versioned payload builders + SSR content routes), `core/` (metrics,
-audit logging, a singleton helper), `data/` (grid, climate fetch/cache, grading/
-scoring, places/cities, the derived-payload store), `notifications/` (push,
-email, Discord bot + interactions + linking, digest, scheduler), `web/app.py`
-(the actual FastAPI app; `app.py` at the repo root is a one-line re-export
-shim). `paths.py` resolves every filesystem location (climate parquet cache,
-accounts DB, logs, bundled `cities.json`/`cities_flavor.json`) from the repo
-root — don't reintroduce `__file__`-relative paths in a module, it breaks the
-moment a module moves. `deploy/entrypoint.sh` is the container entrypoint
-(alembic migrate-then-serve, with a Swarm-secrets-as-files shim); actual deploy
-orchestration lives in `thermograph-infra`, not here.
+`accounts/` (fastapi-users + alembic), `api/` (payload builders + SSR content
+routes), `core/` (metrics, audit, singleton helper), `data/` (grid, climate
+fetch/cache, grading/scoring, places/cities, derived-payload store),
+`notifications/` (push, email, Discord bot, digest, scheduler), `web/app.py`
+(the real app), `daemon/`.
+
+`paths.py` resolves every filesystem location from the repo root — don't
+reintroduce `__file__`-relative paths in a module; it breaks the moment a module
+moves. `deploy/entrypoint.sh` is the container entrypoint (alembic migrate, then
+serve, with a Swarm-secrets-as-files shim).
+
+## Commits & PRs
+
+Concise and technical. Never mention AI, assistants or automated authorship.
diff --git a/frontend/CLAUDE.md b/frontend/CLAUDE.md
index 17ea9fa..5542457 100644
--- a/frontend/CLAUDE.md
+++ b/frontend/CLAUDE.md
@@ -1,181 +1,72 @@
-# Thermograph frontend — agent instructions
+# frontend/ — agent instructions
-This repo is the **SSR + static-asset service** split out of the `emi/thermograph`
-monorepo (repo-split Stage 7). It has no climate data, no DB, and does no
-polars/compute work — everything comes from the backend's content API over
-HTTP. It is one of four sibling repos in `thermograph-repos/`:
+The **SSR + static-asset service**: server-rendered crawlable pages, the
+interactive tool's SPA shells, and every static asset. No climate data, no DB,
+no compute — everything comes from `backend/`'s content API over HTTP.
-- **`thermograph-backend`** — FastAPI API + grading/scoring/grid compute + DB.
- This repo's only dependency.
-- **`thermograph-frontend`** (this repo) — SSR content pages + the interactive
- tool's SPA shells + every static asset.
-- **`thermograph-infra`** — Terraform, `docker-compose*.yml`, `deploy/` (incl.
- `deploy.sh`, the SOPS secrets vault), Caddy config. No app code.
-- **`thermograph-docs`** — architecture decision records + operator runbooks.
- No code.
+Read the root `CLAUDE.md` first for the branch model, deploy contract and image
+names.
-See `MONOREPO-CUTOVER-PLAN.md` (one level up, in `thermograph-repos/`) for the
-full split status and the remaining gap list before the monorepo can be
-archived.
+## Build & verify
-## What this repo is
+- `make test` — whole suite (`scripts/test.sh`).
+- `make test-unit` — hermetic unit tier: SSR rendering fed committed fixtures,
+ no Docker. **This is the tier CI runs.**
+- `make test-integration` — pulls and runs the real backend image and tests the
+ live contract. Local only; CI does not run it.
+- `make backend-up` / `make backend-down` — a local backend container for dev.
+- `make capture-fixtures` — refresh `tests/fixtures/*.json` from a live backend.
-- **`content.py`** — server-rendered, crawlable pages (climate hub, per-city,
- month, records, glossary, about, privacy) + `robots.txt` + `sitemap.xml`.
- Every route fetches its data from the backend's content API
- (`api_client.py`) instead of computing in-process; `content_payloads.py` on
- the backend owns `page_title`/`canonical_path`/`breadcrumb`/`jsonld`, not
- this repo.
-- **`static/*.js`** — the interactive tool: `app.js` (map/search/graded
- results + inline SVG chart), `calendar.js`/`day.js`/`score.js`/`compare.js`
- (SPA shells served by `app.py`'s `_page()`), `account.js` (auth), `cache.js`
- (IndexedDB response cache + `/cell` bundle prefetch), `shared.js` (format
- helpers shared across views).
-- **`static/style.css`** — the single hand-written stylesheet; all design
- tokens live here (see `DESIGN.md`).
-- **`templates/*.html.j2`** — Jinja templates `content.py` renders.
-- **`content/*.yaml`** — structured SSR copy (glossary, static-page SEO meta),
- loaded by `content_loader.py`. Committed here as a starter copy extracted
- alongside the split; real cross-repo copy vendoring (a pinned
- `thermograph-copy` checkout at build time) is deferred, unbuilt follow-up —
- don't assume it exists.
+## Running it
-## Branch/PR + deploy flow
-
-- Work happens on feature branches; PR into `main`. (The monorepo's
- `dev`→`main`→`release` three-stage promotion and its LAN `deploy-dev.yml`
- path do **not** exist in this repo yet — a documented gap, see the cutover
- plan §4. Don't assume a `dev` branch here.)
-- **`.forgejo/workflows/build-push.yml`** — on push to `dev`/`main`/`release`
- or a `v*.*.*` tag: builds THIS repo's own `Dockerfile` and pushes
- `git.thermograph.org/emi/thermograph-frontend/app`, tagged `sha-<12 hex>`
- (every push) and the semver tag (tag pushes only). Backend publishes its own
- separate image the same way — the two are no longer one shared
- `emi/thermograph/app` image.
-- **`.forgejo/workflows/build.yml`** — push/PR to `main`: proves the Dockerfile
- builds. It is a **build check only**, not a boot/health check — booting the
- real app crashes at import without a reachable backend (see API-version
- section below), so a standalone boot check would need to check out
- `thermograph-backend` too. Not yet built; flagged, not silently skipped.
-- **`.forgejo/workflows/deploy.yml`** — push to `main`: SSH to beta, run
- `SERVICE=frontend FRONTEND_IMAGE_TAG=sha-<12 hex> /opt/thermograph/deploy/deploy.sh`
- (that script lives in `thermograph-infra`). Rolls **only** the frontend
- container (`--no-deps`); backend is untouched and deployed independently by
- its own repo's workflow. `deploy.sh` retries the image pull for ~5 min
- (Forgejo has no cross-workflow `needs:`, so this deploy can race ahead of
- `build-push.yml`) and waits for a healthy backend before declaring the roll
- OK (frontend's boot fetches the IndexNow key from backend — see below).
-- **`.forgejo/workflows/deploy-prod.yml`** — push to `release`: same shape,
- targets prod via its own `PROD_SSH_*` secret set (fully separate from beta's,
- so a beta credential leak can't touch prod). Nothing else deploys to prod;
- there is no release-triggered Terraform apply from this repo.
-- The `SERVICE=frontend` + `FRONTEND_IMAGE_TAG=sha-<12 hex>` pair is the entire
- contract into `thermograph-infra/deploy/deploy.sh`: it persists each
- service's live tag in `deploy/.image-tags.env` (host-side, untracked) so a
- frontend-only roll never disturbs backend's currently-running tag, and
- vice versa.
-
-## How to run / test
-
-There is no `run.sh`/`Makefile` in this repo yet (the monorepo's `make
-lan-run`/`make run`/`make stop`/`venv` targets have no home here — see the
-cutover plan's gap list). Run directly:
+`THERMOGRAPH_API_BASE_INTERNAL` is **required** — `api_client.py` raises at
+import if unset.
```bash
-python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
-THERMOGRAPH_API_BASE_INTERNAL=http://127.0.0.1:8137 \
-THERMOGRAPH_BASE=/thermograph \
+THERMOGRAPH_API_BASE_INTERNAL=http://127.0.0.1:8137 THERMOGRAPH_BASE=/thermograph \
.venv/bin/uvicorn app:app --host 0.0.0.0 --port 8080
```
-`THERMOGRAPH_API_BASE_INTERNAL` is **required** — `api_client.py` raises
-`RuntimeError` at import if it's unset. Other env vars: `THERMOGRAPH_BASE`
-(default `/thermograph`; the Dockerfile sets `/` for the deployed clean-root
-topology), `THERMOGRAPH_API_VERSION` (default `v2`, see below),
-`THERMOGRAPH_API_BASE_PUBLIC` (browser-facing backend origin for asset URLs
-when frontend and backend are cross-origin; empty = same-origin, today's
-default), `THERMOGRAPH_SSR_CACHE_TTL` (default 600s, the content-API response
-cache in `api_client.py`), `THERMOGRAPH_GOOGLE_VERIFY`/`THERMOGRAPH_BING_VERIFY`
-(search-console `` tags).
+Other env: `THERMOGRAPH_BASE` (default `/thermograph`; the Dockerfile sets `/`),
+`THERMOGRAPH_API_VERSION` (default `v2`), `THERMOGRAPH_API_BASE_PUBLIC`
+(browser-facing backend origin when cross-origin; empty = same-origin, today's
+default), `THERMOGRAPH_SSR_CACHE_TTL` (default 600s).
-**Known gap — `tests/conftest.py` does not run standalone in this repo.** It
-still assumes the pre-split layout: a sibling `backend/` checkout with its own
-`tests/conftest.py` (`make_history`/`make_recent`) and an `api.content_payloads`
-module, neither of which exists here. It was extracted verbatim as reference
-material for the real fix (a genuine cross-repo contract-test job, or a
-self-contained set of fixtures owned in this repo) — not yet built. Do **not**
-assume `pytest tests` passes here until that lands; `.forgejo/workflows/build.yml`
-only proves the image builds, for the same reason. There is also no
-`requirements-dev.txt` in this repo yet (pytest/httpx test deps aren't pinned).
+## Layout
-## API-version pinning contract
+- **`content.py`** — SSR routes (climate hub, city, month, records, glossary,
+ about, privacy) + `robots.txt` + `sitemap.xml`. Each fetches from the backend
+ via `api_client.py`; the backend's `content_payloads.py` owns `page_title` /
+ `canonical_path` / `breadcrumb` / `jsonld`, not this domain.
+- **`static/*.js`** — `app.js` (map/search/results + inline SVG chart),
+ `calendar.js`/`day.js`/`score.js`/`compare.js` (SPA shells), `account.js`
+ (auth), `cache.js` (IndexedDB cache + `/cell` bundle prefetch), `shared.js`.
+- **`static/style.css`** — the single hand-written stylesheet; design tokens live
+ here, documented in `DESIGN.md`.
+- **`templates/*.html.j2`**, **`content/*.yaml`** (structured SSR copy, loaded by
+ `content_loader.py`).
-Every backend call in this repo goes through a single pinned constant instead
-of scattered `api/v2/...` literals:
+## Contracts with the backend
-- **Python (SSR):** `api_client.py`'s module-level `API_VERSION` (default
- `"v2"`, overridable via `THERMOGRAPH_API_VERSION`). Every path builder
- (`hub()`, `sitemap()`, `indexnow_key()`, `home()`, `city()`, `city_month()`,
- `city_records()`) reads it.
-- **JS (interactive tool):** `static/account.js`'s exported `API_VERSION`
- constant and the `uv(path)` helper — every fetch in `cache.js`, `account.js`
- itself, etc. is built by wrapping the path in `uv(...)` rather than
- hardcoding a version.
-
-**Bump only in lockstep with a verified backend `/api/version` check** — the
-backend exposes `GET {BASE}/api/version` → `{backend_version, min_frontend,
-payload_ver}` (`web/app.py`'s `API_CONTRACT_VERSION`/`MIN_SUPPORTED_FRONTEND`).
-Before bumping this repo's `API_VERSION`, confirm the target backend's
-`backend_version` actually supports it and its `min_frontend` doesn't already
-exclude the version you're moving *away* from.
-
-**A v3 cutover would work like this:** backend mounts a new `v3 = APIRouter()`
-alongside `v2`, re-registering only the changed handlers (v2 keeps serving
-old clients); backend bumps `API_CONTRACT_VERSION` to `"3"` in that same PR.
-Only once that's deployed and verified does this repo bump `API_VERSION` to
-`"v3"` in both `api_client.py` and `account.js` — one PR, both pins together,
-never one without the other. `v2` (and the `/api`, `/api/v1` aliases) stay
-mounted and working until no client depends on them.
-
-The frontend's own boot is **resilient to the backend being down**: the
-`indexnow_key()` fetch in `content.py`'s `register()` is tried once eagerly
-(so the common case still serves `/.txt` as a plain static route), but a
-failure is caught, logged, and falls back to a lazy per-request lookup
-instead of crashing boot — asynchronous frontend/backend deploys depend on
-this (a briefly-unreachable backend at frontend boot must be survivable, not
-fatal).
-
-## Cross-repo contracts this repo depends on
-
-These are **live contracts with the backend** — a change on either side
-without the other breaks something silently, not loudly:
-
-- **The `/cell` bundle + ETag/`If-None-Match`** — `static/cache.js` fetches
- `/api/v2/cell` once per view-set and slices it for calendar/day/score/compare
- instead of one request per view; conditional refetch is via `If-None-Match`
- against the stored `ETag`, so an unchanged payload costs an empty 304. The
- URL map (which slice serves which view) must stay in sync with
- `thermograph-backend`'s route/payload shape.
-- **`shared.js`'s `pctOrd()` must mirror `thermograph-backend`'s
- `data/grading.py`'s `pct_ordinal()`** byte-for-byte in behavior: floor a
- percentile into `1..99` (never round to 100/0 — "100th percentile" reads as
- measurement error, not "as extreme as it has ever been"). Every percentile
- shown anywhere (Day page, calendar tooltip, chart, city pages, homepage
- strip) goes through one of these two functions; if they diverge, the same
- reading says two different things on two different surfaces.
-- **Unit/region logic** (`format.py`'s `F_COUNTRIES` / `static/units.js`'s
- `F_REGIONS` / backend's `api/content_payloads.py`'s `F_COUNTRIES`) — the set
- of Fahrenheit-using country codes must stay identical across all three;
- backend has a test asserting this.
-
-## Design & visual verification
-
-Design tokens + conventions are documented in `DESIGN.md` (source of truth:
-`static/style.css`). To see a change rendered, see `DESIGN.md`'s `make shots`
-section (`tools/shoot.py`).
+- **API version is pinned in exactly two places** — `api_client.py`'s
+ `API_VERSION` (Python) and `static/account.js`'s exported `API_VERSION` plus the
+ `uv(path)` helper (JS). Never hardcode `api/v2/...` anywhere else. Bump both in
+ one PR, and only after the target backend's `/api/version` confirms it serves
+ that version and its `min_frontend` doesn't exclude the one you're leaving.
+- **`shared.js::pctOrd()` must mirror `backend`'s `data/grading.py::pct_ordinal()`**
+ in behaviour — floor into `1..99`, never 0 or 100. Every percentile on every
+ surface goes through one of the two; if they diverge, the same reading says two
+ different things in two places.
+- **`/cell` bundle + ETag** — `cache.js` fetches `/api/v2/cell` once per view-set
+ and slices it, revalidating with `If-None-Match`. The slice→view map must track
+ the backend's payload shape.
+- **Fahrenheit country set** — `format.py::F_COUNTRIES` and `static/units.js`'s
+ `F_REGIONS` must stay identical to the backend's `F_COUNTRIES`.
+- **Boot must survive an unreachable backend.** `content.py`'s `register()` tries
+ the IndexNow-key fetch once, catches failure, and falls back to a lazy
+ per-request lookup. Asynchronous FE/BE deploys depend on this — don't make it
+ fatal.
## Commits & PRs
-Describe only the substance of the change; concise, technical. Never mention
-AI/Claude/assistants or automated authorship anywhere — no trailers,
-co-authors, emoji, or "as requested"/"per the agent" narration.
+Concise and technical. Never mention AI, assistants or automated authorship.
diff --git a/infra/CLAUDE.md b/infra/CLAUDE.md
index e10bb85..ffa2576 100644
--- a/infra/CLAUDE.md
+++ b/infra/CLAUDE.md
@@ -1,42 +1,58 @@
-# thermograph-infra — agent instructions
+# infra/ — agent instructions
-How and where the already-built Thermograph app images run. This repo owns
-Terraform, the SOPS+age secrets vault, compose files, deploy scripts, host
-provisioning, and the ops cron (DB backup + IndexNow). The app repos
-(`thermograph-backend`, `thermograph-frontend`) own building and testing the
-images; this repo never checks out app source. The old monorepo
-(`emi/thermograph`) is archived — do not point anything at it.
+How and where the already-built app images run: Terraform, the SOPS+age secrets
+vault, compose and Swarm stack files, deploy scripts, host provisioning, mail,
+and the ops cron (DB backup + IndexNow). Read the root `CLAUDE.md` first — it
+owns the branch model, the deploy contract, and which environment runs which
+orchestrator.
-## The four machines
+## The machines
-Same as ACCESS.md: **dev machine** (operator's box, LAN dev server, CI runner),
+Per `ACCESS.md`: **dev** (operator's box — LAN dev server and CI runner),
**prod** (`169.58.46.181`, thermograph.org, `agent` user, passwordless sudo),
-**beta** (`75.119.132.91`, beta.thermograph.org + Forgejo, `agent` user). Hosts'
-`/opt/thermograph` (and LAN's `~/thermograph-dev`) are checkouts of THIS repo.
+**beta** (`75.119.132.91`, beta.thermograph.org + Forgejo + Grafana/Loki,
+`agent` user). All four are on a WireGuard mesh. Hosts' `/opt/thermograph` is a
+checkout of **this monorepo**, not an infra-only repo.
-## Deploys
+## Deploy paths
-- `deploy/deploy.sh` — the one deploy path (prod/beta): resets this repo's
- checkout, renders secrets from the SOPS vault, pulls the per-service image
- tags (`BACKEND_IMAGE_TAG`/`FRONTEND_IMAGE_TAG`, persisted in untracked
- `deploy/.image-tags.env`), rolls the target service, health-checks via
- container healthchecks. Invoked over SSH by the app repos' deploy workflows.
-- `deploy/deploy-dev.sh` — thin LAN-dev wrapper around deploy.sh (dev compose
- overlay, `~/thermograph-dev`, branch `dev`).
-- Branch model: see README — infra `main` is live on prod+beta, `dev` on LAN;
- `release` is currently unused. App code is env-staged via image tags; infra
- is not.
+- **`deploy/deploy.sh`** — the single entry point for beta and prod. Resets the
+ host checkout to `BRANCH` (default `main`), renders secrets, then routes:
+ if `/etc/thermograph/deploy-mode` says `stack` it execs
+ `deploy/stack/deploy-stack.sh`, otherwise it rolls compose services. Per-service
+ tags persist in untracked `deploy/.image-tags.env` (compose) and
+ `deploy/.stack-image-tags.env` (stack) so the two never mix.
+- **`deploy/stack/deploy-stack.sh`** — the Swarm path, live on prod. `backend`
+ rolls web **and** worker; `frontend` rolls frontend; `all` runs a full
+ `docker stack deploy`. Start-first, health-gated, auto-rollback.
+ `STACK_TEST=1` rehearses the whole thing under stack name `thermograph-test`
+ on throwaway volumes and ports `18137`/`18080`.
+- **`deploy/deploy-dev.sh`** — thin LAN-dev wrapper (dev compose overlay,
+ `~/thermograph-dev`). Its CI trigger is inert; see root `CLAUDE.md`.
+- **`Makefile`** — compose orchestration only: `up`, `down`, `db-up`, `dev-up`,
+ `om-up`, `om-backfill`.
+
+Volumes are the reason the compose project name is pinned: compose creates
+`thermograph_pgdata`/`_appdata`/`_applogs`, and the Swarm stack declares those
+same names as `external: true` at the same mount paths.
## Rules
-- **Never run `terraform apply` casually** — no tfstate exists anywhere (never
- persisted), so an apply would attempt full re-provisioning of live hosts.
- Terraform here is executable documentation until state is bootstrapped.
-- Secrets: only via the SOPS vault (`deploy/secrets/*.yaml`, `sops edit` +
- commit + deploy). Never hand-edit `/etc/thermograph.env` on a SOPS-enabled
- host — it's a rendered artifact. `secrets-guard` CI rejects plaintext.
-- The ops cron (`.forgejo/workflows/ops-cron.yml`) is THE prod backup — it uses
- this repo's `PROD_SSH_*` Actions secrets. If you touch it, verify a dump
- actually lands in `agent@prod:~/thermograph-backups/`.
-- Commits/PRs: concise and technical; never mention AI/assistants/automated
- authorship.
+- **Never run `terraform apply` casually.** No tfstate is persisted anywhere, so
+ an apply would attempt full re-provisioning of live hosts. Terraform here is
+ executable documentation until state is bootstrapped.
+- **Secrets: SOPS vault only** (`deploy/secrets/*.yaml`, `sops edit` → commit →
+ deploy). Never hand-edit `/etc/thermograph.env` — it is a rendered artifact.
+ `secrets-guard` CI rejects plaintext. `deploy/secrets/seed-from-live.sh` reads
+ production secrets and is **not** for an agent to run.
+- **The ops cron (`.forgejo/workflows/ops-cron.yml`, at the repo root) is THE
+ prod backup.** It uses `PROD_SSH_*`, not `SSH_*` — an earlier revision reused
+ `SSH_*` and silently dumped beta while prod had no backup at all. If you touch
+ it, verify a dump actually lands in `agent@prod:~/thermograph-backups/`.
+- Shell here runs as root over SSH against live hosts with no test suite in
+ front of it. `shell-lint` CI (pinned shellcheck) is the only guard — keep the
+ tree at zero findings.
+
+## Commits & PRs
+
+Concise and technical. Never mention AI, assistants or automated authorship.
diff --git a/infra/README.md b/infra/README.md
index d6f2e55..fdd28e8 100644
--- a/infra/README.md
+++ b/infra/README.md
@@ -1,46 +1,49 @@
-# thermograph-infra
+# infra/
Infrastructure for [Thermograph](https://thermograph.org): Terraform host
-provisioning, the SOPS+age secrets vault, Docker Swarm/WireGuard networking,
-Forgejo, Caddy, and the deploy scripts that run the already-built app image on
-each host. Extracted from the app monorepo (`emi/thermograph`) — the app repo
-owns building and testing the app; this repo owns running it.
+provisioning, the SOPS+age secrets vault, WireGuard/Swarm networking, Forgejo,
+Caddy, mail, and the deploy scripts that run the already-built app images on each
+host. This is a domain of the `emi/thermograph` monorepo — hosts' `/opt/thermograph`
+is a checkout of the whole monorepo, and `infra/` never builds app source; it only
+runs published images.
- **`terraform/`** — provisions/configures hosts (SSH-driven by default; an
- optional GCP-creating module is scaffolded, no live resources yet) and
- triggers each deploy. See `terraform/README.md`.
-- **`deploy/secrets/`** — the git-native SOPS+age secrets vault (every app
- secret, encrypted at rest, rendered at deploy time). See
- `deploy/secrets/README.md`.
-- **`deploy/swarm/`, `deploy/forgejo/`** — the 3-node WireGuard/Swarm cluster
- that hosts Forgejo (git + CI + registry); does not run the app itself. See
- `ACCESS.md` and the READMEs under each directory.
-- **`deploy/deploy.sh`** — pulls the pinned app image (`IMAGE_TAG`) and rolls
- the compose stack; invoked by Terraform and by the app repo's
- `.forgejo/workflows/deploy.yml` over SSH.
-- **`docker-compose*.yml`, `docker-stack.yml`** — how the app image runs
- (compose in production today; `docker-stack.yml` is a design record for a
- possible future Swarm-based app deploy, not currently live).
+ optional GCP-creating module is scaffolded, no live resources yet). See
+ `terraform/README.md`. No tfstate is persisted anywhere — treat `apply` as
+ executable documentation, not a routine operation.
+- **`deploy/secrets/`** — the git-native SOPS+age secrets vault (every app secret,
+ encrypted at rest, rendered at deploy time). See `deploy/secrets/README.md`.
+- **`deploy/swarm/`, `deploy/forgejo/`** — the WireGuard/Swarm cluster hosting
+ Forgejo (git + CI + registry). See `ACCESS.md`.
+- **`deploy/deploy.sh`** — the single deploy entry point for beta and prod.
+ Takes `SERVICE=backend|frontend|all` plus `BACKEND_IMAGE_TAG`/`FRONTEND_IMAGE_TAG`,
+ resets the host checkout, renders secrets, and routes to the right orchestrator.
+- **`deploy/stack/`** — the **Swarm** path, live on **prod**:
+ `thermograph-stack.yml` (db, web, worker, lake, daemon, frontend, autoscaler,
+ autoscaler-lake), `deploy-stack.sh`, `autoscale.sh`, and the LB. Rolling updates
+ are start-first, health-gated, with auto-rollback. `STACK_TEST=1` rehearses the
+ whole stack on throwaway volumes and ports.
+- **`docker-compose*.yml`** — the **compose** path, live on **beta** and LAN dev
+ (db, backend, lake, daemon, frontend). `docker-compose.dev.yml` is the LAN
+ overlay; `docker-compose.openmeteo.yml` is the self-hosted Open-Meteo overlay.
-The app's own source, `Dockerfile`, and build/test CI stay in the app repo —
-this repo never checks out app source; hosts only pull tagged images from the
-registry. See `ACCESS.md` for host access and the Swarm/Forgejo topology, and
-`terraform/README.md` for the day-to-day `plan`/`apply` workflow.
+Which path a host takes is decided by `/etc/thermograph/deploy-mode`: the string
+`stack` makes `deploy.sh` exec `deploy/stack/deploy-stack.sh`; anything else is
+compose. The workflows never need to know which mode a host runs.
## Branches & how changes reach each environment
-- **`main`** — what **prod and beta** run: their `/opt/thermograph` checkouts
- `git reset --hard origin/main` at the start of every deploy (`deploy/deploy.sh`).
- A merge to `main` reaches those hosts on the next app deploy (or a by-hand
- `deploy.sh` run); there is no separate infra deploy trigger.
-- **`dev`** — what **LAN dev** runs: `~/thermograph-dev` resets to it via
- `deploy/deploy-dev.sh`. Keep it fast-forwarded to `main` (infra changes are not
- environment-staged today; the branches exist so LAN dev *can* trail or lead
- when needed).
-- **`release`** — currently consumed by nothing (prod tracks `main`, not
- `release`). It exists to mirror the app repos' dev→main→release promotion
- shape if per-environment infra staging is ever wanted; until then, treat
- `main` as live-everywhere.
+- **`main`** — what **prod and beta** run. `infra-sync.yml` fires on a push to
+ `main` touching `infra/**`, fast-forwards each host's `/opt/thermograph` checkout
+ and re-renders `/etc/thermograph.env` from the vault. It deliberately does
+ **not** roll any service: image tags are the app domains' axis, not infra's. A
+ compose or stack change that must recreate containers takes effect on the next
+ app deploy, or a by-hand `SERVICE=all … deploy/deploy.sh`.
+- **`dev`** — what LAN dev would run via `deploy/deploy-dev.sh`. The CI trigger
+ for this is currently inert (the LAN box still holds a split-era checkout); use
+ `make dev-up` locally.
+- **`release`** — consumed by app deploys only. Both hosts track infra via `main`;
+ prod's *app images* are staged by `release`, but its checkout follows `main`.
-Note the asymmetry with the app repos: app code IS environment-staged
-(dev→main→release maps to LAN→beta→prod via image tags), infra is not.
+Note the asymmetry with the app domains: app code IS environment-staged
+(`dev`→`main`→`release` maps to LAN→beta→prod via image tags); infra is not.
diff --git a/infra/docker-compose.yml b/infra/docker-compose.yml
index 2b13bd5..228dff2 100644
--- a/infra/docker-compose.yml
+++ b/infra/docker-compose.yml
@@ -46,8 +46,10 @@ services:
# to an exact minor (e.g. 2.17.2-pg18) before any host of this stack could ever
# replicate with another — a floating tag risks two hosts landing on different
# extension minors, which blocks a physical replica and risks compressed-chunk
- # corruption on restore. docker-stack.yml (the Swarm interim stack) REQUIRES an
- # exact pin for exactly this reason; use the SAME tag on both once you set one.
+ # corruption on restore. The Swarm path (deploy/stack/) enforces this a
+ # different way: deploy-stack.sh resolves TIMESCALEDB_IMAGE to the digest of
+ # whatever db is ALREADY running -- including this compose stack's
+ # thermograph-db-1 -- so the image under an existing volume can never drift.
image: timescale/timescaledb:${TIMESCALEDB_TAG:-latest-pg18}
environment:
POSTGRES_USER: thermograph
diff --git a/infra/docker-stack.yml b/infra/docker-stack.yml
deleted file mode 100644
index 437148b..0000000
--- a/infra/docker-stack.yml
+++ /dev/null
@@ -1,173 +0,0 @@
-# Docker Swarm stack for the hop-1 interim cutover — see
-# thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md (which this file implements) and
-# thermograph-docs/architecture/repo-topology-and-infrastructure.md §6/§9.
-#
-# Distinct from docker-compose.yml (today's plain-compose deploy, unaffected by
-# this file): Swarm pulls a pre-built image from the registry (IMAGE_TAG) rather
-# than building in place, and needs Swarm-specific keys (deploy:, networks:,
-# secrets:, no host-published ports on the overlay-fronted services).
-#
-# docker stack deploy -c docker-stack.yml thermograph
-#
-# `docker stack deploy` only interpolates from the invoking shell's environment,
-# not a .env file — export the required vars first (or `set -a; . ./stack.env;
-# set +a; docker stack deploy ...`). Required shell vars: IMAGE_TAG (the tag CI
-# pushed to the registry), TIMESCALEDB_TAG (an EXACT pinned minor — see the db
-# service below, hazard #7). POSTGRES_PASSWORD is NOT a shell var here: unlike
-# docker-compose.yml (which interpolates it into THERMOGRAPH_DATABASE_URL at
-# compose time), this file reads it as the `postgres_password` Swarm secret —
-# POSTGRES_PASSWORD_FILE for db (native support in the postgres/timescaledb
-# image), and the chunk-3 entrypoint shim builds THERMOGRAPH_DATABASE_URL for
-# app/worker from the same secret file at container start. All `postgres_password`
-# / `thermograph_*` secrets below must already exist in the Swarm (`docker secret
-# create -`) before the first deploy — Track B step 7 in
-# thermograph-docs/runbooks/implementation-handoff.md.
-#
-# Migrations are NOT run inline by app/worker here (RUN_MIGRATIONS=0) — the
-# runbook's Stage F brings the schema to head via a one-shot task BEFORE scaling
-# app up, so multiple replicas never race Alembic and nothing runs DDL against a
-# still-read-only standby mid-cutover (hazard #9):
-#
-# docker run --rm --network thermograph_internal \
-# -e THERMOGRAPH_DATABASE_URL=... migrate
-
-services:
- db:
- # Pin the EXACT TimescaleDB minor — never latest-pg18 here. A floating tag can
- # give two hosts different extension minors, which blocks a physical replica
- # (a newer .so over an older catalog won't start) and risks compressed-chunk
- # corruption on timescaledb_post_restore() (hazard #7). Get the exact X.Y.Z
- # from the source database: SELECT extversion FROM pg_extension WHERE
- # extname='timescaledb' — use the SAME tag docker-compose.yml's TIMESCALEDB_TAG
- # is pinned to, on every host that could ever replicate with this one.
- image: timescale/timescaledb:${TIMESCALEDB_TAG:?set TIMESCALEDB_TAG to an exact pinned minor, e.g. 2.17.2-pg18 -- not latest-pg18}
- environment:
- POSTGRES_USER: thermograph
- POSTGRES_PASSWORD_FILE: /run/secrets/postgres_password
- POSTGRES_DB: thermograph
- DB_MEMORY: ${DB_MEMORY:-8g}
- volumes:
- - pgdata:/var/lib/postgresql
- - ./deploy/db/init:/docker-entrypoint-initdb.d
- networks:
- - internal
- secrets:
- - postgres_password
- deploy:
- # Pin to the labelled DB node so replication/IO never crosses the slow WG
- # uplink (hazard #14): `docker node update --label-add db=true ` once,
- # on whichever node holds pgdata (Track B).
- placement:
- constraints: ["node.labels.db == true"]
- resources:
- limits:
- cpus: "${DB_CPUS:-2}"
- memory: ${DB_MEMORY:-8g}
- restart_policy:
- condition: on-failure
- # No published port: reachable only as db:5432 on the `internal` overlay —
- # Postgres must never be public (design doc §6).
-
- # Stateless, freely-replicable web tier. ROLE=web means _should_run_notifier()
- # is False unconditionally (web/app.py) — it never starts the notifier or the
- # worker scheduler even if it would otherwise win leader election.
- app:
- image: ${IMAGE_TAG:?set IMAGE_TAG to the image CI pushed, e.g. forge.example/thermograph/app:sha-abc123}
- environment:
- THERMOGRAPH_BASE: /
- PORT: 8137
- WORKERS: ${WORKERS:-4}
- THERMOGRAPH_ROLE: web
- RUN_MIGRATIONS: "0"
- volumes:
- - appdata:/app/data
- - applogs:/app/logs
- networks:
- - internal
- secrets:
- - postgres_password
- - thermograph_auth_secret
- - thermograph_vapid_private_key
- - thermograph_vapid_public_key
- deploy:
- replicas: 1 # single replica this hop -- homepage.json now lives in
- # Postgres (chunk 4), but the notifier/scheduler split (chunk
- # 2/5) is what actually lets this go multi-replica in Phase 2
- resources:
- limits:
- cpus: "${APP_CPUS:-4}"
- restart_policy:
- condition: on-failure
- update_config:
- order: start-first # new task must pass the image's HEALTHCHECK (now
- # GET /healthz) before the old one is stopped
- # No `ports:` — 127.0.0.1:8137:8137 (docker-compose.yml) has no Swarm
- # equivalent: Swarm's routing mesh publishes on 0.0.0.0, which would expose
- # the plaintext app un-fronted (hazard #6). Only the host's Caddy reaches
- # `app`, over the `internal` overlay network — see the Caddyfile templates'
- # health-gated reverse_proxy.
-
- # Owns the notifier + worker scheduler (chunks 1/2/5). Exactly one replica —
- # the leader-election guard is belt-and-suspenders, not a substitute for it.
- worker:
- image: ${IMAGE_TAG:?set IMAGE_TAG to the image CI pushed, e.g. forge.example/thermograph/app:sha-abc123}
- environment:
- THERMOGRAPH_BASE: /
- WORKERS: "1"
- THERMOGRAPH_ROLE: worker
- # Cluster-wide advisory lock, not the flock: the flock only arbitrates
- # workers on ONE host, so under Swarm it guards nothing across replicas.
- THERMOGRAPH_SINGLETON_PG: "1"
- RUN_MIGRATIONS: "0"
- volumes:
- - appdata:/app/data
- - applogs:/app/logs
- networks:
- - internal
- secrets:
- - postgres_password
- - thermograph_auth_secret
- - thermograph_vapid_private_key
- - thermograph_vapid_public_key
- - thermograph_discord_webhook
- deploy:
- replicas: 1
- resources:
- limits:
- cpus: "${WORKER_CPUS:-1}"
- restart_policy:
- condition: on-failure
- # No `ports:` — the worker serves no public traffic; GET /healthz on its own
- # container port is for the image's own HEALTHCHECK (Swarm task health) only.
-
-networks:
- internal:
- driver: overlay
- driver_opts:
- # VXLAN-over-WireGuard double encapsulation needs a lower MTU than the
- # ~1450 overlay default, or large payloads (a /cell bundle, the homepage
- # feed) silently stall while small packets (health checks) keep passing
- # (hazard #13). Validate with a real large-payload transfer, not just a
- # ping, once the mesh is up (Track B).
- com.docker.network.driver.mtu: "1370"
-
-volumes:
- pgdata: {}
- appdata: {}
- applogs: {}
-
-# Declared here, provisioned externally (docker secret create - < file) —
-# never by this stack, and never committed. See Stage 0 of the cutover runbook
-# for which values are continuity-critical (AUTH_SECRET, VAPID, POSTGRES_PASSWORD
-# must be the EXISTING live values, not freshly generated ones).
-secrets:
- postgres_password:
- external: true
- thermograph_auth_secret:
- external: true
- thermograph_vapid_private_key:
- external: true
- thermograph_vapid_public_key:
- external: true
- thermograph_discord_webhook:
- external: true
diff --git a/observability/CLAUDE.md b/observability/CLAUDE.md
index dcca01e..1485024 100644
--- a/observability/CLAUDE.md
+++ b/observability/CLAUDE.md
@@ -1,55 +1,52 @@
-# thermograph-observability — agent instructions
+# observability/ — agent instructions
-The **logging/metrics stack** for the Thermograph fleet: Loki + Grafana on beta,
-with a Grafana Alloy agent on every node (prod, beta, dev) shipping container and
-app logs over the WireGuard mesh. Grafana is fronted by beta's Caddy at
-`dashboard.thermograph.org` (Google SSO, pre-provisioned users only).
+The **logging stack** for the fleet: Loki + Grafana on beta, with a Grafana Alloy
+agent on every node (prod, beta, dev) shipping container and app logs over the
+WireGuard mesh. Grafana is fronted by beta's Caddy at
+**`dashboard.thermograph.org`** (Google SSO, pre-provisioned users only) — use
+that hostname everywhere, never `grafana.thermograph.org`.
-This is **operational infra config, not application code** — there is no build.
-It deploys by hand: `docker compose up -d` on beta for Loki+Grafana, and the
-Alloy agent (`alloy/docker-compose.agent.yml`) on each node. See the README for
-the full per-node procedure.
+This is operational config, not application code: there is **no build** and no
+deploy automation. It ships by hand — `docker compose up -d` on beta for
+Loki+Grafana, and `alloy/docker-compose.agent.yml` on each node. Read the root
+`CLAUDE.md` first; the branch model there applies here too.
## Layout
-- `docker-compose.yml` — the Loki + Grafana stack (runs on beta).
-- `loki/config.yml` — Loki config (mesh-only, filesystem storage).
-- `grafana/provisioning/` — datasource (Loki) + dashboard provider, auto-loaded
- at startup. `grafana/dashboards/*.json` — the dashboards themselves.
-- `grafana/provisioning/alerting/` — the alert rules, the Discord contact point
- and the notification policy, also auto-loaded at startup. Same rule as the
- dashboards: **the repo is the only durable path**, UI edits get overwritten.
- Alerts go to Discord `#ops-alerts`, never email — beta's Grafana relays SMTP
- through prod's Postfix, so email dies exactly when prod does. Every rule is
- LogQL (there is no Prometheus anywhere in the fleet). Thresholds were derived
- from real Loki data and the working is in the comments beside each rule —
- re-derive before changing a number rather than guessing.
-- `alloy/config.alloy` + `alloy/docker-compose.agent.yml` — the per-node log
- shipper. The node name comes from `ALLOY_NODE` (set per host).
-- `caddy-grafana.conf` — the Caddy vhost for `dashboard.thermograph.org` (lives
- in beta's Caddy config, kept here for reference).
-- `.env.example` — copy to `.env` on beta before `docker compose up`. `.env` is
- gitignored; never commit real OAuth secrets or the admin password.
+- `docker-compose.yml` — the Loki + Grafana stack (beta).
+- `loki/config.yml` — mesh-only, filesystem storage.
+- `grafana/provisioning/` — datasource + dashboard provider, auto-loaded at
+ startup. `grafana/dashboards/*.json` — the dashboards.
+- `grafana/provisioning/alerting/` — alert rules, the Discord contact point, and
+ the notification policy.
+- `alloy/config.alloy` + `alloy/docker-compose.agent.yml` — the per-node shipper.
+ Node name comes from `ALLOY_NODE`, set per host.
+- `caddy-grafana.conf` — the beta Caddy vhost, kept here for reference.
+- `.env.example` — copy to `.env` on beta. `.env` is gitignored; never commit
+ OAuth secrets or the admin password.
-## Conventions
+## Rules
+- **The repo is the only durable path.** Dashboards and alerting are provisioned
+ from this directory at startup; edits made through Grafana's UI or API are
+ overwritten on the next provision. Changes go through a PR.
- **Every artifact that ships to a node is CI-validated**
- (`.forgejo/workflows/observability-validate.yml`): both compose files parse, all
- dashboard JSON is valid, and the Loki/provisioning YAML parses. Keep new
- dashboards as valid JSON and new config as valid YAML or CI fails. The alerting
- config gets a stricter third step — every rule's `condition` must name a refId
- that exists, every policy must route to a receiver that exists, and a literal
- Discord webhook URL in the repo is a hard failure. (A rule pointing at a missing
- refId is valid YAML, provisions cleanly, and then never fires; that is precisely
- the silent-no-op this whole domain exists to prevent.) The Alloy config is still
- not CI-validated (needs the `alloy` binary).
-- The public hostname is **`dashboard.thermograph.org`** everywhere — not
- `grafana.thermograph.org`. Match it in any new comment/config.
-- Secrets (OAuth client id/secret, admin password) live only in the host `.env`,
- never in the repo.
+ (`.forgejo/workflows/observability-validate.yml`, at the repo root): both
+ compose files parse, all dashboard JSON is valid, the Loki and provisioning
+ YAML parse, and the Alloy config is checked with the pinned `alloy` binary
+ (v1.9.1, matching what the fleet runs).
+- **Alerting gets a stricter third step.** Every rule's `condition` must name a
+ refId that exists, every policy must route to a receiver that exists, and a
+ literal Discord webhook URL in the repo is a hard failure. A rule pointing at a
+ missing refId is valid YAML, provisions cleanly, and then never fires — that
+ silent no-op is what this check exists to prevent.
+- **Alerts go to Discord `#ops-alerts`, never email.** Beta's Grafana relays SMTP
+ through prod's Postfix, so email dies exactly when prod does.
+- **Every rule is LogQL** — there is no Prometheus anywhere in the fleet.
+ Thresholds were derived from real Loki data and the working is in the comments
+ beside each rule; re-derive before changing a number rather than guessing.
+- Secrets (OAuth client id/secret, admin password) live only in the host `.env`.
-## Branching
+## Commits & PRs
-Single `main` branch, no protection. Open a PR against `main`; the validator
-gates it. Deploying the change to beta / the nodes is a separate manual step
-(this repo has no deploy automation — a known gap).
+Concise and technical. Never mention AI, assistants or automated authorship.