thermograph/infra/DEPLOY.md
Emi Griffith e4693dce58
All checks were successful
secrets-guard / encrypted (pull_request) Successful in 8s
shell-lint / shellcheck (pull_request) Successful in 10s
PR build (required check) / changes (pull_request) Successful in 16s
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
observability: add the estate's first alerting; supervise Postfix
The estate had zero alert rules. The only contact point was Grafana's factory
default pointing at the literal string <example@email.com>, and beta's
untracked compose override routed Grafana's SMTP at prod's Postfix — so
alerts, had any existed, would have been delivered by the box most likely to
be on fire. Prod served 86 5xx in 24h including five /healthz failures and
nobody was told.

Alerting (deploys to beta, which is where Grafana runs):
- 12 Loki-based rules. There is no Prometheus in this estate and Loki is the
  only datasource, so every rule is log-derived.
- Thresholds come from a 24h backtest that happens to contain a real ~20min
  prod outage at 04:20Z. Absolute counts, not ratios: prod runs ~7 req/min, so
  one bad request is 1.4% and a ratio alert would scream all night. The 5xx
  burst rule fires on the outage's 45 and 22 buckets and on nothing else in
  the day; the largest benign bucket all day was 5.
- Routed to a new private #ops-alerts channel, not to any existing channel —
  #weather-events, #announcements and #prod are product surfaces that notify
  real subscribers.
- AlertingWatchdog is a dead-man's switch; its value is its absence.
- Delivery is proven, not assumed: the rules were provisioned into a throwaway
  Grafana against beta's live Loki and a real alert arrived in Discord. This
  matters because GET /api/v1/provisioning/contact-points returns [REDACTED]
  for the URL — a contact point holding an uninterpolated env var looks
  perfectly healthy and pages nobody. The only proof is a message arriving.

CI gains a structural check, because the existing one only proves YAML parses:
an alert rule whose condition names a missing refId is valid YAML, provisions
cleanly, and never fires. It also hard-fails on a literal Discord webhook in
the repo. Verified against all three breakages deliberately introduced.

Postfix supervision:
- The 13h outage was a boot-ordering race, not a Docker renumbering: postfix
  started at 08:09:49, wg0 came up at :51, postfix fataled at :52 on a missing
  docker_gwbridge address, and dockerd did not finish starting until 08:10:16.
  Stock postfix@.service is ordered only After=network-online.target and ships
  no Restart=, so one lost race became a permanent outage.
- An ExecStartPre gate now blocks up to 60s until every inet_interfaces
  address actually exists, which absorbs the transient case inside a single
  start attempt. That makes bounded retry correct: 5 attempts in 600s, then
  failed — a genuinely broken config reaches a visible failed state in ~100s
  instead of re-fataling every 15s forever.
- A 5-minute watchdog timer retries indefinitely and runs reset-failed, so
  "failed" still self-heals. Worst case is ~5 minutes, not 13 hours.
- Wants=, not Requires=: a dockerd failure must not take down the loopback and
  mesh listeners that do not depend on Docker at all.

Health checks must read config with `postconf -c`, never postmulti/postqueue/
postfix — those three RESOLVE inet_interfaces and so fatal precisely when an
address is missing, which made the first version of this check report
status=ok bound=0/0. A health check that fails open is worse than none.

DEPLOY.md carries the monitoring contract, including that systemctl is-active
postfix is a known-false signal: postfix.service is a wrapper whose
ExecStart=/bin/true, so it reports active forever while the real postfix@-
instance is failed with zero listeners. Reproduced live.

Known gap, documented not fixed: Alloy ships Docker stdout, Caddy files and
app JSONL, not journald — so no Postfix line reaches Loki and the mail rules
cannot fire until loki.source.journal is added.

Claude-Session: https://claude.ai/code/session_0182KTMrsTHJc3TcewCatJFY
2026-07-24 13:19:19 -07:00

494 lines
27 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Deploying Thermograph to a prod VPS
> **Prod and beta are now provisioned by Terraform** (`terraform/README.md`) and
> run as a **`docker compose` stack** (backend + frontend + Postgres/TimescaleDB;
> repo-split Stage 4 split the single "app" service in two), not the
> manual venv + systemd model this document originally described. Current shape:
> prod = new box `169.58.46.181`, branch `release`, deployed with `terraform
> apply`; beta = old box `75.119.132.91`, branch `main`, deployed when `main` is
> pushed (`.forgejo/workflows/deploy.yml` SSHes in and runs `deploy/deploy.sh`,
> which `docker compose pull`s the image `build-push.yml` already pushed for
> that commit and `up`s it -- repo-split Stage 6's registry-pull cutover;
> `deploy.sh` no longer builds in place). The **manual walkthrough below
> (`provision.sh`, `thermograph.service`, the venv) is legacy** — kept as a
> from-scratch, no-Terraform reference; the process model is compose, so the
> systemd/venv specifics no longer match how prod/beta actually run.
Pipeline (beta, on `main`): **push to `main` on Forgejo → `build-push.yml` builds
+ pushes the image, tagged by SHA → Forgejo Actions SSHes to beta →
`deploy/deploy.sh` (`git reset` + `docker login` + `docker compose pull && up`,
schema migration runs in the app entrypoint) → Caddy fronts it with Let's Encrypt
TLS.** Prod (on `release`) is deployed with `terraform apply`, not a push
trigger — its `remote-exec` provisioner does the same pull-instead-of-build.
(GitHub is retired/archived — Forgejo, self-hosted at `git.thermograph.org` /
`10.10.0.2:3080` over the WireGuard mesh, is now the sole git host and CI.)
Files that make this work:
| File | Where it lives in prod | Purpose |
|------|------------------------|---------|
| `.forgejo/workflows/deploy.yml` | Forgejo | CI job that SSHes in and runs the deploy script (beta, on `main`) |
| `.forgejo/workflows/build-push.yml` | Forgejo | builds + pushes the SHA-tagged image `deploy.sh`/Terraform pull |
| `deploy/deploy.sh` | `/opt/thermograph/deploy/deploy.sh` | `git reset` + `docker login` + `docker compose pull && up` + health check + warm |
| `docker-compose.yml` | `/opt/thermograph/docker-compose.yml` | app + Postgres/TimescaleDB services (the process model) |
| `deploy/thermograph.env.example` | `/etc/thermograph.env` | Postgres password, VAPID/auth secrets, `WORKERS`, sizing, base path |
| `deploy/Caddyfile` | `/etc/caddy/Caddyfile` | reverse proxy + automatic HTTPS |
| `terraform/` | run from your machine | provisions the box + brings the stack up (replaces `provision.sh`) |
| `deploy/provision.sh`, `deploy/thermograph.service` | *(legacy)* | pre-compose venv+systemd bootstrap — superseded by Terraform |
The app serves on **loopback only**; Caddy is the only thing exposed to the
internet (ports 80/443). The parquet cache in `data/cache/` lives inside the
`/opt/thermograph` checkout and is gitignored, so `git pull` never touches it.
**Any deploy host needs a `/etc/hosts` entry for the registry.**
`git.thermograph.org`'s public DNS resolves to beta's public IP, and beta's
own Caddy rejects `/v2/*` (the registry API) from anything outside the
WireGuard mesh (10.10.0.0/24) — confirmed live: `docker login`/`pull` from
prod failed with a 403 until adding `10.10.0.2 git.thermograph.org` to
`/etc/hosts` (10.10.0.2 is beta's mesh IP; every deploy host is already on
the mesh, this is purely a DNS-routing gap, not a connectivity one). Beta
itself doesn't need this — it's the box Forgejo runs on, so `git.thermograph.org`
just resolves to itself either way. `build-push.yml`'s own header comment
documents the same requirement for the CI runner host.
**Current production layout** (`deploy/Caddyfile`): the app owns
**`thermograph.org`** at its root (`THERMOGRAPH_BASE=/`, so uvicorn serves `/`,
`/calendar`, `/api/v2/…` with no prefix), and **`emigriffith.dev`** serves a
static portfolio at its root with `emigriffith.dev/thermograph*` permanently
redirecting to `thermograph.org` (prefix stripped). Both domains' `A` records
point at the same VPS; Caddy provisions a separate cert for each.
> **Note:** `deploy.sh` (the CI deploy) only pulls code, reinstalls deps, and
> restarts the app service — it does **not** touch Caddy or `/etc/thermograph.env`.
> After changing `deploy/Caddyfile` or the env, copy it onto the VPS and reload
> Caddy / restart the service by hand (see Part 2 step 4 and Part 3).
---
## Part 1 — Keys: how to manage them
You need **two** separate SSH keypairs. Don't reuse one for both.
### Key A — your login key (you ↔ VPS), you probably already have this
This is how *you* SSH in to run `provision.sh`. Nothing to set up here beyond
your existing access.
### Key B — the CI deploy key (Forgejo Actions ↔ VPS), you create this
A dedicated key whose **private half** lives in a Forgejo Actions secret and
whose **public half** is authorized on the VPS's deploy user. It should only
be able to log in as `deploy` and run the deploy script.
Generate it on your machine (no passphrase — CI can't type one):
```bash
ssh-keygen -t ed25519 -C "thermograph-ci" -f ~/.ssh/thermograph_ci -N ""
```
That makes two files:
- `~/.ssh/thermograph_ci`**private** key → goes into the Forgejo Actions secret `SSH_KEY`
- `~/.ssh/thermograph_ci.pub`**public** key → goes on the VPS
Install the **public** key on the VPS (as the `deploy` user):
```bash
ssh-copy-id -i ~/.ssh/thermograph_ci.pub deploy@YOUR_VPS_IP
# or manually: append the .pub line to /home/deploy/.ssh/authorized_keys
```
Put the **private** key into Forgejo → repo **Settings → Actions → Secrets →
Add Secret**. Set all of these:
| Secret | Value |
|--------|-------|
| `SSH_KEY` | full contents of `~/.ssh/thermograph_ci` (the private key, incl. BEGIN/END lines) |
| `SSH_HOST` | VPS IP or hostname |
| `SSH_USER` | `deploy` |
| `SSH_PORT` | `22` (or your custom SSH port) |
No `tea` CLI is installed for a command-line alternative; use the web UI, or
the REST API with a personal access token:
`POST /api/v1/repos/{owner}/{repo}/actions/secrets/{name}` (body `{"data": "..."}`).
**Why I can't do this for you:** Forgejo secrets are write-only and require
your Forgejo auth; the VPS's `authorized_keys` requires your SSH access. Both
are credentials only you should hold.
### Deploy key for cloning a private repo (optional)
If the Forgejo repo is **private**, the VPS also needs to *pull* from it. Two
options:
- Make `provision.sh`'s `REPO_URL` an **HTTPS** URL with an embedded personal
access token (what the live VPS checkout actually uses today — see its
`origin` remote), or
- Add the `deploy` user's own key (`ssh-keygen` on the VPS) as a **read-only
Deploy Key** in the repo (Settings → Deploy keys). This is separate from Key B.
If the repo is **public**, skip this — `git pull` needs no auth.
---
## Part 2 — Install (once, on the VPS)
1. **Create the Forgejo repo and push** (this project isn't versioned yet):
```bash
cd ~/Code/Thermograph
git init && git add -A && git commit -m "Initial commit"
# create an empty private repo on the Forgejo instance first (web UI or
# POST /api/v1/user/repos), then:
git remote add origin https://<forgejo-host>/OWNER/thermograph.git
git push -u origin main
```
2. **Point DNS** for each domain you serve (`thermograph.org` for the app,
`emigriffith.dev` for the portfolio) at the VPS IP with an `A` record. Let's
Encrypt won't issue a cert until the name resolves to the box.
3. **Provision the box.** Copy `deploy/provision.sh` over (or clone the repo)
and edit `REPO_URL` at the top, then:
```bash
REPO_URL="https://<forgejo-host>/OWNER/thermograph.git" bash deploy/provision.sh
```
It installs Python + git + Caddy, creates the `deploy` user, clones to
`/opt/thermograph`, installs the systemd unit + env file, grants the deploy
user a password-less `systemctl restart thermograph` (so CI can restart it),
and starts the service.
4. **Install the Caddy config** and reload. `deploy/Caddyfile` already carries the
real domains (`thermograph.org` + `emigriffith.dev`); edit them if yours differ,
then copy it into place, validate, and reload:
```bash
sudo cp /opt/thermograph/deploy/Caddyfile /etc/caddy/Caddyfile
sudo caddy validate --config /etc/caddy/Caddyfile # syntax check before reload
sudo systemctl reload caddy # Caddy fetches each domain's TLS cert automatically
```
Do this again any time `deploy/Caddyfile` changes — the CI deploy does not ship
it. Set the app's base path to match in `/etc/thermograph.env`
(`THERMOGRAPH_BASE=/` for the app-at-root layout), then `sudo systemctl restart
thermograph`.
5. **Install Key B** (Part 1) so CI can log in.
6. **Verify:** open `https://YOUR_DOMAIN/` — you should get a clean padlock.
---
## Part 3 — Day-to-day
- **Deploy:** just `git push` to `main`. Actions runs `deploy.sh` and
health-checks the service. Watch it under the repo's **Actions** tab on
Forgejo (or query `/api/v1/repos/{owner}/{repo}/actions/tasks`).
- **Manual deploy / rollback:** `ssh deploy@vps '/opt/thermograph/deploy/deploy.sh'`
(set `BRANCH=` or check out a tag first to roll back).
- **Logs:** `journalctl -u thermograph -f` (app), `/var/log/caddy/thermograph.log` (proxy).
- **Restart:** `sudo systemctl restart thermograph`.
- **Change port / base path:** edit `/etc/thermograph.env`, then `sudo systemctl
restart thermograph`. Changing the base path usually pairs with a Caddy change
(proxy target / redirects) — update `/etc/caddy/Caddyfile` and reload Caddy too.
- **Change routing / domains:** edit `/etc/caddy/Caddyfile` (or copy the repo's),
`sudo caddy validate --config /etc/caddy/Caddyfile`, then `sudo systemctl reload
caddy`. The CI deploy never ships the Caddyfile, so this is always by hand.
## Security notes
- App is bound to `127.0.0.1` — not reachable except through Caddy.
- Firewall: allow only `22`, `80`, `443` (e.g. `ufw allow 22,80,443/tcp`).
- The CI sudoers rule is scoped to exactly `systemctl restart/status thermograph`.
- Rotate Key B by regenerating it and updating the `SSH_KEY` secret +
`authorized_keys`.
## Accounts & notifications
Accounts, subscriptions, and notifications are **authoritative and not
regenerable** — back them up. On **prod/beta they live in Postgres/TimescaleDB**
(the compose `db` service, on the `pgdata` volume) alongside the climate record;
on **LAN dev** they fall back to SQLite (`data/accounts.sqlite`) when
`THERMOGRAPH_DATABASE_URL` isn't a Postgres URL.
- **Back up the Postgres volume** (e.g. `pg_dump` the `thermograph` database) as
part of the VPS backup routine — losing it wipes all accounts and alerts. (On
a SQLite LAN dev box, back up `data/accounts.sqlite` and its `-wal`/`-shm`
sidecars instead.)
- **Multiple workers, one notifier (leader election).** Prod runs several uvicorn
workers (`WORKERS`, currently 8 on prod / 4 on beta). The subscription
evaluator and the recurring scheduler must run in exactly one, so the workers
elect a single leader — a host-local `flock` (`THERMOGRAPH_SINGLETON_LOCK`,
set by the compose app service to `/app/data/notifier.lock`), or a cluster-wide
Postgres advisory lock (`THERMOGRAPH_SINGLETON_PG=1`) if the notifier must be
single across *hosts*. See `backend/core/singleton.py`. `THERMOGRAPH_ROLE`
(`web`/`worker`/`all`, default `all`) further restricts which processes may own
it, so a stateless web tier can scale without scaling notifiers.
- **Env vars** (all optional, sensible defaults):
- `THERMOGRAPH_COOKIE_SECURE` — set to `1` when serving over HTTPS (the VPS/TLS
deploy) so the session cookie is `Secure`. Leave unset for plain-HTTP LAN, where
a `Secure` cookie would never be sent.
- `THERMOGRAPH_SESSION_TTL_DAYS` — login/cookie lifetime (default 30).
- `THERMOGRAPH_NOTIFY_INTERVAL` — seconds between evaluation passes (default 900).
- `THERMOGRAPH_NOTIFY_ARCHIVE_FETCHES` — max missing-archive fetches per pass
(default 4). A newly-subscribed cell with no cached ~45-year archive is fetched
once (then only read); this caps how many such fetches one pass may do so a burst
of new subscriptions can't exhaust the archive quota.
- `THERMOGRAPH_ENABLE_NOTIFIER` — set to `0` to disable the background evaluator.
- `THERMOGRAPH_AUTH_SECRET` — signing secret for (future) email reset/verify
tokens; unused today, a per-boot random value is used if unset.
### Web Push (VAPID)
Push notifications are signed with a VAPID keypair. If `THERMOGRAPH_VAPID_PRIVATE_KEY`
/ `THERMOGRAPH_VAPID_PUBLIC_KEY` are unset, the app generates a pair into
**`data/vapid.json`** on first run. That's fine **only if the file persists** — the
key the browser subscribed with must keep matching the key the server signs with. If
the keys ever change (regenerated `data/vapid.json`, a non-persistent data dir), every
existing subscription silently stops delivering: the push service returns 401/403 and
the app records it (`logs/errors/*.jsonl`, tagged `"phase":"push"`) but the `/push/test`
call still returns 202 with `sent:0, failed:N`. To be safe, **pin the keys** in
`/etc/thermograph.env` (see `deploy/thermograph.env.example` for the generate command).
After changing keys, users must toggle alerts off/on to re-subscribe.
Diagnose a "test says sent but nothing arrives": check the `/push/test` response
(`failed:N` ⇒ delivery rejected) or `tail logs/errors/*.jsonl | grep push`, and
`docker compose logs backend | grep -i push` for the exact status code.
## Homepage "unusual right now" feed
`backend/api/homepage.py` sweeps the tracked cities, grades each one's latest
observation from the parquet cache, and writes the ranked result to
`data/homepage.json` for the homepage template. It runs on the notifier's timer
(hourly guard) and at the tail of `warm_cities.py`, so a deploy repopulates it.
**It spends a small amount of forecast quota, and has to.** Nothing else keeps
the recent/forecast cache current for the tracked cities: `warm_cities.py` skips
any city whose archive is already cached, so it never re-fetches their forecast
bundle, and the notifier only touches cells somebody has subscribed to. Reading
that cache without refreshing it made the strip present days-old readings as
"right now". Each pass therefore tops up the **stalest** cells first, capped:
- `THERMOGRAPH_HOMEPAGE_REFRESH` — recent/forecast fetches per pass (default 40,
one per cell). At the hourly cadence that is ~960/day. Set `0` to disable, in
which case the strip only ranks cities something else happened to refresh.
- `THERMOGRAPH_HOMEPAGE_CITIES` — how many cities the strip ranks over (default
250). The strip shows ~12 cards; ranking over a smaller set that is genuinely
fresh beats ranking over every cached city when most of those are stale.
Readings older than a day are dropped rather than ranked, so an empty strip means
the cache went cold — not that the weather is unremarkable.
## Outbound email
The app never talks to a mail provider directly. It speaks plain SMTP to
**Postfix running as a send-only null client on `127.0.0.1:25`**
(`deploy/provision-mail.sh`, run once as root). `backend/mailer.py` is the only
code that sends, over stdlib `smtplib` — no new dependency.
Why route through a local MTA rather than a provider's API:
- **Delivery policy stays swappable.** Direct-to-MX or relayed through a
transactional provider is a Postfix setting; the app's config is just
"localhost:25" either way, so switching needs no code change or redeploy.
- **Postfix queues and retries.** A handler hands the message off in
microseconds; a slow or briefly-down upstream can't stall a request or lose a
signup.
- **Nothing new is exposed.** `inet_interfaces = loopback-only`, so the box
accepts mail from itself and nothing else. A loopback socket is not a
filesystem write, so the hardened unit (`ProtectSystem=full`,
`ReadWritePaths=…`) needs no change.
**Default is safe:** `THERMOGRAPH_MAIL_BACKEND` defaults to `console` — it logs
the message and sends nothing — so LAN dev and the test suite exercise the whole
signup path with no mail server and no risk of mailing a real person. Production
opts in with `=smtp`.
### Postfix unit topology — and the check that lies
Debian/Ubuntu ships Postfix as **three** systemd units, and only one of them
does anything:
| Unit | What it is | Does it bind sockets? |
|---|---|---|
| `postfix.service` | umbrella wrapper: `Type=oneshot`, `ExecStart=/bin/true`, `RemainAfterExit=yes` | **No.** Reports `active (exited)` forever. |
| `postfix@.service` | the template all instances are built from | n/a |
| `postfix@-.service` | the instance for `/etc/postfix` — the real MTA | **Yes.** This is the one that matters. |
> **`systemctl is-active postfix` is not a health check.** It returned `active`
> for the entire 13-hour mail outage of 2026-07-24 while `postfix@-.service` was
> `failed` and nothing was listening on `:25`. Reproduced deliberately during
> supervision testing: with the instance in `failed` and zero listeners,
> `is-active postfix` still said `active`. **Always query `postfix@-.service`.**
Two consequences worth internalising:
- Restart the **instance** (`systemctl restart 'postfix@-'`), not the umbrella,
for a `main.cf` change to take effect.
- **Postfix's own tools lie in the one state you most need to detect.** Any
command that *resolves* `inet_interfaces``postmulti`, `postqueue`,
`postfix` — fatals when one of the listed addresses is missing from the host.
A checker built on them returns empty and reads as "nothing wrong". Use plain
`postconf -c /etc/postfix -h inet_interfaces`, which only reads `main.cf`.
### Supervision (`deploy/provision-mail-supervision.sh`)
Installed by `provision-mail.sh`, and safe to run on its own — it touches
nothing in `main.cf`. Everything it writes lives outside `/opt/thermograph`, so
`deploy.sh`'s `git reset --hard` cannot erase it; re-running the script is what
makes a rebuilt box match.
- **Drop-in** `/etc/systemd/system/postfix@.service.d/10-wait-for-interfaces.conf`
(a drop-in, never an edit to the packaged unit, which apt replaces):
orders Postfix `After=docker.service wg-quick@wg0.service`, adds
`Restart=on-failure` / `RestartSec=15s`, and bounds retries at
`StartLimitBurst=5` / `StartLimitIntervalSec=600`.
- **Start gate** `/usr/local/sbin/postfix-wait-interfaces` — an `ExecStartPre`
that blocks (≤60s) until every address in `inet_interfaces` actually exists.
Ordering makes the boot race unlikely; this makes it deterministic, and it is
what makes a *bounded* start limit safe: a late interface costs seconds inside
one attempt instead of burning the retry budget.
- **Health check** `/usr/local/sbin/postfix-health` — the single definition of
"mail works" (see below).
- **Watchdog** `/usr/local/sbin/postfix-watchdog` + `postfix-watchdog.timer`
every 5 minutes. `Restart=` gives up after 5 attempts *by design*, so that a
genuinely broken config lands in `failed` loudly instead of re-fataling every
15s forever; the watchdog is what stops "failed" meaning "dead until a human
notices". It runs `reset-failed` (clearing the start-limit counter, which
otherwise makes plain `systemctl start` refuse) and retries, forever, at a
cadence 20× slower than `RestartSec`. Net effect: **bounded fast retry for
transients, slow unbounded retry for everything else**, worst-case recovery
~5 minutes instead of 13 hours.
`sudo touch /etc/postfix/maintenance` stops the watchdog fighting a deliberate
`systemctl stop postfix@-`. Remove it when done.
### Monitoring specification
For the Grafana/Discord alerting and any future `mail_health` tool. Alerts must
**not** route over this box's own SMTP — prod's MTA cannot be the transport for
"prod's MTA is down".
`postfix-health` is the contract. Exit `0` healthy, `1` unhealthy, and one
logfmt line on stdout (LogQL: `| logfmt`):
```
mail_health status=ok unit=postfix@-.service unit_state=active bound=3/3 \
banner=ok queue=2 queue_real=0 oldest_deferred_s=0 reason=none
```
It asserts, in order: the **instance** unit is active; every address in
`inet_interfaces` is really bound on `:25`; none of them is **missing from the
host** (the latent state that stays invisible until the next reboot and then
kills all mail); a live `220` greeting comes back, not merely an open socket;
**no public address is bound**; and queue depth/age with null-sender bounces
excluded (they are undeliverable by design — this host accepts no inbound mail —
and would otherwise hold an alert open forever).
Alert on:
| Severity | Condition | Why |
|---|---|---|
| **Page** | `mail_health status=fail` twice consecutively (≥10 min) | mail is down |
| **Page** | any `postfix_watchdog action=RECOVERY_FAILED` | self-heal exhausted |
| **Page** | **no `postfix_watchdog` line at all for 15 min** | see below |
| Warn | `status=fail` where `reason` contains only `address-missing-from-host` | still serving, but the next reboot kills all mail |
| Warn | `queue_real > 0` and `oldest_deferred_s > 3600` | mail accepted but not leaving |
The **heartbeat-absence** rule is the one that is easy to leave out and the one
that would have caught this outage. Every other check is an *active probe* and
therefore only reports while something is alive to run it; a dead box, dead
timer, or dead log shipper produces silence, and silence renders identically to
health on a dashboard. Alert on the absence of the 5-minute heartbeat, or the
monitoring inherits the same blind spot the umbrella unit had.
> **Blocker for the alerting agent:** none of this reaches Loki today. Alloy
> ships Docker container stdout, Caddy files and the app's JSONL — **not the
> systemd journal** — and Postfix runs on the host. Until
> `observability/alloy/config.alloy` gains a `loki.source.journal` (plus
> `/var/log/journal` and `/etc/machine-id` mounted read-only in
> `alloy/docker-compose.agent.yml`), no Postfix or watchdog line is queryable
> and none of the rules above can fire. That change is untested — validate it
> before relying on it. Until then the check is still reachable synchronously
> via Centralis `run_on_host` running `postfix-health`, which is also the
> natural implementation of a `mail_health` tool.
### Deliverability (do this before mailing real subscribers)
Mail from a bare VPS IP is very often junked no matter how Postfix is configured,
because the IP has no sending reputation. Either:
- **Relay through a provider** (recommended) — `RELAYHOST` + credentials in
`provision-mail.sh`. They handle SPF/DKIM alignment and reputation.
- **Direct to MX** — then you own the DNS work: an **SPF** TXT record, **DKIM**
(opendkim) with the public key published, a **DMARC** TXT record, and a
**PTR / reverse-DNS** record on the VPS IP pointing at `mail.thermograph.org`.
Missing rDNS alone is enough for Gmail and Outlook to junk everything.
Check placement with <https://www.mail-tester.com>, which scores all four at once.
- **Env vars** (all in `deploy/thermograph.env.example`):
`THERMOGRAPH_MAIL_BACKEND` (`console` | `smtp` | `disabled`),
`THERMOGRAPH_SMTP_HOST` / `_PORT` / `_USER` / `_PASSWORD` / `_STARTTLS`,
`THERMOGRAPH_MAIL_FROM`, `THERMOGRAPH_MAIL_REPLY_TO`.
- **`THERMOGRAPH_AUTH_SECRET` must be set to a fixed value** before any
confirmation or password-reset link is mailed. It currently defaults to a
per-boot random value, so every outstanding link would break on restart.
- **Digest signups** land in the `pending_digest` table in the accounts DB
(Postgres on prod/beta; authoritative, back it up). The form ships ahead of
delivery on purpose:
addresses are collected now and confirmed once SMTP is live.
- **Logs:** `journalctl -u postfix -f`; queue with `mailq`.
## SEO content pages
Crawlable, server-rendered pages (Jinja2) sit alongside the interactive tool and
link into it: `/climate` (hub), `/climate/<slug>` (per-city), `/climate/<slug>/<month>`,
`/climate/<slug>/records`, `/glossary`, `/about`, plus `/robots.txt` and
`/sitemap.xml`. The routable city set is `backend/cities.json` (~1000 metros: top
~500 global + ~250 core-English + ~250 extended-English/high-proficiency),
regenerated with `python gen_cities.py [n_global] [n_english] [n_extended]`.
- A **values filter** in `gen_cities.py` drops cities in countries that criminalize
LGBTQ+ people (plus China/Pakistan/Indonesia by choice) via `EXCLUDE_CC`, except a
hand-kept list of notable/tourist/high-English hubs in `KEEP_SLUGS` (Lagos, Nairobi,
Cairo, Dubai, KL, Shanghai×5-from-China, Karachi, Jakarta, …). Excluded slots are
backfilled from the next-ranked non-excluded cities, so the total stays ~1000. Edit
either set and regenerate to change the policy.
- **Per-city blurbs** live in `backend/cities_flavor.json` (a short Wikipedia summary
per city, CC BY-SA, attributed on the page). `python gen_flavor.py` fetches
Wikipedia's free REST summary API and validates each match by coordinates; it's
**incremental** by default (only fetches cities missing a blurb, and prunes ones no
longer in cities.json) — pass `--full` to rebuild. ~94% of cities get a blurb; the
rest render without one.
- **Notable weather events** are a small **hand-curated** list in `backend/city_events.py`
(`{slug: {text, url}}`) — one verified iconic event per city (Katrina/New Orleans,
Harvey/Houston, the 1952 Great Smog/London, …). Auto-sourcing these from search
proved unreliable (wrong matches), so they're edited by hand; a city not in the list
simply shows its flavor blurb. Add entries freely — a test checks the slugs are valid.
- **Archive warming runs automatically on every deploy.** `deploy.sh` (prod) and
`deploy-dev.sh` (dev) launch `warm_cities.py` detached in the background after the
health check, so the ~750 city pages serve from cache and a crawl can't burst the
archive API quota. It's idempotent (skips already-cached cells), so only the first
deploy does the full ~25-min warm; later deploys just top up new cities. A page
also self-heals (fetches its archive once) if hit before warming finishes. To run
it by hand: `cd backend && python warm_cities.py --pace 2`. Logs: prod
`logs/warm-cities.log`; dev `journalctl --user -u thermograph-warm-cities`.
- **Submit the sitemap** in Google Search Console: `https://thermograph.org/sitemap.xml`.
- `jinja2` is a new dependency (already in `requirements.txt`).
## Monitoring
Fleet-wide logs and dashboards live in a **separate project** — a central
Grafana + Loki stack fed by a Grafana Alloy agent on every node, over the
WireGuard mesh (`thermograph-observability` on Forgejo; UI at
**`https://dashboard.thermograph.org`**, Google SSO). It ingests every
container's stdout, Caddy's access logs, and the app's structured JSON logs
(`logs/{errors,access,audit}/*.jsonl`). This replaced the old SSH-tailed
`scripts/dashboard.py`.
The app still exposes raw counters at the gated `GET /api/v2/metrics` route
(loopback-only; refuses any request carrying a proxy `X-Forwarded-*` header, so
it's never reachable through Caddy). It reports inbound requests per endpoint,
outbound calls per upstream source, and background-daemon heartbeats
(e.g. the subscription notifier). Read it over an SSH tunnel with
`THERMOGRAPH_METRICS_TOKEN` (set in `/etc/thermograph.env`) as the `token=`
query param — handy for a quick check without opening Grafana.