thermograph/observability/CLAUDE.md
Emi Griffith e4693dce58
All checks were successful
secrets-guard / encrypted (pull_request) Successful in 8s
shell-lint / shellcheck (pull_request) Successful in 10s
PR build (required check) / changes (pull_request) Successful in 16s
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
observability: add the estate's first alerting; supervise Postfix
The estate had zero alert rules. The only contact point was Grafana's factory
default pointing at the literal string <example@email.com>, and beta's
untracked compose override routed Grafana's SMTP at prod's Postfix — so
alerts, had any existed, would have been delivered by the box most likely to
be on fire. Prod served 86 5xx in 24h including five /healthz failures and
nobody was told.

Alerting (deploys to beta, which is where Grafana runs):
- 12 Loki-based rules. There is no Prometheus in this estate and Loki is the
  only datasource, so every rule is log-derived.
- Thresholds come from a 24h backtest that happens to contain a real ~20min
  prod outage at 04:20Z. Absolute counts, not ratios: prod runs ~7 req/min, so
  one bad request is 1.4% and a ratio alert would scream all night. The 5xx
  burst rule fires on the outage's 45 and 22 buckets and on nothing else in
  the day; the largest benign bucket all day was 5.
- Routed to a new private #ops-alerts channel, not to any existing channel —
  #weather-events, #announcements and #prod are product surfaces that notify
  real subscribers.
- AlertingWatchdog is a dead-man's switch; its value is its absence.
- Delivery is proven, not assumed: the rules were provisioned into a throwaway
  Grafana against beta's live Loki and a real alert arrived in Discord. This
  matters because GET /api/v1/provisioning/contact-points returns [REDACTED]
  for the URL — a contact point holding an uninterpolated env var looks
  perfectly healthy and pages nobody. The only proof is a message arriving.

CI gains a structural check, because the existing one only proves YAML parses:
an alert rule whose condition names a missing refId is valid YAML, provisions
cleanly, and never fires. It also hard-fails on a literal Discord webhook in
the repo. Verified against all three breakages deliberately introduced.

Postfix supervision:
- The 13h outage was a boot-ordering race, not a Docker renumbering: postfix
  started at 08:09:49, wg0 came up at :51, postfix fataled at :52 on a missing
  docker_gwbridge address, and dockerd did not finish starting until 08:10:16.
  Stock postfix@.service is ordered only After=network-online.target and ships
  no Restart=, so one lost race became a permanent outage.
- An ExecStartPre gate now blocks up to 60s until every inet_interfaces
  address actually exists, which absorbs the transient case inside a single
  start attempt. That makes bounded retry correct: 5 attempts in 600s, then
  failed — a genuinely broken config reaches a visible failed state in ~100s
  instead of re-fataling every 15s forever.
- A 5-minute watchdog timer retries indefinitely and runs reset-failed, so
  "failed" still self-heals. Worst case is ~5 minutes, not 13 hours.
- Wants=, not Requires=: a dockerd failure must not take down the loopback and
  mesh listeners that do not depend on Docker at all.

Health checks must read config with `postconf -c`, never postmulti/postqueue/
postfix — those three RESOLVE inet_interfaces and so fatal precisely when an
address is missing, which made the first version of this check report
status=ok bound=0/0. A health check that fails open is worse than none.

DEPLOY.md carries the monitoring contract, including that systemctl is-active
postfix is a known-false signal: postfix.service is a wrapper whose
ExecStart=/bin/true, so it reports active forever while the real postfix@-
instance is failed with zero listeners. Reproduced live.

Known gap, documented not fixed: Alloy ships Docker stdout, Caddy files and
app JSONL, not journald — so no Postfix line reaches Loki and the mail rules
cannot fire until loki.source.journal is added.

Claude-Session: https://claude.ai/code/session_0182KTMrsTHJc3TcewCatJFY
2026-07-24 13:19:19 -07:00

3.1 KiB

thermograph-observability — agent instructions

The logging/metrics stack for the Thermograph fleet: Loki + Grafana on beta, with a Grafana Alloy agent on every node (prod, beta, dev) shipping container and app logs over the WireGuard mesh. Grafana is fronted by beta's Caddy at dashboard.thermograph.org (Google SSO, pre-provisioned users only).

This is operational infra config, not application code — there is no build. It deploys by hand: docker compose up -d on beta for Loki+Grafana, and the Alloy agent (alloy/docker-compose.agent.yml) on each node. See the README for the full per-node procedure.

Layout

  • docker-compose.yml — the Loki + Grafana stack (runs on beta).
  • loki/config.yml — Loki config (mesh-only, filesystem storage).
  • grafana/provisioning/ — datasource (Loki) + dashboard provider, auto-loaded at startup. grafana/dashboards/*.json — the dashboards themselves.
  • grafana/provisioning/alerting/ — the alert rules, the Discord contact point and the notification policy, also auto-loaded at startup. Same rule as the dashboards: the repo is the only durable path, UI edits get overwritten. Alerts go to Discord #ops-alerts, never email — beta's Grafana relays SMTP through prod's Postfix, so email dies exactly when prod does. Every rule is LogQL (there is no Prometheus anywhere in the fleet). Thresholds were derived from real Loki data and the working is in the comments beside each rule — re-derive before changing a number rather than guessing.
  • alloy/config.alloy + alloy/docker-compose.agent.yml — the per-node log shipper. The node name comes from ALLOY_NODE (set per host).
  • caddy-grafana.conf — the Caddy vhost for dashboard.thermograph.org (lives in beta's Caddy config, kept here for reference).
  • .env.example — copy to .env on beta before docker compose up. .env is gitignored; never commit real OAuth secrets or the admin password.

Conventions

  • Every artifact that ships to a node is CI-validated (.forgejo/workflows/observability-validate.yml): both compose files parse, all dashboard JSON is valid, and the Loki/provisioning YAML parses. Keep new dashboards as valid JSON and new config as valid YAML or CI fails. The alerting config gets a stricter third step — every rule's condition must name a refId that exists, every policy must route to a receiver that exists, and a literal Discord webhook URL in the repo is a hard failure. (A rule pointing at a missing refId is valid YAML, provisions cleanly, and then never fires; that is precisely the silent-no-op this whole domain exists to prevent.) The Alloy config is still not CI-validated (needs the alloy binary).
  • The public hostname is dashboard.thermograph.org everywhere — not grafana.thermograph.org. Match it in any new comment/config.
  • Secrets (OAuth client id/secret, admin password) live only in the host .env, never in the repo.

Branching

Single main branch, no protection. Open a PR against main; the validator gates it. Deploying the change to beta / the nodes is a separate manual step (this repo has no deploy automation — a known gap).