thermograph/observability/grafana/provisioning/alerting/contact-points.yml
Emi Griffith e4693dce58
All checks were successful
secrets-guard / encrypted (pull_request) Successful in 8s
shell-lint / shellcheck (pull_request) Successful in 10s
PR build (required check) / changes (pull_request) Successful in 16s
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
observability: add the estate's first alerting; supervise Postfix
The estate had zero alert rules. The only contact point was Grafana's factory
default pointing at the literal string <example@email.com>, and beta's
untracked compose override routed Grafana's SMTP at prod's Postfix — so
alerts, had any existed, would have been delivered by the box most likely to
be on fire. Prod served 86 5xx in 24h including five /healthz failures and
nobody was told.

Alerting (deploys to beta, which is where Grafana runs):
- 12 Loki-based rules. There is no Prometheus in this estate and Loki is the
  only datasource, so every rule is log-derived.
- Thresholds come from a 24h backtest that happens to contain a real ~20min
  prod outage at 04:20Z. Absolute counts, not ratios: prod runs ~7 req/min, so
  one bad request is 1.4% and a ratio alert would scream all night. The 5xx
  burst rule fires on the outage's 45 and 22 buckets and on nothing else in
  the day; the largest benign bucket all day was 5.
- Routed to a new private #ops-alerts channel, not to any existing channel —
  #weather-events, #announcements and #prod are product surfaces that notify
  real subscribers.
- AlertingWatchdog is a dead-man's switch; its value is its absence.
- Delivery is proven, not assumed: the rules were provisioned into a throwaway
  Grafana against beta's live Loki and a real alert arrived in Discord. This
  matters because GET /api/v1/provisioning/contact-points returns [REDACTED]
  for the URL — a contact point holding an uninterpolated env var looks
  perfectly healthy and pages nobody. The only proof is a message arriving.

CI gains a structural check, because the existing one only proves YAML parses:
an alert rule whose condition names a missing refId is valid YAML, provisions
cleanly, and never fires. It also hard-fails on a literal Discord webhook in
the repo. Verified against all three breakages deliberately introduced.

Postfix supervision:
- The 13h outage was a boot-ordering race, not a Docker renumbering: postfix
  started at 08:09:49, wg0 came up at :51, postfix fataled at :52 on a missing
  docker_gwbridge address, and dockerd did not finish starting until 08:10:16.
  Stock postfix@.service is ordered only After=network-online.target and ships
  no Restart=, so one lost race became a permanent outage.
- An ExecStartPre gate now blocks up to 60s until every inet_interfaces
  address actually exists, which absorbs the transient case inside a single
  start attempt. That makes bounded retry correct: 5 attempts in 600s, then
  failed — a genuinely broken config reaches a visible failed state in ~100s
  instead of re-fataling every 15s forever.
- A 5-minute watchdog timer retries indefinitely and runs reset-failed, so
  "failed" still self-heals. Worst case is ~5 minutes, not 13 hours.
- Wants=, not Requires=: a dockerd failure must not take down the loopback and
  mesh listeners that do not depend on Docker at all.

Health checks must read config with `postconf -c`, never postmulti/postqueue/
postfix — those three RESOLVE inet_interfaces and so fatal precisely when an
address is missing, which made the first version of this check report
status=ok bound=0/0. A health check that fails open is worse than none.

DEPLOY.md carries the monitoring contract, including that systemctl is-active
postfix is a known-false signal: postfix.service is a wrapper whose
ExecStart=/bin/true, so it reports active forever while the real postfix@-
instance is failed with zero listeners. Reproduced live.

Known gap, documented not fixed: Alloy ships Docker stdout, Caddy files and
app JSONL, not journald — so no Postfix line reaches Loki and the mail rules
cannot fire until loki.source.journal is added.

Claude-Session: https://claude.ai/code/session_0182KTMrsTHJc3TcewCatJFY
2026-07-24 13:19:19 -07:00

64 lines
3.4 KiB
YAML

# Where alerts go. ONE destination: a Discord webhook posting into the private
# #ops-alerts channel (Owners category) of the Thermograph.org server.
#
# WHY NOT EMAIL. The Grafana factory default routes to the literal string
# "<example@email.com>", i.e. nowhere. Fixing it by pointing at a real mailbox
# would still be wrong: beta's Grafana relays SMTP through prod's Postfix at
# 10.10.0.1:25 (see the live docker-compose.override.yml on beta), so email
# alerts travel *through the box most likely to be on fire* and are silently
# lost whenever Postfix is down — which is exactly when you need them. Discord
# is off-estate: it works when prod is dead, and it works from a phone.
#
# THE WEBHOOK URL IS A SECRET. Anyone holding it can post into the channel, so
# it is NOT in this repo. It is read from the Grafana process environment, which
# docker-compose.yml feeds from beta's gitignored .env (see .env.example).
# Grafana expands $VAR / $__env{VAR} when it reads provisioning files; if the
# variable is unset the contact point ends up with a literal "$DISCORD_..."
# string and every notification fails — see README "Verify alerting".
#
# To rotate: make a new webhook on the #ops-alerts channel, replace the value in
# beta's .env, `docker restart observability-grafana-1`, delete the old webhook
# in Discord.
apiVersion: 1
contactPoints:
- orgId: 1
name: thermograph-ops-discord
receivers:
- uid: tg_ops_discord
type: discord
# Send a green "resolved" message too — a page you never see close is a
# page you stop trusting.
disableResolveMessage: false
settings:
url: $DISCORD_ALERT_WEBHOOK_URL
use_discord_username: false
title: '{{ if eq .Status "firing" }}[FIRING]{{ else }}[RESOLVED]{{ end }} {{ .CommonLabels.alertname }}'
# Deliberately plain. A template error here breaks EVERY notification
# silently, so this uses only functions verified against Grafana
# 11.6.1 via /api/alertmanager/grafana/config/api/v1/receivers/test.
message: |-
{{ range .Alerts }}**severity:** {{ .Labels.severity }} · **host:** {{ .Labels.host }}
{{ .Annotations.summary }}
{{ .Annotations.description }}
`value: {{ .ValueString }}`
{{ end }}
<https://dashboard.thermograph.org/alerting/list>
# --- The factory-default contact point -------------------------------------------
# Grafana ships a built-in receiver "grafana-default-email" whose only integration
# is `email receiver` -> "<example@email.com>". It is created by Grafana's default
# Alertmanager config, NOT by provisioning, and its uid is the empty string:
#
# GET /api/v1/provisioning/contact-points
# -> [{"name":"email receiver","type":"email","settings":{...},"uid":""}]
#
# Because `deleteContactPoints` matches on uid, provisioning CANNOT remove it —
# DELETE /api/v1/provisioning/contact-points/ with an empty uid 404s (verified
# against 11.6.1). What notification-policies.yml *does* do is stop anything ever
# routing to it, which makes it inert: it can no longer receive anything.
#
# Deleting the row itself is cosmetic and manual, and only works once no policy
# references it (i.e. after this config is live) — either the UI (Alerting ->
# Contact points -> grafana-default-email -> Delete) or the Alertmanager config
# API. The README's "The leftover default contact point" has the exact command.