All checks were successful
secrets-guard / encrypted (pull_request) Successful in 8s
shell-lint / shellcheck (pull_request) Successful in 10s
PR build (required check) / changes (pull_request) Successful in 16s
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
The estate had zero alert rules. The only contact point was Grafana's factory default pointing at the literal string <example@email.com>, and beta's untracked compose override routed Grafana's SMTP at prod's Postfix — so alerts, had any existed, would have been delivered by the box most likely to be on fire. Prod served 86 5xx in 24h including five /healthz failures and nobody was told. Alerting (deploys to beta, which is where Grafana runs): - 12 Loki-based rules. There is no Prometheus in this estate and Loki is the only datasource, so every rule is log-derived. - Thresholds come from a 24h backtest that happens to contain a real ~20min prod outage at 04:20Z. Absolute counts, not ratios: prod runs ~7 req/min, so one bad request is 1.4% and a ratio alert would scream all night. The 5xx burst rule fires on the outage's 45 and 22 buckets and on nothing else in the day; the largest benign bucket all day was 5. - Routed to a new private #ops-alerts channel, not to any existing channel — #weather-events, #announcements and #prod are product surfaces that notify real subscribers. - AlertingWatchdog is a dead-man's switch; its value is its absence. - Delivery is proven, not assumed: the rules were provisioned into a throwaway Grafana against beta's live Loki and a real alert arrived in Discord. This matters because GET /api/v1/provisioning/contact-points returns [REDACTED] for the URL — a contact point holding an uninterpolated env var looks perfectly healthy and pages nobody. The only proof is a message arriving. CI gains a structural check, because the existing one only proves YAML parses: an alert rule whose condition names a missing refId is valid YAML, provisions cleanly, and never fires. It also hard-fails on a literal Discord webhook in the repo. Verified against all three breakages deliberately introduced. Postfix supervision: - The 13h outage was a boot-ordering race, not a Docker renumbering: postfix started at 08:09:49, wg0 came up at :51, postfix fataled at :52 on a missing docker_gwbridge address, and dockerd did not finish starting until 08:10:16. Stock postfix@.service is ordered only After=network-online.target and ships no Restart=, so one lost race became a permanent outage. - An ExecStartPre gate now blocks up to 60s until every inet_interfaces address actually exists, which absorbs the transient case inside a single start attempt. That makes bounded retry correct: 5 attempts in 600s, then failed — a genuinely broken config reaches a visible failed state in ~100s instead of re-fataling every 15s forever. - A 5-minute watchdog timer retries indefinitely and runs reset-failed, so "failed" still self-heals. Worst case is ~5 minutes, not 13 hours. - Wants=, not Requires=: a dockerd failure must not take down the loopback and mesh listeners that do not depend on Docker at all. Health checks must read config with `postconf -c`, never postmulti/postqueue/ postfix — those three RESOLVE inet_interfaces and so fatal precisely when an address is missing, which made the first version of this check report status=ok bound=0/0. A health check that fails open is worse than none. DEPLOY.md carries the monitoring contract, including that systemctl is-active postfix is a known-false signal: postfix.service is a wrapper whose ExecStart=/bin/true, so it reports active forever while the real postfix@- instance is failed with zero listeners. Reproduced live. Known gap, documented not fixed: Alloy ships Docker stdout, Caddy files and app JSONL, not journald — so no Postfix line reaches Loki and the mail rules cannot fire until loki.source.journal is added. Claude-Session: https://claude.ai/code/session_0182KTMrsTHJc3TcewCatJFY
64 lines
2.4 KiB
YAML
64 lines
2.4 KiB
YAML
# The routing tree. Provisioning `policies:` REPLACES Grafana's root policy, so
|
|
# applying this file is what actually severs the factory default
|
|
# (`receiver: grafana-default-email` -> "<example@email.com>") and points the
|
|
# whole estate at Discord. Everything lands in the same #ops-alerts channel; the
|
|
# severity branches differ only in how *insistent* they are.
|
|
#
|
|
# Timing rationale, for a solo operator on a phone:
|
|
# critical — 10s wait so a burst arrives as one message, re-nag hourly while
|
|
# it is still broken. You want to be pestered about a dead prod.
|
|
# warning — 2m wait (lets a flapping thing settle), re-nag twice a day.
|
|
# Not worth waking up for; worth not forgetting.
|
|
# watchdog — the dead-man's-switch. Fires forever by design, so it is throttled
|
|
# to one message per day. Its VALUE IS ITS ABSENCE: if #ops-alerts
|
|
# is silent for >24h, Grafana/Loki/the mesh is down and no other
|
|
# alert in this file can reach you either.
|
|
#
|
|
# There is deliberately no catch-all branch: an alert whose `severity` label is
|
|
# missing or unrecognised matches none of the routes and falls through to the
|
|
# ROOT receiver, which is also Discord. Nothing can be silently swallowed by a
|
|
# typo in a label — it just arrives on the default (4h) cadence instead.
|
|
apiVersion: 1
|
|
|
|
policies:
|
|
- orgId: 1
|
|
receiver: thermograph-ops-discord
|
|
# Collapse by rule + host so a whole-node outage (every prod rule firing at
|
|
# once) still arrives as a handful of messages, not thirty.
|
|
group_by:
|
|
- alertname
|
|
- host
|
|
group_wait: 30s
|
|
group_interval: 5m
|
|
repeat_interval: 4h
|
|
|
|
routes:
|
|
- receiver: thermograph-ops-discord
|
|
object_matchers:
|
|
- ['severity', '=', 'watchdog']
|
|
group_wait: 0s
|
|
group_interval: 24h
|
|
repeat_interval: 24h
|
|
continue: false
|
|
|
|
- receiver: thermograph-ops-discord
|
|
object_matchers:
|
|
- ['severity', '=', 'critical']
|
|
group_by:
|
|
- alertname
|
|
- host
|
|
group_wait: 10s
|
|
group_interval: 5m
|
|
repeat_interval: 1h
|
|
continue: false
|
|
|
|
- receiver: thermograph-ops-discord
|
|
object_matchers:
|
|
- ['severity', '=', 'warning']
|
|
group_by:
|
|
- alertname
|
|
- host
|
|
group_wait: 2m
|
|
group_interval: 10m
|
|
repeat_interval: 12h
|
|
continue: false
|