The estate had zero alert rules. The only contact point was Grafana's factory default pointing at the literal string <example@email.com>, and beta's untracked compose override routed Grafana's SMTP at prod's Postfix — so alerts, had any existed, would have been delivered by the box most likely to be on fire. Prod served 86 5xx in 24h including five /healthz failures and nobody was told. Alerting (deploys to beta, which is where Grafana runs): - 12 Loki-based rules. There is no Prometheus in this estate and Loki is the only datasource, so every rule is log-derived. - Thresholds come from a 24h backtest that happens to contain a real ~20min prod outage at 04:20Z. Absolute counts, not ratios: prod runs ~7 req/min, so one bad request is 1.4% and a ratio alert would scream all night. The 5xx burst rule fires on the outage's 45 and 22 buckets and on nothing else in the day; the largest benign bucket all day was 5. - Routed to a new private #ops-alerts channel, not to any existing channel — #weather-events, #announcements and #prod are product surfaces that notify real subscribers. - AlertingWatchdog is a dead-man's switch; its value is its absence. - Delivery is proven, not assumed: the rules were provisioned into a throwaway Grafana against beta's live Loki and a real alert arrived in Discord. This matters because GET /api/v1/provisioning/contact-points returns [REDACTED] for the URL — a contact point holding an uninterpolated env var looks perfectly healthy and pages nobody. The only proof is a message arriving. CI gains a structural check, because the existing one only proves YAML parses: an alert rule whose condition names a missing refId is valid YAML, provisions cleanly, and never fires. It also hard-fails on a literal Discord webhook in the repo. Verified against all three breakages deliberately introduced. Postfix supervision: - The 13h outage was a boot-ordering race, not a Docker renumbering: postfix started at 08:09:49, wg0 came up at :51, postfix fataled at :52 on a missing docker_gwbridge address, and dockerd did not finish starting until 08:10:16. Stock postfix@.service is ordered only After=network-online.target and ships no Restart=, so one lost race became a permanent outage. - An ExecStartPre gate now blocks up to 60s until every inet_interfaces address actually exists, which absorbs the transient case inside a single start attempt. That makes bounded retry correct: 5 attempts in 600s, then failed — a genuinely broken config reaches a visible failed state in ~100s instead of re-fataling every 15s forever. - A 5-minute watchdog timer retries indefinitely and runs reset-failed, so "failed" still self-heals. Worst case is ~5 minutes, not 13 hours. - Wants=, not Requires=: a dockerd failure must not take down the loopback and mesh listeners that do not depend on Docker at all. Health checks must read config with `postconf -c`, never postmulti/postqueue/ postfix — those three RESOLVE inet_interfaces and so fatal precisely when an address is missing, which made the first version of this check report status=ok bound=0/0. A health check that fails open is worse than none. DEPLOY.md carries the monitoring contract, including that systemctl is-active postfix is a known-false signal: postfix.service is a wrapper whose ExecStart=/bin/true, so it reports active forever while the real postfix@- instance is failed with zero listeners. Reproduced live. Known gap, documented not fixed: Alloy ships Docker stdout, Caddy files and app JSONL, not journald — so no Postfix line reaches Loki and the mail rules cannot fire until loki.source.journal is added. Claude-Session: https://claude.ai/code/session_0182KTMrsTHJc3TcewCatJFY
3.1 KiB
thermograph-observability — agent instructions
The logging/metrics stack for the Thermograph fleet: Loki + Grafana on beta,
with a Grafana Alloy agent on every node (prod, beta, dev) shipping container and
app logs over the WireGuard mesh. Grafana is fronted by beta's Caddy at
dashboard.thermograph.org (Google SSO, pre-provisioned users only).
This is operational infra config, not application code — there is no build.
It deploys by hand: docker compose up -d on beta for Loki+Grafana, and the
Alloy agent (alloy/docker-compose.agent.yml) on each node. See the README for
the full per-node procedure.
Layout
docker-compose.yml— the Loki + Grafana stack (runs on beta).loki/config.yml— Loki config (mesh-only, filesystem storage).grafana/provisioning/— datasource (Loki) + dashboard provider, auto-loaded at startup.grafana/dashboards/*.json— the dashboards themselves.grafana/provisioning/alerting/— the alert rules, the Discord contact point and the notification policy, also auto-loaded at startup. Same rule as the dashboards: the repo is the only durable path, UI edits get overwritten. Alerts go to Discord#ops-alerts, never email — beta's Grafana relays SMTP through prod's Postfix, so email dies exactly when prod does. Every rule is LogQL (there is no Prometheus anywhere in the fleet). Thresholds were derived from real Loki data and the working is in the comments beside each rule — re-derive before changing a number rather than guessing.alloy/config.alloy+alloy/docker-compose.agent.yml— the per-node log shipper. The node name comes fromALLOY_NODE(set per host).caddy-grafana.conf— the Caddy vhost fordashboard.thermograph.org(lives in beta's Caddy config, kept here for reference)..env.example— copy to.envon beta beforedocker compose up..envis gitignored; never commit real OAuth secrets or the admin password.
Conventions
- Every artifact that ships to a node is CI-validated
(
.forgejo/workflows/observability-validate.yml): both compose files parse, all dashboard JSON is valid, and the Loki/provisioning YAML parses. Keep new dashboards as valid JSON and new config as valid YAML or CI fails. The alerting config gets a stricter third step — every rule'sconditionmust name a refId that exists, every policy must route to a receiver that exists, and a literal Discord webhook URL in the repo is a hard failure. (A rule pointing at a missing refId is valid YAML, provisions cleanly, and then never fires; that is precisely the silent-no-op this whole domain exists to prevent.) The Alloy config is still not CI-validated (needs thealloybinary). - The public hostname is
dashboard.thermograph.orgeverywhere — notgrafana.thermograph.org. Match it in any new comment/config. - Secrets (OAuth client id/secret, admin password) live only in the host
.env, never in the repo.
Branching
Single main branch, no protection. Open a PR against main; the validator
gates it. Deploying the change to beta / the nodes is a separate manual step
(this repo has no deploy automation — a known gap).