All checks were successful
secrets-guard / encrypted (pull_request) Successful in 8s
shell-lint / shellcheck (pull_request) Successful in 10s
PR build (required check) / changes (pull_request) Successful in 16s
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
The estate had zero alert rules. The only contact point was Grafana's factory default pointing at the literal string <example@email.com>, and beta's untracked compose override routed Grafana's SMTP at prod's Postfix — so alerts, had any existed, would have been delivered by the box most likely to be on fire. Prod served 86 5xx in 24h including five /healthz failures and nobody was told. Alerting (deploys to beta, which is where Grafana runs): - 12 Loki-based rules. There is no Prometheus in this estate and Loki is the only datasource, so every rule is log-derived. - Thresholds come from a 24h backtest that happens to contain a real ~20min prod outage at 04:20Z. Absolute counts, not ratios: prod runs ~7 req/min, so one bad request is 1.4% and a ratio alert would scream all night. The 5xx burst rule fires on the outage's 45 and 22 buckets and on nothing else in the day; the largest benign bucket all day was 5. - Routed to a new private #ops-alerts channel, not to any existing channel — #weather-events, #announcements and #prod are product surfaces that notify real subscribers. - AlertingWatchdog is a dead-man's switch; its value is its absence. - Delivery is proven, not assumed: the rules were provisioned into a throwaway Grafana against beta's live Loki and a real alert arrived in Discord. This matters because GET /api/v1/provisioning/contact-points returns [REDACTED] for the URL — a contact point holding an uninterpolated env var looks perfectly healthy and pages nobody. The only proof is a message arriving. CI gains a structural check, because the existing one only proves YAML parses: an alert rule whose condition names a missing refId is valid YAML, provisions cleanly, and never fires. It also hard-fails on a literal Discord webhook in the repo. Verified against all three breakages deliberately introduced. Postfix supervision: - The 13h outage was a boot-ordering race, not a Docker renumbering: postfix started at 08:09:49, wg0 came up at :51, postfix fataled at :52 on a missing docker_gwbridge address, and dockerd did not finish starting until 08:10:16. Stock postfix@.service is ordered only After=network-online.target and ships no Restart=, so one lost race became a permanent outage. - An ExecStartPre gate now blocks up to 60s until every inet_interfaces address actually exists, which absorbs the transient case inside a single start attempt. That makes bounded retry correct: 5 attempts in 600s, then failed — a genuinely broken config reaches a visible failed state in ~100s instead of re-fataling every 15s forever. - A 5-minute watchdog timer retries indefinitely and runs reset-failed, so "failed" still self-heals. Worst case is ~5 minutes, not 13 hours. - Wants=, not Requires=: a dockerd failure must not take down the loopback and mesh listeners that do not depend on Docker at all. Health checks must read config with `postconf -c`, never postmulti/postqueue/ postfix — those three RESOLVE inet_interfaces and so fatal precisely when an address is missing, which made the first version of this check report status=ok bound=0/0. A health check that fails open is worse than none. DEPLOY.md carries the monitoring contract, including that systemctl is-active postfix is a known-false signal: postfix.service is a wrapper whose ExecStart=/bin/true, so it reports active forever while the real postfix@- instance is failed with zero listeners. Reproduced live. Known gap, documented not fixed: Alloy ships Docker stdout, Caddy files and app JSONL, not journald — so no Postfix line reaches Loki and the mail rules cannot fire until loki.source.journal is added. Claude-Session: https://claude.ai/code/session_0182KTMrsTHJc3TcewCatJFY
55 lines
3.1 KiB
Markdown
55 lines
3.1 KiB
Markdown
# thermograph-observability — agent instructions
|
|
|
|
The **logging/metrics stack** for the Thermograph fleet: Loki + Grafana on beta,
|
|
with a Grafana Alloy agent on every node (prod, beta, dev) shipping container and
|
|
app logs over the WireGuard mesh. Grafana is fronted by beta's Caddy at
|
|
`dashboard.thermograph.org` (Google SSO, pre-provisioned users only).
|
|
|
|
This is **operational infra config, not application code** — there is no build.
|
|
It deploys by hand: `docker compose up -d` on beta for Loki+Grafana, and the
|
|
Alloy agent (`alloy/docker-compose.agent.yml`) on each node. See the README for
|
|
the full per-node procedure.
|
|
|
|
## Layout
|
|
|
|
- `docker-compose.yml` — the Loki + Grafana stack (runs on beta).
|
|
- `loki/config.yml` — Loki config (mesh-only, filesystem storage).
|
|
- `grafana/provisioning/` — datasource (Loki) + dashboard provider, auto-loaded
|
|
at startup. `grafana/dashboards/*.json` — the dashboards themselves.
|
|
- `grafana/provisioning/alerting/` — the alert rules, the Discord contact point
|
|
and the notification policy, also auto-loaded at startup. Same rule as the
|
|
dashboards: **the repo is the only durable path**, UI edits get overwritten.
|
|
Alerts go to Discord `#ops-alerts`, never email — beta's Grafana relays SMTP
|
|
through prod's Postfix, so email dies exactly when prod does. Every rule is
|
|
LogQL (there is no Prometheus anywhere in the fleet). Thresholds were derived
|
|
from real Loki data and the working is in the comments beside each rule —
|
|
re-derive before changing a number rather than guessing.
|
|
- `alloy/config.alloy` + `alloy/docker-compose.agent.yml` — the per-node log
|
|
shipper. The node name comes from `ALLOY_NODE` (set per host).
|
|
- `caddy-grafana.conf` — the Caddy vhost for `dashboard.thermograph.org` (lives
|
|
in beta's Caddy config, kept here for reference).
|
|
- `.env.example` — copy to `.env` on beta before `docker compose up`. `.env` is
|
|
gitignored; never commit real OAuth secrets or the admin password.
|
|
|
|
## Conventions
|
|
|
|
- **Every artifact that ships to a node is CI-validated**
|
|
(`.forgejo/workflows/observability-validate.yml`): both compose files parse, all
|
|
dashboard JSON is valid, and the Loki/provisioning YAML parses. Keep new
|
|
dashboards as valid JSON and new config as valid YAML or CI fails. The alerting
|
|
config gets a stricter third step — every rule's `condition` must name a refId
|
|
that exists, every policy must route to a receiver that exists, and a literal
|
|
Discord webhook URL in the repo is a hard failure. (A rule pointing at a missing
|
|
refId is valid YAML, provisions cleanly, and then never fires; that is precisely
|
|
the silent-no-op this whole domain exists to prevent.) The Alloy config is still
|
|
not CI-validated (needs the `alloy` binary).
|
|
- The public hostname is **`dashboard.thermograph.org`** everywhere — not
|
|
`grafana.thermograph.org`. Match it in any new comment/config.
|
|
- Secrets (OAuth client id/secret, admin password) live only in the host `.env`,
|
|
never in the repo.
|
|
|
|
## Branching
|
|
|
|
Single `main` branch, no protection. Open a PR against `main`; the validator
|
|
gates it. Deploying the change to beta / the nodes is a separate manual step
|
|
(this repo has no deploy automation — a known gap).
|