thermograph/observability/CLAUDE.md

52 lines
2.8 KiB
Markdown

# observability/ — agent instructions
The **logging stack** for the fleet: Loki + Grafana on beta, with a Grafana Alloy
agent on every node (prod, beta, dev) shipping container and app logs over the
WireGuard mesh. Grafana is fronted by beta's Caddy at
**`dashboard.thermograph.org`** (Google SSO, pre-provisioned users only) — use
that hostname everywhere, never `grafana.thermograph.org`.
This is operational config, not application code: there is **no build** and no
deploy automation. It ships by hand — `docker compose up -d` on beta for
Loki+Grafana, and `alloy/docker-compose.agent.yml` on each node. Read the root
`CLAUDE.md` first; the branch model there applies here too.
## Layout
- `docker-compose.yml` — the Loki + Grafana stack (beta).
- `loki/config.yml` — mesh-only, filesystem storage.
- `grafana/provisioning/` — datasource + dashboard provider, auto-loaded at
startup. `grafana/dashboards/*.json` — the dashboards.
- `grafana/provisioning/alerting/` — alert rules, the Discord contact point, and
the notification policy.
- `alloy/config.alloy` + `alloy/docker-compose.agent.yml` — the per-node shipper.
Node name comes from `ALLOY_NODE`, set per host.
- `caddy-grafana.conf` — the beta Caddy vhost, kept here for reference.
- `.env.example` — copy to `.env` on beta. `.env` is gitignored; never commit
OAuth secrets or the admin password.
## Rules
- **The repo is the only durable path.** Dashboards and alerting are provisioned
from this directory at startup; edits made through Grafana's UI or API are
overwritten on the next provision. Changes go through a PR.
- **Every artifact that ships to a node is CI-validated**
(`.forgejo/workflows/observability-validate.yml`, at the repo root): both
compose files parse, all dashboard JSON is valid, the Loki and provisioning
YAML parse, and the Alloy config is checked with the pinned `alloy` binary
(v1.9.1, matching what the fleet runs).
- **Alerting gets a stricter third step.** Every rule's `condition` must name a
refId that exists, every policy must route to a receiver that exists, and a
literal Discord webhook URL in the repo is a hard failure. A rule pointing at a
missing refId is valid YAML, provisions cleanly, and then never fires — that
silent no-op is what this check exists to prevent.
- **Alerts go to Discord `#ops-alerts`, never email.** Beta's Grafana relays SMTP
through prod's Postfix, so email dies exactly when prod does.
- **Every rule is LogQL** — there is no Prometheus anywhere in the fleet.
Thresholds were derived from real Loki data and the working is in the comments
beside each rule; re-derive before changing a number rather than guessing.
- Secrets (OAuth client id/secret, admin password) live only in the host `.env`.
## Commits & PRs
Concise and technical. Never mention AI, assistants or automated authorship.