# observability/ — agent instructions The **logging stack** for the fleet: Loki + Grafana on **vps1** — the monitoring host, mesh `10.10.0.2`, and NOT the beta *environment* (that runs on vps2 now; any doc that still says "beta" meaning this box is stale) — with a Grafana Alloy agent on every node shipping container and app logs over the WireGuard mesh. vps1 also hosts Forgejo, the portfolio site, and the **dev** environment (mesh-only, no public DNS). vps2 runs **prod and beta** as two Swarm stacks. Grafana is fronted by vps1's Caddy at **`dashboard.thermograph.org`** (Google SSO, pre-provisioned users only) — use that hostname everywhere, never `grafana.thermograph.org`. This is operational config, not application code: there is **no build** and no deploy automation. It ships by hand — `docker compose up -d` on vps1 for Loki+Grafana, and `alloy/docker-compose.agent.yml` on each node. Read the root `CLAUDE.md` first; the branch model there applies here too. ## Layout - `docker-compose.yml` — the Loki + Grafana stack (vps1). - `loki/config.yml` — mesh-only, filesystem storage. - `grafana/provisioning/` — datasource + dashboard provider, auto-loaded at startup. `grafana/dashboards/*.json` — the dashboards. - `grafana/provisioning/alerting/` — alert rules, the Discord contact point, and the notification policy. - `alloy/config.alloy` + `alloy/docker-compose.agent.yml` (+ `alloy/docker-compose.agent.beta.yml`, a vps2-only overlay adding beta's `applogs` mount) — the per-node shipper. `ALLOY_NODE` is the MACHINE (`vps1`|`vps2`|`desktop`) — that's a semantic change from when the fleet had one environment per node; `host` (`prod`|`beta`|`dev`), the ENVIRONMENT label every dashboard/alert slices on, is now derived per source instead, since vps2 alone carries two of them. - `caddy-grafana.conf` — the vps1 Caddy vhost, kept here for reference. - `.env.example` — copy to `.env` on vps1. `.env` is gitignored; never commit OAuth secrets or the admin password. ## Rules - **The repo is the only durable path.** Dashboards and alerting are provisioned from this directory at startup; edits made through Grafana's UI or API are overwritten on the next provision. Changes go through a PR. - **Every artifact that ships to a node is CI-validated** (`.forgejo/workflows/observability-validate.yml`, at the repo root): both compose files parse, all dashboard JSON is valid, the Loki and provisioning YAML parse, and the Alloy config is checked with the pinned `alloy` binary (v1.9.1, matching what the fleet runs). - **Alerting gets a stricter third step.** Every rule's `condition` must name a refId that exists, every policy must route to a receiver that exists, and a literal Discord webhook URL in the repo is a hard failure. A rule pointing at a missing refId is valid YAML, provisions cleanly, and then never fires — that silent no-op is what this check exists to prevent. - **Alerts go to Discord `#ops-alerts`, never email.** vps1's Grafana relays SMTP through vps2's Postfix, so email dies exactly when prod (on vps2) does. - **Every rule is LogQL** — there is no Prometheus anywhere in the fleet. Thresholds were derived from real Loki data and the working is in the comments beside each rule; re-derive before changing a number rather than guessing. - Secrets (OAuth client id/secret, admin password) live only in the host `.env`. ## Commits & PRs Concise and technical. Never mention AI, assistants or automated authorship.