3.5 KiB
observability/ — agent instructions
The logging stack for the fleet: Loki + Grafana on vps1 — the monitoring
host, mesh 10.10.0.2, and NOT the beta environment (that runs on vps2 now;
any doc that still says "beta" meaning this box is stale) — with a Grafana
Alloy agent on every node shipping container and app logs over the WireGuard
mesh. vps1 also hosts Forgejo, the portfolio site, and the dev environment
(mesh-only, no public DNS). vps2 runs prod and beta as two Swarm stacks.
Grafana is fronted by vps1's Caddy at dashboard.thermograph.org (Google
SSO, pre-provisioned users only) — use that hostname everywhere, never
grafana.thermograph.org.
This is operational config, not application code: there is no build and no
deploy automation. It ships by hand — docker compose up -d on vps1 for
Loki+Grafana, and alloy/docker-compose.agent.yml on each node. Read the root
CLAUDE.md first; the branch model there applies here too.
Layout
docker-compose.yml— the Loki + Grafana stack (vps1).loki/config.yml— mesh-only, filesystem storage.grafana/provisioning/— datasource + dashboard provider, auto-loaded at startup.grafana/dashboards/*.json— the dashboards.grafana/provisioning/alerting/— alert rules, the Discord contact point, and the notification policy.alloy/config.alloy+alloy/docker-compose.agent.yml(+alloy/docker-compose.agent.beta.yml, a vps2-only overlay adding beta'sapplogsmount) — the per-node shipper.ALLOY_NODEis the MACHINE (vps1|vps2|desktop) — that's a semantic change from when the fleet had one environment per node;host(prod|beta|dev), the ENVIRONMENT label every dashboard/alert slices on, is now derived per source instead, since vps2 alone carries two of them.caddy-grafana.conf— the vps1 Caddy vhost, kept here for reference..env.example— copy to.envon vps1..envis gitignored; never commit OAuth secrets or the admin password.
Rules
- The repo is the only durable path. Dashboards and alerting are provisioned from this directory at startup; edits made through Grafana's UI or API are overwritten on the next provision. Changes go through a PR.
- Every artifact that ships to a node is CI-validated
(
.forgejo/workflows/observability-validate.yml, at the repo root): both compose files parse, all dashboard JSON is valid, the Loki and provisioning YAML parse, and the Alloy config is checked with the pinnedalloybinary (v1.9.1, matching what the fleet runs). - Alerting gets a stricter third step. Every rule's
conditionmust name a refId that exists, every policy must route to a receiver that exists, and a literal Discord webhook URL in the repo is a hard failure. A rule pointing at a missing refId is valid YAML, provisions cleanly, and then never fires — that silent no-op is what this check exists to prevent. - Alerts go to Discord
#ops-alerts, never email. vps1's Grafana relays SMTP through vps2's Postfix, so email dies exactly when prod (on vps2) does. - Every rule is LogQL — there is no Prometheus anywhere in the fleet. Thresholds were derived from real Loki data and the working is in the comments beside each rule; re-derive before changing a number rather than guessing.
- Secrets (OAuth client id/secret, admin password) live only in the host
.env.
Commits & PRs
Concise and technical. Never mention AI, assistants or automated authorship.