thermograph/observability/CLAUDE.md
emi d138f00a20
Some checks failed
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Failing after 6s
secrets-guard / encrypted (push) Successful in 6s
shell-lint / shellcheck (push) Successful in 13s
Validate observability stack / validate (push) Successful in 17s
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103)
2026-07-26 06:56:38 +00:00

61 lines
3.5 KiB
Markdown

# observability/ — agent instructions
The **logging stack** for the fleet: Loki + Grafana on **vps1** — the monitoring
host, mesh `10.10.0.2`, and NOT the beta *environment* (that runs on vps2 now;
any doc that still says "beta" meaning this box is stale) — with a Grafana
Alloy agent on every node shipping container and app logs over the WireGuard
mesh. vps1 also hosts Forgejo, the portfolio site, and the **dev** environment
(mesh-only, no public DNS). vps2 runs **prod and beta** as two Swarm stacks.
Grafana is fronted by vps1's Caddy at **`dashboard.thermograph.org`** (Google
SSO, pre-provisioned users only) — use that hostname everywhere, never
`grafana.thermograph.org`.
This is operational config, not application code: there is **no build** and no
deploy automation. It ships by hand — `docker compose up -d` on vps1 for
Loki+Grafana, and `alloy/docker-compose.agent.yml` on each node. Read the root
`CLAUDE.md` first; the branch model there applies here too.
## Layout
- `docker-compose.yml` — the Loki + Grafana stack (vps1).
- `loki/config.yml` — mesh-only, filesystem storage.
- `grafana/provisioning/` — datasource + dashboard provider, auto-loaded at
startup. `grafana/dashboards/*.json` — the dashboards.
- `grafana/provisioning/alerting/` — alert rules, the Discord contact point, and
the notification policy.
- `alloy/config.alloy` + `alloy/docker-compose.agent.yml` (+
`alloy/docker-compose.agent.beta.yml`, a vps2-only overlay adding beta's
`applogs` mount) — the per-node shipper. `ALLOY_NODE` is the MACHINE
(`vps1`|`vps2`|`desktop`) — that's a semantic change from when the fleet had
one environment per node; `host` (`prod`|`beta`|`dev`), the ENVIRONMENT
label every dashboard/alert slices on, is now derived per source instead,
since vps2 alone carries two of them.
- `caddy-grafana.conf` — the vps1 Caddy vhost, kept here for reference.
- `.env.example` — copy to `.env` on vps1. `.env` is gitignored; never commit
OAuth secrets or the admin password.
## Rules
- **The repo is the only durable path.** Dashboards and alerting are provisioned
from this directory at startup; edits made through Grafana's UI or API are
overwritten on the next provision. Changes go through a PR.
- **Every artifact that ships to a node is CI-validated**
(`.forgejo/workflows/observability-validate.yml`, at the repo root): both
compose files parse, all dashboard JSON is valid, the Loki and provisioning
YAML parse, and the Alloy config is checked with the pinned `alloy` binary
(v1.9.1, matching what the fleet runs).
- **Alerting gets a stricter third step.** Every rule's `condition` must name a
refId that exists, every policy must route to a receiver that exists, and a
literal Discord webhook URL in the repo is a hard failure. A rule pointing at a
missing refId is valid YAML, provisions cleanly, and then never fires — that
silent no-op is what this check exists to prevent.
- **Alerts go to Discord `#ops-alerts`, never email.** vps1's Grafana relays
SMTP through vps2's Postfix, so email dies exactly when prod (on vps2) does.
- **Every rule is LogQL** — there is no Prometheus anywhere in the fleet.
Thresholds were derived from real Loki data and the working is in the comments
beside each rule; re-derive before changing a number rather than guessing.
- Secrets (OAuth client id/secret, admin password) live only in the host `.env`.
## Commits & PRs
Concise and technical. Never mention AI, assistants or automated authorship.