thermograph/observability/CLAUDE.md
Emi Griffith c98512cfcc
All checks were successful
secrets-guard / encrypted (pull_request) Successful in 10s
PR build (required check) / changes (pull_request) Successful in 17s
shell-lint / shellcheck (pull_request) Successful in 14s
PR build (required check) / validate-observability (pull_request) Successful in 44s
PR build (required check) / build-frontend (pull_request) Successful in 2m13s
PR build (required check) / build-backend (pull_request) Successful in 2m31s
PR build (required check) / gate (pull_request) Successful in 3s
docs: rewrite the agent context layer to match the live system
The four domain CLAUDE.md files still described the pre-monorepo split-repo
topology, and several statements were the exact inverse of current reality:
infra/CLAUDE.md told a reader that emi/thermograph is archived and must not be
pointed at, when that is the live repo; backend/ and frontend/ both claimed
there is no Makefile and no requirements-dev.txt when all four files exist;
frontend/ described itself as one of four sibling repos; observability/ claimed
a single unprotected main branch.

These files are read before every change, so a stale one is a correctness
problem rather than a documentation one. Rewritten against the tree:

- Root CLAUDE.md now owns the cross-cutting truth once — branch model, the
  SERVICE + *_IMAGE_TAG deploy contract, image names, which orchestrator each
  environment runs, and that everything is a PR. Domain files carry only what
  differs and are capped at ~70 lines.
- Records that prod runs Swarm from deploy/stack/thermograph-stack.yml while
  beta and LAN dev run compose, routed by /etc/thermograph/deploy-mode.
- Documents the *-deploy-dev.yml workflows as inert rather than leaving a
  reader to discover it.

Deletes infra/docker-stack.yml. It defined db/app/worker, matched nothing that
deploys, and was referenced only by prose describing it as a future design
record — while the real 8-service prod stack lives under deploy/stack/. A file
that looks authoritative and affects nothing is the worst case for a reader
asked to change the prod stack.

infra/README.md claimed "compose in production today" and that Swarm was "not
currently live"; corrected, and infra-sync.yml's existence is now recorded
instead of "there is no separate infra deploy trigger". The compose file's
timescale-pin comment now points at how deploy-stack.sh actually resolves the
digest from the running container.
2026-07-24 20:59:44 -07:00

2.8 KiB

observability/ — agent instructions

The logging stack for the fleet: Loki + Grafana on beta, with a Grafana Alloy agent on every node (prod, beta, dev) shipping container and app logs over the WireGuard mesh. Grafana is fronted by beta's Caddy at dashboard.thermograph.org (Google SSO, pre-provisioned users only) — use that hostname everywhere, never grafana.thermograph.org.

This is operational config, not application code: there is no build and no deploy automation. It ships by hand — docker compose up -d on beta for Loki+Grafana, and alloy/docker-compose.agent.yml on each node. Read the root CLAUDE.md first; the branch model there applies here too.

Layout

  • docker-compose.yml — the Loki + Grafana stack (beta).
  • loki/config.yml — mesh-only, filesystem storage.
  • grafana/provisioning/ — datasource + dashboard provider, auto-loaded at startup. grafana/dashboards/*.json — the dashboards.
  • grafana/provisioning/alerting/ — alert rules, the Discord contact point, and the notification policy.
  • alloy/config.alloy + alloy/docker-compose.agent.yml — the per-node shipper. Node name comes from ALLOY_NODE, set per host.
  • caddy-grafana.conf — the beta Caddy vhost, kept here for reference.
  • .env.example — copy to .env on beta. .env is gitignored; never commit OAuth secrets or the admin password.

Rules

  • The repo is the only durable path. Dashboards and alerting are provisioned from this directory at startup; edits made through Grafana's UI or API are overwritten on the next provision. Changes go through a PR.
  • Every artifact that ships to a node is CI-validated (.forgejo/workflows/observability-validate.yml, at the repo root): both compose files parse, all dashboard JSON is valid, the Loki and provisioning YAML parse, and the Alloy config is checked with the pinned alloy binary (v1.9.1, matching what the fleet runs).
  • Alerting gets a stricter third step. Every rule's condition must name a refId that exists, every policy must route to a receiver that exists, and a literal Discord webhook URL in the repo is a hard failure. A rule pointing at a missing refId is valid YAML, provisions cleanly, and then never fires — that silent no-op is what this check exists to prevent.
  • Alerts go to Discord #ops-alerts, never email. Beta's Grafana relays SMTP through prod's Postfix, so email dies exactly when prod does.
  • Every rule is LogQL — there is no Prometheus anywhere in the fleet. Thresholds were derived from real Loki data and the working is in the comments beside each rule; re-derive before changing a number rather than guessing.
  • Secrets (OAuth client id/secret, admin password) live only in the host .env.

Commits & PRs

Concise and technical. Never mention AI, assistants or automated authorship.