All checks were successful
secrets-guard / encrypted (pull_request) Successful in 10s
PR build (required check) / changes (pull_request) Successful in 17s
shell-lint / shellcheck (pull_request) Successful in 14s
PR build (required check) / validate-observability (pull_request) Successful in 44s
PR build (required check) / build-frontend (pull_request) Successful in 2m13s
PR build (required check) / build-backend (pull_request) Successful in 2m31s
PR build (required check) / gate (pull_request) Successful in 3s
The four domain CLAUDE.md files still described the pre-monorepo split-repo topology, and several statements were the exact inverse of current reality: infra/CLAUDE.md told a reader that emi/thermograph is archived and must not be pointed at, when that is the live repo; backend/ and frontend/ both claimed there is no Makefile and no requirements-dev.txt when all four files exist; frontend/ described itself as one of four sibling repos; observability/ claimed a single unprotected main branch. These files are read before every change, so a stale one is a correctness problem rather than a documentation one. Rewritten against the tree: - Root CLAUDE.md now owns the cross-cutting truth once — branch model, the SERVICE + *_IMAGE_TAG deploy contract, image names, which orchestrator each environment runs, and that everything is a PR. Domain files carry only what differs and are capped at ~70 lines. - Records that prod runs Swarm from deploy/stack/thermograph-stack.yml while beta and LAN dev run compose, routed by /etc/thermograph/deploy-mode. - Documents the *-deploy-dev.yml workflows as inert rather than leaving a reader to discover it. Deletes infra/docker-stack.yml. It defined db/app/worker, matched nothing that deploys, and was referenced only by prose describing it as a future design record — while the real 8-service prod stack lives under deploy/stack/. A file that looks authoritative and affects nothing is the worst case for a reader asked to change the prod stack. infra/README.md claimed "compose in production today" and that Swarm was "not currently live"; corrected, and infra-sync.yml's existence is now recorded instead of "there is no separate infra deploy trigger". The compose file's timescale-pin comment now points at how deploy-stack.sh actually resolves the digest from the running container.
52 lines
2.8 KiB
Markdown
52 lines
2.8 KiB
Markdown
# observability/ — agent instructions
|
|
|
|
The **logging stack** for the fleet: Loki + Grafana on beta, with a Grafana Alloy
|
|
agent on every node (prod, beta, dev) shipping container and app logs over the
|
|
WireGuard mesh. Grafana is fronted by beta's Caddy at
|
|
**`dashboard.thermograph.org`** (Google SSO, pre-provisioned users only) — use
|
|
that hostname everywhere, never `grafana.thermograph.org`.
|
|
|
|
This is operational config, not application code: there is **no build** and no
|
|
deploy automation. It ships by hand — `docker compose up -d` on beta for
|
|
Loki+Grafana, and `alloy/docker-compose.agent.yml` on each node. Read the root
|
|
`CLAUDE.md` first; the branch model there applies here too.
|
|
|
|
## Layout
|
|
|
|
- `docker-compose.yml` — the Loki + Grafana stack (beta).
|
|
- `loki/config.yml` — mesh-only, filesystem storage.
|
|
- `grafana/provisioning/` — datasource + dashboard provider, auto-loaded at
|
|
startup. `grafana/dashboards/*.json` — the dashboards.
|
|
- `grafana/provisioning/alerting/` — alert rules, the Discord contact point, and
|
|
the notification policy.
|
|
- `alloy/config.alloy` + `alloy/docker-compose.agent.yml` — the per-node shipper.
|
|
Node name comes from `ALLOY_NODE`, set per host.
|
|
- `caddy-grafana.conf` — the beta Caddy vhost, kept here for reference.
|
|
- `.env.example` — copy to `.env` on beta. `.env` is gitignored; never commit
|
|
OAuth secrets or the admin password.
|
|
|
|
## Rules
|
|
|
|
- **The repo is the only durable path.** Dashboards and alerting are provisioned
|
|
from this directory at startup; edits made through Grafana's UI or API are
|
|
overwritten on the next provision. Changes go through a PR.
|
|
- **Every artifact that ships to a node is CI-validated**
|
|
(`.forgejo/workflows/observability-validate.yml`, at the repo root): both
|
|
compose files parse, all dashboard JSON is valid, the Loki and provisioning
|
|
YAML parse, and the Alloy config is checked with the pinned `alloy` binary
|
|
(v1.9.1, matching what the fleet runs).
|
|
- **Alerting gets a stricter third step.** Every rule's `condition` must name a
|
|
refId that exists, every policy must route to a receiver that exists, and a
|
|
literal Discord webhook URL in the repo is a hard failure. A rule pointing at a
|
|
missing refId is valid YAML, provisions cleanly, and then never fires — that
|
|
silent no-op is what this check exists to prevent.
|
|
- **Alerts go to Discord `#ops-alerts`, never email.** Beta's Grafana relays SMTP
|
|
through prod's Postfix, so email dies exactly when prod does.
|
|
- **Every rule is LogQL** — there is no Prometheus anywhere in the fleet.
|
|
Thresholds were derived from real Loki data and the working is in the comments
|
|
beside each rule; re-derive before changing a number rather than guessing.
|
|
- Secrets (OAuth client id/secret, admin password) live only in the host `.env`.
|
|
|
|
## Commits & PRs
|
|
|
|
Concise and technical. Never mention AI, assistants or automated authorship.
|