|
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / validate-observability (pull_request) Successful in 23s
PR build (required check) / gate (pull_request) Successful in 2s
thermograph-observability (and thermograph-backend/frontend/infra) are archived on Forgejo now that the monorepo cutover has landed, so this domain's work lands here in observability/ instead of a standalone repo PR. Alloy (observability/alloy/config.alloy), deployed from /opt/observability on prod and beta (confirmed via docker inspect mounts on the live hosts; the monorepo's own observability/ copy isn't what's running yet -- that's a separate, not-yet-done cutover step for this domain, flagged below): - Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from the 250ms/250ms PollingFileWatcher default to 2s/10s, and local.file_match's glob-rescan sync_period from 10s to 1m. - Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task churn (~every 90s) stops restarting tailers on every refresh. - Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*, thermograph-lb, thermograph_worker) via a relabel rule; the worker's stdout is pure /healthz noise, already superseded by its own access/*.jsonl. - Fix loki.source.docker double-relabeling: it took both the already- relabelled target output AND relabel_rules, running every rule twice per entry and making the new drop rules a no-op on the second pass. Pass raw discovery targets + relabel_rules instead. - Add a loki.process stage for Caddy logs that drops well-known crawler/bot lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot, petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older (beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth addressing separately), so this is the shipper-side equivalent. Loki (observability/loki/config.yml): 30 low-rate streams were hitting the 30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98% full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream limits sized for a three-node hobby fleet -- their absence was the source of the sporadic 429s (Loki was falling back to its stricter multi-tenant defaults). CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no syntax coverage. Download the pinned v1.9.1 alloy binary (matches observability/alloy/docker-compose.agent.yml) straight from its GitHub release and run `alloy validate` (minus the one line covering a known gap where validate doesn't recognize the top-level livedebugging singleton block, unlike alloy run). Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl, the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile is the documented manual-apply fallback, updated in step to avoid drift): - health_interval 5s -> 15s: a quarter of the active-healthcheck polling load for the same restart-safety guarantee. - A `filter` format log encoder deleting request>headers, request>tls, and resp_headers (the full-header/TLS-block serialization was the main driver of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the file itself doesn't grow unbounded between logrotate runs. - Strip the query string entirely from logged request URIs via a `regexp` field filter -- Caddy's default logger recorded request.uri *including* the query string, so every `?q=<search text>` sat in Loki next to the client IP for the full 30-day retention. Confirmed empirically (via `caddy validate` and a scratch `caddy run` against beta's real 2.6.2 binary) that keeping request>headers>User-Agent alongside a parent-level header delete isn't achievable in this Caddy version -- deleting the parent drops the whole subtree regardless of a child override, and `log_append` (which would let it be hoisted out first) needs a newer Caddy than beta runs. Traded away rather than block on it. Deferred (host-level, not in any repo, applied directly and separately): /etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts -- config written, but NOT restarting dockerd on either: prod is the fleet's sole Swarm manager with Live Restore disabled, so a daemon restart there would both interrupt Swarm's control plane and actually stop/restart every container on the host (no live-restore to keep them up), a real production outage window, not a blip. Left for a planned maintenance window instead. A 14-day retention cron for the app's applogs JSONL volume is applied on both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete rather than logrotate, since the app already self-dates one file per day per category and logrotate's rotate-in-place model doesn't fit files it doesn't own the naming of. Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf, dashboard.thermograph.org's Caddy block) is untouched by any of this. |
||
|---|---|---|
| .. | ||
| alloy | ||
| grafana | ||
| loki | ||
| .env.example | ||
| .gitignore | ||
| caddy-grafana.conf | ||
| CLAUDE.md | ||
| docker-compose.yml | ||
| README.md | ||
Thermograph — observability
Fleet-wide log aggregation for Thermograph: a central
Grafana + Loki stack, fed by a Grafana Alloy agent on every node, all
over the private WireGuard mesh. This replaces the old in-repo
scripts/dashboard.py (an SSH-tailed, single-host, metrics-endpoint view) with a
persistent, queryable, fleet-wide log store and UI.
prod (10.10.0.1) beta (10.10.0.2) desktop / LAN dev (10.10.0.3)
┌────────────┐ ┌──────────────────┐ ┌────────────┐
│ Alloy agent│──wg0──┐ │ Loki ◀── Alloy │ ┌──wg0│ Alloy agent│
└────────────┘ └───▶│ Grafana (Caddy) │◀──┘ └────────────┘
docker+caddy+app └──────────────────┘ docker+caddy+app
logs central store + UI logs
- Loki stores logs (filesystem, 30-day retention). Listens on the mesh IP
10.10.0.2:3100— never public. - Grafana is the UI, fronted by beta's Caddy at
dashboard.thermograph.org, login via Google SSO (with a break-glass local admin); Loki is not exposed. - Alloy runs on each node and ships three sources to Loki: every Docker
container's stdout/stderr, Caddy's host access logs, and the app's structured
JSON logs (
errors/access/audit*.jsonl, parsed solevel/tag/phasebecome labels). Every line is tagged with its node (host = prod|beta|dev).
Layout
| Path | What |
|---|---|
docker-compose.yml |
the central Loki + Grafana stack (runs on beta) |
loki/config.yml |
Loki config (filesystem, retention) |
grafana/provisioning/ |
auto-wired Loki datasource + dashboard provider |
grafana/dashboards/thermograph-logs.json |
the fleet-logs dashboard |
alloy/config.alloy |
the per-node collector config |
alloy/docker-compose.agent.yml |
runs the Alloy agent on a node |
caddy-grafana.conf |
beta Caddy site block for the Grafana UI |
Deploy
1. Central stack (on beta)
# on beta (75.119.132.91):
git clone <this repo> thermograph-observability && cd thermograph-observability
cp .env.example .env # set GF_SECURITY_ADMIN_PASSWORD
docker compose up -d # Loki on 10.10.0.2:3100 + 127.0.0.1:3100; Grafana on 127.0.0.1:3000
Expose the Grafana UI: point dashboard.thermograph.org at beta, append
caddy-grafana.conf to /etc/caddy/Caddyfile, systemctl reload caddy.
Google SSO (Grafana + Forgejo)
Both the dashboard and Forgejo log in with Google. Create one Google Cloud OAuth 2.0 Client (type: Web application) and register both redirect URIs:
- Grafana:
https://dashboard.thermograph.org/login/google - Forgejo:
https://git.thermograph.org/user/oauth2/google/callback
Grafana: put the client id/secret in .env, set OAUTH_ENABLED=true, and
docker compose up -d. Login is locked to pre-provisioned accounts
(allow_sign_up=false), so provision your email as an admin once:
# on beta, after the stack is up (uses the break-glass admin creds from .env):
curl -s -u admin:"$GF_SECURITY_ADMIN_PASSWORD" -H 'Content-Type: application/json' \
-X POST http://127.0.0.1:3000/api/admin/users \
-d '{"name":"you","login":"you@gmail.com","email":"you@gmail.com","password":"'"$(openssl rand -base64 24)"'"}'
# then make that user a Grafana admin (org role Admin + server admin) via the UI
# or /api/org/users; the Google login matches on email.
Forgejo: add the Google auth source (once, as a Forgejo admin):
docker exec -u git <forgejo-container> forgejo admin auth add-oauth \
--name google --provider openidConnect \
--key "<CLIENT_ID>" --secret "<CLIENT_SECRET>" \
--auto-discover-url https://accounts.google.com/.well-known/openid-configuration
2. Alloy agent (on every node — prod, beta, desktop)
cd thermograph-observability/alloy
# Every node (beta included) pushes to Loki over the mesh IP — a containerized
# agent can't reach the host's 127.0.0.1, but it can reach beta's wg0 address:
# beta: LOKI_URL=http://10.10.0.2:3100/loki/api/v1/push ALLOY_NODE=beta
# prod: LOKI_URL=http://10.10.0.2:3100/loki/api/v1/push ALLOY_NODE=prod
# desktop: LOKI_URL=http://10.10.0.2:3100/loki/api/v1/push ALLOY_NODE=dev
ALLOY_NODE=prod LOKI_URL=http://10.10.0.2:3100/loki/api/v1/push \
docker compose -f docker-compose.agent.yml up -d
On the LAN dev node the app's log volume is thermograph-dev_applogs (not
thermograph_applogs) — edit the two references in docker-compose.agent.yml
there. The mesh must already be up (see the main repo's deploy/swarm/ /
INFRA.md); Swarm is not required — the agents talk plain HTTP over wg0.
Verify
# From any mesh node, confirm Loki is receiving from every host:
curl -s 'http://10.10.0.2:3100/loki/api/v1/label/host/values' # -> {"data":["beta","dev","prod"]}
# Then open Grafana → the "Thermograph — Fleet Logs" dashboard.
The dashboard
Thermograph — Fleet Logs (auto-provisioned), with a Node selector:
log volume by service, app error rate by tag, upstream 429 count, Caddy 5xx
count, a per-node notifier-liveness indicator (from the heartbeat log), recent app errors, and a live all-container
tail. Edit it in Grafana and re-export the JSON here to version a change.
Notes / follow-ups
- Loki is mesh-only by design; only Grafana (with auth) is public.
- Alloy reads the Docker socket read-only to discover containers.
- Retention is 30 days on beta's disk — bump
loki/config.ymlif you want more. - The app emits a
tag=heartbeatline for the subscription notifier every ~15m; the dashboard's "Notifier alive?" panel turns red if a node misses it. Extend the sameaudit.log_heartbeat()to any future background daemon.