All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / validate-observability (pull_request) Successful in 23s
PR build (required check) / gate (pull_request) Successful in 2s
thermograph-observability (and thermograph-backend/frontend/infra) are archived on Forgejo now that the monorepo cutover has landed, so this domain's work lands here in observability/ instead of a standalone repo PR. Alloy (observability/alloy/config.alloy), deployed from /opt/observability on prod and beta (confirmed via docker inspect mounts on the live hosts; the monorepo's own observability/ copy isn't what's running yet -- that's a separate, not-yet-done cutover step for this domain, flagged below): - Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from the 250ms/250ms PollingFileWatcher default to 2s/10s, and local.file_match's glob-rescan sync_period from 10s to 1m. - Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task churn (~every 90s) stops restarting tailers on every refresh. - Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*, thermograph-lb, thermograph_worker) via a relabel rule; the worker's stdout is pure /healthz noise, already superseded by its own access/*.jsonl. - Fix loki.source.docker double-relabeling: it took both the already- relabelled target output AND relabel_rules, running every rule twice per entry and making the new drop rules a no-op on the second pass. Pass raw discovery targets + relabel_rules instead. - Add a loki.process stage for Caddy logs that drops well-known crawler/bot lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot, petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older (beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth addressing separately), so this is the shipper-side equivalent. Loki (observability/loki/config.yml): 30 low-rate streams were hitting the 30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98% full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream limits sized for a three-node hobby fleet -- their absence was the source of the sporadic 429s (Loki was falling back to its stricter multi-tenant defaults). CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no syntax coverage. Download the pinned v1.9.1 alloy binary (matches observability/alloy/docker-compose.agent.yml) straight from its GitHub release and run `alloy validate` (minus the one line covering a known gap where validate doesn't recognize the top-level livedebugging singleton block, unlike alloy run). Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl, the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile is the documented manual-apply fallback, updated in step to avoid drift): - health_interval 5s -> 15s: a quarter of the active-healthcheck polling load for the same restart-safety guarantee. - A `filter` format log encoder deleting request>headers, request>tls, and resp_headers (the full-header/TLS-block serialization was the main driver of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the file itself doesn't grow unbounded between logrotate runs. - Strip the query string entirely from logged request URIs via a `regexp` field filter -- Caddy's default logger recorded request.uri *including* the query string, so every `?q=<search text>` sat in Loki next to the client IP for the full 30-day retention. Confirmed empirically (via `caddy validate` and a scratch `caddy run` against beta's real 2.6.2 binary) that keeping request>headers>User-Agent alongside a parent-level header delete isn't achievable in this Caddy version -- deleting the parent drops the whole subtree regardless of a child override, and `log_append` (which would let it be hoisted out first) needs a newer Caddy than beta runs. Traded away rather than block on it. Deferred (host-level, not in any repo, applied directly and separately): /etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts -- config written, but NOT restarting dockerd on either: prod is the fleet's sole Swarm manager with Live Restore disabled, so a daemon restart there would both interrupt Swarm's control plane and actually stop/restart every container on the host (no live-restore to keep them up), a real production outage window, not a blip. Left for a planned maintenance window instead. A 14-day retention cron for the app's applogs JSONL volume is applied on both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete rather than logrotate, since the app already self-dates one file per day per category and logrotate's rotate-in-place model doesn't fit files it doesn't own the naming of. Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf, dashboard.thermograph.org's Caddy block) is untouched by any of this.
140 lines
5.4 KiB
Text
140 lines
5.4 KiB
Text
// Grafana Alloy — the log collector that runs on EVERY node (prod, beta, LAN
|
|
// dev). It gathers three sources and ships them to the central Loki on beta:
|
|
//
|
|
// 1. Every Docker container's stdout/stderr (app, db, and on beta forgejo)
|
|
// via the Docker socket.
|
|
// 2. Caddy's host access logs (/var/log/caddy/*.log) — the reverse proxy runs
|
|
// on the host, not in a container, so its logs aren't in Docker.
|
|
// 3. The app's structured JSON logs (errors/access/audit *.jsonl) from the
|
|
// `applogs` Docker volume, parsed so `level`/`tag`/`phase` become labels.
|
|
//
|
|
// Per-node settings come from the environment (see docker-compose.agent.yml):
|
|
// ALLOY_NODE — this node's name label (prod | beta | dev)
|
|
// LOKI_URL — where to push (http://10.10.0.2:3100/loki/api/v1/push over wg0)
|
|
|
|
livedebugging { enabled = false }
|
|
|
|
// --- 1. All Docker container logs ------------------------------------------------
|
|
// refresh_interval defaults to 60s; Swarm task churn reshuffles the target set on
|
|
// roughly that cadence, which restarts tailers ~every 90s and was costing ~11% of
|
|
// Alloy's own CPU in tailer restarts alone. 5m is still fast enough to pick up a
|
|
// real deploy without paying that churn cost.
|
|
discovery.docker "containers" {
|
|
host = "unix:///var/run/docker.sock"
|
|
refresh_interval = "5m"
|
|
}
|
|
|
|
// Turn Docker metadata into tidy labels: `container` (short name) and `service`
|
|
// (the compose service, e.g. app/db). Drop noise/duplicate containers so Loki never
|
|
// ingests them: Alloy itself (loop), the autoscaler, throwaway `thermograph-test_*`
|
|
// stacks, the loopback LB bridge (`thermograph-lb`, pure plumbing, nothing to debug
|
|
// from its logs), and the app's own worker (`thermograph_worker`'s stdout is 100%
|
|
// `/healthz` poll noise — the app's `access/*.jsonl` under source #3 is a strict
|
|
// superset of anything useful it logs).
|
|
discovery.relabel "containers" {
|
|
targets = discovery.docker.containers.targets
|
|
|
|
rule {
|
|
source_labels = ["__meta_docker_container_name"]
|
|
regex = "/(.*)"
|
|
target_label = "container"
|
|
}
|
|
rule {
|
|
source_labels = ["__meta_docker_container_label_com_docker_compose_service"]
|
|
target_label = "service"
|
|
}
|
|
rule {
|
|
source_labels = ["container"]
|
|
regex = "(alloy|autoscaler|thermograph-test_.*|thermograph-lb|thermograph_worker).*"
|
|
action = "drop"
|
|
}
|
|
}
|
|
|
|
// Pass RAW targets here (not discovery.relabel.containers.output) alongside
|
|
// relabel_rules: loki.source.docker applies relabel_rules itself, once, using the
|
|
// __meta_docker_* metadata it still holds at collection time. Passing the
|
|
// already-relabelled output *and* relabel_rules ran the same rules twice per log
|
|
// entry, and made the drop rules above a no-op on the second pass since the
|
|
// __meta_docker_* labels are already gone from the pre-relabelled output.
|
|
loki.source.docker "containers" {
|
|
host = "unix:///var/run/docker.sock"
|
|
targets = discovery.docker.containers.targets
|
|
forward_to = [loki.write.central.receiver]
|
|
relabel_rules = discovery.relabel.containers.rules
|
|
labels = { job = "docker" }
|
|
}
|
|
|
|
// --- 2. Caddy host access logs ---------------------------------------------------
|
|
// sync_period (glob rescan) defaults to 10s; 1m is plenty for a log file that only
|
|
// appears/rotates on the order of hours.
|
|
local.file_match "caddy" {
|
|
path_targets = [{ __path__ = "/var/log/caddy/*.log", job = "caddy" }]
|
|
sync_period = "1m"
|
|
}
|
|
|
|
loki.source.file "caddy" {
|
|
targets = local.file_match.caddy.targets
|
|
forward_to = [loki.process.caddy.receiver]
|
|
|
|
// PollingFileWatcher defaults (250ms/250ms) stat every tailed file 4x/second
|
|
// forever. Caddy's access log doesn't need sub-second latency into Loki.
|
|
file_watch {
|
|
min_poll_frequency = "2s"
|
|
max_poll_frequency = "10s"
|
|
}
|
|
}
|
|
|
|
// Drop well-known crawler/bot traffic before it hits Loki. Caddy itself can't do
|
|
// this cheaply (log_skip needs Caddy >= 2.7; both hosts run older Caddy, and an
|
|
// upgrade is out of scope here), so filter it at the shipper instead.
|
|
loki.process "caddy" {
|
|
forward_to = [loki.write.central.receiver]
|
|
|
|
stage.drop {
|
|
expression = "(?i)(semrushbot|claudebot|ahrefsbot|yandexbot|bytespider|mj12bot|petalbot)"
|
|
drop_counter_reason = "crawler"
|
|
}
|
|
}
|
|
|
|
// --- 3. App structured JSON logs (errors / access / audit) -----------------------
|
|
// Mounted read-only from the app's `applogs` volume at /applogs (see the agent
|
|
// compose). Lift `level`/`tag`/`phase` out of the JSON so they're queryable.
|
|
local.file_match "app_jsonl" {
|
|
path_targets = [{ __path__ = "/applogs/**/*.jsonl", job = "app-json" }]
|
|
sync_period = "1m"
|
|
}
|
|
|
|
loki.source.file "app_jsonl" {
|
|
targets = local.file_match.app_jsonl.targets
|
|
forward_to = [loki.process.app_jsonl.receiver]
|
|
|
|
file_watch {
|
|
min_poll_frequency = "2s"
|
|
max_poll_frequency = "10s"
|
|
}
|
|
}
|
|
|
|
loki.process "app_jsonl" {
|
|
forward_to = [loki.write.central.receiver]
|
|
|
|
stage.json {
|
|
expressions = { level = "level", tag = "tag", phase = "phase", status = "status" }
|
|
}
|
|
// A record with no explicit level: an error-folder line is an error, else info.
|
|
stage.static_labels {
|
|
values = { source = "app" }
|
|
}
|
|
stage.labels {
|
|
values = { level = "", tag = "", phase = "" }
|
|
}
|
|
}
|
|
|
|
// --- Ship to the central Loki over the WireGuard mesh ----------------------------
|
|
loki.write "central" {
|
|
endpoint {
|
|
url = sys.env("LOKI_URL")
|
|
}
|
|
// Every line from this node is stamped with its node name, so one Grafana
|
|
// view can slice prod vs beta vs dev.
|
|
external_labels = { host = sys.env("ALLOY_NODE") }
|
|
}
|