Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / validate-observability (pull_request) Successful in 23s
PR build (required check) / gate (pull_request) Successful in 2s

thermograph-observability (and thermograph-backend/frontend/infra) are
archived on Forgejo now that the monorepo cutover has landed, so this domain's
work lands here in observability/ instead of a standalone repo PR.

Alloy (observability/alloy/config.alloy), deployed from /opt/observability on
prod and beta (confirmed via docker inspect mounts on the live hosts; the
monorepo's own observability/ copy isn't what's running yet -- that's a
separate, not-yet-done cutover step for this domain, flagged below):

- Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from
  the 250ms/250ms PollingFileWatcher default to 2s/10s, and
  local.file_match's glob-rescan sync_period from 10s to 1m.
- Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task
  churn (~every 90s) stops restarting tailers on every refresh.
- Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*,
  thermograph-lb, thermograph_worker) via a relabel rule; the worker's
  stdout is pure /healthz noise, already superseded by its own
  access/*.jsonl.
- Fix loki.source.docker double-relabeling: it took both the already-
  relabelled target output AND relabel_rules, running every rule twice per
  entry and making the new drop rules a no-op on the second pass. Pass raw
  discovery targets + relabel_rules instead.
- Add a loki.process stage for Caddy logs that drops well-known crawler/bot
  lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot,
  petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older
  (beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth
  addressing separately), so this is the shipper-side equivalent.

Loki (observability/loki/config.yml): 30 low-rate streams were hitting the
30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98%
full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set
chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream
limits sized for a three-node hobby fleet -- their absence was the source
of the sporadic 429s (Loki was falling back to its stricter multi-tenant
defaults).

CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no
syntax coverage. Download the pinned v1.9.1 alloy binary (matches
observability/alloy/docker-compose.agent.yml) straight from its GitHub
release and run `alloy validate` (minus the one line covering a known gap
where validate doesn't recognize the top-level livedebugging singleton
block, unlike alloy run).

Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl,
the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose
deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile
is the documented manual-apply fallback, updated in step to avoid drift):

- health_interval 5s -> 15s: a quarter of the active-healthcheck polling
  load for the same restart-safety guarantee.
- A `filter` format log encoder deleting request>headers, request>tls, and
  resp_headers (the full-header/TLS-block serialization was the main driver
  of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the
  file itself doesn't grow unbounded between logrotate runs.
- Strip the query string entirely from logged request URIs via a `regexp`
  field filter -- Caddy's default logger recorded request.uri *including*
  the query string, so every `?q=<search text>` sat in Loki next to the
  client IP for the full 30-day retention. Confirmed empirically (via
  `caddy validate` and a scratch `caddy run` against beta's real 2.6.2
  binary) that keeping request>headers>User-Agent alongside a parent-level
  header delete isn't achievable in this Caddy version -- deleting the
  parent drops the whole subtree regardless of a child override, and
  `log_append` (which would let it be hoisted out first) needs a newer
  Caddy than beta runs. Traded away rather than block on it.

Deferred (host-level, not in any repo, applied directly and separately):
/etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts --
config written, but NOT restarting dockerd on either: prod is the fleet's
sole Swarm manager with Live Restore disabled, so a daemon restart there
would both interrupt Swarm's control plane and actually stop/restart every
container on the host (no live-restore to keep them up), a real production
outage window, not a blip. Left for a planned maintenance window instead.
A 14-day retention cron for the app's applogs JSONL volume is applied on
both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete
rather than logrotate, since the app already self-dates one file per day
per category and logrotate's rotate-in-place model doesn't fit files it
doesn't own the naming of.

Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf,
dashboard.thermograph.org's Caddy block) is untouched by any of this.
This commit is contained in:
Emi Griffith 2026-07-23 21:34:06 -07:00
parent eaaa4f1ad9
commit 9c0f493d68
5 changed files with 146 additions and 14 deletions

View file

@ -12,10 +12,13 @@ name: Validate observability stack
# path-filtered to the domain, and workflow_call lets pr-build.yml reuse this # path-filtered to the domain, and workflow_call lets pr-build.yml reuse this
# as the domain's PR check under its single `gate` required status. # as the domain's PR check under its single `gate` required status.
# #
# Not validated here: the Alloy config (observability/alloy/config.alloy, an # The Alloy config (observability/alloy/config.alloy, an HCL-like format) IS
# HCL-like format, needs the `alloy` binary) and docker-compose *schema* # validated below -- the static `alloy` binary is downloaded straight from its
# (unknown-key checks, which would need the docker CLI). Add either if the # GitHub release (pinned to the same v1.9.1 the fleet runs; see
# churn warrants the extra tooling. # observability/alloy/docker-compose.agent.yml), no docker CLI needed.
#
# Not validated here: docker-compose *schema* (unknown-key checks, which would
# need the docker CLI). Add it if the churn warrants the extra tooling.
on: on:
workflow_call: {} workflow_call: {}
@ -70,3 +73,21 @@ jobs:
print("INVALID YAML", f, "-", e); bad = True print("INVALID YAML", f, "-", e); bad = True
sys.exit(1 if bad else 0) sys.exit(1 if bad else 0)
PY PY
- name: Alloy config parses and validates (components, relabel/process wiring)
run: |
set -euo pipefail
apt-get update -qq && apt-get install -y -qq unzip >/dev/null
curl -sSL -o /tmp/alloy.zip \
https://github.com/grafana/alloy/releases/download/v1.9.1/alloy-linux-amd64.zip
unzip -q /tmp/alloy.zip -d /tmp/alloybin
chmod +x /tmp/alloybin/alloy-linux-amd64
# `alloy validate` builds the real component graph (catches bad field
# names, dangling forward_to/receiver refs, malformed relabel/process
# stages) but doesn't recognize the top-level `livedebugging {}` singleton
# block as a component -- it only knows named components, even though
# `alloy run` loads that block fine. Known gap in the `validate`
# subcommand, not a config error, so strip that one line before checking.
grep -v '^livedebugging ' observability/alloy/config.alloy > /tmp/config-for-validate.alloy
/tmp/alloybin/alloy-linux-amd64 validate /tmp/config-for-validate.alloy

View file

@ -31,10 +31,12 @@ thermograph.org {
# Active health check on the same cheap /healthz route each container's own # Active health check on the same cheap /healthz route each container's own
# HEALTHCHECK uses (Dockerfile) — so a deploy that's still restarting/booting # HEALTHCHECK uses (Dockerfile) — so a deploy that's still restarting/booting
# never gets proxied into (a reload alone has no gate, hop-1 runbook hazard #10). # never gets proxied into (a reload alone has no gate, hop-1 runbook hazard #10).
# 15s (was 5s): plenty responsive for a process that only restarts on a deploy,
# and a quarter of the polling load.
handle @backend_paths { handle @backend_paths {
reverse_proxy 127.0.0.1:8137 { reverse_proxy 127.0.0.1:8137 {
health_uri /healthz health_uri /healthz
health_interval 5s health_interval 15s
health_timeout 3s health_timeout 3s
health_status 2xx health_status 2xx
} }
@ -43,14 +45,35 @@ thermograph.org {
handle { handle {
reverse_proxy 127.0.0.1:8080 { reverse_proxy 127.0.0.1:8080 {
health_uri /healthz health_uri /healthz
health_interval 5s health_interval 15s
health_timeout 3s health_timeout 3s
health_status 2xx health_status 2xx
} }
} }
# Access-log hygiene: the default JSON encoder serializes full request headers,
# the TLS block, and response headers on every line (measured ~1,133B/line) --
# strip those with the `filter` format encoder. Also strip the query string from
# the logged URI: Caddy's default logger records request.uri *including* the
# query string, so every `?q=<search text>` a visitor typed sat in Loki next to
# their client IP for the full 30-day retention -- a real privacy leak, not just
# noise. Bot/crawler skipping stays out of here: `log_skip` needs Caddy >= 2.7
# and an upgrade is out of scope, so that's handled downstream in Alloy's
# loki.process "caddy" stage instead (see observability/alloy/config.alloy).
log { log {
output file /var/log/caddy/thermograph.log output file /var/log/caddy/thermograph.log {
roll_size 20MiB
roll_keep 5
}
format filter {
wrap json
fields {
request>headers delete
request>tls delete
resp_headers delete
request>uri regexp \?.* ""
}
}
} }
} }

View file

@ -12,6 +12,16 @@
# #
# NOTE: the repo's deploy/Caddyfile additionally serves the emigriffith.dev portfolio # NOTE: the repo's deploy/Caddyfile additionally serves the emigriffith.dev portfolio
# and legacy redirects; those are host-specific and intentionally not templated here. # and legacy redirects; those are host-specific and intentionally not templated here.
#
# Access-log hygiene: the default JSON encoder serializes full request headers, the
# TLS block, and response headers on every line (measured ~1,133B/line on prod,
# ~922B on beta) -- strip those with the `filter` format encoder. Also strip the
# query string from the logged URI: Caddy's default logger records request.uri
# *including* the query string, so every `?q=<search text>` a visitor typed sat in
# Loki next to their client IP for the full 30-day retention -- a real privacy
# leak, not just noise. Bot/crawler skipping stays out of here: `log_skip` needs
# Caddy >= 2.7 and an upgrade is out of scope, so that's handled downstream in
# Alloy's loki.process "caddy" stage instead (see observability/alloy/config.alloy).
${domain} { ${domain} {
encode zstd gzip encode zstd gzip
@ -21,10 +31,12 @@ ${domain} {
# HEALTHCHECK uses (Dockerfile) — so a `docker compose up -d --build` deploy # HEALTHCHECK uses (Dockerfile) — so a `docker compose up -d --build` deploy
# that's still restarting/booting never gets proxied into (a reload alone has # that's still restarting/booting never gets proxied into (a reload alone has
# no gate, hop-1 runbook hazard #10). health_uri is relative to the upstream. # no gate, hop-1 runbook hazard #10). health_uri is relative to the upstream.
# 15s (was 5s): plenty responsive for an active health check against a process
# that only restarts on a deploy, and a quarter of the polling load.
handle @backend_paths { handle @backend_paths {
reverse_proxy 127.0.0.1:${port} { reverse_proxy 127.0.0.1:${port} {
health_uri /healthz health_uri /healthz
health_interval 5s health_interval 15s
health_timeout 3s health_timeout 3s
health_status 2xx health_status 2xx
} }
@ -33,13 +45,25 @@ ${domain} {
handle { handle {
reverse_proxy 127.0.0.1:${frontend_port} { reverse_proxy 127.0.0.1:${frontend_port} {
health_uri /healthz health_uri /healthz
health_interval 5s health_interval 15s
health_timeout 3s health_timeout 3s
health_status 2xx health_status 2xx
} }
} }
log { log {
output file /var/log/caddy/thermograph.log output file /var/log/caddy/thermograph.log {
roll_size 20MiB
roll_keep 5
}
format filter {
wrap json
fields {
request>headers delete
request>tls delete
resp_headers delete
request>uri regexp \?.* ""
}
}
} }
} }

View file

@ -15,12 +15,22 @@
livedebugging { enabled = false } livedebugging { enabled = false }
// --- 1. All Docker container logs ------------------------------------------------ // --- 1. All Docker container logs ------------------------------------------------
// refresh_interval defaults to 60s; Swarm task churn reshuffles the target set on
// roughly that cadence, which restarts tailers ~every 90s and was costing ~11% of
// Alloy's own CPU in tailer restarts alone. 5m is still fast enough to pick up a
// real deploy without paying that churn cost.
discovery.docker "containers" { discovery.docker "containers" {
host = "unix:///var/run/docker.sock" host = "unix:///var/run/docker.sock"
refresh_interval = "5m"
} }
// Turn Docker metadata into tidy labels: `container` (short name) and `service` // Turn Docker metadata into tidy labels: `container` (short name) and `service`
// (the compose service, e.g. app/db). Drop Alloy's own container to avoid a loop. // (the compose service, e.g. app/db). Drop noise/duplicate containers so Loki never
// ingests them: Alloy itself (loop), the autoscaler, throwaway `thermograph-test_*`
// stacks, the loopback LB bridge (`thermograph-lb`, pure plumbing, nothing to debug
// from its logs), and the app's own worker (`thermograph_worker`'s stdout is 100%
// `/healthz` poll noise — the app's `access/*.jsonl` under source #3 is a strict
// superset of anything useful it logs).
discovery.relabel "containers" { discovery.relabel "containers" {
targets = discovery.docker.containers.targets targets = discovery.docker.containers.targets
@ -35,27 +45,55 @@ discovery.relabel "containers" {
} }
rule { rule {
source_labels = ["container"] source_labels = ["container"]
regex = ".*alloy.*" regex = "(alloy|autoscaler|thermograph-test_.*|thermograph-lb|thermograph_worker).*"
action = "drop" action = "drop"
} }
} }
// Pass RAW targets here (not discovery.relabel.containers.output) alongside
// relabel_rules: loki.source.docker applies relabel_rules itself, once, using the
// __meta_docker_* metadata it still holds at collection time. Passing the
// already-relabelled output *and* relabel_rules ran the same rules twice per log
// entry, and made the drop rules above a no-op on the second pass since the
// __meta_docker_* labels are already gone from the pre-relabelled output.
loki.source.docker "containers" { loki.source.docker "containers" {
host = "unix:///var/run/docker.sock" host = "unix:///var/run/docker.sock"
targets = discovery.relabel.containers.output targets = discovery.docker.containers.targets
forward_to = [loki.write.central.receiver] forward_to = [loki.write.central.receiver]
relabel_rules = discovery.relabel.containers.rules relabel_rules = discovery.relabel.containers.rules
labels = { job = "docker" } labels = { job = "docker" }
} }
// --- 2. Caddy host access logs --------------------------------------------------- // --- 2. Caddy host access logs ---------------------------------------------------
// sync_period (glob rescan) defaults to 10s; 1m is plenty for a log file that only
// appears/rotates on the order of hours.
local.file_match "caddy" { local.file_match "caddy" {
path_targets = [{ __path__ = "/var/log/caddy/*.log", job = "caddy" }] path_targets = [{ __path__ = "/var/log/caddy/*.log", job = "caddy" }]
sync_period = "1m"
} }
loki.source.file "caddy" { loki.source.file "caddy" {
targets = local.file_match.caddy.targets targets = local.file_match.caddy.targets
forward_to = [loki.process.caddy.receiver]
// PollingFileWatcher defaults (250ms/250ms) stat every tailed file 4x/second
// forever. Caddy's access log doesn't need sub-second latency into Loki.
file_watch {
min_poll_frequency = "2s"
max_poll_frequency = "10s"
}
}
// Drop well-known crawler/bot traffic before it hits Loki. Caddy itself can't do
// this cheaply (log_skip needs Caddy >= 2.7; both hosts run older Caddy, and an
// upgrade is out of scope here), so filter it at the shipper instead.
loki.process "caddy" {
forward_to = [loki.write.central.receiver] forward_to = [loki.write.central.receiver]
stage.drop {
expression = "(?i)(semrushbot|claudebot|ahrefsbot|yandexbot|bytespider|mj12bot|petalbot)"
drop_counter_reason = "crawler"
}
} }
// --- 3. App structured JSON logs (errors / access / audit) ----------------------- // --- 3. App structured JSON logs (errors / access / audit) -----------------------
@ -63,11 +101,17 @@ loki.source.file "caddy" {
// compose). Lift `level`/`tag`/`phase` out of the JSON so they're queryable. // compose). Lift `level`/`tag`/`phase` out of the JSON so they're queryable.
local.file_match "app_jsonl" { local.file_match "app_jsonl" {
path_targets = [{ __path__ = "/applogs/**/*.jsonl", job = "app-json" }] path_targets = [{ __path__ = "/applogs/**/*.jsonl", job = "app-json" }]
sync_period = "1m"
} }
loki.source.file "app_jsonl" { loki.source.file "app_jsonl" {
targets = local.file_match.app_jsonl.targets targets = local.file_match.app_jsonl.targets
forward_to = [loki.process.app_jsonl.receiver] forward_to = [loki.process.app_jsonl.receiver]
file_watch {
min_poll_frequency = "2s"
max_poll_frequency = "10s"
}
} }
loki.process "app_jsonl" { loki.process "app_jsonl" {

View file

@ -32,6 +32,17 @@ schema_config:
prefix: index_ prefix: index_
period: 24h period: 24h
# 30 low-rate streams almost always hit chunk_idle_period (30m default) long before
# they'd ever fill chunk_target_size, so chunks were flushed 1.98% full on average
# (972 near-empty chunks for 30MB of actual log data). Give idle streams far longer
# to accumulate before an idle flush, and cap age/size so a chunk still can't grow
# unbounded.
ingester:
chunk_idle_period: 2h
max_chunk_age: 12h
chunk_target_size: 1572864
chunk_encoding: snappy
limits_config: limits_config:
# A hobby fleet's volume is tiny; keep 30 days and cap ingestion generously. # A hobby fleet's volume is tiny; keep 30 days and cap ingestion generously.
retention_period: 720h retention_period: 720h
@ -40,6 +51,15 @@ limits_config:
max_query_series: 5000 max_query_series: 5000
allow_structured_metadata: true allow_structured_metadata: true
volume_enabled: true volume_enabled: true
# No explicit ingestion/stream limits meant Loki fell back to its (much stricter)
# built-in defaults, which is why 25 pushes came back HTTP 429 with nothing in
# this file explaining why. These are sized for a three-node hobby fleet, not the
# multi-tenant defaults.
ingestion_rate_mb: 8
ingestion_burst_size_mb: 16
per_stream_rate_limit: 3MB
per_stream_rate_limit_burst: 10MB
max_global_streams_per_user: 1000
compactor: compactor:
working_directory: /loki/compactor working_directory: /loki/compactor