Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping #36

Merged
admin_emi merged 1 commit from infra/logging-optimizations into dev 2026-07-24 04:37:42 +00:00
Owner

What changed

Alloy (observability/alloy/config.alloy):

  • file_watch poll interval 250ms/250ms -> 2s/10s on both loki.source.file blocks; local.file_match sync_period 10s -> 1m.
  • discovery.docker refresh_interval 60s -> 5m (Swarm task churn was restarting tailers ~every 90s).
  • New drop rule for noise/duplicate containers: alloy, autoscaler, thermograph-test_*, thermograph-lb, thermograph_worker.
  • Fixed loki.source.docker double-relabeling (it took both the pre-relabelled target output and relabel_rules, running every rule twice and making the new drop rules a no-op on the second pass).
  • Added a loki.process "caddy" stage that drops known crawler/bot lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot, petalbot) -- the Caddy-side log_skip needs >= 2.7 and both hosts are older.

Loki (observability/loki/config.yml): chunk_idle_period 30m -> 2h, max_chunk_age 12h, explicit chunk_target_size/chunk_encoding; explicit ingestion/stream limits (ingestion_rate_mb, per_stream_rate_limit, max_global_streams_per_user, etc.) sized for a 3-node hobby fleet -- their absence was the source of intermittent 429s.

CI (.forgejo/workflows/observability-validate.yml): added real syntax coverage for alloy/config.alloy -- downloads the pinned v1.9.1 alloy binary from its GitHub release and runs alloy validate (there was none before).

Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl + infra/deploy/Caddyfile): health_interval 5s -> 15s; a filter format log encoder dropping request>headers, request>tls, resp_headers (the main driver of ~1,133B/line on prod, ~922B on beta) plus roll_size/roll_keep; and a regexp field filter that strips the query string entirely from logged request URIs (Caddy's default logger recorded it verbatim -- a real privacy leak, since every ?q=<search text> sat in Loki next to the client IP for 30 days).

Where the true source of truth was found

  • thermograph-observability is archived (confirmed via the Forgejo API -- along with thermograph-backend/frontend/infra) now that the monorepo cutover (see CUTOVER-NOTES.md) has landed. It can still be cloned/viewed but not pushed to, so this domain's work lands here in observability/ instead of a standalone-repo PR as originally planned.
  • Deployed Alloy config: confirmed via docker inspect on prod that the live alloy container's config bind-mount source is /opt/observability/alloy/config.alloy -- a checkout of the now-archived split repo, not this monorepo's observability/ copy. That means this PR fixes the canonical/going-forward copy, but taking effect on hosts needs a follow-up cutover of the observability stack's own deploy (analogous to the already-completed app-side /opt/thermograph swap) -- flagged, not done here, since it's a host-side change outside this PR's scope.
  • Deployed Caddyfile: confirmed on both prod and beta that the live /etc/caddy/Caddyfile's own header says "RENDERED BY TERRAFORM for <domain>", matching Caddyfile.tftpl's template header -- that's the true source, not a hand-edited file. infra/deploy/Caddyfile is a documented manual-apply fallback (see infra/DEPLOY.md) and was updated in step to avoid the two drifting further.

Validated

  • alloy validate (real v1.9.1 binary, downloaded from the GitHub release) against the edited config: passes. (One pre-existing, unrelated tool gap: alloy validate doesn't recognize the top-level livedebugging {} singleton block as a component, even though alloy run loads it fine -- confirmed this is not something my change introduced by validating the original file too. Worked around in both the local check and the new CI step by stripping that one line before validating.)
  • loki -verify-config (real v3.5.1 image, matching what's deployed) against the edited config: passes.
  • caddy validate against the rendered template (via terraform console + the real caddy binary on beta, matching what's actually installed there) and against the edited infra/deploy/Caddyfile: both pass.
  • Query-string stripping was verified against a live scratch caddy run on beta (a temporary process on an unused port, not the real service) -- confirmed the uri field in the resulting JSON log line drops the query string entirely, and confirmed empirically that keeping request>headers>User-Agent alongside a parent-level request>headers delete is not achievable on beta's actual Caddy version (2.6.2): deleting the parent drops the whole subtree regardless of a child override, and the log_append directive that would let User-Agent be hoisted out first isn't available before Caddy 2.8ish. Traded away rather than block on it -- noted in the code comment.
  • Also discovered (worth a flag): prod actually runs Caddy v2.11.4, not 2.6.2 as assumed going in -- only beta is still on 2.6.2. Left both on the Alloy-side crawler drop for uniformity/simplicity per the task, but the version skew itself is worth resolving separately.

Host-level changes applied directly (not in any repo, per scope)

  • /etc/docker/daemon.json ({"log-driver": "json-file", "log-opts": {"max-size": "10m", "max-file": "3"}}) written on both prod and beta. Not restarting dockerd on either -- checked first: prod is the fleet's sole Swarm manager (Managers: 1, this is the Leader) with Live Restore disabled, so a daemon restart there would both interrupt Swarm's control plane and actually stop/restart every container on the host (no live-restore to keep them running) -- a real outage window, not a blip. Deferring both restarts to a planned maintenance window; flagging this explicitly rather than forcing it.
  • /etc/cron.d/thermograph-applogs-retention on both hosts: find /var/lib/docker/volumes/thermograph_applogs/_data -name '*.jsonl' -mtime +14 -delete, daily at 03:00. Used plain cron rather than logrotate (logrotate is set up and active on both hosts for other logs, but the app already self-dates one file per category per day -- e.g. access-2026-07-24.jsonl -- so logrotate's rotate-in-place model doesn't fit files it doesn't own the naming of; a straight find -delete is the correct tool here).
  • Checked service health on both hosts before and after all of the above: all docker service ls replicas 1/1 on prod, and thermograph.org, beta.thermograph.org, git.thermograph.org all return 200, dashboard.thermograph.org its normal 302 (Google OAuth redirect) -- nothing was disrupted.

Grafana

Not touched. No changes to observability/docker-compose.yml, observability/grafana/, observability/caddy-grafana.conf, or dashboard.thermograph.org's Caddy block. Confirmed via git diff --stat that only the five files listed above changed.

Deferred / flagged

  • Docker daemon restarts on both hosts (see above) -- config is in place, restart is not.
  • The observability stack's own host-side cutover (repointing /opt/observability from the archived split repo to track this monorepo's observability/ subtree) -- needed before any of the Alloy/Loki changes here actually take effect in production. Out of scope for this PR.
  • Caddy version skew (prod 2.11.4 vs beta 2.6.2) -- worth its own investigation/alignment pass.
## What changed **Alloy** (`observability/alloy/config.alloy`): - `file_watch` poll interval 250ms/250ms -> 2s/10s on both `loki.source.file` blocks; `local.file_match` `sync_period` 10s -> 1m. - `discovery.docker` `refresh_interval` 60s -> 5m (Swarm task churn was restarting tailers ~every 90s). - New drop rule for noise/duplicate containers: `alloy`, `autoscaler`, `thermograph-test_*`, `thermograph-lb`, `thermograph_worker`. - Fixed `loki.source.docker` double-relabeling (it took both the pre-relabelled target output *and* `relabel_rules`, running every rule twice and making the new drop rules a no-op on the second pass). - Added a `loki.process "caddy"` stage that drops known crawler/bot lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot, petalbot) -- the Caddy-side `log_skip` needs >= 2.7 and both hosts are older. **Loki** (`observability/loki/config.yml`): `chunk_idle_period` 30m -> 2h, `max_chunk_age` 12h, explicit `chunk_target_size`/`chunk_encoding`; explicit ingestion/stream limits (`ingestion_rate_mb`, `per_stream_rate_limit`, `max_global_streams_per_user`, etc.) sized for a 3-node hobby fleet -- their absence was the source of intermittent 429s. **CI** (`.forgejo/workflows/observability-validate.yml`): added real syntax coverage for `alloy/config.alloy` -- downloads the pinned v1.9.1 `alloy` binary from its GitHub release and runs `alloy validate` (there was none before). **Caddy** (`infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl` + `infra/deploy/Caddyfile`): `health_interval` 5s -> 15s; a `filter` format log encoder dropping `request>headers`, `request>tls`, `resp_headers` (the main driver of ~1,133B/line on prod, ~922B on beta) plus `roll_size`/`roll_keep`; and a `regexp` field filter that strips the query string entirely from logged request URIs (Caddy's default logger recorded it verbatim -- a real privacy leak, since every `?q=<search text>` sat in Loki next to the client IP for 30 days). ## Where the true source of truth was found - **thermograph-observability is archived** (confirmed via the Forgejo API -- along with thermograph-backend/frontend/infra) now that the monorepo cutover (see `CUTOVER-NOTES.md`) has landed. It can still be cloned/viewed but not pushed to, so this domain's work lands here in `observability/` instead of a standalone-repo PR as originally planned. - **Deployed Alloy config**: confirmed via `docker inspect` on prod that the live `alloy` container's config bind-mount source is `/opt/observability/alloy/config.alloy` -- a checkout of the now-archived split repo, **not** this monorepo's `observability/` copy. That means this PR fixes the canonical/going-forward copy, but taking effect on hosts needs a follow-up cutover of the observability stack's own deploy (analogous to the already-completed app-side `/opt/thermograph` swap) -- flagged, not done here, since it's a host-side change outside this PR's scope. - **Deployed Caddyfile**: confirmed on both prod and beta that the live `/etc/caddy/Caddyfile`'s own header says "RENDERED BY TERRAFORM for `<domain>`", matching `Caddyfile.tftpl`'s template header -- that's the true source, not a hand-edited file. `infra/deploy/Caddyfile` is a documented manual-apply fallback (see `infra/DEPLOY.md`) and was updated in step to avoid the two drifting further. ## Validated - `alloy validate` (real v1.9.1 binary, downloaded from the GitHub release) against the edited config: passes. (One pre-existing, unrelated tool gap: `alloy validate` doesn't recognize the top-level `livedebugging {}` singleton block as a component, even though `alloy run` loads it fine -- confirmed this is not something my change introduced by validating the original file too. Worked around in both the local check and the new CI step by stripping that one line before validating.) - `loki -verify-config` (real v3.5.1 image, matching what's deployed) against the edited config: passes. - `caddy validate` against the rendered template (via `terraform console` + the real `caddy` binary on beta, matching what's actually installed there) and against the edited `infra/deploy/Caddyfile`: both pass. - Query-string stripping was verified against a live scratch `caddy run` on beta (a temporary process on an unused port, not the real service) -- confirmed the `uri` field in the resulting JSON log line drops the query string entirely, and confirmed empirically that keeping `request>headers>User-Agent` alongside a parent-level `request>headers delete` is **not achievable** on beta's actual Caddy version (2.6.2): deleting the parent drops the whole subtree regardless of a child override, and the `log_append` directive that would let User-Agent be hoisted out first isn't available before Caddy 2.8ish. Traded away rather than block on it -- noted in the code comment. - Also discovered (worth a flag): **prod actually runs Caddy v2.11.4**, not 2.6.2 as assumed going in -- only beta is still on 2.6.2. Left both on the Alloy-side crawler drop for uniformity/simplicity per the task, but the version skew itself is worth resolving separately. ## Host-level changes applied directly (not in any repo, per scope) - `/etc/docker/daemon.json` (`{"log-driver": "json-file", "log-opts": {"max-size": "10m", "max-file": "3"}}`) written on **both** prod and beta. **Not restarting dockerd on either** -- checked first: prod is the fleet's sole Swarm manager (`Managers: 1`, this is the Leader) with **Live Restore disabled**, so a daemon restart there would both interrupt Swarm's control plane and actually stop/restart every container on the host (no live-restore to keep them running) -- a real outage window, not a blip. Deferring both restarts to a planned maintenance window; flagging this explicitly rather than forcing it. - `/etc/cron.d/thermograph-applogs-retention` on both hosts: `find /var/lib/docker/volumes/thermograph_applogs/_data -name '*.jsonl' -mtime +14 -delete`, daily at 03:00. Used plain cron rather than logrotate (logrotate *is* set up and active on both hosts for other logs, but the app already self-dates one file per category per day -- e.g. `access-2026-07-24.jsonl` -- so logrotate's rotate-in-place model doesn't fit files it doesn't own the naming of; a straight `find -delete` is the correct tool here). - Checked service health on both hosts before and after all of the above: all `docker service ls` replicas 1/1 on prod, and `thermograph.org`, `beta.thermograph.org`, `git.thermograph.org` all return 200, `dashboard.thermograph.org` its normal 302 (Google OAuth redirect) -- nothing was disrupted. ## Grafana Not touched. No changes to `observability/docker-compose.yml`, `observability/grafana/`, `observability/caddy-grafana.conf`, or `dashboard.thermograph.org`'s Caddy block. Confirmed via `git diff --stat` that only the five files listed above changed. ## Deferred / flagged - Docker daemon restarts on both hosts (see above) -- config is in place, restart is not. - The observability stack's own host-side cutover (repointing `/opt/observability` from the archived split repo to track this monorepo's `observability/` subtree) -- needed before any of the Alloy/Loki changes here actually take effect in production. Out of scope for this PR. - Caddy version skew (prod 2.11.4 vs beta 2.6.2) -- worth its own investigation/alignment pass.
admin_emi added 1 commit 2026-07-24 04:35:28 +00:00
Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / validate-observability (pull_request) Successful in 23s
PR build (required check) / gate (pull_request) Successful in 2s
9c0f493d68
thermograph-observability (and thermograph-backend/frontend/infra) are
archived on Forgejo now that the monorepo cutover has landed, so this domain's
work lands here in observability/ instead of a standalone repo PR.

Alloy (observability/alloy/config.alloy), deployed from /opt/observability on
prod and beta (confirmed via docker inspect mounts on the live hosts; the
monorepo's own observability/ copy isn't what's running yet -- that's a
separate, not-yet-done cutover step for this domain, flagged below):

- Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from
  the 250ms/250ms PollingFileWatcher default to 2s/10s, and
  local.file_match's glob-rescan sync_period from 10s to 1m.
- Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task
  churn (~every 90s) stops restarting tailers on every refresh.
- Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*,
  thermograph-lb, thermograph_worker) via a relabel rule; the worker's
  stdout is pure /healthz noise, already superseded by its own
  access/*.jsonl.
- Fix loki.source.docker double-relabeling: it took both the already-
  relabelled target output AND relabel_rules, running every rule twice per
  entry and making the new drop rules a no-op on the second pass. Pass raw
  discovery targets + relabel_rules instead.
- Add a loki.process stage for Caddy logs that drops well-known crawler/bot
  lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot,
  petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older
  (beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth
  addressing separately), so this is the shipper-side equivalent.

Loki (observability/loki/config.yml): 30 low-rate streams were hitting the
30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98%
full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set
chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream
limits sized for a three-node hobby fleet -- their absence was the source
of the sporadic 429s (Loki was falling back to its stricter multi-tenant
defaults).

CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no
syntax coverage. Download the pinned v1.9.1 alloy binary (matches
observability/alloy/docker-compose.agent.yml) straight from its GitHub
release and run `alloy validate` (minus the one line covering a known gap
where validate doesn't recognize the top-level livedebugging singleton
block, unlike alloy run).

Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl,
the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose
deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile
is the documented manual-apply fallback, updated in step to avoid drift):

- health_interval 5s -> 15s: a quarter of the active-healthcheck polling
  load for the same restart-safety guarantee.
- A `filter` format log encoder deleting request>headers, request>tls, and
  resp_headers (the full-header/TLS-block serialization was the main driver
  of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the
  file itself doesn't grow unbounded between logrotate runs.
- Strip the query string entirely from logged request URIs via a `regexp`
  field filter -- Caddy's default logger recorded request.uri *including*
  the query string, so every `?q=<search text>` sat in Loki next to the
  client IP for the full 30-day retention. Confirmed empirically (via
  `caddy validate` and a scratch `caddy run` against beta's real 2.6.2
  binary) that keeping request>headers>User-Agent alongside a parent-level
  header delete isn't achievable in this Caddy version -- deleting the
  parent drops the whole subtree regardless of a child override, and
  `log_append` (which would let it be hoisted out first) needs a newer
  Caddy than beta runs. Traded away rather than block on it.

Deferred (host-level, not in any repo, applied directly and separately):
/etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts --
config written, but NOT restarting dockerd on either: prod is the fleet's
sole Swarm manager with Live Restore disabled, so a daemon restart there
would both interrupt Swarm's control plane and actually stop/restart every
container on the host (no live-restore to keep them up), a real production
outage window, not a blip. Left for a planned maintenance window instead.
A 14-day retention cron for the app's applogs JSONL volume is applied on
both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete
rather than logrotate, since the app already self-dates one file per day
per category and logrotate's rotate-in-place model doesn't fit files it
doesn't own the naming of.

Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf,
dashboard.thermograph.org's Caddy block) is untouched by any of this.
admin_emi merged commit 0862395dda into dev 2026-07-24 04:37:42 +00:00
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference: Jinemi/thermograph#36
No description provided.