thermograph/.forgejo/workflows/observability-validate.yml

94 lines
4 KiB
YAML
Raw Normal View History

name: Validate observability stack
# The observability DOMAIN deploys by hand (`docker compose up -d` on beta +
# the Alloy agent on each node) with no build step, so nothing caught a
# malformed compose file, a broken dashboard JSON, or an unparseable config
# until it failed live on the host. This is that missing guard: it parses
# every YAML/JSON artifact that ships to a node. It does NOT deploy, and
# deliberately needs no docker CLI (the node:20-bookworm job image doesn't
# ship one) -- a YAML parse catches the real breakage without it.
#
# Monorepo port: paths are prefixed observability/, the push trigger is
# path-filtered to the domain, and workflow_call lets pr-build.yml reuse this
# as the domain's PR check under its single `gate` required status.
#
Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping thermograph-observability (and thermograph-backend/frontend/infra) are archived on Forgejo now that the monorepo cutover has landed, so this domain's work lands here in observability/ instead of a standalone repo PR. Alloy (observability/alloy/config.alloy), deployed from /opt/observability on prod and beta (confirmed via docker inspect mounts on the live hosts; the monorepo's own observability/ copy isn't what's running yet -- that's a separate, not-yet-done cutover step for this domain, flagged below): - Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from the 250ms/250ms PollingFileWatcher default to 2s/10s, and local.file_match's glob-rescan sync_period from 10s to 1m. - Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task churn (~every 90s) stops restarting tailers on every refresh. - Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*, thermograph-lb, thermograph_worker) via a relabel rule; the worker's stdout is pure /healthz noise, already superseded by its own access/*.jsonl. - Fix loki.source.docker double-relabeling: it took both the already- relabelled target output AND relabel_rules, running every rule twice per entry and making the new drop rules a no-op on the second pass. Pass raw discovery targets + relabel_rules instead. - Add a loki.process stage for Caddy logs that drops well-known crawler/bot lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot, petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older (beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth addressing separately), so this is the shipper-side equivalent. Loki (observability/loki/config.yml): 30 low-rate streams were hitting the 30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98% full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream limits sized for a three-node hobby fleet -- their absence was the source of the sporadic 429s (Loki was falling back to its stricter multi-tenant defaults). CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no syntax coverage. Download the pinned v1.9.1 alloy binary (matches observability/alloy/docker-compose.agent.yml) straight from its GitHub release and run `alloy validate` (minus the one line covering a known gap where validate doesn't recognize the top-level livedebugging singleton block, unlike alloy run). Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl, the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile is the documented manual-apply fallback, updated in step to avoid drift): - health_interval 5s -> 15s: a quarter of the active-healthcheck polling load for the same restart-safety guarantee. - A `filter` format log encoder deleting request>headers, request>tls, and resp_headers (the full-header/TLS-block serialization was the main driver of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the file itself doesn't grow unbounded between logrotate runs. - Strip the query string entirely from logged request URIs via a `regexp` field filter -- Caddy's default logger recorded request.uri *including* the query string, so every `?q=<search text>` sat in Loki next to the client IP for the full 30-day retention. Confirmed empirically (via `caddy validate` and a scratch `caddy run` against beta's real 2.6.2 binary) that keeping request>headers>User-Agent alongside a parent-level header delete isn't achievable in this Caddy version -- deleting the parent drops the whole subtree regardless of a child override, and `log_append` (which would let it be hoisted out first) needs a newer Caddy than beta runs. Traded away rather than block on it. Deferred (host-level, not in any repo, applied directly and separately): /etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts -- config written, but NOT restarting dockerd on either: prod is the fleet's sole Swarm manager with Live Restore disabled, so a daemon restart there would both interrupt Swarm's control plane and actually stop/restart every container on the host (no live-restore to keep them up), a real production outage window, not a blip. Left for a planned maintenance window instead. A 14-day retention cron for the app's applogs JSONL volume is applied on both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete rather than logrotate, since the app already self-dates one file per day per category and logrotate's rotate-in-place model doesn't fit files it doesn't own the naming of. Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf, dashboard.thermograph.org's Caddy block) is untouched by any of this.
2026-07-24 04:34:06 +00:00
# The Alloy config (observability/alloy/config.alloy, an HCL-like format) IS
# validated below -- the static `alloy` binary is downloaded straight from its
# GitHub release (pinned to the same v1.9.1 the fleet runs; see
# observability/alloy/docker-compose.agent.yml), no docker CLI needed.
#
# Not validated here: docker-compose *schema* (unknown-key checks, which would
# need the docker CLI). Add it if the churn warrants the extra tooling.
on:
workflow_call: {}
push:
branches: [main, dev]
paths: ['observability/**']
jobs:
validate:
runs-on: docker
steps:
- uses: actions/checkout@v4
- name: Dashboards are valid JSON
run: |
set -euo pipefail
python3 - <<'PY'
import json, pathlib, sys
bad = False
files = sorted(pathlib.Path("observability/grafana/dashboards").glob("*.json"))
if not files:
print("no dashboards found"); sys.exit(1)
for f in files:
try:
json.loads(f.read_text()); print("OK", f)
except Exception as e: # noqa: BLE001
print("INVALID JSON", f, "-", e); bad = True
sys.exit(1 if bad else 0)
PY
- name: YAML artifacts parse (compose + loki + grafana provisioning)
run: |
set -euo pipefail
apt-get update -qq && apt-get install -y -qq python3-yaml >/dev/null
python3 - <<'PY'
import pathlib, sys
import yaml
bad = False
root = pathlib.Path("observability")
targets = [
root / "docker-compose.yml",
root / "alloy/docker-compose.agent.yml",
root / "loki/config.yml",
*sorted((root / "grafana/provisioning").rglob("*.yml")),
]
for f in targets:
if not f.exists():
print("MISSING", f); bad = True; continue
try:
yaml.safe_load(f.read_text()); print("OK", f)
except Exception as e: # noqa: BLE001
print("INVALID YAML", f, "-", e); bad = True
sys.exit(1 if bad else 0)
PY
Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping thermograph-observability (and thermograph-backend/frontend/infra) are archived on Forgejo now that the monorepo cutover has landed, so this domain's work lands here in observability/ instead of a standalone repo PR. Alloy (observability/alloy/config.alloy), deployed from /opt/observability on prod and beta (confirmed via docker inspect mounts on the live hosts; the monorepo's own observability/ copy isn't what's running yet -- that's a separate, not-yet-done cutover step for this domain, flagged below): - Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from the 250ms/250ms PollingFileWatcher default to 2s/10s, and local.file_match's glob-rescan sync_period from 10s to 1m. - Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task churn (~every 90s) stops restarting tailers on every refresh. - Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*, thermograph-lb, thermograph_worker) via a relabel rule; the worker's stdout is pure /healthz noise, already superseded by its own access/*.jsonl. - Fix loki.source.docker double-relabeling: it took both the already- relabelled target output AND relabel_rules, running every rule twice per entry and making the new drop rules a no-op on the second pass. Pass raw discovery targets + relabel_rules instead. - Add a loki.process stage for Caddy logs that drops well-known crawler/bot lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot, petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older (beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth addressing separately), so this is the shipper-side equivalent. Loki (observability/loki/config.yml): 30 low-rate streams were hitting the 30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98% full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream limits sized for a three-node hobby fleet -- their absence was the source of the sporadic 429s (Loki was falling back to its stricter multi-tenant defaults). CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no syntax coverage. Download the pinned v1.9.1 alloy binary (matches observability/alloy/docker-compose.agent.yml) straight from its GitHub release and run `alloy validate` (minus the one line covering a known gap where validate doesn't recognize the top-level livedebugging singleton block, unlike alloy run). Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl, the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile is the documented manual-apply fallback, updated in step to avoid drift): - health_interval 5s -> 15s: a quarter of the active-healthcheck polling load for the same restart-safety guarantee. - A `filter` format log encoder deleting request>headers, request>tls, and resp_headers (the full-header/TLS-block serialization was the main driver of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the file itself doesn't grow unbounded between logrotate runs. - Strip the query string entirely from logged request URIs via a `regexp` field filter -- Caddy's default logger recorded request.uri *including* the query string, so every `?q=<search text>` sat in Loki next to the client IP for the full 30-day retention. Confirmed empirically (via `caddy validate` and a scratch `caddy run` against beta's real 2.6.2 binary) that keeping request>headers>User-Agent alongside a parent-level header delete isn't achievable in this Caddy version -- deleting the parent drops the whole subtree regardless of a child override, and `log_append` (which would let it be hoisted out first) needs a newer Caddy than beta runs. Traded away rather than block on it. Deferred (host-level, not in any repo, applied directly and separately): /etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts -- config written, but NOT restarting dockerd on either: prod is the fleet's sole Swarm manager with Live Restore disabled, so a daemon restart there would both interrupt Swarm's control plane and actually stop/restart every container on the host (no live-restore to keep them up), a real production outage window, not a blip. Left for a planned maintenance window instead. A 14-day retention cron for the app's applogs JSONL volume is applied on both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete rather than logrotate, since the app already self-dates one file per day per category and logrotate's rotate-in-place model doesn't fit files it doesn't own the naming of. Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf, dashboard.thermograph.org's Caddy block) is untouched by any of this.
2026-07-24 04:34:06 +00:00
- name: Alloy config parses and validates (components, relabel/process wiring)
run: |
set -euo pipefail
apt-get update -qq && apt-get install -y -qq unzip >/dev/null
curl -sSL -o /tmp/alloy.zip \
https://github.com/grafana/alloy/releases/download/v1.9.1/alloy-linux-amd64.zip
unzip -q /tmp/alloy.zip -d /tmp/alloybin
chmod +x /tmp/alloybin/alloy-linux-amd64
# `alloy validate` builds the real component graph (catches bad field
# names, dangling forward_to/receiver refs, malformed relabel/process
# stages) but doesn't recognize the top-level `livedebugging {}` singleton
# block as a component -- it only knows named components, even though
# `alloy run` loads that block fine. Known gap in the `validate`
# subcommand, not a config error, so strip that one line before checking.
grep -v '^livedebugging ' observability/alloy/config.alloy > /tmp/config-for-validate.alloy
/tmp/alloybin/alloy-linux-amd64 validate /tmp/config-for-validate.alloy