thermograph-observability (and thermograph-backend/frontend/infra) are
archived on Forgejo now that the monorepo cutover has landed, so this domain's
work lands here in observability/ instead of a standalone repo PR.
Alloy (observability/alloy/config.alloy), deployed from /opt/observability on
prod and beta (confirmed via docker inspect mounts on the live hosts; the
monorepo's own observability/ copy isn't what's running yet -- that's a
separate, not-yet-done cutover step for this domain, flagged below):
- Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from
the 250ms/250ms PollingFileWatcher default to 2s/10s, and
local.file_match's glob-rescan sync_period from 10s to 1m.
- Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task
churn (~every 90s) stops restarting tailers on every refresh.
- Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*,
thermograph-lb, thermograph_worker) via a relabel rule; the worker's
stdout is pure /healthz noise, already superseded by its own
access/*.jsonl.
- Fix loki.source.docker double-relabeling: it took both the already-
relabelled target output AND relabel_rules, running every rule twice per
entry and making the new drop rules a no-op on the second pass. Pass raw
discovery targets + relabel_rules instead.
- Add a loki.process stage for Caddy logs that drops well-known crawler/bot
lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot,
petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older
(beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth
addressing separately), so this is the shipper-side equivalent.
Loki (observability/loki/config.yml): 30 low-rate streams were hitting the
30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98%
full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set
chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream
limits sized for a three-node hobby fleet -- their absence was the source
of the sporadic 429s (Loki was falling back to its stricter multi-tenant
defaults).
CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no
syntax coverage. Download the pinned v1.9.1 alloy binary (matches
observability/alloy/docker-compose.agent.yml) straight from its GitHub
release and run `alloy validate` (minus the one line covering a known gap
where validate doesn't recognize the top-level livedebugging singleton
block, unlike alloy run).
Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl,
the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose
deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile
is the documented manual-apply fallback, updated in step to avoid drift):
- health_interval 5s -> 15s: a quarter of the active-healthcheck polling
load for the same restart-safety guarantee.
- A `filter` format log encoder deleting request>headers, request>tls, and
resp_headers (the full-header/TLS-block serialization was the main driver
of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the
file itself doesn't grow unbounded between logrotate runs.
- Strip the query string entirely from logged request URIs via a `regexp`
field filter -- Caddy's default logger recorded request.uri *including*
the query string, so every `?q=<search text>` sat in Loki next to the
client IP for the full 30-day retention. Confirmed empirically (via
`caddy validate` and a scratch `caddy run` against beta's real 2.6.2
binary) that keeping request>headers>User-Agent alongside a parent-level
header delete isn't achievable in this Caddy version -- deleting the
parent drops the whole subtree regardless of a child override, and
`log_append` (which would let it be hoisted out first) needs a newer
Caddy than beta runs. Traded away rather than block on it.
Deferred (host-level, not in any repo, applied directly and separately):
/etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts --
config written, but NOT restarting dockerd on either: prod is the fleet's
sole Swarm manager with Live Restore disabled, so a daemon restart there
would both interrupt Swarm's control plane and actually stop/restart every
container on the host (no live-restore to keep them up), a real production
outage window, not a blip. Left for a planned maintenance window instead.
A 14-day retention cron for the app's applogs JSONL volume is applied on
both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete
rather than logrotate, since the app already self-dates one file per day
per category and logrotate's rotate-in-place model doesn't fit files it
doesn't own the naming of.
Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf,
dashboard.thermograph.org's Caddy block) is untouched by any of this.
main took two direct PRs (#19 shell-lint, #21 the thermograph-daemon Go
service) while dev accumulated the ERA5 lake stack; both sides added compose/
stack services and touched the same seams. Resolutions keep both worlds:
- compose + Swarm stack: lake AND daemon are sibling services; backend env
carries THERMOGRAPH_LAKE_URL and THERMOGRAPH_INTERNAL_TOKEN.
- deploy.sh: backend deploys roll (backend lake daemon); all adds frontend.
- requirements.txt: websockets/apscheduler stay removed (moved to the Go
daemon), duckdb stays (the lake role's engine).
Merged tree: backend suite green, both compose configs validate.
Adds .forgejo/workflows/shell-lint.yml (pinned shellcheck v0.11.0 + sha256, -x,
default severity, fail on any finding, not path-filtered) and drives all 26
scripts to zero findings.
Two defects shellcheck cannot see:
render-secrets.sh left DECRYPTED vault contents in /tmp whenever a sops decrypt
failed -- the caller's set -e aborted the function before either cleanup ran.
Now removed on every exit path, with `|| return 1` on both sops calls so a
decrypt failure can never write a partial /etc/thermograph.env regardless of the
caller's shell options. Explicitly not a `trap ... RETURN`: such a trap set in a
sourced function persists into the caller's shell and re-fires when the caller's
next `.`/source completes, where the function-local tmp is unset -- fatal and
silent under deploy.sh's set -u. The file now records that reasoning.
autoscale.sh ran `set -eu` without pipefail while piping docker stats into awk,
so a failed left side was swallowed and the loop autoscaled on empty input.
Promoted to pipefail with a missed sample treated as a skip; verified busybox ash
in docker:27-cli supports it.
Also: capture-fixtures.sh's `jq . || cat` ran cat after jq had consumed stdin,
silently writing truncated fixtures; deploy.sh/deploy-stack.sh `# shellcheck
source=` paths corrected for the monorepo layout.
Backend runs its full hermetic suite, frontend its unit tier, inside the
just-built image -- the dev-only feature both app repos carried in their own
build.yml (deleted here in favor of the root workflow).
One root workflow set replaces the four repos' copies (deleted -- root-only
is where Forgejo reads them, and dead copies are a trap): per-domain
build-push with explicit image paths (emi/thermograph/backend|frontend; the
old github.repository-derived path collides in a monorepo), path-filtered
per-domain beta/prod/dev deploys, a domain-input reusable build check, a
single always-reporting PR gate (path-filtered required checks deadlock
auto-merge), a new infra-sync pipeline (host checkout + secrets render on
infra/** pushes), and ports of secrets-guard / ops-cron /
observability-validate to monorepo paths.