thermograph/infra/deploy/Caddyfile

121 lines
4.5 KiB
Text
Raw Normal View History

# /etc/caddy/Caddyfile on the VPS.
# Each domain's A/AAAA record must already point at this VPS — Caddy provisions a
# Let's Encrypt cert on first request and auto-renews. Nothing else to do for TLS
# (just make sure ports 80 and 443 are open).
#
# Layout:
# thermograph.org/* -> the Thermograph app (path-split across
# backend/frontend -- see below)
# emigriffith.dev/ -> static portfolio site (served straight from disk)
# emigriffith.dev/thermograph* -> permanent redirect to thermograph.org (the app moved)
#
# Thermograph now owns thermograph.org's root, so both services run with
# THERMOGRAPH_BASE=/ (see /etc/thermograph.env) — pages, assets and API all sit
# at "/" with no sub-path prefix. Repo-split Stage 4: backend and frontend are
# two containers now (docker-compose.yml), each on its own loopback port --
# Caddy path-splits directly to whichever owns a given path, so the browser
# still sees one apparent origin. Repo-split Stage 7a flipped which side owns
# the enumerated list: frontend now owns everything (content pages, the
# interactive tool's SPA shells, every static asset, the dynamic IndexNow key
# file) except the short, stable set below, which mirrors backend/web/app.py's
# own routing exactly (a single catch-all proxy to frontend for everything
# else) -- unlike frontend's paths, backend's don't grow every time a new
# static asset filename is added. A gap in this list still just degrades to
# "one extra hop" through backend's own proxy fallback, never a 404.
thermograph.org {
encode zstd gzip
@backend_paths path /api/* /digest /discord/interactions
# Active health check on the same cheap /healthz route each container's own
Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235) * Split web/worker duties with THERMOGRAPH_ROLE Background work (the subscription notifier) is welded to the same process that serves requests, so scaling the web tier to N replicas would also scale notifier instances unless something restricts it further than leader election alone. Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process behavior). Every replica runs the same image; ROLE only gates whether a process is allowed to own the notifier at all, layered on top of the existing leader election: web replicas never start it even if they'd win leader election, worker replicas start it if they win. The decision is pulled into _should_run_notifier() so it's unit-testable without booting the full app (DB init, places index, neighbor warmer). Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE) so a worker replica - which serves no real traffic - still has something Swarm can health-check. * Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate Three changes toward the hop-1 interim cutover, all inert until Track B stands up the platform: docker-stack.yml: the Swarm stack file for the interim cutover, distinct from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a pre-built image (IMAGE_TAG) instead of building in place; app/worker publish no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted), reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard needs a smaller MTU or large payloads silently stall); db is placement- pinned to a labelled node; app/worker skip inline migrations (RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing that ever runs Alembic; secrets are real Swarm secrets mounted at /run/secrets, read by the entrypoint shim rather than plain env vars. TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default latest-pg18, today's behavior unchanged), wired through Terraform (timescaledb_tag, default "latest-pg18") so it can actually be pinned to an exact minor without hand-editing the host - required before any host of the stack could replicate with another (a floating tag risks mismatched extension minors, which blocks a physical replica and risks compressed- chunk corruption on restore). Caddy active health-gate: both the Terraform-rendered Caddyfile and the live deploy/Caddyfile now health-check the app on the same cheap /healthz route its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage, so it's cheap enough for a tight interval and works identically for a worker replica, which serves no public traffic at all) - Caddy won't forward into a container that's still booting or unhealthy. Verified live: built and booted the real image via docker compose - both containers report healthy via the new /healthz-based HEALTHCHECK, and GET / still renders the full SSR homepage unchanged. Both Caddyfiles validated with the real caddy binary. docker-stack.yml validated with docker compose config (required-var guards fire with clear messages; secrets correctly mount at /run/secrets/<name>, matching the entrypoint shim's mapping). docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside the existing openmeteo overlay. terraform validate + fmt clean.
2026-07-21 00:39:48 +00:00
# HEALTHCHECK uses (Dockerfile) — so a deploy that's still restarting/booting
# never gets proxied into (a reload alone has no gate, hop-1 runbook hazard #10).
Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping thermograph-observability (and thermograph-backend/frontend/infra) are archived on Forgejo now that the monorepo cutover has landed, so this domain's work lands here in observability/ instead of a standalone repo PR. Alloy (observability/alloy/config.alloy), deployed from /opt/observability on prod and beta (confirmed via docker inspect mounts on the live hosts; the monorepo's own observability/ copy isn't what's running yet -- that's a separate, not-yet-done cutover step for this domain, flagged below): - Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from the 250ms/250ms PollingFileWatcher default to 2s/10s, and local.file_match's glob-rescan sync_period from 10s to 1m. - Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task churn (~every 90s) stops restarting tailers on every refresh. - Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*, thermograph-lb, thermograph_worker) via a relabel rule; the worker's stdout is pure /healthz noise, already superseded by its own access/*.jsonl. - Fix loki.source.docker double-relabeling: it took both the already- relabelled target output AND relabel_rules, running every rule twice per entry and making the new drop rules a no-op on the second pass. Pass raw discovery targets + relabel_rules instead. - Add a loki.process stage for Caddy logs that drops well-known crawler/bot lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot, petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older (beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth addressing separately), so this is the shipper-side equivalent. Loki (observability/loki/config.yml): 30 low-rate streams were hitting the 30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98% full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream limits sized for a three-node hobby fleet -- their absence was the source of the sporadic 429s (Loki was falling back to its stricter multi-tenant defaults). CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no syntax coverage. Download the pinned v1.9.1 alloy binary (matches observability/alloy/docker-compose.agent.yml) straight from its GitHub release and run `alloy validate` (minus the one line covering a known gap where validate doesn't recognize the top-level livedebugging singleton block, unlike alloy run). Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl, the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile is the documented manual-apply fallback, updated in step to avoid drift): - health_interval 5s -> 15s: a quarter of the active-healthcheck polling load for the same restart-safety guarantee. - A `filter` format log encoder deleting request>headers, request>tls, and resp_headers (the full-header/TLS-block serialization was the main driver of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the file itself doesn't grow unbounded between logrotate runs. - Strip the query string entirely from logged request URIs via a `regexp` field filter -- Caddy's default logger recorded request.uri *including* the query string, so every `?q=<search text>` sat in Loki next to the client IP for the full 30-day retention. Confirmed empirically (via `caddy validate` and a scratch `caddy run` against beta's real 2.6.2 binary) that keeping request>headers>User-Agent alongside a parent-level header delete isn't achievable in this Caddy version -- deleting the parent drops the whole subtree regardless of a child override, and `log_append` (which would let it be hoisted out first) needs a newer Caddy than beta runs. Traded away rather than block on it. Deferred (host-level, not in any repo, applied directly and separately): /etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts -- config written, but NOT restarting dockerd on either: prod is the fleet's sole Swarm manager with Live Restore disabled, so a daemon restart there would both interrupt Swarm's control plane and actually stop/restart every container on the host (no live-restore to keep them up), a real production outage window, not a blip. Left for a planned maintenance window instead. A 14-day retention cron for the app's applogs JSONL volume is applied on both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete rather than logrotate, since the app already self-dates one file per day per category and logrotate's rotate-in-place model doesn't fit files it doesn't own the naming of. Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf, dashboard.thermograph.org's Caddy block) is untouched by any of this.
2026-07-24 04:34:06 +00:00
# 15s (was 5s): plenty responsive for a process that only restarts on a deploy,
# and a quarter of the polling load.
handle @backend_paths {
reverse_proxy 127.0.0.1:8137 {
health_uri /healthz
Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping thermograph-observability (and thermograph-backend/frontend/infra) are archived on Forgejo now that the monorepo cutover has landed, so this domain's work lands here in observability/ instead of a standalone repo PR. Alloy (observability/alloy/config.alloy), deployed from /opt/observability on prod and beta (confirmed via docker inspect mounts on the live hosts; the monorepo's own observability/ copy isn't what's running yet -- that's a separate, not-yet-done cutover step for this domain, flagged below): - Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from the 250ms/250ms PollingFileWatcher default to 2s/10s, and local.file_match's glob-rescan sync_period from 10s to 1m. - Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task churn (~every 90s) stops restarting tailers on every refresh. - Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*, thermograph-lb, thermograph_worker) via a relabel rule; the worker's stdout is pure /healthz noise, already superseded by its own access/*.jsonl. - Fix loki.source.docker double-relabeling: it took both the already- relabelled target output AND relabel_rules, running every rule twice per entry and making the new drop rules a no-op on the second pass. Pass raw discovery targets + relabel_rules instead. - Add a loki.process stage for Caddy logs that drops well-known crawler/bot lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot, petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older (beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth addressing separately), so this is the shipper-side equivalent. Loki (observability/loki/config.yml): 30 low-rate streams were hitting the 30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98% full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream limits sized for a three-node hobby fleet -- their absence was the source of the sporadic 429s (Loki was falling back to its stricter multi-tenant defaults). CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no syntax coverage. Download the pinned v1.9.1 alloy binary (matches observability/alloy/docker-compose.agent.yml) straight from its GitHub release and run `alloy validate` (minus the one line covering a known gap where validate doesn't recognize the top-level livedebugging singleton block, unlike alloy run). Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl, the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile is the documented manual-apply fallback, updated in step to avoid drift): - health_interval 5s -> 15s: a quarter of the active-healthcheck polling load for the same restart-safety guarantee. - A `filter` format log encoder deleting request>headers, request>tls, and resp_headers (the full-header/TLS-block serialization was the main driver of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the file itself doesn't grow unbounded between logrotate runs. - Strip the query string entirely from logged request URIs via a `regexp` field filter -- Caddy's default logger recorded request.uri *including* the query string, so every `?q=<search text>` sat in Loki next to the client IP for the full 30-day retention. Confirmed empirically (via `caddy validate` and a scratch `caddy run` against beta's real 2.6.2 binary) that keeping request>headers>User-Agent alongside a parent-level header delete isn't achievable in this Caddy version -- deleting the parent drops the whole subtree regardless of a child override, and `log_append` (which would let it be hoisted out first) needs a newer Caddy than beta runs. Traded away rather than block on it. Deferred (host-level, not in any repo, applied directly and separately): /etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts -- config written, but NOT restarting dockerd on either: prod is the fleet's sole Swarm manager with Live Restore disabled, so a daemon restart there would both interrupt Swarm's control plane and actually stop/restart every container on the host (no live-restore to keep them up), a real production outage window, not a blip. Left for a planned maintenance window instead. A 14-day retention cron for the app's applogs JSONL volume is applied on both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete rather than logrotate, since the app already self-dates one file per day per category and logrotate's rotate-in-place model doesn't fit files it doesn't own the naming of. Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf, dashboard.thermograph.org's Caddy block) is untouched by any of this.
2026-07-24 04:34:06 +00:00
health_interval 15s
health_timeout 3s
health_status 2xx
}
}
handle {
reverse_proxy 127.0.0.1:8080 {
health_uri /healthz
Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping thermograph-observability (and thermograph-backend/frontend/infra) are archived on Forgejo now that the monorepo cutover has landed, so this domain's work lands here in observability/ instead of a standalone repo PR. Alloy (observability/alloy/config.alloy), deployed from /opt/observability on prod and beta (confirmed via docker inspect mounts on the live hosts; the monorepo's own observability/ copy isn't what's running yet -- that's a separate, not-yet-done cutover step for this domain, flagged below): - Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from the 250ms/250ms PollingFileWatcher default to 2s/10s, and local.file_match's glob-rescan sync_period from 10s to 1m. - Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task churn (~every 90s) stops restarting tailers on every refresh. - Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*, thermograph-lb, thermograph_worker) via a relabel rule; the worker's stdout is pure /healthz noise, already superseded by its own access/*.jsonl. - Fix loki.source.docker double-relabeling: it took both the already- relabelled target output AND relabel_rules, running every rule twice per entry and making the new drop rules a no-op on the second pass. Pass raw discovery targets + relabel_rules instead. - Add a loki.process stage for Caddy logs that drops well-known crawler/bot lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot, petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older (beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth addressing separately), so this is the shipper-side equivalent. Loki (observability/loki/config.yml): 30 low-rate streams were hitting the 30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98% full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream limits sized for a three-node hobby fleet -- their absence was the source of the sporadic 429s (Loki was falling back to its stricter multi-tenant defaults). CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no syntax coverage. Download the pinned v1.9.1 alloy binary (matches observability/alloy/docker-compose.agent.yml) straight from its GitHub release and run `alloy validate` (minus the one line covering a known gap where validate doesn't recognize the top-level livedebugging singleton block, unlike alloy run). Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl, the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile is the documented manual-apply fallback, updated in step to avoid drift): - health_interval 5s -> 15s: a quarter of the active-healthcheck polling load for the same restart-safety guarantee. - A `filter` format log encoder deleting request>headers, request>tls, and resp_headers (the full-header/TLS-block serialization was the main driver of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the file itself doesn't grow unbounded between logrotate runs. - Strip the query string entirely from logged request URIs via a `regexp` field filter -- Caddy's default logger recorded request.uri *including* the query string, so every `?q=<search text>` sat in Loki next to the client IP for the full 30-day retention. Confirmed empirically (via `caddy validate` and a scratch `caddy run` against beta's real 2.6.2 binary) that keeping request>headers>User-Agent alongside a parent-level header delete isn't achievable in this Caddy version -- deleting the parent drops the whole subtree regardless of a child override, and `log_append` (which would let it be hoisted out first) needs a newer Caddy than beta runs. Traded away rather than block on it. Deferred (host-level, not in any repo, applied directly and separately): /etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts -- config written, but NOT restarting dockerd on either: prod is the fleet's sole Swarm manager with Live Restore disabled, so a daemon restart there would both interrupt Swarm's control plane and actually stop/restart every container on the host (no live-restore to keep them up), a real production outage window, not a blip. Left for a planned maintenance window instead. A 14-day retention cron for the app's applogs JSONL volume is applied on both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete rather than logrotate, since the app already self-dates one file per day per category and logrotate's rotate-in-place model doesn't fit files it doesn't own the naming of. Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf, dashboard.thermograph.org's Caddy block) is untouched by any of this.
2026-07-24 04:34:06 +00:00
health_interval 15s
health_timeout 3s
health_status 2xx
}
Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235) * Split web/worker duties with THERMOGRAPH_ROLE Background work (the subscription notifier) is welded to the same process that serves requests, so scaling the web tier to N replicas would also scale notifier instances unless something restricts it further than leader election alone. Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process behavior). Every replica runs the same image; ROLE only gates whether a process is allowed to own the notifier at all, layered on top of the existing leader election: web replicas never start it even if they'd win leader election, worker replicas start it if they win. The decision is pulled into _should_run_notifier() so it's unit-testable without booting the full app (DB init, places index, neighbor warmer). Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE) so a worker replica - which serves no real traffic - still has something Swarm can health-check. * Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate Three changes toward the hop-1 interim cutover, all inert until Track B stands up the platform: docker-stack.yml: the Swarm stack file for the interim cutover, distinct from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a pre-built image (IMAGE_TAG) instead of building in place; app/worker publish no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted), reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard needs a smaller MTU or large payloads silently stall); db is placement- pinned to a labelled node; app/worker skip inline migrations (RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing that ever runs Alembic; secrets are real Swarm secrets mounted at /run/secrets, read by the entrypoint shim rather than plain env vars. TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default latest-pg18, today's behavior unchanged), wired through Terraform (timescaledb_tag, default "latest-pg18") so it can actually be pinned to an exact minor without hand-editing the host - required before any host of the stack could replicate with another (a floating tag risks mismatched extension minors, which blocks a physical replica and risks compressed- chunk corruption on restore). Caddy active health-gate: both the Terraform-rendered Caddyfile and the live deploy/Caddyfile now health-check the app on the same cheap /healthz route its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage, so it's cheap enough for a tight interval and works identically for a worker replica, which serves no public traffic at all) - Caddy won't forward into a container that's still booting or unhealthy. Verified live: built and booted the real image via docker compose - both containers report healthy via the new /healthz-based HEALTHCHECK, and GET / still renders the full SSR homepage unchanged. Both Caddyfiles validated with the real caddy binary. docker-stack.yml validated with docker compose config (required-var guards fire with clear messages; secrets correctly mount at /run/secrets/<name>, matching the entrypoint shim's mapping). docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside the existing openmeteo overlay. terraform validate + fmt clean.
2026-07-21 00:39:48 +00:00
}
Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping thermograph-observability (and thermograph-backend/frontend/infra) are archived on Forgejo now that the monorepo cutover has landed, so this domain's work lands here in observability/ instead of a standalone repo PR. Alloy (observability/alloy/config.alloy), deployed from /opt/observability on prod and beta (confirmed via docker inspect mounts on the live hosts; the monorepo's own observability/ copy isn't what's running yet -- that's a separate, not-yet-done cutover step for this domain, flagged below): - Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from the 250ms/250ms PollingFileWatcher default to 2s/10s, and local.file_match's glob-rescan sync_period from 10s to 1m. - Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task churn (~every 90s) stops restarting tailers on every refresh. - Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*, thermograph-lb, thermograph_worker) via a relabel rule; the worker's stdout is pure /healthz noise, already superseded by its own access/*.jsonl. - Fix loki.source.docker double-relabeling: it took both the already- relabelled target output AND relabel_rules, running every rule twice per entry and making the new drop rules a no-op on the second pass. Pass raw discovery targets + relabel_rules instead. - Add a loki.process stage for Caddy logs that drops well-known crawler/bot lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot, petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older (beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth addressing separately), so this is the shipper-side equivalent. Loki (observability/loki/config.yml): 30 low-rate streams were hitting the 30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98% full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream limits sized for a three-node hobby fleet -- their absence was the source of the sporadic 429s (Loki was falling back to its stricter multi-tenant defaults). CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no syntax coverage. Download the pinned v1.9.1 alloy binary (matches observability/alloy/docker-compose.agent.yml) straight from its GitHub release and run `alloy validate` (minus the one line covering a known gap where validate doesn't recognize the top-level livedebugging singleton block, unlike alloy run). Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl, the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile is the documented manual-apply fallback, updated in step to avoid drift): - health_interval 5s -> 15s: a quarter of the active-healthcheck polling load for the same restart-safety guarantee. - A `filter` format log encoder deleting request>headers, request>tls, and resp_headers (the full-header/TLS-block serialization was the main driver of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the file itself doesn't grow unbounded between logrotate runs. - Strip the query string entirely from logged request URIs via a `regexp` field filter -- Caddy's default logger recorded request.uri *including* the query string, so every `?q=<search text>` sat in Loki next to the client IP for the full 30-day retention. Confirmed empirically (via `caddy validate` and a scratch `caddy run` against beta's real 2.6.2 binary) that keeping request>headers>User-Agent alongside a parent-level header delete isn't achievable in this Caddy version -- deleting the parent drops the whole subtree regardless of a child override, and `log_append` (which would let it be hoisted out first) needs a newer Caddy than beta runs. Traded away rather than block on it. Deferred (host-level, not in any repo, applied directly and separately): /etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts -- config written, but NOT restarting dockerd on either: prod is the fleet's sole Swarm manager with Live Restore disabled, so a daemon restart there would both interrupt Swarm's control plane and actually stop/restart every container on the host (no live-restore to keep them up), a real production outage window, not a blip. Left for a planned maintenance window instead. A 14-day retention cron for the app's applogs JSONL volume is applied on both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete rather than logrotate, since the app already self-dates one file per day per category and logrotate's rotate-in-place model doesn't fit files it doesn't own the naming of. Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf, dashboard.thermograph.org's Caddy block) is untouched by any of this.
2026-07-24 04:34:06 +00:00
# Access-log hygiene: the default JSON encoder serializes full request headers,
# the TLS block, and response headers on every line (measured ~1,133B/line) --
# strip those with the `filter` format encoder. Also strip the query string from
# the logged URI: Caddy's default logger records request.uri *including* the
# query string, so every `?q=<search text>` a visitor typed sat in Loki next to
# their client IP for the full 30-day retention -- a real privacy leak, not just
# noise. Bot/crawler skipping stays out of here: `log_skip` needs Caddy >= 2.7
# and an upgrade is out of scope, so that's handled downstream in Alloy's
# loki.process "caddy" stage instead (see observability/alloy/config.alloy).
log {
Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping thermograph-observability (and thermograph-backend/frontend/infra) are archived on Forgejo now that the monorepo cutover has landed, so this domain's work lands here in observability/ instead of a standalone repo PR. Alloy (observability/alloy/config.alloy), deployed from /opt/observability on prod and beta (confirmed via docker inspect mounts on the live hosts; the monorepo's own observability/ copy isn't what's running yet -- that's a separate, not-yet-done cutover step for this domain, flagged below): - Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from the 250ms/250ms PollingFileWatcher default to 2s/10s, and local.file_match's glob-rescan sync_period from 10s to 1m. - Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task churn (~every 90s) stops restarting tailers on every refresh. - Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*, thermograph-lb, thermograph_worker) via a relabel rule; the worker's stdout is pure /healthz noise, already superseded by its own access/*.jsonl. - Fix loki.source.docker double-relabeling: it took both the already- relabelled target output AND relabel_rules, running every rule twice per entry and making the new drop rules a no-op on the second pass. Pass raw discovery targets + relabel_rules instead. - Add a loki.process stage for Caddy logs that drops well-known crawler/bot lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot, petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older (beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth addressing separately), so this is the shipper-side equivalent. Loki (observability/loki/config.yml): 30 low-rate streams were hitting the 30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98% full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream limits sized for a three-node hobby fleet -- their absence was the source of the sporadic 429s (Loki was falling back to its stricter multi-tenant defaults). CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no syntax coverage. Download the pinned v1.9.1 alloy binary (matches observability/alloy/docker-compose.agent.yml) straight from its GitHub release and run `alloy validate` (minus the one line covering a known gap where validate doesn't recognize the top-level livedebugging singleton block, unlike alloy run). Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl, the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile is the documented manual-apply fallback, updated in step to avoid drift): - health_interval 5s -> 15s: a quarter of the active-healthcheck polling load for the same restart-safety guarantee. - A `filter` format log encoder deleting request>headers, request>tls, and resp_headers (the full-header/TLS-block serialization was the main driver of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the file itself doesn't grow unbounded between logrotate runs. - Strip the query string entirely from logged request URIs via a `regexp` field filter -- Caddy's default logger recorded request.uri *including* the query string, so every `?q=<search text>` sat in Loki next to the client IP for the full 30-day retention. Confirmed empirically (via `caddy validate` and a scratch `caddy run` against beta's real 2.6.2 binary) that keeping request>headers>User-Agent alongside a parent-level header delete isn't achievable in this Caddy version -- deleting the parent drops the whole subtree regardless of a child override, and `log_append` (which would let it be hoisted out first) needs a newer Caddy than beta runs. Traded away rather than block on it. Deferred (host-level, not in any repo, applied directly and separately): /etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts -- config written, but NOT restarting dockerd on either: prod is the fleet's sole Swarm manager with Live Restore disabled, so a daemon restart there would both interrupt Swarm's control plane and actually stop/restart every container on the host (no live-restore to keep them up), a real production outage window, not a blip. Left for a planned maintenance window instead. A 14-day retention cron for the app's applogs JSONL volume is applied on both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete rather than logrotate, since the app already self-dates one file per day per category and logrotate's rotate-in-place model doesn't fit files it doesn't own the naming of. Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf, dashboard.thermograph.org's Caddy block) is untouched by any of this.
2026-07-24 04:34:06 +00:00
output file /var/log/caddy/thermograph.log {
roll_size 20MiB
roll_keep 5
}
format filter {
wrap json
fields {
request>headers delete
request>tls delete
resp_headers delete
request>uri regexp \?.* ""
}
}
}
}
emigriffith.dev {
encode zstd gzip
# Thermograph moved to its own domain. Send the old sub-path there with a
# permanent redirect, stripping the /thermograph prefix so deep links map
# straight across (…/thermograph/calendar -> thermograph.org/calendar). The
# bare /thermograph (no trailing slash) goes to the new root.
handle_path /thermograph/* {
redir https://thermograph.org{uri} permanent
}
handle /thermograph {
redir https://thermograph.org/ permanent
}
# Portfolio at the root. Point `root` at the built static site (for the Astro
# portfolio that's its `dist/` output). file_server serves index.html for
# directories and returns a real 404 for missing paths.
handle {
root * /var/www/emigriffith
file_server
}
log {
output file /var/log/caddy/emigriffith.log
}
}
# Old bookmarks to the raw IP (the pre-domain URL) would otherwise get bounced to
# HTTPS-on-the-IP, which has no cert and fails. Redirect them to the portfolio domain.
http://75.119.132.91 {
redir https://emigriffith.dev{uri} permanent
}
# Optional: redirect www -> apex for either domain. Add the www CNAME/A record
# first, then uncomment the matching block.
# www.emigriffith.dev {
# redir https://emigriffith.dev{uri} permanent
# }
# www.thermograph.org {
# redir https://thermograph.org{uri} permanent
# }