thermograph/infra
Emi Griffith 9c0f493d68
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / validate-observability (pull_request) Successful in 23s
PR build (required check) / gate (pull_request) Successful in 2s
Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping
thermograph-observability (and thermograph-backend/frontend/infra) are
archived on Forgejo now that the monorepo cutover has landed, so this domain's
work lands here in observability/ instead of a standalone repo PR.

Alloy (observability/alloy/config.alloy), deployed from /opt/observability on
prod and beta (confirmed via docker inspect mounts on the live hosts; the
monorepo's own observability/ copy isn't what's running yet -- that's a
separate, not-yet-done cutover step for this domain, flagged below):

- Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from
  the 250ms/250ms PollingFileWatcher default to 2s/10s, and
  local.file_match's glob-rescan sync_period from 10s to 1m.
- Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task
  churn (~every 90s) stops restarting tailers on every refresh.
- Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*,
  thermograph-lb, thermograph_worker) via a relabel rule; the worker's
  stdout is pure /healthz noise, already superseded by its own
  access/*.jsonl.
- Fix loki.source.docker double-relabeling: it took both the already-
  relabelled target output AND relabel_rules, running every rule twice per
  entry and making the new drop rules a no-op on the second pass. Pass raw
  discovery targets + relabel_rules instead.
- Add a loki.process stage for Caddy logs that drops well-known crawler/bot
  lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot,
  petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older
  (beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth
  addressing separately), so this is the shipper-side equivalent.

Loki (observability/loki/config.yml): 30 low-rate streams were hitting the
30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98%
full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set
chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream
limits sized for a three-node hobby fleet -- their absence was the source
of the sporadic 429s (Loki was falling back to its stricter multi-tenant
defaults).

CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no
syntax coverage. Download the pinned v1.9.1 alloy binary (matches
observability/alloy/docker-compose.agent.yml) straight from its GitHub
release and run `alloy validate` (minus the one line covering a known gap
where validate doesn't recognize the top-level livedebugging singleton
block, unlike alloy run).

Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl,
the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose
deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile
is the documented manual-apply fallback, updated in step to avoid drift):

- health_interval 5s -> 15s: a quarter of the active-healthcheck polling
  load for the same restart-safety guarantee.
- A `filter` format log encoder deleting request>headers, request>tls, and
  resp_headers (the full-header/TLS-block serialization was the main driver
  of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the
  file itself doesn't grow unbounded between logrotate runs.
- Strip the query string entirely from logged request URIs via a `regexp`
  field filter -- Caddy's default logger recorded request.uri *including*
  the query string, so every `?q=<search text>` sat in Loki next to the
  client IP for the full 30-day retention. Confirmed empirically (via
  `caddy validate` and a scratch `caddy run` against beta's real 2.6.2
  binary) that keeping request>headers>User-Agent alongside a parent-level
  header delete isn't achievable in this Caddy version -- deleting the
  parent drops the whole subtree regardless of a child override, and
  `log_append` (which would let it be hoisted out first) needs a newer
  Caddy than beta runs. Traded away rather than block on it.

Deferred (host-level, not in any repo, applied directly and separately):
/etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts --
config written, but NOT restarting dockerd on either: prod is the fleet's
sole Swarm manager with Live Restore disabled, so a daemon restart there
would both interrupt Swarm's control plane and actually stop/restart every
container on the host (no live-restore to keep them up), a real production
outage window, not a blip. Left for a planned maintenance window instead.
A 14-day retention cron for the app's applogs JSONL volume is applied on
both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete
rather than logrotate, since the app already self-dates one file per day
per category and logrotate's rotate-in-place model doesn't fit files it
doesn't own the naming of.

Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf,
dashboard.thermograph.org's Caddy block) is untouched by any of this.
2026-07-23 21:34:06 -07:00
..
.claude/skills/key-gaps Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
deploy Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping 2026-07-23 21:34:06 -07:00
lake-iceberg Iceberg conversion container for the ERA5 lake (infra/lake-iceberg) (#24) 2026-07-23 22:56:27 +00:00
terraform Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping 2026-07-23 21:34:06 -07:00
.env.example daemon: move the Discord gateway and scheduler out of the web process into Go (#21) 2026-07-23 22:49:54 +00:00
.gitignore Iceberg conversion container for the ERA5 lake (infra/lake-iceberg) (#24) 2026-07-23 22:56:27 +00:00
.sops.yaml Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
ACCESS.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
CLAUDE.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
DEPLOY-DEV.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
DEPLOY.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
docker-compose.dev.yml Reconcile: merge main (shellcheck guard, Go daemon) into dev 2026-07-23 17:06:52 -07:00
docker-compose.openmeteo.yml Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
docker-compose.yml Reconcile: merge main (shellcheck guard, Go daemon) into dev 2026-07-23 17:06:52 -07:00
docker-stack.yml Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
Makefile Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
README.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00

thermograph-infra

Infrastructure for Thermograph: Terraform host provisioning, the SOPS+age secrets vault, Docker Swarm/WireGuard networking, Forgejo, Caddy, and the deploy scripts that run the already-built app image on each host. Extracted from the app monorepo (emi/thermograph) — the app repo owns building and testing the app; this repo owns running it.

  • terraform/ — provisions/configures hosts (SSH-driven by default; an optional GCP-creating module is scaffolded, no live resources yet) and triggers each deploy. See terraform/README.md.
  • deploy/secrets/ — the git-native SOPS+age secrets vault (every app secret, encrypted at rest, rendered at deploy time). See deploy/secrets/README.md.
  • deploy/swarm/, deploy/forgejo/ — the 3-node WireGuard/Swarm cluster that hosts Forgejo (git + CI + registry); does not run the app itself. See ACCESS.md and the READMEs under each directory.
  • deploy/deploy.sh — pulls the pinned app image (IMAGE_TAG) and rolls the compose stack; invoked by Terraform and by the app repo's .forgejo/workflows/deploy.yml over SSH.
  • docker-compose*.yml, docker-stack.yml — how the app image runs (compose in production today; docker-stack.yml is a design record for a possible future Swarm-based app deploy, not currently live).

The app's own source, Dockerfile, and build/test CI stay in the app repo — this repo never checks out app source; hosts only pull tagged images from the registry. See ACCESS.md for host access and the Swarm/Forgejo topology, and terraform/README.md for the day-to-day plan/apply workflow.

Branches & how changes reach each environment

  • main — what prod and beta run: their /opt/thermograph checkouts git reset --hard origin/main at the start of every deploy (deploy/deploy.sh). A merge to main reaches those hosts on the next app deploy (or a by-hand deploy.sh run); there is no separate infra deploy trigger.
  • dev — what LAN dev runs: ~/thermograph-dev resets to it via deploy/deploy-dev.sh. Keep it fast-forwarded to main (infra changes are not environment-staged today; the branches exist so LAN dev can trail or lead when needed).
  • release — currently consumed by nothing (prod tracks main, not release). It exists to mirror the app repos' dev→main→release promotion shape if per-environment infra staging is ever wanted; until then, treat main as live-everywhere.

Note the asymmetry with the app repos: app code IS environment-staged (dev→main→release maps to LAN→beta→prod via image tags), infra is not.