Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
# /etc/caddy/Caddyfile — RENDERED BY TERRAFORM for ${domain}.
|
|
|
|
|
# Caddy provisions a Let's Encrypt cert on first request and auto-renews; the domain's
|
|
|
|
|
# A/AAAA record must already point at this host and ports 80/443 must be open.
|
|
|
|
|
#
|
2026-07-21 20:01:30 +00:00
|
|
|
# The app owns the domain root and runs with THERMOGRAPH_BASE=/. Repo-split
|
|
|
|
|
# Stage 4: backend (127.0.0.1:${port}) and frontend (127.0.0.1:${frontend_port})
|
2026-07-21 22:48:59 +00:00
|
|
|
# are two containers now -- path-split directly to whichever owns a given
|
|
|
|
|
# path. Repo-split Stage 7a flipped which side owns the enumerated list:
|
|
|
|
|
# backend owns the short, stable set below (mirrors backend/web/app.py's own
|
|
|
|
|
# routing exactly -- a single catch-all proxy to frontend for everything
|
|
|
|
|
# else), so the two never disagree about who owns what.
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
#
|
|
|
|
|
# NOTE: the repo's deploy/Caddyfile additionally serves the emigriffith.dev portfolio
|
|
|
|
|
# and legacy redirects; those are host-specific and intentionally not templated here.
|
Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping
thermograph-observability (and thermograph-backend/frontend/infra) are
archived on Forgejo now that the monorepo cutover has landed, so this domain's
work lands here in observability/ instead of a standalone repo PR.
Alloy (observability/alloy/config.alloy), deployed from /opt/observability on
prod and beta (confirmed via docker inspect mounts on the live hosts; the
monorepo's own observability/ copy isn't what's running yet -- that's a
separate, not-yet-done cutover step for this domain, flagged below):
- Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from
the 250ms/250ms PollingFileWatcher default to 2s/10s, and
local.file_match's glob-rescan sync_period from 10s to 1m.
- Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task
churn (~every 90s) stops restarting tailers on every refresh.
- Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*,
thermograph-lb, thermograph_worker) via a relabel rule; the worker's
stdout is pure /healthz noise, already superseded by its own
access/*.jsonl.
- Fix loki.source.docker double-relabeling: it took both the already-
relabelled target output AND relabel_rules, running every rule twice per
entry and making the new drop rules a no-op on the second pass. Pass raw
discovery targets + relabel_rules instead.
- Add a loki.process stage for Caddy logs that drops well-known crawler/bot
lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot,
petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older
(beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth
addressing separately), so this is the shipper-side equivalent.
Loki (observability/loki/config.yml): 30 low-rate streams were hitting the
30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98%
full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set
chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream
limits sized for a three-node hobby fleet -- their absence was the source
of the sporadic 429s (Loki was falling back to its stricter multi-tenant
defaults).
CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no
syntax coverage. Download the pinned v1.9.1 alloy binary (matches
observability/alloy/docker-compose.agent.yml) straight from its GitHub
release and run `alloy validate` (minus the one line covering a known gap
where validate doesn't recognize the top-level livedebugging singleton
block, unlike alloy run).
Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl,
the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose
deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile
is the documented manual-apply fallback, updated in step to avoid drift):
- health_interval 5s -> 15s: a quarter of the active-healthcheck polling
load for the same restart-safety guarantee.
- A `filter` format log encoder deleting request>headers, request>tls, and
resp_headers (the full-header/TLS-block serialization was the main driver
of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the
file itself doesn't grow unbounded between logrotate runs.
- Strip the query string entirely from logged request URIs via a `regexp`
field filter -- Caddy's default logger recorded request.uri *including*
the query string, so every `?q=<search text>` sat in Loki next to the
client IP for the full 30-day retention. Confirmed empirically (via
`caddy validate` and a scratch `caddy run` against beta's real 2.6.2
binary) that keeping request>headers>User-Agent alongside a parent-level
header delete isn't achievable in this Caddy version -- deleting the
parent drops the whole subtree regardless of a child override, and
`log_append` (which would let it be hoisted out first) needs a newer
Caddy than beta runs. Traded away rather than block on it.
Deferred (host-level, not in any repo, applied directly and separately):
/etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts --
config written, but NOT restarting dockerd on either: prod is the fleet's
sole Swarm manager with Live Restore disabled, so a daemon restart there
would both interrupt Swarm's control plane and actually stop/restart every
container on the host (no live-restore to keep them up), a real production
outage window, not a blip. Left for a planned maintenance window instead.
A 14-day retention cron for the app's applogs JSONL volume is applied on
both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete
rather than logrotate, since the app already self-dates one file per day
per category and logrotate's rotate-in-place model doesn't fit files it
doesn't own the naming of.
Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf,
dashboard.thermograph.org's Caddy block) is untouched by any of this.
2026-07-24 04:34:06 +00:00
|
|
|
#
|
|
|
|
|
# Access-log hygiene: the default JSON encoder serializes full request headers, the
|
|
|
|
|
# TLS block, and response headers on every line (measured ~1,133B/line on prod,
|
|
|
|
|
# ~922B on beta) -- strip those with the `filter` format encoder. Also strip the
|
|
|
|
|
# query string from the logged URI: Caddy's default logger records request.uri
|
|
|
|
|
# *including* the query string, so every `?q=<search text>` a visitor typed sat in
|
|
|
|
|
# Loki next to their client IP for the full 30-day retention -- a real privacy
|
|
|
|
|
# leak, not just noise. Bot/crawler skipping stays out of here: `log_skip` needs
|
|
|
|
|
# Caddy >= 2.7 and an upgrade is out of scope, so that's handled downstream in
|
|
|
|
|
# Alloy's loki.process "caddy" stage instead (see observability/alloy/config.alloy).
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
${domain} {
|
|
|
|
|
encode zstd gzip
|
|
|
|
|
|
2026-07-21 22:48:59 +00:00
|
|
|
@backend_paths path /api/* /digest /discord/interactions
|
2026-07-21 20:01:30 +00:00
|
|
|
|
|
|
|
|
# Active health check on the same cheap /healthz route each container's own
|
Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235)
* Split web/worker duties with THERMOGRAPH_ROLE
Background work (the subscription notifier) is welded to the same process
that serves requests, so scaling the web tier to N replicas would also scale
notifier instances unless something restricts it further than leader
election alone.
Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process
behavior). Every replica runs the same image; ROLE only gates whether a
process is allowed to own the notifier at all, layered on top of the
existing leader election: web replicas never start it even if they'd win
leader election, worker replicas start it if they win. The decision is
pulled into _should_run_notifier() so it's unit-testable without booting the
full app (DB init, places index, neighbor warmer).
Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE)
so a worker replica - which serves no real traffic - still has something
Swarm can health-check.
* Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate
Three changes toward the hop-1 interim cutover, all inert until Track B
stands up the platform:
docker-stack.yml: the Swarm stack file for the interim cutover, distinct
from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a
pre-built image (IMAGE_TAG) instead of building in place; app/worker publish
no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing
mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted),
reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard
needs a smaller MTU or large payloads silently stall); db is placement-
pinned to a labelled node; app/worker skip inline migrations
(RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing
that ever runs Alembic; secrets are real Swarm secrets mounted at
/run/secrets, read by the entrypoint shim rather than plain env vars.
TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default
latest-pg18, today's behavior unchanged), wired through Terraform
(timescaledb_tag, default "latest-pg18") so it can actually be pinned to an
exact minor without hand-editing the host - required before any host of the
stack could replicate with another (a floating tag risks mismatched
extension minors, which blocks a physical replica and risks compressed-
chunk corruption on restore).
Caddy active health-gate: both the Terraform-rendered Caddyfile and the live
deploy/Caddyfile now health-check the app on the same cheap /healthz route
its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage,
so it's cheap enough for a tight interval and works identically for a
worker replica, which serves no public traffic at all) - Caddy won't
forward into a container that's still booting or unhealthy.
Verified live: built and booted the real image via docker compose - both
containers report healthy via the new /healthz-based HEALTHCHECK, and GET /
still renders the full SSR homepage unchanged. Both Caddyfiles validated
with the real caddy binary. docker-stack.yml validated with docker compose
config (required-var guards fire with clear messages; secrets correctly
mount at /run/secrets/<name>, matching the entrypoint shim's mapping).
docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside
the existing openmeteo overlay. terraform validate + fmt clean.
2026-07-21 00:39:48 +00:00
|
|
|
# HEALTHCHECK uses (Dockerfile) — so a `docker compose up -d --build` deploy
|
|
|
|
|
# that's still restarting/booting never gets proxied into (a reload alone has
|
|
|
|
|
# no gate, hop-1 runbook hazard #10). health_uri is relative to the upstream.
|
Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping
thermograph-observability (and thermograph-backend/frontend/infra) are
archived on Forgejo now that the monorepo cutover has landed, so this domain's
work lands here in observability/ instead of a standalone repo PR.
Alloy (observability/alloy/config.alloy), deployed from /opt/observability on
prod and beta (confirmed via docker inspect mounts on the live hosts; the
monorepo's own observability/ copy isn't what's running yet -- that's a
separate, not-yet-done cutover step for this domain, flagged below):
- Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from
the 250ms/250ms PollingFileWatcher default to 2s/10s, and
local.file_match's glob-rescan sync_period from 10s to 1m.
- Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task
churn (~every 90s) stops restarting tailers on every refresh.
- Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*,
thermograph-lb, thermograph_worker) via a relabel rule; the worker's
stdout is pure /healthz noise, already superseded by its own
access/*.jsonl.
- Fix loki.source.docker double-relabeling: it took both the already-
relabelled target output AND relabel_rules, running every rule twice per
entry and making the new drop rules a no-op on the second pass. Pass raw
discovery targets + relabel_rules instead.
- Add a loki.process stage for Caddy logs that drops well-known crawler/bot
lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot,
petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older
(beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth
addressing separately), so this is the shipper-side equivalent.
Loki (observability/loki/config.yml): 30 low-rate streams were hitting the
30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98%
full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set
chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream
limits sized for a three-node hobby fleet -- their absence was the source
of the sporadic 429s (Loki was falling back to its stricter multi-tenant
defaults).
CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no
syntax coverage. Download the pinned v1.9.1 alloy binary (matches
observability/alloy/docker-compose.agent.yml) straight from its GitHub
release and run `alloy validate` (minus the one line covering a known gap
where validate doesn't recognize the top-level livedebugging singleton
block, unlike alloy run).
Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl,
the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose
deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile
is the documented manual-apply fallback, updated in step to avoid drift):
- health_interval 5s -> 15s: a quarter of the active-healthcheck polling
load for the same restart-safety guarantee.
- A `filter` format log encoder deleting request>headers, request>tls, and
resp_headers (the full-header/TLS-block serialization was the main driver
of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the
file itself doesn't grow unbounded between logrotate runs.
- Strip the query string entirely from logged request URIs via a `regexp`
field filter -- Caddy's default logger recorded request.uri *including*
the query string, so every `?q=<search text>` sat in Loki next to the
client IP for the full 30-day retention. Confirmed empirically (via
`caddy validate` and a scratch `caddy run` against beta's real 2.6.2
binary) that keeping request>headers>User-Agent alongside a parent-level
header delete isn't achievable in this Caddy version -- deleting the
parent drops the whole subtree regardless of a child override, and
`log_append` (which would let it be hoisted out first) needs a newer
Caddy than beta runs. Traded away rather than block on it.
Deferred (host-level, not in any repo, applied directly and separately):
/etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts --
config written, but NOT restarting dockerd on either: prod is the fleet's
sole Swarm manager with Live Restore disabled, so a daemon restart there
would both interrupt Swarm's control plane and actually stop/restart every
container on the host (no live-restore to keep them up), a real production
outage window, not a blip. Left for a planned maintenance window instead.
A 14-day retention cron for the app's applogs JSONL volume is applied on
both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete
rather than logrotate, since the app already self-dates one file per day
per category and logrotate's rotate-in-place model doesn't fit files it
doesn't own the naming of.
Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf,
dashboard.thermograph.org's Caddy block) is untouched by any of this.
2026-07-24 04:34:06 +00:00
|
|
|
# 15s (was 5s): plenty responsive for an active health check against a process
|
|
|
|
|
# that only restarts on a deploy, and a quarter of the polling load.
|
2026-07-21 22:48:59 +00:00
|
|
|
handle @backend_paths {
|
|
|
|
|
reverse_proxy 127.0.0.1:${port} {
|
2026-07-21 20:01:30 +00:00
|
|
|
health_uri /healthz
|
Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping
thermograph-observability (and thermograph-backend/frontend/infra) are
archived on Forgejo now that the monorepo cutover has landed, so this domain's
work lands here in observability/ instead of a standalone repo PR.
Alloy (observability/alloy/config.alloy), deployed from /opt/observability on
prod and beta (confirmed via docker inspect mounts on the live hosts; the
monorepo's own observability/ copy isn't what's running yet -- that's a
separate, not-yet-done cutover step for this domain, flagged below):
- Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from
the 250ms/250ms PollingFileWatcher default to 2s/10s, and
local.file_match's glob-rescan sync_period from 10s to 1m.
- Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task
churn (~every 90s) stops restarting tailers on every refresh.
- Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*,
thermograph-lb, thermograph_worker) via a relabel rule; the worker's
stdout is pure /healthz noise, already superseded by its own
access/*.jsonl.
- Fix loki.source.docker double-relabeling: it took both the already-
relabelled target output AND relabel_rules, running every rule twice per
entry and making the new drop rules a no-op on the second pass. Pass raw
discovery targets + relabel_rules instead.
- Add a loki.process stage for Caddy logs that drops well-known crawler/bot
lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot,
petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older
(beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth
addressing separately), so this is the shipper-side equivalent.
Loki (observability/loki/config.yml): 30 low-rate streams were hitting the
30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98%
full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set
chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream
limits sized for a three-node hobby fleet -- their absence was the source
of the sporadic 429s (Loki was falling back to its stricter multi-tenant
defaults).
CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no
syntax coverage. Download the pinned v1.9.1 alloy binary (matches
observability/alloy/docker-compose.agent.yml) straight from its GitHub
release and run `alloy validate` (minus the one line covering a known gap
where validate doesn't recognize the top-level livedebugging singleton
block, unlike alloy run).
Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl,
the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose
deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile
is the documented manual-apply fallback, updated in step to avoid drift):
- health_interval 5s -> 15s: a quarter of the active-healthcheck polling
load for the same restart-safety guarantee.
- A `filter` format log encoder deleting request>headers, request>tls, and
resp_headers (the full-header/TLS-block serialization was the main driver
of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the
file itself doesn't grow unbounded between logrotate runs.
- Strip the query string entirely from logged request URIs via a `regexp`
field filter -- Caddy's default logger recorded request.uri *including*
the query string, so every `?q=<search text>` sat in Loki next to the
client IP for the full 30-day retention. Confirmed empirically (via
`caddy validate` and a scratch `caddy run` against beta's real 2.6.2
binary) that keeping request>headers>User-Agent alongside a parent-level
header delete isn't achievable in this Caddy version -- deleting the
parent drops the whole subtree regardless of a child override, and
`log_append` (which would let it be hoisted out first) needs a newer
Caddy than beta runs. Traded away rather than block on it.
Deferred (host-level, not in any repo, applied directly and separately):
/etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts --
config written, but NOT restarting dockerd on either: prod is the fleet's
sole Swarm manager with Live Restore disabled, so a daemon restart there
would both interrupt Swarm's control plane and actually stop/restart every
container on the host (no live-restore to keep them up), a real production
outage window, not a blip. Left for a planned maintenance window instead.
A 14-day retention cron for the app's applogs JSONL volume is applied on
both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete
rather than logrotate, since the app already self-dates one file per day
per category and logrotate's rotate-in-place model doesn't fit files it
doesn't own the naming of.
Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf,
dashboard.thermograph.org's Caddy block) is untouched by any of this.
2026-07-24 04:34:06 +00:00
|
|
|
health_interval 15s
|
2026-07-21 20:01:30 +00:00
|
|
|
health_timeout 3s
|
|
|
|
|
health_status 2xx
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
handle {
|
2026-07-21 22:48:59 +00:00
|
|
|
reverse_proxy 127.0.0.1:${frontend_port} {
|
2026-07-21 20:01:30 +00:00
|
|
|
health_uri /healthz
|
Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping
thermograph-observability (and thermograph-backend/frontend/infra) are
archived on Forgejo now that the monorepo cutover has landed, so this domain's
work lands here in observability/ instead of a standalone repo PR.
Alloy (observability/alloy/config.alloy), deployed from /opt/observability on
prod and beta (confirmed via docker inspect mounts on the live hosts; the
monorepo's own observability/ copy isn't what's running yet -- that's a
separate, not-yet-done cutover step for this domain, flagged below):
- Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from
the 250ms/250ms PollingFileWatcher default to 2s/10s, and
local.file_match's glob-rescan sync_period from 10s to 1m.
- Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task
churn (~every 90s) stops restarting tailers on every refresh.
- Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*,
thermograph-lb, thermograph_worker) via a relabel rule; the worker's
stdout is pure /healthz noise, already superseded by its own
access/*.jsonl.
- Fix loki.source.docker double-relabeling: it took both the already-
relabelled target output AND relabel_rules, running every rule twice per
entry and making the new drop rules a no-op on the second pass. Pass raw
discovery targets + relabel_rules instead.
- Add a loki.process stage for Caddy logs that drops well-known crawler/bot
lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot,
petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older
(beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth
addressing separately), so this is the shipper-side equivalent.
Loki (observability/loki/config.yml): 30 low-rate streams were hitting the
30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98%
full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set
chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream
limits sized for a three-node hobby fleet -- their absence was the source
of the sporadic 429s (Loki was falling back to its stricter multi-tenant
defaults).
CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no
syntax coverage. Download the pinned v1.9.1 alloy binary (matches
observability/alloy/docker-compose.agent.yml) straight from its GitHub
release and run `alloy validate` (minus the one line covering a known gap
where validate doesn't recognize the top-level livedebugging singleton
block, unlike alloy run).
Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl,
the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose
deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile
is the documented manual-apply fallback, updated in step to avoid drift):
- health_interval 5s -> 15s: a quarter of the active-healthcheck polling
load for the same restart-safety guarantee.
- A `filter` format log encoder deleting request>headers, request>tls, and
resp_headers (the full-header/TLS-block serialization was the main driver
of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the
file itself doesn't grow unbounded between logrotate runs.
- Strip the query string entirely from logged request URIs via a `regexp`
field filter -- Caddy's default logger recorded request.uri *including*
the query string, so every `?q=<search text>` sat in Loki next to the
client IP for the full 30-day retention. Confirmed empirically (via
`caddy validate` and a scratch `caddy run` against beta's real 2.6.2
binary) that keeping request>headers>User-Agent alongside a parent-level
header delete isn't achievable in this Caddy version -- deleting the
parent drops the whole subtree regardless of a child override, and
`log_append` (which would let it be hoisted out first) needs a newer
Caddy than beta runs. Traded away rather than block on it.
Deferred (host-level, not in any repo, applied directly and separately):
/etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts --
config written, but NOT restarting dockerd on either: prod is the fleet's
sole Swarm manager with Live Restore disabled, so a daemon restart there
would both interrupt Swarm's control plane and actually stop/restart every
container on the host (no live-restore to keep them up), a real production
outage window, not a blip. Left for a planned maintenance window instead.
A 14-day retention cron for the app's applogs JSONL volume is applied on
both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete
rather than logrotate, since the app already self-dates one file per day
per category and logrotate's rotate-in-place model doesn't fit files it
doesn't own the naming of.
Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf,
dashboard.thermograph.org's Caddy block) is untouched by any of this.
2026-07-24 04:34:06 +00:00
|
|
|
health_interval 15s
|
2026-07-21 20:01:30 +00:00
|
|
|
health_timeout 3s
|
|
|
|
|
health_status 2xx
|
|
|
|
|
}
|
Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235)
* Split web/worker duties with THERMOGRAPH_ROLE
Background work (the subscription notifier) is welded to the same process
that serves requests, so scaling the web tier to N replicas would also scale
notifier instances unless something restricts it further than leader
election alone.
Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process
behavior). Every replica runs the same image; ROLE only gates whether a
process is allowed to own the notifier at all, layered on top of the
existing leader election: web replicas never start it even if they'd win
leader election, worker replicas start it if they win. The decision is
pulled into _should_run_notifier() so it's unit-testable without booting the
full app (DB init, places index, neighbor warmer).
Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE)
so a worker replica - which serves no real traffic - still has something
Swarm can health-check.
* Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate
Three changes toward the hop-1 interim cutover, all inert until Track B
stands up the platform:
docker-stack.yml: the Swarm stack file for the interim cutover, distinct
from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a
pre-built image (IMAGE_TAG) instead of building in place; app/worker publish
no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing
mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted),
reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard
needs a smaller MTU or large payloads silently stall); db is placement-
pinned to a labelled node; app/worker skip inline migrations
(RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing
that ever runs Alembic; secrets are real Swarm secrets mounted at
/run/secrets, read by the entrypoint shim rather than plain env vars.
TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default
latest-pg18, today's behavior unchanged), wired through Terraform
(timescaledb_tag, default "latest-pg18") so it can actually be pinned to an
exact minor without hand-editing the host - required before any host of the
stack could replicate with another (a floating tag risks mismatched
extension minors, which blocks a physical replica and risks compressed-
chunk corruption on restore).
Caddy active health-gate: both the Terraform-rendered Caddyfile and the live
deploy/Caddyfile now health-check the app on the same cheap /healthz route
its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage,
so it's cheap enough for a tight interval and works identically for a
worker replica, which serves no public traffic at all) - Caddy won't
forward into a container that's still booting or unhealthy.
Verified live: built and booted the real image via docker compose - both
containers report healthy via the new /healthz-based HEALTHCHECK, and GET /
still renders the full SSR homepage unchanged. Both Caddyfiles validated
with the real caddy binary. docker-stack.yml validated with docker compose
config (required-var guards fire with clear messages; secrets correctly
mount at /run/secrets/<name>, matching the entrypoint shim's mapping).
docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside
the existing openmeteo overlay. terraform validate + fmt clean.
2026-07-21 00:39:48 +00:00
|
|
|
}
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
|
|
|
|
|
log {
|
Log hygiene: Alloy CPU, Loki chunks/limits, Caddy field-stripping
thermograph-observability (and thermograph-backend/frontend/infra) are
archived on Forgejo now that the monorepo cutover has landed, so this domain's
work lands here in observability/ instead of a standalone repo PR.
Alloy (observability/alloy/config.alloy), deployed from /opt/observability on
prod and beta (confirmed via docker inspect mounts on the live hosts; the
monorepo's own observability/ copy isn't what's running yet -- that's a
separate, not-yet-done cutover step for this domain, flagged below):
- Raise loki.source.file's file_watch poll interval (caddy + app_jsonl) from
the 250ms/250ms PollingFileWatcher default to 2s/10s, and
local.file_match's glob-rescan sync_period from 10s to 1m.
- Raise discovery.docker's refresh_interval from 60s to 5m so Swarm task
churn (~every 90s) stops restarting tailers on every refresh.
- Drop noise/duplicate containers (alloy, autoscaler, thermograph-test_*,
thermograph-lb, thermograph_worker) via a relabel rule; the worker's
stdout is pure /healthz noise, already superseded by its own
access/*.jsonl.
- Fix loki.source.docker double-relabeling: it took both the already-
relabelled target output AND relabel_rules, running every rule twice per
entry and making the new drop rules a no-op on the second pass. Pass raw
discovery targets + relabel_rules instead.
- Add a loki.process stage for Caddy logs that drops well-known crawler/bot
lines (semrushbot, claudebot, ahrefsbot, yandexbot, bytespider, mj12bot,
petalbot) -- Caddy's own log_skip needs >= 2.7 and both hosts are older
(beta confirmed 2.6.2; prod is actually 2.11.4, a version-skew worth
addressing separately), so this is the shipper-side equivalent.
Loki (observability/loki/config.yml): 30 low-rate streams were hitting the
30m chunk_idle_period long before chunk_target_size, flushing chunks 1.98%
full on average. Raise chunk_idle_period to 2h, max_chunk_age to 12h, set
chunk_target_size/chunk_encoding explicitly. Add explicit ingestion/stream
limits sized for a three-node hobby fleet -- their absence was the source
of the sporadic 429s (Loki was falling back to its stricter multi-tenant
defaults).
CI (.forgejo/workflows/observability-validate.yml): the Alloy config had no
syntax coverage. Download the pinned v1.9.1 alloy binary (matches
observability/alloy/docker-compose.agent.yml) straight from its GitHub
release and run `alloy validate` (minus the one line covering a known gap
where validate doesn't recognize the top-level livedebugging singleton
block, unlike alloy run).
Caddy (infra/terraform/modules/thermograph-host/templates/Caddyfile.tftpl,
the true source for /etc/caddy/Caddyfile -- confirmed live on prod, whose
deployed file's own header says "RENDERED BY TERRAFORM"; infra/deploy/Caddyfile
is the documented manual-apply fallback, updated in step to avoid drift):
- health_interval 5s -> 15s: a quarter of the active-healthcheck polling
load for the same restart-safety guarantee.
- A `filter` format log encoder deleting request>headers, request>tls, and
resp_headers (the full-header/TLS-block serialization was the main driver
of ~1,133B/line on prod, ~922B on beta), plus roll_size/roll_keep so the
file itself doesn't grow unbounded between logrotate runs.
- Strip the query string entirely from logged request URIs via a `regexp`
field filter -- Caddy's default logger recorded request.uri *including*
the query string, so every `?q=<search text>` sat in Loki next to the
client IP for the full 30-day retention. Confirmed empirically (via
`caddy validate` and a scratch `caddy run` against beta's real 2.6.2
binary) that keeping request>headers>User-Agent alongside a parent-level
header delete isn't achievable in this Caddy version -- deleting the
parent drops the whole subtree regardless of a child override, and
`log_append` (which would let it be hoisted out first) needs a newer
Caddy than beta runs. Traded away rather than block on it.
Deferred (host-level, not in any repo, applied directly and separately):
/etc/docker/daemon.json (10m/3-file json-file log caps) on both hosts --
config written, but NOT restarting dockerd on either: prod is the fleet's
sole Swarm manager with Live Restore disabled, so a daemon restart there
would both interrupt Swarm's control plane and actually stop/restart every
container on the host (no live-restore to keep them up), a real production
outage window, not a blip. Left for a planned maintenance window instead.
A 14-day retention cron for the app's applogs JSONL volume is applied on
both hosts (/etc/cron.d/thermograph-applogs-retention) -- plain find+delete
rather than logrotate, since the app already self-dates one file per day
per category and logrotate's rotate-in-place model doesn't fit files it
doesn't own the naming of.
Grafana itself (docker-compose.yml, grafana/, caddy-grafana.conf,
dashboard.thermograph.org's Caddy block) is untouched by any of this.
2026-07-24 04:34:06 +00:00
|
|
|
output file /var/log/caddy/thermograph.log {
|
|
|
|
|
roll_size 20MiB
|
|
|
|
|
roll_keep 5
|
|
|
|
|
}
|
|
|
|
|
format filter {
|
|
|
|
|
wrap json
|
|
|
|
|
fields {
|
|
|
|
|
request>headers delete
|
|
|
|
|
request>tls delete
|
|
|
|
|
resp_headers delete
|
|
|
|
|
request>uri regexp \?.* ""
|
|
|
|
|
}
|
|
|
|
|
}
|
Add Terraform to provision the VPS hosts (compose keeps running the app) (#223)
Terraform config under terraform/ manages the two existing VPS hosts and hands the
app to docker-compose, with local state:
- prod: the new 48GB/12-core VPS (release branch, thermograph.org), sized larger.
- beta: the old VPS 75.119.132.91 (main branch, testing tier), no public domain.
- The LAN dev box stays on deploy/deploy-dev.sh (dev branch) — out of Terraform.
A reusable module (modules/thermograph-host) SSHes each host to install docker/
compose/ufw (+ Caddy when a domain is set), sync the checkout to the host's branch,
render /etc/thermograph.env from Terraform variables (secrets pushed via provisioner
content, never on local disk), `docker compose up -d`, and health-check. Named
volumes are preserved on re-apply, so the Postgres data is never recreated.
Container resources are now env-driven in docker-compose.yml (APP_CPUS/DB_CPUS/
DB_MEMORY/WORKERS) with unchanged defaults, so Terraform can size each host.
2026-07-20 07:42:15 +00:00
|
|
|
}
|
|
|
|
|
}
|