thermograph/observability/alloy/config.alloy
Emi Griffith 4e97d8e5dc
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
infra: split the estate into vps1/vps2, moving beta next to prod and dev onto vps1
Environments stop being machines. Until now each box WAS an environment --
"beta" named both a deploy target and a host, "the desktop" named both the
operator's computer and the dev server -- so every path could assume one
environment per host and hardcode /opt/thermograph, /etc/thermograph.env and
ports 8137/8080. That assumption ends here:

  vps1  75.119.132.91  Forgejo, Grafana/Loki, the portfolio site, and DEV
                       (own Postgres, mesh-only on 10.10.0.2:8137)
  vps2  169.58.46.181  PROD and BETA as two Swarm stacks sharing one
                       TimescaleDB instance, plus Centralis, Postfix, backups
  desktop              AI model hosting + flex Swarm capacity, no environment

Nothing live has moved; the ordered cutover is in
infra/deploy/RUNBOOK-vps1-vps2-cutover.md.

deploy/env-topology.sh is the single source of truth: env -> host, checkout,
branch, deploy mode, stack name, env file, LB ports, DB role/database, service
prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing
distinguishing a beta deploy from a prod one -- deploy.sh refuses to run when it
disagrees with the checkout it was invoked from. The host-wide secrets-env and
deploy-mode markers survive only as a fallback for a single-environment box.

Beta gets its own stack file rather than an overlay (a merge cannot REMOVE the
db service, and beta having no database of its own is the point). Its services
are prefixed beta-*: Swarm registers a service's short name as a DNS alias on
every network it joins, so two stacks both calling a service `web` on the shared
data network would let prod's frontend resolve a beta task. Prod's stack, env
and LB config are untouched.

One Postgres, two databases with two roles: deploy/db/provision-env-db.sh
creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT
revoked from PUBLIC) plus a <role>_ro for ad-hoc queries, and refuses to run for
the environment that owns the instance -- doing so would demote prod's bootstrap
superuser.

Fixes that co-residency would otherwise have broken silently:

- CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and
  also "the Forgejo box" because those shared a machine; that conflation is what
  once had the prod backup dumping beta.
- The nightly backup dumps BOTH application databases to separate off-box
  prefixes, and fails loudly on a missing one instead of skipping it.
- Alloy stopped deriving the `host` label from the node -- on vps2 that would
  have filed every beta line as prod, feeding prod's alert rules with beta's
  traffic. It is now derived per source, with a new `node` label for the machine,
  and beta's log volume is mounted via a vps2-only overlay.
- ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derived their SSH
  target and env-file path from the topology instead of hardcoding beta to
  75.119.132.91 -- which is vps1's address now.
- Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on
  vps1) instead of 0.0.0.0: correct for a home LAN box, a public exposure of
  unreviewed branches on a VPS.

Terraform's hosts variable is keyed by machine with a nested environments map,
and main.tf flattens (host, environment) pairs so two environments on one box
cannot share a checkout. Caddy's single reference config is split per host.

Docs, onboarding and the runbooks are updated throughout, including the security
rationales that co-residency makes false: separate SSH credentials no longer put
a host boundary between beta and prod, and the boundary that remains is the
database and the filesystem.
2026-07-25 15:01:29 -07:00

273 lines
11 KiB
Text

// Grafana Alloy — the log collector that runs on EVERY node (vps1, vps2). It
// gathers three sources and ships them to the central Loki on vps1:
//
// 1. Every Docker container's stdout/stderr (app, db, and on vps1 forgejo +
// the monitoring stack itself) via the Docker socket.
// 2. Caddy's host access logs (/var/log/caddy/*.log) — the reverse proxy runs
// on the host, not in a container, so its logs aren't in Docker.
// 3. The app's structured JSON logs (errors/access/audit *.jsonl) from the
// `applogs` Docker volumes, parsed so `level`/`tag`/`phase` become labels.
//
// ---------------------------------------------------------------------------
// `host` IS THE ENVIRONMENT, AND IT IS NO LONGER THE NODE
// ---------------------------------------------------------------------------
// Every query, dashboard and alert in this estate slices on host="prod" |
// "beta" | "dev". That label used to be an EXTERNAL label — one value stamped
// on everything the agent shipped — which was exact while one node ran exactly
// one environment.
//
// vps2 now runs prod AND beta. A single external label there would have marked
// every beta container, every beta access log and every beta JSON line as
// `prod`: not a cosmetic problem, but wrong data feeding the prod alert rules,
// and beta's own alerts (NotifierHeartbeatMissingBeta) firing on prod's
// traffic or never firing at all.
//
// So `host` is now derived PER SOURCE — from the container's stack prefix, from
// the access-log filename, from which log volume a JSON line came out of — and
// ALLOY_ENV supplies the fallback for anything on the node that isn't
// environment-specific (Forgejo, Grafana, Loki, the portfolio site). The node
// itself is still labelled, as `node`, because "which machine" is a real
// question too — it is just a different question from "which environment".
//
// Per-node settings come from the environment (see docker-compose.agent.yml):
// ALLOY_NODE — the MACHINE this agent runs on (vps1 | vps2 | desktop)
// ALLOY_ENV — the default environment label for logs on this node that
// aren't attributable to a specific stack (vps1 -> dev,
// vps2 -> prod)
// LOKI_URL — where to push (http://10.10.0.2:3100/loki/api/v1/push over wg0)
livedebugging { enabled = false }
// --- 1. All Docker container logs ------------------------------------------------
// refresh_interval defaults to 60s; Swarm task churn reshuffles the target set on
// roughly that cadence, which restarts tailers ~every 90s and was costing ~11% of
// Alloy's own CPU in tailer restarts alone. 5m is still fast enough to pick up a
// real deploy without paying that churn cost.
discovery.docker "containers" {
host = "unix:///var/run/docker.sock"
refresh_interval = "5m"
}
// Turn Docker metadata into tidy labels: `container` (short name) and `service`
// (the compose service, e.g. app/db). Drop noise/duplicate containers so Loki never
// ingests them: Alloy itself (loop), the autoscaler, throwaway `thermograph-test_*`
// stacks, the loopback LB bridge (`thermograph-lb`, pure plumbing, nothing to debug
// from its logs), and the app's own worker (`thermograph_worker`'s stdout is 100%
// `/healthz` poll noise — the app's `access/*.jsonl` under source #3 is a strict
// superset of anything useful it logs).
discovery.relabel "containers" {
targets = discovery.docker.containers.targets
rule {
source_labels = ["__meta_docker_container_name"]
regex = "/(.*)"
target_label = "container"
}
rule {
source_labels = ["__meta_docker_container_label_com_docker_compose_service"]
target_label = "service"
}
// --- environment attribution -------------------------------------------
// Relabel rules are applied in order and a later write to the same target
// wins, so this is "default, then override with anything more specific".
//
// The regexes are fully anchored (Prometheus relabel semantics), which is
// what keeps `/thermograph_web.1.x` and `/thermograph-beta_beta-web.1.x`
// from matching each other: prod's stack prefix ends in an underscore,
// beta's in `-beta_`.
rule {
source_labels = ["__meta_docker_container_name"]
regex = ".*"
target_label = "host"
replacement = sys.env("ALLOY_ENV")
}
rule {
source_labels = ["__meta_docker_container_name"]
regex = "/thermograph_.*"
target_label = "host"
replacement = "prod"
}
rule {
source_labels = ["__meta_docker_container_name"]
regex = "/thermograph-beta_.*"
target_label = "host"
replacement = "beta"
}
rule {
source_labels = ["__meta_docker_container_name"]
regex = "/thermograph-dev[-_].*"
target_label = "host"
replacement = "dev"
}
// Drop noise/duplicate containers, per environment. Beta's equivalents are
// listed explicitly: its LB is `thermograph-beta-lb` (which does NOT match
// `thermograph-lb.*`) and its worker is `thermograph-beta_beta-worker`, so
// omitting them would have quietly re-admitted exactly the healthz-poll
// noise the prod entries exist to exclude.
rule {
source_labels = ["container"]
regex = "(alloy|autoscaler|thermograph-test_.*|thermograph-lb|thermograph-beta-lb|thermograph_worker|thermograph-beta_beta-worker).*"
action = "drop"
}
}
// Pass RAW targets here (not discovery.relabel.containers.output) alongside
// relabel_rules: loki.source.docker applies relabel_rules itself, once, using the
// __meta_docker_* metadata it still holds at collection time. Passing the
// already-relabelled output *and* relabel_rules ran the same rules twice per log
// entry, and made the drop rules above a no-op on the second pass since the
// __meta_docker_* labels are already gone from the pre-relabelled output.
loki.source.docker "containers" {
host = "unix:///var/run/docker.sock"
targets = discovery.docker.containers.targets
forward_to = [loki.write.central.receiver]
relabel_rules = discovery.relabel.containers.rules
labels = { job = "docker" }
}
// --- 2. Caddy host access logs ---------------------------------------------------
// sync_period (glob rescan) defaults to 10s; 1m is plenty for a log file that only
// appears/rotates on the order of hours.
local.file_match "caddy" {
path_targets = [{ __path__ = "/var/log/caddy/*.log", job = "caddy" }]
sync_period = "1m"
}
loki.source.file "caddy" {
targets = local.file_match.caddy.targets
forward_to = [loki.process.caddy.receiver]
// PollingFileWatcher defaults (250ms/250ms) stat every tailed file 4x/second
// forever. Caddy's access log doesn't need sub-second latency into Loki.
file_watch {
min_poll_frequency = "2s"
max_poll_frequency = "10s"
}
}
// Drop well-known crawler/bot traffic before it hits Loki. Caddy itself can't do
// this cheaply (log_skip needs Caddy >= 2.7; both hosts run older Caddy, and an
// upgrade is out of scope here), so filter it at the shipper instead.
loki.process "caddy" {
forward_to = [loki.write.central.receiver]
stage.drop {
expression = "(?i)(semrushbot|claudebot|ahrefsbot|yandexbot|bytespider|mj12bot|petalbot)"
drop_counter_reason = "crawler"
}
// Attribute each access log to an environment BY FILENAME. On vps2 one
// Caddy fronts both thermograph.org and beta.thermograph.org, so a single
// node-wide label would file every beta request under prod.
//
// By filename rather than by parsing `request.host` out of the JSON: the
// site block that writes the file is the same place the hostname is
// declared (see infra/deploy/Caddyfile.vps2), one label costs nothing to
// evaluate, and it keeps working for a non-JSON log format.
stage.static_labels {
values = { host = sys.env("ALLOY_ENV") }
}
stage.match {
selector = "{filename=~\".*/beta[.]log\"}"
stage.static_labels {
values = { host = "beta" }
}
}
stage.match {
selector = "{filename=~\".*/thermograph[.]log\"}"
stage.static_labels {
values = { host = "prod" }
}
}
}
// --- 3. App structured JSON logs (errors / access / audit) -----------------------
// Mounted read-only from the app's `applogs` volumes (see the agent compose).
// Lift `level`/`tag`/`phase` out of the JSON so they're queryable.
//
// TWO MOUNTS, NOT ONE. Each environment has its own applogs volume, and on vps2
// both exist side by side:
//
// /applogs this node's primary environment (vps2 -> prod, vps1 -> dev)
// /applogs-beta beta's, mounted on vps2 only
//
// A single mount was structurally unable to carry both: whichever volume the
// agent bound, the other environment's structured logs would never reach Loki
// at all — silently, since a missing file source is not an error. Every alert
// built on the JSON stream (the notifier heartbeats in particular) reads this.
local.file_match "app_jsonl" {
path_targets = [{ __path__ = "/applogs/**/*.jsonl", job = "app-json" }]
sync_period = "1m"
}
loki.source.file "app_jsonl" {
targets = local.file_match.app_jsonl.targets
forward_to = [loki.process.app_jsonl.receiver]
file_watch {
min_poll_frequency = "2s"
max_poll_frequency = "10s"
}
}
loki.process "app_jsonl" {
forward_to = [loki.write.central.receiver]
stage.json {
expressions = { level = "level", tag = "tag", phase = "phase", status = "status" }
}
// A record with no explicit level: an error-folder line is an error, else info.
stage.static_labels {
values = { source = "app", host = sys.env("ALLOY_ENV") }
}
stage.labels {
values = { level = "", tag = "", phase = "" }
}
}
// Beta's structured logs. Absent on vps1 (the mount simply isn't there), where
// this matches nothing and costs nothing.
local.file_match "app_jsonl_beta" {
path_targets = [{ __path__ = "/applogs-beta/**/*.jsonl", job = "app-json" }]
sync_period = "1m"
}
loki.source.file "app_jsonl_beta" {
targets = local.file_match.app_jsonl_beta.targets
forward_to = [loki.process.app_jsonl_beta.receiver]
file_watch {
min_poll_frequency = "2s"
max_poll_frequency = "10s"
}
}
loki.process "app_jsonl_beta" {
forward_to = [loki.write.central.receiver]
stage.json {
expressions = { level = "level", tag = "tag", phase = "phase", status = "status" }
}
stage.static_labels {
values = { source = "app", host = "beta" }
}
stage.labels {
values = { level = "", tag = "", phase = "" }
}
}
// --- Ship to the central Loki over the WireGuard mesh ----------------------------
loki.write "central" {
endpoint {
url = sys.env("LOKI_URL")
}
// `node` — the MACHINE, stamped on everything this agent ships. `host` is
// deliberately NOT here: it is the ENVIRONMENT, set per source above,
// because vps2 carries two of them and an external label cannot vary per
// stream. Setting host here would override that work on every line.
external_labels = { node = sys.env("ALLOY_NODE") }
}