2026-07-26 06:56:38 +00:00
|
|
|
// Grafana Alloy — the log collector that runs on EVERY node (vps1, vps2). It
|
|
|
|
|
// gathers three sources and ships them to the central Loki on vps1:
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
//
|
2026-07-26 06:56:38 +00:00
|
|
|
// 1. Every Docker container's stdout/stderr (app, db, and on vps1 forgejo +
|
|
|
|
|
// the monitoring stack itself) via the Docker socket.
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
// 2. Caddy's host access logs (/var/log/caddy/*.log) — the reverse proxy runs
|
|
|
|
|
// on the host, not in a container, so its logs aren't in Docker.
|
|
|
|
|
// 3. The app's structured JSON logs (errors/access/audit *.jsonl) from the
|
2026-07-26 06:56:38 +00:00
|
|
|
// `applogs` Docker volumes, parsed so `level`/`tag`/`phase` become labels.
|
|
|
|
|
//
|
|
|
|
|
// ---------------------------------------------------------------------------
|
|
|
|
|
// `host` IS THE ENVIRONMENT, AND IT IS NO LONGER THE NODE
|
|
|
|
|
// ---------------------------------------------------------------------------
|
|
|
|
|
// Every query, dashboard and alert in this estate slices on host="prod" |
|
|
|
|
|
// "beta" | "dev". That label used to be an EXTERNAL label — one value stamped
|
|
|
|
|
// on everything the agent shipped — which was exact while one node ran exactly
|
|
|
|
|
// one environment.
|
|
|
|
|
//
|
|
|
|
|
// vps2 now runs prod AND beta. A single external label there would have marked
|
|
|
|
|
// every beta container, every beta access log and every beta JSON line as
|
|
|
|
|
// `prod`: not a cosmetic problem, but wrong data feeding the prod alert rules,
|
|
|
|
|
// and beta's own alerts (NotifierHeartbeatMissingBeta) firing on prod's
|
|
|
|
|
// traffic or never firing at all.
|
|
|
|
|
//
|
|
|
|
|
// So `host` is now derived PER SOURCE — from the container's stack prefix, from
|
|
|
|
|
// the access-log filename, from which log volume a JSON line came out of — and
|
|
|
|
|
// ALLOY_ENV supplies the fallback for anything on the node that isn't
|
|
|
|
|
// environment-specific (Forgejo, Grafana, Loki, the portfolio site). The node
|
|
|
|
|
// itself is still labelled, as `node`, because "which machine" is a real
|
|
|
|
|
// question too — it is just a different question from "which environment".
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
//
|
|
|
|
|
// Per-node settings come from the environment (see docker-compose.agent.yml):
|
2026-07-26 06:56:38 +00:00
|
|
|
// ALLOY_NODE — the MACHINE this agent runs on (vps1 | vps2 | desktop)
|
|
|
|
|
// ALLOY_ENV — the default environment label for logs on this node that
|
|
|
|
|
// aren't attributable to a specific stack (vps1 -> dev,
|
|
|
|
|
// vps2 -> prod)
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
// LOKI_URL — where to push (http://10.10.0.2:3100/loki/api/v1/push over wg0)
|
|
|
|
|
|
|
|
|
|
livedebugging { enabled = false }
|
|
|
|
|
|
|
|
|
|
// --- 1. All Docker container logs ------------------------------------------------
|
2026-07-24 04:37:41 +00:00
|
|
|
// refresh_interval defaults to 60s; Swarm task churn reshuffles the target set on
|
|
|
|
|
// roughly that cadence, which restarts tailers ~every 90s and was costing ~11% of
|
|
|
|
|
// Alloy's own CPU in tailer restarts alone. 5m is still fast enough to pick up a
|
|
|
|
|
// real deploy without paying that churn cost.
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
discovery.docker "containers" {
|
2026-07-24 04:37:41 +00:00
|
|
|
host = "unix:///var/run/docker.sock"
|
|
|
|
|
refresh_interval = "5m"
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
// Turn Docker metadata into tidy labels: `container` (short name) and `service`
|
2026-07-24 04:37:41 +00:00
|
|
|
// (the compose service, e.g. app/db). Drop noise/duplicate containers so Loki never
|
|
|
|
|
// ingests them: Alloy itself (loop), the autoscaler, throwaway `thermograph-test_*`
|
|
|
|
|
// stacks, the loopback LB bridge (`thermograph-lb`, pure plumbing, nothing to debug
|
|
|
|
|
// from its logs), and the app's own worker (`thermograph_worker`'s stdout is 100%
|
|
|
|
|
// `/healthz` poll noise — the app's `access/*.jsonl` under source #3 is a strict
|
|
|
|
|
// superset of anything useful it logs).
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
discovery.relabel "containers" {
|
|
|
|
|
targets = discovery.docker.containers.targets
|
|
|
|
|
|
|
|
|
|
rule {
|
|
|
|
|
source_labels = ["__meta_docker_container_name"]
|
|
|
|
|
regex = "/(.*)"
|
|
|
|
|
target_label = "container"
|
|
|
|
|
}
|
|
|
|
|
rule {
|
|
|
|
|
source_labels = ["__meta_docker_container_label_com_docker_compose_service"]
|
|
|
|
|
target_label = "service"
|
|
|
|
|
}
|
2026-07-26 06:56:38 +00:00
|
|
|
|
|
|
|
|
// --- environment attribution -------------------------------------------
|
|
|
|
|
// Relabel rules are applied in order and a later write to the same target
|
|
|
|
|
// wins, so this is "default, then override with anything more specific".
|
|
|
|
|
//
|
|
|
|
|
// The regexes are fully anchored (Prometheus relabel semantics), which is
|
|
|
|
|
// what keeps `/thermograph_web.1.x` and `/thermograph-beta_beta-web.1.x`
|
|
|
|
|
// from matching each other: prod's stack prefix ends in an underscore,
|
|
|
|
|
// beta's in `-beta_`.
|
|
|
|
|
rule {
|
|
|
|
|
source_labels = ["__meta_docker_container_name"]
|
|
|
|
|
regex = ".*"
|
|
|
|
|
target_label = "host"
|
|
|
|
|
replacement = sys.env("ALLOY_ENV")
|
|
|
|
|
}
|
|
|
|
|
rule {
|
|
|
|
|
source_labels = ["__meta_docker_container_name"]
|
|
|
|
|
regex = "/thermograph_.*"
|
|
|
|
|
target_label = "host"
|
|
|
|
|
replacement = "prod"
|
|
|
|
|
}
|
|
|
|
|
rule {
|
|
|
|
|
source_labels = ["__meta_docker_container_name"]
|
|
|
|
|
regex = "/thermograph-beta_.*"
|
|
|
|
|
target_label = "host"
|
|
|
|
|
replacement = "beta"
|
|
|
|
|
}
|
|
|
|
|
rule {
|
|
|
|
|
source_labels = ["__meta_docker_container_name"]
|
|
|
|
|
regex = "/thermograph-dev[-_].*"
|
|
|
|
|
target_label = "host"
|
|
|
|
|
replacement = "dev"
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
// Drop noise/duplicate containers, per environment. Beta's equivalents are
|
|
|
|
|
// listed explicitly: its LB is `thermograph-beta-lb` (which does NOT match
|
|
|
|
|
// `thermograph-lb.*`) and its worker is `thermograph-beta_beta-worker`, so
|
|
|
|
|
// omitting them would have quietly re-admitted exactly the healthz-poll
|
|
|
|
|
// noise the prod entries exist to exclude.
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
rule {
|
|
|
|
|
source_labels = ["container"]
|
2026-07-26 06:56:38 +00:00
|
|
|
regex = "(alloy|autoscaler|thermograph-test_.*|thermograph-lb|thermograph-beta-lb|thermograph_worker|thermograph-beta_beta-worker).*"
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
action = "drop"
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
2026-07-24 04:37:41 +00:00
|
|
|
// Pass RAW targets here (not discovery.relabel.containers.output) alongside
|
|
|
|
|
// relabel_rules: loki.source.docker applies relabel_rules itself, once, using the
|
|
|
|
|
// __meta_docker_* metadata it still holds at collection time. Passing the
|
|
|
|
|
// already-relabelled output *and* relabel_rules ran the same rules twice per log
|
|
|
|
|
// entry, and made the drop rules above a no-op on the second pass since the
|
|
|
|
|
// __meta_docker_* labels are already gone from the pre-relabelled output.
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
loki.source.docker "containers" {
|
|
|
|
|
host = "unix:///var/run/docker.sock"
|
2026-07-24 04:37:41 +00:00
|
|
|
targets = discovery.docker.containers.targets
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
forward_to = [loki.write.central.receiver]
|
|
|
|
|
relabel_rules = discovery.relabel.containers.rules
|
|
|
|
|
labels = { job = "docker" }
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
// --- 2. Caddy host access logs ---------------------------------------------------
|
2026-07-24 04:37:41 +00:00
|
|
|
// sync_period (glob rescan) defaults to 10s; 1m is plenty for a log file that only
|
|
|
|
|
// appears/rotates on the order of hours.
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
local.file_match "caddy" {
|
|
|
|
|
path_targets = [{ __path__ = "/var/log/caddy/*.log", job = "caddy" }]
|
2026-07-24 04:37:41 +00:00
|
|
|
sync_period = "1m"
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
loki.source.file "caddy" {
|
|
|
|
|
targets = local.file_match.caddy.targets
|
2026-07-24 04:37:41 +00:00
|
|
|
forward_to = [loki.process.caddy.receiver]
|
|
|
|
|
|
|
|
|
|
// PollingFileWatcher defaults (250ms/250ms) stat every tailed file 4x/second
|
|
|
|
|
// forever. Caddy's access log doesn't need sub-second latency into Loki.
|
|
|
|
|
file_watch {
|
|
|
|
|
min_poll_frequency = "2s"
|
|
|
|
|
max_poll_frequency = "10s"
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
// Drop well-known crawler/bot traffic before it hits Loki. Caddy itself can't do
|
|
|
|
|
// this cheaply (log_skip needs Caddy >= 2.7; both hosts run older Caddy, and an
|
|
|
|
|
// upgrade is out of scope here), so filter it at the shipper instead.
|
|
|
|
|
loki.process "caddy" {
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
forward_to = [loki.write.central.receiver]
|
2026-07-24 04:37:41 +00:00
|
|
|
|
|
|
|
|
stage.drop {
|
|
|
|
|
expression = "(?i)(semrushbot|claudebot|ahrefsbot|yandexbot|bytespider|mj12bot|petalbot)"
|
|
|
|
|
drop_counter_reason = "crawler"
|
|
|
|
|
}
|
2026-07-26 06:56:38 +00:00
|
|
|
|
|
|
|
|
// Attribute each access log to an environment BY FILENAME. On vps2 one
|
|
|
|
|
// Caddy fronts both thermograph.org and beta.thermograph.org, so a single
|
|
|
|
|
// node-wide label would file every beta request under prod.
|
|
|
|
|
//
|
|
|
|
|
// By filename rather than by parsing `request.host` out of the JSON: the
|
|
|
|
|
// site block that writes the file is the same place the hostname is
|
|
|
|
|
// declared (see infra/deploy/Caddyfile.vps2), one label costs nothing to
|
|
|
|
|
// evaluate, and it keeps working for a non-JSON log format.
|
|
|
|
|
stage.static_labels {
|
|
|
|
|
values = { host = sys.env("ALLOY_ENV") }
|
|
|
|
|
}
|
|
|
|
|
stage.match {
|
|
|
|
|
selector = "{filename=~\".*/beta[.]log\"}"
|
|
|
|
|
|
|
|
|
|
stage.static_labels {
|
|
|
|
|
values = { host = "beta" }
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
stage.match {
|
|
|
|
|
selector = "{filename=~\".*/thermograph[.]log\"}"
|
|
|
|
|
|
|
|
|
|
stage.static_labels {
|
|
|
|
|
values = { host = "prod" }
|
|
|
|
|
}
|
|
|
|
|
}
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
// --- 3. App structured JSON logs (errors / access / audit) -----------------------
|
2026-07-26 06:56:38 +00:00
|
|
|
// Mounted read-only from the app's `applogs` volumes (see the agent compose).
|
|
|
|
|
// Lift `level`/`tag`/`phase` out of the JSON so they're queryable.
|
|
|
|
|
//
|
|
|
|
|
// TWO MOUNTS, NOT ONE. Each environment has its own applogs volume, and on vps2
|
|
|
|
|
// both exist side by side:
|
|
|
|
|
//
|
|
|
|
|
// /applogs this node's primary environment (vps2 -> prod, vps1 -> dev)
|
|
|
|
|
// /applogs-beta beta's, mounted on vps2 only
|
|
|
|
|
//
|
|
|
|
|
// A single mount was structurally unable to carry both: whichever volume the
|
|
|
|
|
// agent bound, the other environment's structured logs would never reach Loki
|
|
|
|
|
// at all — silently, since a missing file source is not an error. Every alert
|
|
|
|
|
// built on the JSON stream (the notifier heartbeats in particular) reads this.
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
local.file_match "app_jsonl" {
|
|
|
|
|
path_targets = [{ __path__ = "/applogs/**/*.jsonl", job = "app-json" }]
|
2026-07-24 04:37:41 +00:00
|
|
|
sync_period = "1m"
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
loki.source.file "app_jsonl" {
|
|
|
|
|
targets = local.file_match.app_jsonl.targets
|
|
|
|
|
forward_to = [loki.process.app_jsonl.receiver]
|
2026-07-24 04:37:41 +00:00
|
|
|
|
|
|
|
|
file_watch {
|
|
|
|
|
min_poll_frequency = "2s"
|
|
|
|
|
max_poll_frequency = "10s"
|
|
|
|
|
}
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
loki.process "app_jsonl" {
|
|
|
|
|
forward_to = [loki.write.central.receiver]
|
|
|
|
|
|
|
|
|
|
stage.json {
|
|
|
|
|
expressions = { level = "level", tag = "tag", phase = "phase", status = "status" }
|
|
|
|
|
}
|
|
|
|
|
// A record with no explicit level: an error-folder line is an error, else info.
|
|
|
|
|
stage.static_labels {
|
2026-07-26 06:56:38 +00:00
|
|
|
values = { source = "app", host = sys.env("ALLOY_ENV") }
|
|
|
|
|
}
|
|
|
|
|
stage.labels {
|
|
|
|
|
values = { level = "", tag = "", phase = "" }
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
// Beta's structured logs. Absent on vps1 (the mount simply isn't there), where
|
|
|
|
|
// this matches nothing and costs nothing.
|
|
|
|
|
local.file_match "app_jsonl_beta" {
|
|
|
|
|
path_targets = [{ __path__ = "/applogs-beta/**/*.jsonl", job = "app-json" }]
|
|
|
|
|
sync_period = "1m"
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
loki.source.file "app_jsonl_beta" {
|
|
|
|
|
targets = local.file_match.app_jsonl_beta.targets
|
|
|
|
|
forward_to = [loki.process.app_jsonl_beta.receiver]
|
|
|
|
|
|
|
|
|
|
file_watch {
|
|
|
|
|
min_poll_frequency = "2s"
|
|
|
|
|
max_poll_frequency = "10s"
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
loki.process "app_jsonl_beta" {
|
|
|
|
|
forward_to = [loki.write.central.receiver]
|
|
|
|
|
|
|
|
|
|
stage.json {
|
|
|
|
|
expressions = { level = "level", tag = "tag", phase = "phase", status = "status" }
|
|
|
|
|
}
|
|
|
|
|
stage.static_labels {
|
|
|
|
|
values = { source = "app", host = "beta" }
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
}
|
|
|
|
|
stage.labels {
|
|
|
|
|
values = { level = "", tag = "", phase = "" }
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
// --- Ship to the central Loki over the WireGuard mesh ----------------------------
|
|
|
|
|
loki.write "central" {
|
|
|
|
|
endpoint {
|
|
|
|
|
url = sys.env("LOKI_URL")
|
|
|
|
|
}
|
2026-07-26 06:56:38 +00:00
|
|
|
// `node` — the MACHINE, stamped on everything this agent ships. `host` is
|
|
|
|
|
// deliberately NOT here: it is the ENVIRONMENT, set per source above,
|
|
|
|
|
// because vps2 carries two of them and an external label cannot vary per
|
|
|
|
|
// stream. Setting host here would override that work on every line.
|
|
|
|
|
external_labels = { node = sys.env("ALLOY_NODE") }
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
}
|