2026-07-26 06:56:38 +00:00
|
|
|
# The Alloy log-collector agent. Runs on EACH node (vps1, vps2) — one per node
|
|
|
|
|
# — reading that node's Docker containers, Caddy logs, and the app's JSON logs,
|
|
|
|
|
# and shipping them to the central Loki on vps1 over the mesh.
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
#
|
|
|
|
|
# Per-node config comes from the environment (an .env file next to this, or
|
|
|
|
|
# exported before `up`):
|
2026-07-26 06:56:38 +00:00
|
|
|
# ALLOY_NODE the MACHINE: vps1 | vps2 | desktop
|
|
|
|
|
# ALLOY_ENV the default ENVIRONMENT label for logs on this node that
|
|
|
|
|
# aren't attributable to a specific stack (Forgejo, Grafana,
|
|
|
|
|
# the portfolio site): vps1 -> dev, vps2 -> prod
|
|
|
|
|
# APPLOGS_VOLUME this node's primary applogs volume
|
2026-07-21 16:24:26 +00:00
|
|
|
# LOKI_URL push endpoint. All nodes push over the mesh:
|
|
|
|
|
# http://10.10.0.2:3100/loki/api/v1/push
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
#
|
2026-07-26 06:56:38 +00:00
|
|
|
# Deploy on vps1 (dev + Forgejo + the monitoring stack itself):
|
|
|
|
|
# ALLOY_NODE=vps1 ALLOY_ENV=dev APPLOGS_VOLUME=thermograph-dev_applogs \
|
|
|
|
|
# LOKI_URL=http://10.10.0.2:3100/loki/api/v1/push \
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
# docker compose -f docker-compose.agent.yml up -d
|
|
|
|
|
#
|
2026-07-26 06:56:38 +00:00
|
|
|
# Deploy on vps2 (prod + beta) — note the SECOND file, which adds beta's log
|
|
|
|
|
# volume:
|
|
|
|
|
# ALLOY_NODE=vps2 ALLOY_ENV=prod APPLOGS_VOLUME=thermograph_applogs \
|
|
|
|
|
# BETA_APPLOGS_VOLUME=thermograph-beta_applogs \
|
|
|
|
|
# LOKI_URL=http://10.10.0.2:3100/loki/api/v1/push \
|
|
|
|
|
# docker compose -f docker-compose.agent.yml -f docker-compose.agent.beta.yml up -d
|
|
|
|
|
#
|
|
|
|
|
# WHY THE VOLUME NAME IS A VARIABLE NOW, AND WHY BETA IS AN OVERLAY. The name
|
|
|
|
|
# used to be the literal `thermograph_applogs`, with a comment telling you to
|
|
|
|
|
# hand-edit two lines on the dev node. That was survivable while one node ran
|
|
|
|
|
# one environment. vps2 runs two, each with its own applogs volume, and a single
|
|
|
|
|
# hardcoded mount there would ship one environment's structured logs and
|
|
|
|
|
# silently drop the other's.
|
|
|
|
|
#
|
|
|
|
|
# Beta's mount is a separate overlay file rather than a second variable in this
|
|
|
|
|
# one because the second mount is labelled host="beta" unconditionally. Pointing
|
|
|
|
|
# it at dev's volume on vps1 just to satisfy compose would ship every dev log
|
|
|
|
|
# line a second time labelled as beta — inventing an environment's worth of
|
|
|
|
|
# fake data on a box beta does not run on.
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
|
|
|
|
|
services:
|
|
|
|
|
alloy:
|
|
|
|
|
image: grafana/alloy:v1.9.1
|
|
|
|
|
command:
|
|
|
|
|
- run
|
|
|
|
|
- /etc/alloy/config.alloy
|
|
|
|
|
- --storage.path=/var/lib/alloy/data
|
|
|
|
|
- --server.http.listen-addr=0.0.0.0:12345
|
|
|
|
|
environment:
|
2026-07-26 06:56:38 +00:00
|
|
|
ALLOY_NODE: ${ALLOY_NODE:?set ALLOY_NODE (vps1|vps2|desktop) — the MACHINE}
|
|
|
|
|
ALLOY_ENV: ${ALLOY_ENV:?set ALLOY_ENV (dev|beta|prod) — this node's default environment label}
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
LOKI_URL: ${LOKI_URL:?set LOKI_URL}
|
|
|
|
|
volumes:
|
|
|
|
|
- ./config.alloy:/etc/alloy/config.alloy:ro
|
|
|
|
|
# Read-only Docker socket: discover + tail every container's stdout/stderr.
|
|
|
|
|
- /var/run/docker.sock:/var/run/docker.sock:ro
|
|
|
|
|
# Caddy's host access logs (runs on the host, not a container).
|
|
|
|
|
- /var/log/caddy:/var/log/caddy:ro
|
|
|
|
|
# The app's structured JSON logs, straight off its named volume.
|
2026-07-26 06:56:38 +00:00
|
|
|
- applogs:/applogs:ro
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
- alloy_data:/var/lib/alloy/data
|
|
|
|
|
restart: unless-stopped
|
|
|
|
|
|
|
|
|
|
volumes:
|
|
|
|
|
# The app's existing log volume (created by the app's own stack). external:true
|
2026-07-26 06:56:38 +00:00
|
|
|
# means compose references it, never creates or deletes it. The NAME varies by
|
|
|
|
|
# node — see the header — while the mount path stays /applogs, which is what
|
|
|
|
|
# lets one config.alloy serve every node.
|
|
|
|
|
applogs:
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
external: true
|
2026-07-26 06:56:38 +00:00
|
|
|
name: ${APPLOGS_VOLUME:?set APPLOGS_VOLUME (e.g. thermograph_applogs on vps2, thermograph-dev_applogs on vps1)}
|
Fleet log aggregation: Grafana + Loki + Alloy over the WireGuard mesh
A central Grafana + Loki stack (on beta) fed by a Grafana Alloy agent on
every node, replacing the old SSH-tailed single-host scripts/dashboard.py.
- docker-compose.yml: central Loki (mesh-only :3100) + Grafana (Caddy-fronted)
- loki/config.yml: single-binary, filesystem storage, 30-day retention
- alloy/config.alloy + docker-compose.agent.yml: per-node collector — every
container's stdout/stderr via the Docker socket, Caddy host logs, and the
app's structured JSON logs (errors/access/audit), each line tagged by node
- grafana/: auto-provisioned Loki datasource + a fleet-logs dashboard
(volume by service, error rate, upstream 429s, Caddy 5xx, notifier liveness,
live tail), with a per-node selector
- caddy-grafana.conf, README, .env.example
2026-07-21 16:11:52 +00:00
|
|
|
alloy_data: {}
|