2026-07-23 05:11:33 +00:00
|
|
|
name: Validate observability stack
|
|
|
|
|
|
2026-07-26 06:56:38 +00:00
|
|
|
# The observability DOMAIN deploys by hand (`docker compose up -d` on vps1 —
|
|
|
|
|
# the monitoring host, not the beta environment, which runs on vps2 now — plus
|
2026-07-23 05:11:33 +00:00
|
|
|
# the Alloy agent on each node) with no build step, so nothing caught a
|
|
|
|
|
# malformed compose file, a broken dashboard JSON, or an unparseable config
|
|
|
|
|
# until it failed live on the host. This is that missing guard: it parses
|
|
|
|
|
# every YAML/JSON artifact that ships to a node. It does NOT deploy, and
|
|
|
|
|
# deliberately needs no docker CLI (the node:20-bookworm job image doesn't
|
|
|
|
|
# ship one) -- a YAML parse catches the real breakage without it.
|
|
|
|
|
#
|
|
|
|
|
# Monorepo port: paths are prefixed observability/, the push trigger is
|
|
|
|
|
# path-filtered to the domain, and workflow_call lets pr-build.yml reuse this
|
|
|
|
|
# as the domain's PR check under its single `gate` required status.
|
|
|
|
|
#
|
2026-07-24 04:37:41 +00:00
|
|
|
# The Alloy config (observability/alloy/config.alloy, an HCL-like format) IS
|
|
|
|
|
# validated below -- the static `alloy` binary is downloaded straight from its
|
|
|
|
|
# GitHub release (pinned to the same v1.9.1 the fleet runs; see
|
|
|
|
|
# observability/alloy/docker-compose.agent.yml), no docker CLI needed.
|
|
|
|
|
#
|
observability: add the estate's first alerting; supervise Postfix
The estate had zero alert rules. The only contact point was Grafana's factory
default pointing at the literal string <example@email.com>, and beta's
untracked compose override routed Grafana's SMTP at prod's Postfix — so
alerts, had any existed, would have been delivered by the box most likely to
be on fire. Prod served 86 5xx in 24h including five /healthz failures and
nobody was told.
Alerting (deploys to beta, which is where Grafana runs):
- 12 Loki-based rules. There is no Prometheus in this estate and Loki is the
only datasource, so every rule is log-derived.
- Thresholds come from a 24h backtest that happens to contain a real ~20min
prod outage at 04:20Z. Absolute counts, not ratios: prod runs ~7 req/min, so
one bad request is 1.4% and a ratio alert would scream all night. The 5xx
burst rule fires on the outage's 45 and 22 buckets and on nothing else in
the day; the largest benign bucket all day was 5.
- Routed to a new private #ops-alerts channel, not to any existing channel —
#weather-events, #announcements and #prod are product surfaces that notify
real subscribers.
- AlertingWatchdog is a dead-man's switch; its value is its absence.
- Delivery is proven, not assumed: the rules were provisioned into a throwaway
Grafana against beta's live Loki and a real alert arrived in Discord. This
matters because GET /api/v1/provisioning/contact-points returns [REDACTED]
for the URL — a contact point holding an uninterpolated env var looks
perfectly healthy and pages nobody. The only proof is a message arriving.
CI gains a structural check, because the existing one only proves YAML parses:
an alert rule whose condition names a missing refId is valid YAML, provisions
cleanly, and never fires. It also hard-fails on a literal Discord webhook in
the repo. Verified against all three breakages deliberately introduced.
Postfix supervision:
- The 13h outage was a boot-ordering race, not a Docker renumbering: postfix
started at 08:09:49, wg0 came up at :51, postfix fataled at :52 on a missing
docker_gwbridge address, and dockerd did not finish starting until 08:10:16.
Stock postfix@.service is ordered only After=network-online.target and ships
no Restart=, so one lost race became a permanent outage.
- An ExecStartPre gate now blocks up to 60s until every inet_interfaces
address actually exists, which absorbs the transient case inside a single
start attempt. That makes bounded retry correct: 5 attempts in 600s, then
failed — a genuinely broken config reaches a visible failed state in ~100s
instead of re-fataling every 15s forever.
- A 5-minute watchdog timer retries indefinitely and runs reset-failed, so
"failed" still self-heals. Worst case is ~5 minutes, not 13 hours.
- Wants=, not Requires=: a dockerd failure must not take down the loopback and
mesh listeners that do not depend on Docker at all.
Health checks must read config with `postconf -c`, never postmulti/postqueue/
postfix — those three RESOLVE inet_interfaces and so fatal precisely when an
address is missing, which made the first version of this check report
status=ok bound=0/0. A health check that fails open is worse than none.
DEPLOY.md carries the monitoring contract, including that systemctl is-active
postfix is a known-false signal: postfix.service is a wrapper whose
ExecStart=/bin/true, so it reports active forever while the real postfix@-
instance is failed with zero listeners. Reproduced live.
Known gap, documented not fixed: Alloy ships Docker stdout, Caddy files and
app JSONL, not journald — so no Postfix line reaches Loki and the mail rules
cannot fire until loki.source.journal is added.
Claude-Session: https://claude.ai/code/session_0182KTMrsTHJc3TcewCatJFY
2026-07-24 20:19:19 +00:00
|
|
|
# The alerting config gets a third, stricter step. A YAML parse is not enough
|
|
|
|
|
# for it: an alert rule whose `condition` names a refId that does not exist is
|
|
|
|
|
# perfectly valid YAML, provisions without complaint, and then never fires --
|
|
|
|
|
# a paging path that looks healthy and reaches nobody, which is the exact
|
|
|
|
|
# failure this whole config exists to end. That step also refuses a literal
|
|
|
|
|
# Discord webhook in the repo (it must stay an env reference).
|
|
|
|
|
#
|
2026-07-24 04:37:41 +00:00
|
|
|
# Not validated here: docker-compose *schema* (unknown-key checks, which would
|
|
|
|
|
# need the docker CLI). Add it if the churn warrants the extra tooling.
|
2026-07-23 05:11:33 +00:00
|
|
|
|
|
|
|
|
on:
|
|
|
|
|
workflow_call: {}
|
|
|
|
|
push:
|
|
|
|
|
branches: [main, dev]
|
|
|
|
|
paths: ['observability/**']
|
|
|
|
|
|
|
|
|
|
jobs:
|
|
|
|
|
validate:
|
|
|
|
|
runs-on: docker
|
|
|
|
|
steps:
|
|
|
|
|
- uses: actions/checkout@v4
|
|
|
|
|
|
|
|
|
|
- name: Dashboards are valid JSON
|
|
|
|
|
run: |
|
|
|
|
|
set -euo pipefail
|
|
|
|
|
python3 - <<'PY'
|
|
|
|
|
import json, pathlib, sys
|
|
|
|
|
bad = False
|
|
|
|
|
files = sorted(pathlib.Path("observability/grafana/dashboards").glob("*.json"))
|
|
|
|
|
if not files:
|
|
|
|
|
print("no dashboards found"); sys.exit(1)
|
|
|
|
|
for f in files:
|
|
|
|
|
try:
|
|
|
|
|
json.loads(f.read_text()); print("OK", f)
|
|
|
|
|
except Exception as e: # noqa: BLE001
|
|
|
|
|
print("INVALID JSON", f, "-", e); bad = True
|
|
|
|
|
sys.exit(1 if bad else 0)
|
|
|
|
|
PY
|
|
|
|
|
|
|
|
|
|
- name: YAML artifacts parse (compose + loki + grafana provisioning)
|
|
|
|
|
run: |
|
|
|
|
|
set -euo pipefail
|
|
|
|
|
apt-get update -qq && apt-get install -y -qq python3-yaml >/dev/null
|
|
|
|
|
python3 - <<'PY'
|
|
|
|
|
import pathlib, sys
|
|
|
|
|
import yaml
|
|
|
|
|
bad = False
|
|
|
|
|
root = pathlib.Path("observability")
|
|
|
|
|
targets = [
|
|
|
|
|
root / "docker-compose.yml",
|
|
|
|
|
root / "alloy/docker-compose.agent.yml",
|
|
|
|
|
root / "loki/config.yml",
|
|
|
|
|
*sorted((root / "grafana/provisioning").rglob("*.yml")),
|
|
|
|
|
]
|
|
|
|
|
for f in targets:
|
|
|
|
|
if not f.exists():
|
|
|
|
|
print("MISSING", f); bad = True; continue
|
|
|
|
|
try:
|
|
|
|
|
yaml.safe_load(f.read_text()); print("OK", f)
|
|
|
|
|
except Exception as e: # noqa: BLE001
|
|
|
|
|
print("INVALID YAML", f, "-", e); bad = True
|
|
|
|
|
sys.exit(1 if bad else 0)
|
|
|
|
|
PY
|
2026-07-24 04:37:41 +00:00
|
|
|
|
|
|
|
|
- name: Alloy config parses and validates (components, relabel/process wiring)
|
|
|
|
|
run: |
|
|
|
|
|
set -euo pipefail
|
|
|
|
|
apt-get update -qq && apt-get install -y -qq unzip >/dev/null
|
|
|
|
|
curl -sSL -o /tmp/alloy.zip \
|
|
|
|
|
https://github.com/grafana/alloy/releases/download/v1.9.1/alloy-linux-amd64.zip
|
|
|
|
|
unzip -q /tmp/alloy.zip -d /tmp/alloybin
|
|
|
|
|
chmod +x /tmp/alloybin/alloy-linux-amd64
|
|
|
|
|
|
|
|
|
|
# `alloy validate` builds the real component graph (catches bad field
|
|
|
|
|
# names, dangling forward_to/receiver refs, malformed relabel/process
|
|
|
|
|
# stages) but doesn't recognize the top-level `livedebugging {}` singleton
|
|
|
|
|
# block as a component -- it only knows named components, even though
|
|
|
|
|
# `alloy run` loads that block fine. Known gap in the `validate`
|
|
|
|
|
# subcommand, not a config error, so strip that one line before checking.
|
|
|
|
|
grep -v '^livedebugging ' observability/alloy/config.alloy > /tmp/config-for-validate.alloy
|
|
|
|
|
/tmp/alloybin/alloy-linux-amd64 validate /tmp/config-for-validate.alloy
|
observability: add the estate's first alerting; supervise Postfix
The estate had zero alert rules. The only contact point was Grafana's factory
default pointing at the literal string <example@email.com>, and beta's
untracked compose override routed Grafana's SMTP at prod's Postfix — so
alerts, had any existed, would have been delivered by the box most likely to
be on fire. Prod served 86 5xx in 24h including five /healthz failures and
nobody was told.
Alerting (deploys to beta, which is where Grafana runs):
- 12 Loki-based rules. There is no Prometheus in this estate and Loki is the
only datasource, so every rule is log-derived.
- Thresholds come from a 24h backtest that happens to contain a real ~20min
prod outage at 04:20Z. Absolute counts, not ratios: prod runs ~7 req/min, so
one bad request is 1.4% and a ratio alert would scream all night. The 5xx
burst rule fires on the outage's 45 and 22 buckets and on nothing else in
the day; the largest benign bucket all day was 5.
- Routed to a new private #ops-alerts channel, not to any existing channel —
#weather-events, #announcements and #prod are product surfaces that notify
real subscribers.
- AlertingWatchdog is a dead-man's switch; its value is its absence.
- Delivery is proven, not assumed: the rules were provisioned into a throwaway
Grafana against beta's live Loki and a real alert arrived in Discord. This
matters because GET /api/v1/provisioning/contact-points returns [REDACTED]
for the URL — a contact point holding an uninterpolated env var looks
perfectly healthy and pages nobody. The only proof is a message arriving.
CI gains a structural check, because the existing one only proves YAML parses:
an alert rule whose condition names a missing refId is valid YAML, provisions
cleanly, and never fires. It also hard-fails on a literal Discord webhook in
the repo. Verified against all three breakages deliberately introduced.
Postfix supervision:
- The 13h outage was a boot-ordering race, not a Docker renumbering: postfix
started at 08:09:49, wg0 came up at :51, postfix fataled at :52 on a missing
docker_gwbridge address, and dockerd did not finish starting until 08:10:16.
Stock postfix@.service is ordered only After=network-online.target and ships
no Restart=, so one lost race became a permanent outage.
- An ExecStartPre gate now blocks up to 60s until every inet_interfaces
address actually exists, which absorbs the transient case inside a single
start attempt. That makes bounded retry correct: 5 attempts in 600s, then
failed — a genuinely broken config reaches a visible failed state in ~100s
instead of re-fataling every 15s forever.
- A 5-minute watchdog timer retries indefinitely and runs reset-failed, so
"failed" still self-heals. Worst case is ~5 minutes, not 13 hours.
- Wants=, not Requires=: a dockerd failure must not take down the loopback and
mesh listeners that do not depend on Docker at all.
Health checks must read config with `postconf -c`, never postmulti/postqueue/
postfix — those three RESOLVE inet_interfaces and so fatal precisely when an
address is missing, which made the first version of this check report
status=ok bound=0/0. A health check that fails open is worse than none.
DEPLOY.md carries the monitoring contract, including that systemctl is-active
postfix is a known-false signal: postfix.service is a wrapper whose
ExecStart=/bin/true, so it reports active forever while the real postfix@-
instance is failed with zero listeners. Reproduced live.
Known gap, documented not fixed: Alloy ships Docker stdout, Caddy files and
app JSONL, not journald — so no Postfix line reaches Loki and the mail rules
cannot fire until loki.source.journal is added.
Claude-Session: https://claude.ai/code/session_0182KTMrsTHJc3TcewCatJFY
2026-07-24 20:19:19 +00:00
|
|
|
|
|
|
|
|
- name: Alerting config is structurally sound
|
|
|
|
|
run: |
|
|
|
|
|
set -euo pipefail
|
|
|
|
|
apt-get update -qq && apt-get install -y -qq python3-yaml >/dev/null
|
|
|
|
|
python3 - <<'PY'
|
|
|
|
|
import pathlib, re, sys
|
|
|
|
|
import yaml
|
|
|
|
|
|
|
|
|
|
d = pathlib.Path("observability/grafana/provisioning/alerting")
|
|
|
|
|
if not d.is_dir():
|
|
|
|
|
print("no alerting provisioning yet — nothing to check"); sys.exit(0)
|
|
|
|
|
|
|
|
|
|
docs, bad = {}, []
|
|
|
|
|
for f in sorted(list(d.glob("*.yml")) + list(d.glob("*.yaml"))):
|
|
|
|
|
docs[f] = yaml.safe_load(f.read_text()) or {}
|
|
|
|
|
print("loaded", f)
|
|
|
|
|
|
|
|
|
|
receivers, referenced = set(), set()
|
|
|
|
|
|
|
|
|
|
for f, doc in docs.items():
|
|
|
|
|
# --- contact points ---------------------------------------------------
|
|
|
|
|
for cp in doc.get("contactPoints") or []:
|
|
|
|
|
receivers.add(cp.get("name"))
|
|
|
|
|
for r in cp.get("receivers") or []:
|
|
|
|
|
url = str((r.get("settings") or {}).get("url", ""))
|
|
|
|
|
# A real webhook in the repo is a credential leak: anyone with
|
|
|
|
|
# it can post into the ops channel. It must stay an env ref.
|
|
|
|
|
if re.search(r"https://(discord|discordapp)\.com/api/webhooks/", url):
|
|
|
|
|
bad.append(f"{f}: contact point '{cp.get('name')}' has a LITERAL Discord webhook URL — use $DISCORD_ALERT_WEBHOOK_URL")
|
|
|
|
|
elif not url.startswith("$"):
|
|
|
|
|
bad.append(f"{f}: contact point '{cp.get('name')}' url is neither an env reference nor empty: {url!r}")
|
|
|
|
|
|
|
|
|
|
# --- notification policies --------------------------------------------
|
|
|
|
|
def walk(route, where):
|
|
|
|
|
if route.get("receiver"):
|
|
|
|
|
referenced.add(route["receiver"])
|
|
|
|
|
for child in route.get("routes") or []:
|
|
|
|
|
walk(child, where)
|
|
|
|
|
|
|
|
|
|
for pol in doc.get("policies") or []:
|
|
|
|
|
walk(pol, f)
|
|
|
|
|
|
|
|
|
|
# --- alert rules ------------------------------------------------------
|
|
|
|
|
for g in doc.get("groups") or []:
|
|
|
|
|
for rule in g.get("rules") or []:
|
|
|
|
|
title = rule.get("title", "<untitled>")
|
|
|
|
|
where = f"{f}: rule '{title}'"
|
|
|
|
|
refs = {q.get("refId") for q in rule.get("data") or []}
|
|
|
|
|
if not rule.get("uid"):
|
|
|
|
|
bad.append(f"{where}: missing uid (provisioning needs a stable one)")
|
|
|
|
|
if rule.get("condition") not in refs:
|
|
|
|
|
bad.append(f"{where}: condition '{rule.get('condition')}' is not one of {sorted(refs)} — this rule can never fire")
|
|
|
|
|
if not (rule.get("labels") or {}).get("severity"):
|
|
|
|
|
bad.append(f"{where}: no severity label — the notification policy routes on it")
|
|
|
|
|
if not (rule.get("annotations") or {}).get("summary"):
|
|
|
|
|
bad.append(f"{where}: no summary annotation — the Discord message would be blank")
|
|
|
|
|
for q in rule.get("data") or []:
|
|
|
|
|
model = q.get("model") or {}
|
|
|
|
|
upstream = model.get("expression")
|
|
|
|
|
if upstream and upstream not in refs:
|
|
|
|
|
bad.append(f"{where}: query '{q.get('refId')}' reads refId '{upstream}' which does not exist")
|
|
|
|
|
|
|
|
|
|
missing = referenced - receivers
|
|
|
|
|
if missing:
|
|
|
|
|
bad.append(f"notification policy routes to undefined receiver(s): {sorted(missing)}")
|
|
|
|
|
|
|
|
|
|
for b in bad:
|
|
|
|
|
print("::error::" + b)
|
|
|
|
|
print(f"checked {len(docs)} file(s): {len(receivers)} contact point(s), {len(bad)} problem(s)")
|
|
|
|
|
sys.exit(1 if bad else 0)
|
|
|
|
|
PY
|