thermograph/.forgejo/workflows/observability-validate.yml
emi d138f00a20
Some checks failed
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Failing after 6s
secrets-guard / encrypted (push) Successful in 6s
shell-lint / shellcheck (push) Successful in 13s
Validate observability stack / validate (push) Successful in 17s
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103)
2026-07-26 06:56:38 +00:00

173 lines
8.3 KiB
YAML

name: Validate observability stack
# The observability DOMAIN deploys by hand (`docker compose up -d` on vps1 —
# the monitoring host, not the beta environment, which runs on vps2 now — plus
# the Alloy agent on each node) with no build step, so nothing caught a
# malformed compose file, a broken dashboard JSON, or an unparseable config
# until it failed live on the host. This is that missing guard: it parses
# every YAML/JSON artifact that ships to a node. It does NOT deploy, and
# deliberately needs no docker CLI (the node:20-bookworm job image doesn't
# ship one) -- a YAML parse catches the real breakage without it.
#
# Monorepo port: paths are prefixed observability/, the push trigger is
# path-filtered to the domain, and workflow_call lets pr-build.yml reuse this
# as the domain's PR check under its single `gate` required status.
#
# The Alloy config (observability/alloy/config.alloy, an HCL-like format) IS
# validated below -- the static `alloy` binary is downloaded straight from its
# GitHub release (pinned to the same v1.9.1 the fleet runs; see
# observability/alloy/docker-compose.agent.yml), no docker CLI needed.
#
# The alerting config gets a third, stricter step. A YAML parse is not enough
# for it: an alert rule whose `condition` names a refId that does not exist is
# perfectly valid YAML, provisions without complaint, and then never fires --
# a paging path that looks healthy and reaches nobody, which is the exact
# failure this whole config exists to end. That step also refuses a literal
# Discord webhook in the repo (it must stay an env reference).
#
# Not validated here: docker-compose *schema* (unknown-key checks, which would
# need the docker CLI). Add it if the churn warrants the extra tooling.
on:
workflow_call: {}
push:
branches: [main, dev]
paths: ['observability/**']
jobs:
validate:
runs-on: docker
steps:
- uses: actions/checkout@v4
- name: Dashboards are valid JSON
run: |
set -euo pipefail
python3 - <<'PY'
import json, pathlib, sys
bad = False
files = sorted(pathlib.Path("observability/grafana/dashboards").glob("*.json"))
if not files:
print("no dashboards found"); sys.exit(1)
for f in files:
try:
json.loads(f.read_text()); print("OK", f)
except Exception as e: # noqa: BLE001
print("INVALID JSON", f, "-", e); bad = True
sys.exit(1 if bad else 0)
PY
- name: YAML artifacts parse (compose + loki + grafana provisioning)
run: |
set -euo pipefail
apt-get update -qq && apt-get install -y -qq python3-yaml >/dev/null
python3 - <<'PY'
import pathlib, sys
import yaml
bad = False
root = pathlib.Path("observability")
targets = [
root / "docker-compose.yml",
root / "alloy/docker-compose.agent.yml",
root / "loki/config.yml",
*sorted((root / "grafana/provisioning").rglob("*.yml")),
]
for f in targets:
if not f.exists():
print("MISSING", f); bad = True; continue
try:
yaml.safe_load(f.read_text()); print("OK", f)
except Exception as e: # noqa: BLE001
print("INVALID YAML", f, "-", e); bad = True
sys.exit(1 if bad else 0)
PY
- name: Alloy config parses and validates (components, relabel/process wiring)
run: |
set -euo pipefail
apt-get update -qq && apt-get install -y -qq unzip >/dev/null
curl -sSL -o /tmp/alloy.zip \
https://github.com/grafana/alloy/releases/download/v1.9.1/alloy-linux-amd64.zip
unzip -q /tmp/alloy.zip -d /tmp/alloybin
chmod +x /tmp/alloybin/alloy-linux-amd64
# `alloy validate` builds the real component graph (catches bad field
# names, dangling forward_to/receiver refs, malformed relabel/process
# stages) but doesn't recognize the top-level `livedebugging {}` singleton
# block as a component -- it only knows named components, even though
# `alloy run` loads that block fine. Known gap in the `validate`
# subcommand, not a config error, so strip that one line before checking.
grep -v '^livedebugging ' observability/alloy/config.alloy > /tmp/config-for-validate.alloy
/tmp/alloybin/alloy-linux-amd64 validate /tmp/config-for-validate.alloy
- name: Alerting config is structurally sound
run: |
set -euo pipefail
apt-get update -qq && apt-get install -y -qq python3-yaml >/dev/null
python3 - <<'PY'
import pathlib, re, sys
import yaml
d = pathlib.Path("observability/grafana/provisioning/alerting")
if not d.is_dir():
print("no alerting provisioning yet — nothing to check"); sys.exit(0)
docs, bad = {}, []
for f in sorted(list(d.glob("*.yml")) + list(d.glob("*.yaml"))):
docs[f] = yaml.safe_load(f.read_text()) or {}
print("loaded", f)
receivers, referenced = set(), set()
for f, doc in docs.items():
# --- contact points ---------------------------------------------------
for cp in doc.get("contactPoints") or []:
receivers.add(cp.get("name"))
for r in cp.get("receivers") or []:
url = str((r.get("settings") or {}).get("url", ""))
# A real webhook in the repo is a credential leak: anyone with
# it can post into the ops channel. It must stay an env ref.
if re.search(r"https://(discord|discordapp)\.com/api/webhooks/", url):
bad.append(f"{f}: contact point '{cp.get('name')}' has a LITERAL Discord webhook URL — use $DISCORD_ALERT_WEBHOOK_URL")
elif not url.startswith("$"):
bad.append(f"{f}: contact point '{cp.get('name')}' url is neither an env reference nor empty: {url!r}")
# --- notification policies --------------------------------------------
def walk(route, where):
if route.get("receiver"):
referenced.add(route["receiver"])
for child in route.get("routes") or []:
walk(child, where)
for pol in doc.get("policies") or []:
walk(pol, f)
# --- alert rules ------------------------------------------------------
for g in doc.get("groups") or []:
for rule in g.get("rules") or []:
title = rule.get("title", "<untitled>")
where = f"{f}: rule '{title}'"
refs = {q.get("refId") for q in rule.get("data") or []}
if not rule.get("uid"):
bad.append(f"{where}: missing uid (provisioning needs a stable one)")
if rule.get("condition") not in refs:
bad.append(f"{where}: condition '{rule.get('condition')}' is not one of {sorted(refs)} — this rule can never fire")
if not (rule.get("labels") or {}).get("severity"):
bad.append(f"{where}: no severity label — the notification policy routes on it")
if not (rule.get("annotations") or {}).get("summary"):
bad.append(f"{where}: no summary annotation — the Discord message would be blank")
for q in rule.get("data") or []:
model = q.get("model") or {}
upstream = model.get("expression")
if upstream and upstream not in refs:
bad.append(f"{where}: query '{q.get('refId')}' reads refId '{upstream}' which does not exist")
missing = referenced - receivers
if missing:
bad.append(f"notification policy routes to undefined receiver(s): {sorted(missing)}")
for b in bad:
print("::error::" + b)
print(f"checked {len(docs)} file(s): {len(receivers)} contact point(s), {len(bad)} problem(s)")
sys.exit(1 if bad else 0)
PY