thermograph/.claude/hooks/README.md
emi d138f00a20
Some checks failed
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Failing after 6s
secrets-guard / encrypted (push) Successful in 6s
shell-lint / shellcheck (push) Successful in 13s
Validate observability stack / validate (push) Successful in 17s
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
infra: split the estate into vps1/vps2 — beta joins prod, dev gets a home (#103)
2026-07-26 06:56:38 +00:00

92 lines
4.6 KiB
Markdown

# .claude/hooks — enforcement, not advice
`CLAUDE.md` files can only ask. These hooks are the part that enforces, and they
travel with the repo so a session started anywhere in this checkout gets them.
They are wired up in `.claude/settings.json`.
| Hook | Event | What it does |
|---|---|---|
| `prod-guard.sh` | PreToolUse | Live-host commands: reads run unprompted, mutations ask first |
| `secrets-guard.sh` | PreToolUse | Denies any direct write to `infra/deploy/secrets/*.yaml` |
| `lint-after-edit.sh` | PostToolUse | shellchecks an edited `*.sh` and feeds findings back immediately |
## Why prod-guard exists
`agent` has passwordless sudo on vps1 and vps2, and the operator's global settings
allow `Bash(ssh prod:*)` outright. Nothing about `ssh prod 'docker service rm …'`
goes through a pull request, so CI cannot see it — before this hook, any session in
any directory could delete production with no prompt.
**It classifies by allowlist, not blocklist.** Only commands positively recognised as
read-only are allowed through; everything else asks. A blocklist of dangerous verbs
is wrong by construction — the first destructive verb nobody thought of sails
straight through. Being wrong in the "ask" direction costs a keystroke; being wrong
the other way costs production.
Beta is guarded as strictly as prod — more strictly than the names alone suggest,
now: vps2 (`169.58.46.181`) runs **both** prod and beta as separate Swarm stacks on
the same manager, sharing one TimescaleDB instance, so a "beta" command there has
prod's blast radius (see `infra/deploy/env-topology.sh`). vps1 (`75.119.132.91`) is
the other guarded host: Forgejo (git + CI + registry) and Grafana/Loki live there,
plus the `dev` environment, so a destructive command there takes out git, CI and the
registry at once. Mesh IPs did not move in the vps1/vps2 split (vps1 is still
`10.10.0.2`, vps2 is still `10.10.0.1`); what moved is which environment runs where
`HOST_TOKEN_RE` in `prod-guard.sh` recognises `vps1`/`vps2` alongside the
environment names (`prod`/`beta`) and the raw IPs, so a command naming the host
either way still gets caught. The LAN dev box is deliberately unguarded.
## Changing the classifier
`is_readonly()` in `prod-guard.sh` splits a command on `&&`, `||`, `;` and `|`, then
checks each segment's leading word. Redirection to a file, command substitution,
`sed -i`, and mutating `docker`/`git`/`systemctl` subcommands all count as writes.
If you add a command to the allowlist, **re-run the test matrix**. There is a
non-obvious failure mode this already hit once: the loop is fed by
`printf '%s\n' | sed`, and dropping that trailing newline makes `read` return
non-zero on the only line, so the loop body never runs and *every* command
classifies as read-only — a silent, total bypass that still looks like it works.
Test by piping a payload straight in:
```bash
echo '{"tool_name":"Bash","tool_input":{"command":"ssh prod '\''rm -rf /'\''"}}' \
| .claude/hooks/prod-guard.sh
# expect: permissionDecision "ask"
```
Exit 0 with no output means "no opinion" and the normal permission flow proceeds.
That is also what happens if `jq` is missing or the script errors, so a bug here
fails toward the prompt rather than toward silent execution.
## lint-after-edit needs shellcheck on PATH
It exits quietly when shellcheck is absent — a missing local tool must not break
editing, and `shell-lint` CI is still the backstop. v0.11.0 (the version CI pins) is
installed at `~/.local/bin/shellcheck`; if you move machines, reinstall it:
```bash
VER=v0.11.0
curl -fsSL "https://github.com/koalaman/shellcheck/releases/download/${VER}/shellcheck-${VER}.linux.x86_64.tar.xz" \
| tar -xJ --strip-components=1 -C ~/.local/bin "shellcheck-${VER}/shellcheck"
```
Keep it matched to `shell-lint.yml`'s pin. A drifted shellcheck grows *new* warnings
and fails CI on an unrelated push, which is why that workflow pins the version and
its sha256 rather than using `apt-get install`.
## The global copy
`~/.claude/settings.json` runs `prod-guard.sh` from `~/.claude/hooks/` as well, so a
session started **outside** this repo is covered too — otherwise the guard would be
trivially bypassed by working from `~/Code`. That copy is a mirror; this one is
canonical. If you change the classifier here, copy it over:
```bash
cp .claude/hooks/prod-guard.sh ~/.claude/hooks/prod-guard.sh
```
The blanket `Bash(ssh prod:*)` / `Bash(ssh beta:*)` allow rules were removed from
global settings at the same time, so the hook is the only thing granting live-host
access. If it is missing or broken, nothing allows the command and you get a prompt.