thermograph/.claude/hooks/README.md
Emi Griffith c184fd55ea
All checks were successful
PR build (required check) / changes (pull_request) Successful in 11s
secrets-guard / encrypted (pull_request) Successful in 10s
shell-lint / shellcheck (pull_request) Successful in 14s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 3s
guardrails: silence the intentional SC2016 in prod-guard
shell-lint flagged the literal `$(` / backtick patterns the read-only classifier
searches for inside a command string. Expanding them is exactly what must not
happen — the check exists to refuse commands containing command substitution, so
the single quotes are the point. Same treatment and rationale as the existing
disables in render-secrets.sh and verify-centralis-render.sh.

Verified with the pinned shellcheck v0.11.0 (sha256 matching shell-lint.yml):
all 34 scripts in the tree clean. Hook behaviour unchanged — the classifier
test matrix still allows reads and asks on mutations and compound smuggling.

Also documents that ~/.claude/hooks/ carries a mirror of prod-guard.sh so
sessions started outside this repo are covered, and that the blanket ssh allow
rules were removed from global settings so the hook cannot be out-ranked.
2026-07-24 21:12:53 -07:00

84 lines
3.9 KiB
Markdown

# .claude/hooks — enforcement, not advice
`CLAUDE.md` files can only ask. These hooks are the part that enforces, and they
travel with the repo so a session started anywhere in this checkout gets them.
They are wired up in `.claude/settings.json`.
| Hook | Event | What it does |
|---|---|---|
| `prod-guard.sh` | PreToolUse | Live-host commands: reads run unprompted, mutations ask first |
| `secrets-guard.sh` | PreToolUse | Denies any direct write to `infra/deploy/secrets/*.yaml` |
| `lint-after-edit.sh` | PostToolUse | shellchecks an edited `*.sh` and feeds findings back immediately |
## Why prod-guard exists
`agent` has passwordless sudo on prod and beta, and the operator's global settings
allow `Bash(ssh prod:*)` outright. Nothing about `ssh prod 'docker service rm …'`
goes through a pull request, so CI cannot see it — before this hook, any session in
any directory could delete production with no prompt.
**It classifies by allowlist, not blocklist.** Only commands positively recognised as
read-only are allowed through; everything else asks. A blocklist of dangerous verbs
is wrong by construction — the first destructive verb nobody thought of sails
straight through. Being wrong in the "ask" direction costs a keystroke; being wrong
the other way costs production.
Beta is guarded as strictly as prod: it serves beta.thermograph.org *and* hosts
Forgejo, so a destructive command there takes out git, CI and the registry at once.
The LAN dev box is deliberately unguarded.
## Changing the classifier
`is_readonly()` in `prod-guard.sh` splits a command on `&&`, `||`, `;` and `|`, then
checks each segment's leading word. Redirection to a file, command substitution,
`sed -i`, and mutating `docker`/`git`/`systemctl` subcommands all count as writes.
If you add a command to the allowlist, **re-run the test matrix**. There is a
non-obvious failure mode this already hit once: the loop is fed by
`printf '%s\n' | sed`, and dropping that trailing newline makes `read` return
non-zero on the only line, so the loop body never runs and *every* command
classifies as read-only — a silent, total bypass that still looks like it works.
Test by piping a payload straight in:
```bash
echo '{"tool_name":"Bash","tool_input":{"command":"ssh prod '\''rm -rf /'\''"}}' \
| .claude/hooks/prod-guard.sh
# expect: permissionDecision "ask"
```
Exit 0 with no output means "no opinion" and the normal permission flow proceeds.
That is also what happens if `jq` is missing or the script errors, so a bug here
fails toward the prompt rather than toward silent execution.
## lint-after-edit needs shellcheck on PATH
It exits quietly when shellcheck is absent — a missing local tool must not break
editing, and `shell-lint` CI is still the backstop. v0.11.0 (the version CI pins) is
installed at `~/.local/bin/shellcheck`; if you move machines, reinstall it:
```bash
VER=v0.11.0
curl -fsSL "https://github.com/koalaman/shellcheck/releases/download/${VER}/shellcheck-${VER}.linux.x86_64.tar.xz" \
| tar -xJ --strip-components=1 -C ~/.local/bin "shellcheck-${VER}/shellcheck"
```
Keep it matched to `shell-lint.yml`'s pin. A drifted shellcheck grows *new* warnings
and fails CI on an unrelated push, which is why that workflow pins the version and
its sha256 rather than using `apt-get install`.
## The global copy
`~/.claude/settings.json` runs `prod-guard.sh` from `~/.claude/hooks/` as well, so a
session started **outside** this repo is covered too — otherwise the guard would be
trivially bypassed by working from `~/Code`. That copy is a mirror; this one is
canonical. If you change the classifier here, copy it over:
```bash
cp .claude/hooks/prod-guard.sh ~/.claude/hooks/prod-guard.sh
```
The blanket `Bash(ssh prod:*)` / `Bash(ssh beta:*)` allow rules were removed from
global settings at the same time, so the hook is the only thing granting live-host
access. If it is missing or broken, nothing allows the command and you get a prompt.