promote: main → release (#93)
All checks were successful
secrets-guard / encrypted (push) Successful in 18s
shell-lint / shellcheck (push) Successful in 16s
Build + push images (Forgejo registry) / build-push (backend) (push) Successful in 1m27s
Build + push images (Forgejo registry) / build-push (frontend) (push) Successful in 1m25s
Deploy / deploy (backend) (push) Successful in 1m31s
Deploy / deploy (frontend) (push) Successful in 1m45s
All checks were successful
secrets-guard / encrypted (push) Successful in 18s
shell-lint / shellcheck (push) Successful in 16s
Build + push images (Forgejo registry) / build-push (backend) (push) Successful in 1m27s
Build + push images (Forgejo registry) / build-push (frontend) (push) Successful in 1m25s
Deploy / deploy (backend) (push) Successful in 1m31s
Deploy / deploy (frontend) (push) Successful in 1m45s
Reviewed-on: https://git.thermograph.org/emi/thermograph/pulls/93
This commit is contained in:
commit
90f699ceaf
37 changed files with 1243 additions and 1372 deletions
84
.claude/hooks/README.md
Normal file
84
.claude/hooks/README.md
Normal file
|
|
@ -0,0 +1,84 @@
|
|||
# .claude/hooks — enforcement, not advice
|
||||
|
||||
`CLAUDE.md` files can only ask. These hooks are the part that enforces, and they
|
||||
travel with the repo so a session started anywhere in this checkout gets them.
|
||||
|
||||
They are wired up in `.claude/settings.json`.
|
||||
|
||||
| Hook | Event | What it does |
|
||||
|---|---|---|
|
||||
| `prod-guard.sh` | PreToolUse | Live-host commands: reads run unprompted, mutations ask first |
|
||||
| `secrets-guard.sh` | PreToolUse | Denies any direct write to `infra/deploy/secrets/*.yaml` |
|
||||
| `lint-after-edit.sh` | PostToolUse | shellchecks an edited `*.sh` and feeds findings back immediately |
|
||||
|
||||
## Why prod-guard exists
|
||||
|
||||
`agent` has passwordless sudo on prod and beta, and the operator's global settings
|
||||
allow `Bash(ssh prod:*)` outright. Nothing about `ssh prod 'docker service rm …'`
|
||||
goes through a pull request, so CI cannot see it — before this hook, any session in
|
||||
any directory could delete production with no prompt.
|
||||
|
||||
**It classifies by allowlist, not blocklist.** Only commands positively recognised as
|
||||
read-only are allowed through; everything else asks. A blocklist of dangerous verbs
|
||||
is wrong by construction — the first destructive verb nobody thought of sails
|
||||
straight through. Being wrong in the "ask" direction costs a keystroke; being wrong
|
||||
the other way costs production.
|
||||
|
||||
Beta is guarded as strictly as prod: it serves beta.thermograph.org *and* hosts
|
||||
Forgejo, so a destructive command there takes out git, CI and the registry at once.
|
||||
The LAN dev box is deliberately unguarded.
|
||||
|
||||
## Changing the classifier
|
||||
|
||||
`is_readonly()` in `prod-guard.sh` splits a command on `&&`, `||`, `;` and `|`, then
|
||||
checks each segment's leading word. Redirection to a file, command substitution,
|
||||
`sed -i`, and mutating `docker`/`git`/`systemctl` subcommands all count as writes.
|
||||
|
||||
If you add a command to the allowlist, **re-run the test matrix**. There is a
|
||||
non-obvious failure mode this already hit once: the loop is fed by
|
||||
`printf '%s\n' | sed`, and dropping that trailing newline makes `read` return
|
||||
non-zero on the only line, so the loop body never runs and *every* command
|
||||
classifies as read-only — a silent, total bypass that still looks like it works.
|
||||
|
||||
Test by piping a payload straight in:
|
||||
|
||||
```bash
|
||||
echo '{"tool_name":"Bash","tool_input":{"command":"ssh prod '\''rm -rf /'\''"}}' \
|
||||
| .claude/hooks/prod-guard.sh
|
||||
# expect: permissionDecision "ask"
|
||||
```
|
||||
|
||||
Exit 0 with no output means "no opinion" and the normal permission flow proceeds.
|
||||
That is also what happens if `jq` is missing or the script errors, so a bug here
|
||||
fails toward the prompt rather than toward silent execution.
|
||||
|
||||
## lint-after-edit needs shellcheck on PATH
|
||||
|
||||
It exits quietly when shellcheck is absent — a missing local tool must not break
|
||||
editing, and `shell-lint` CI is still the backstop. v0.11.0 (the version CI pins) is
|
||||
installed at `~/.local/bin/shellcheck`; if you move machines, reinstall it:
|
||||
|
||||
```bash
|
||||
VER=v0.11.0
|
||||
curl -fsSL "https://github.com/koalaman/shellcheck/releases/download/${VER}/shellcheck-${VER}.linux.x86_64.tar.xz" \
|
||||
| tar -xJ --strip-components=1 -C ~/.local/bin "shellcheck-${VER}/shellcheck"
|
||||
```
|
||||
|
||||
Keep it matched to `shell-lint.yml`'s pin. A drifted shellcheck grows *new* warnings
|
||||
and fails CI on an unrelated push, which is why that workflow pins the version and
|
||||
its sha256 rather than using `apt-get install`.
|
||||
|
||||
## The global copy
|
||||
|
||||
`~/.claude/settings.json` runs `prod-guard.sh` from `~/.claude/hooks/` as well, so a
|
||||
session started **outside** this repo is covered too — otherwise the guard would be
|
||||
trivially bypassed by working from `~/Code`. That copy is a mirror; this one is
|
||||
canonical. If you change the classifier here, copy it over:
|
||||
|
||||
```bash
|
||||
cp .claude/hooks/prod-guard.sh ~/.claude/hooks/prod-guard.sh
|
||||
```
|
||||
|
||||
The blanket `Bash(ssh prod:*)` / `Bash(ssh beta:*)` allow rules were removed from
|
||||
global settings at the same time, so the hook is the only thing granting live-host
|
||||
access. If it is missing or broken, nothing allows the command and you get a prompt.
|
||||
34
.claude/hooks/lint-after-edit.sh
Executable file
34
.claude/hooks/lint-after-edit.sh
Executable file
|
|
@ -0,0 +1,34 @@
|
|||
#!/usr/bin/env bash
|
||||
# PostToolUse: shellcheck a shell script the moment it is edited.
|
||||
#
|
||||
# The scripts under infra/deploy/ run as root over SSH against live hosts with no
|
||||
# test suite in front of them — a quoting bug or an unset-variable typo is found on
|
||||
# the host, or not at all. shell-lint.yml already guards this, but only after a
|
||||
# push: a five-minute round trip to learn about a missing quote, by which point the
|
||||
# context that produced it is gone.
|
||||
#
|
||||
# Findings are fed back to the model rather than blocking the edit. The file is
|
||||
# already written; the useful move is to fix it now, in the same turn.
|
||||
#
|
||||
# Matches shell-lint.yml's pinned shellcheck when one is on PATH. If shellcheck is
|
||||
# not installed this exits quietly — CI is still the backstop, and a missing local
|
||||
# tool must not break editing.
|
||||
set -uo pipefail
|
||||
|
||||
payload=$(cat)
|
||||
command -v jq >/dev/null 2>&1 || exit 0
|
||||
command -v shellcheck >/dev/null 2>&1 || exit 0
|
||||
|
||||
path=$(printf '%s' "$payload" | jq -r '.tool_response.filePath // .tool_input.file_path // ""')
|
||||
[ -n "$path" ] || exit 0
|
||||
case "$path" in *.sh) ;; *) exit 0 ;; esac
|
||||
[ -f "$path" ] || exit 0
|
||||
|
||||
if out=$(shellcheck -f gcc "$path" 2>&1); then
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# Keep the feedback small: findings only, capped.
|
||||
out=$(printf '%s' "$out" | head -30)
|
||||
jq -n --arg p "$path" --arg o "$out" \
|
||||
'{decision:"block", reason:("shellcheck findings in \($p) — shell-lint CI will fail on these, so fix them now:\n\($o)")}'
|
||||
156
.claude/hooks/prod-guard.sh
Executable file
156
.claude/hooks/prod-guard.sh
Executable file
|
|
@ -0,0 +1,156 @@
|
|||
#!/usr/bin/env bash
|
||||
# PreToolUse guard for commands that reach a live host.
|
||||
#
|
||||
# Policy: reads stay friction-free, mutations ask first.
|
||||
#
|
||||
# The estate's widest privilege is that `agent` has passwordless sudo on prod and
|
||||
# beta, and the global settings allow `ssh prod:*` outright — so before this hook
|
||||
# existed, any session in any directory could roll, delete or restart production
|
||||
# with no prompt. CI cannot help here: nothing about `ssh prod 'docker service rm'`
|
||||
# goes through a pull request.
|
||||
#
|
||||
# Classification is an ALLOWLIST of read-only commands, not a blocklist of
|
||||
# dangerous ones. A blocklist is wrong by construction: the first destructive verb
|
||||
# nobody thought of sails straight through. Anything this script does not
|
||||
# positively recognise as a read becomes an "ask", which is the safe direction to
|
||||
# be wrong in.
|
||||
#
|
||||
# Exit 0 with no output = "no opinion", and the normal permission flow proceeds.
|
||||
# That is also what happens if this script errors or jq is missing, so a bug here
|
||||
# fails toward the prompt rather than toward silent execution.
|
||||
#
|
||||
# Beta is guarded as strictly as prod: it serves beta.thermograph.org AND hosts
|
||||
# Forgejo, so a destructive command there can take out the git server, CI and the
|
||||
# container registry at once.
|
||||
set -uo pipefail
|
||||
|
||||
payload=$(cat)
|
||||
command -v jq >/dev/null 2>&1 || exit 0
|
||||
|
||||
tool=$(printf '%s' "$payload" | jq -r '.tool_name // ""')
|
||||
|
||||
# An ssh target naming a live host, matched as a whole token after any `user@`.
|
||||
# Backslashes are doubled because awk -v processes escape sequences in the value
|
||||
# before the regex ever sees it; a single \. arrives as a bare dot (and warns).
|
||||
readonly HOST_TOKEN_RE='^(prod|beta|169\\.58\\.46\\.181|75\\.119\\.132\\.91|10\\.10\\.0\\.[12])$'
|
||||
|
||||
emit() { # $1=allow|ask|deny $2=reason
|
||||
jq -n --arg d "$1" --arg r "$2" \
|
||||
'{hookSpecificOutput:{hookEventName:"PreToolUse",permissionDecision:$d,permissionDecisionReason:$r}}'
|
||||
exit 0
|
||||
}
|
||||
|
||||
# --- is this shell command read-only? ----------------------------------------
|
||||
# Returns 0 only when every segment leads with a recognised read-only command.
|
||||
is_readonly() {
|
||||
local cmd="$1" seg w1 w2 stripped
|
||||
|
||||
# Discard the harmless redirections that appear in ordinary read commands,
|
||||
# then treat any surviving redirect as a write.
|
||||
stripped=${cmd//2>&1/}
|
||||
stripped=${stripped//&>\/dev\/null/}
|
||||
stripped=${stripped//2>\/dev\/null/}
|
||||
stripped=${stripped//>\/dev\/null/}
|
||||
stripped=${stripped//> \/dev\/null/}
|
||||
case "$stripped" in *'>'*) return 1 ;; esac
|
||||
|
||||
# Command substitution can hide anything; refuse to reason about it.
|
||||
# shellcheck disable=SC2016 # single quotes are the point: these are literal
|
||||
# `$(` and backtick characters being searched for inside the command text, not
|
||||
# expansions to perform. Expanding them here would defeat the check entirely.
|
||||
case "$cmd" in *'$('*|*'`'*) return 1 ;; esac
|
||||
|
||||
while IFS= read -r seg; do
|
||||
seg="${seg#"${seg%%[![:space:]]*}"}"
|
||||
[ -z "$seg" ] && continue
|
||||
w1=$(printf '%s' "$seg" | awk '{print $1}')
|
||||
w2=$(printf '%s' "$seg" | awk '{print $2}')
|
||||
case "$w1" in
|
||||
cat|ls|head|tail|wc|grep|egrep|fgrep|rg|find|stat|df|du|free|uptime|whoami) ;;
|
||||
id|hostname|date|ps|env|printenv|pwd|which|uname|lsblk|ip|ss|netstat) ;;
|
||||
pg_isready|echo|true|sort|uniq|cut|tr|awk|jq|nproc|readlink|realpath|test) ;;
|
||||
sed) case "$seg" in *' -i'*) return 1 ;; esac ;;
|
||||
curl) case "$seg" in *' -X'*|*' -d'*|*'--data'*|*' -T'*|*' -F'*|*'--upload'*) return 1 ;; esac ;;
|
||||
journalctl) ;;
|
||||
systemctl)
|
||||
case "$w2" in status|show|is-active|is-enabled|is-failed|list-units|cat) ;; *) return 1 ;; esac ;;
|
||||
docker)
|
||||
case "$w2" in
|
||||
ps|logs|inspect|stats|images|version|info|top|port|events|history|diff) ;;
|
||||
service) case "$seg" in *'service ls'*|*'service ps'*|*'service inspect'*|*'service logs'*) ;; *) return 1 ;; esac ;;
|
||||
stack) case "$seg" in *'stack ls'*|*'stack ps'*|*'stack services'*) ;; *) return 1 ;; esac ;;
|
||||
compose) case "$seg" in *'compose ps'*|*'compose logs'*|*'compose config'*|*'compose top'*|*'compose images'*) ;; *) return 1 ;; esac ;;
|
||||
node) case "$seg" in *'node ls'*|*'node inspect'*) ;; *) return 1 ;; esac ;;
|
||||
volume) case "$seg" in *'volume ls'*|*'volume inspect'*) ;; *) return 1 ;; esac ;;
|
||||
*) return 1 ;;
|
||||
esac ;;
|
||||
git)
|
||||
case "$w2" in status|log|diff|show|branch|remote|rev-parse|describe|blame|ls-files|config) ;; *) return 1 ;; esac ;;
|
||||
*) return 1 ;;
|
||||
esac
|
||||
# NOTE: the trailing newline matters. Without it `read` returns non-zero on the
|
||||
# final (only) line, the loop body never runs, and every command classifies as
|
||||
# read-only — a silent false negative that allows anything.
|
||||
done < <(printf '%s\n' "$cmd" | sed 's/&&/\n/g; s/||/\n/g; s/;/\n/g; s/|/\n/g')
|
||||
return 0
|
||||
}
|
||||
|
||||
# --- find an ssh target and return what would run on it -----------------------
|
||||
# Prints "<host>\t<remote command>" when a live host is addressed, nothing
|
||||
# otherwise. The remote command is empty for a bare login shell.
|
||||
ssh_target_of() {
|
||||
printf '%s' "$1" | awk -v re="$HOST_TOKEN_RE" '
|
||||
{
|
||||
for (i = 1; i <= NF; i++) {
|
||||
t = $i; sub(/^[^@]*@/, "", t)
|
||||
if (t ~ re) {
|
||||
out = ""
|
||||
for (j = i + 1; j <= NF; j++) out = out (out == "" ? "" : " ") $j
|
||||
printf "%s\t%s\n", t, out
|
||||
exit
|
||||
}
|
||||
}
|
||||
}'
|
||||
}
|
||||
|
||||
case "$tool" in
|
||||
Bash)
|
||||
cmd=$(printf '%s' "$payload" | jq -r '.tool_input.command // ""')
|
||||
case "$cmd" in *ssh*) ;; *) exit 0 ;; esac
|
||||
found=$(ssh_target_of "$cmd")
|
||||
[ -n "$found" ] || exit 0 # no live host addressed
|
||||
host=${found%%$'\t'*}
|
||||
remote=${found#*$'\t'}
|
||||
remote=$(printf '%s' "$remote" | sed -E "s/^'(.*)'$/\1/; s/^\"(.*)\"$/\1/")
|
||||
[ -z "$remote" ] && exit 0 # bare login shell
|
||||
if is_readonly "$remote"; then
|
||||
emit allow "read-only command on ${host}"
|
||||
fi
|
||||
emit ask "This mutates ${host}, a live host where the agent user has passwordless sudo. Command: ${remote}"
|
||||
;;
|
||||
|
||||
mcp__centralis__run_on_host)
|
||||
env=$(printf '%s' "$payload" | jq -r '.tool_input.env // ""')
|
||||
cmd=$(printf '%s' "$payload" | jq -r '.tool_input.command // ""')
|
||||
case "$env" in prod|beta) ;; *) exit 0 ;; esac
|
||||
if is_readonly "$cmd"; then
|
||||
emit allow "read-only command on ${env}"
|
||||
fi
|
||||
emit ask "run_on_host would mutate ${env}. Command: ${cmd}"
|
||||
;;
|
||||
|
||||
mcp__centralis__sql_query)
|
||||
env=$(printf '%s' "$payload" | jq -r '.tool_input.env // ""')
|
||||
write=$(printf '%s' "$payload" | jq -r '.tool_input.write // false')
|
||||
[ "$write" = "true" ] || exit 0
|
||||
case "$env" in prod|beta) ;; *) exit 0 ;; esac
|
||||
sql=$(printf '%s' "$payload" | jq -r '.tool_input.sql // ""' | tr '\n' ' ' | cut -c1-300)
|
||||
emit ask "Write query against the ${env} database. sql_query escalates to the app role and has no confirmation of its own. SQL: ${sql}"
|
||||
;;
|
||||
|
||||
mcp__centralis__rollback_to|mcp__centralis__promote|mcp__centralis__secrets_rotate)
|
||||
emit ask "${tool#mcp__centralis__} changes production state directly."
|
||||
;;
|
||||
|
||||
*) exit 0 ;;
|
||||
esac
|
||||
34
.claude/hooks/secrets-guard.sh
Executable file
34
.claude/hooks/secrets-guard.sh
Executable file
|
|
@ -0,0 +1,34 @@
|
|||
#!/usr/bin/env bash
|
||||
# PreToolUse guard: never write infra/deploy/secrets/*.yaml directly.
|
||||
#
|
||||
# Every file matching that glob is SOPS-encrypted, and secrets-guard.yml fails the
|
||||
# build if one is not. This is the same rule enforced one step earlier: a direct
|
||||
# Write or Edit produces plaintext, and by the time CI catches it the secret has
|
||||
# already been written to disk and probably committed.
|
||||
#
|
||||
# The supported path is `sops edit <file>`, which decrypts to a temp file, opens an
|
||||
# editor and re-encrypts on save — the plaintext never touches the working tree.
|
||||
#
|
||||
# Denies rather than asks: there is no legitimate reason to hand-write one of these,
|
||||
# so a prompt would only be an invitation to click through.
|
||||
set -uo pipefail
|
||||
|
||||
payload=$(cat)
|
||||
command -v jq >/dev/null 2>&1 || exit 0
|
||||
|
||||
path=$(printf '%s' "$payload" | jq -r '.tool_input.file_path // ""')
|
||||
[ -n "$path" ] || exit 0
|
||||
|
||||
case "$path" in
|
||||
*/infra/deploy/secrets/*.yaml)
|
||||
jq -n --arg p "$path" '{
|
||||
hookSpecificOutput: {
|
||||
hookEventName: "PreToolUse",
|
||||
permissionDecision: "deny",
|
||||
permissionDecisionReason: ("Refusing to write \($p) directly — every file under infra/deploy/secrets/ is SOPS-encrypted and a direct write would land plaintext. Use `sops edit \($p)` instead, which re-encrypts on save. (secrets-guard.yml would fail the build on this anyway, but only after the secret was already on disk.)")
|
||||
}
|
||||
}'
|
||||
exit 0
|
||||
;;
|
||||
*) exit 0 ;;
|
||||
esac
|
||||
42
.claude/settings.json
Normal file
42
.claude/settings.json
Normal file
|
|
@ -0,0 +1,42 @@
|
|||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"hooks": {
|
||||
"PreToolUse": [
|
||||
{
|
||||
"matcher": "Bash|mcp__centralis__run_on_host|mcp__centralis__sql_query|mcp__centralis__rollback_to|mcp__centralis__promote|mcp__centralis__secrets_rotate",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "\"${CLAUDE_PROJECT_DIR:-.}/.claude/hooks/prod-guard.sh\"",
|
||||
"timeout": 10,
|
||||
"statusMessage": "Checking live-host access"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"matcher": "Write|Edit",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "\"${CLAUDE_PROJECT_DIR:-.}/.claude/hooks/secrets-guard.sh\"",
|
||||
"timeout": 10,
|
||||
"statusMessage": "Checking SOPS vault"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"PostToolUse": [
|
||||
{
|
||||
"matcher": "Write|Edit",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "\"${CLAUDE_PROJECT_DIR:-.}/.claude/hooks/lint-after-edit.sh\"",
|
||||
"timeout": 30,
|
||||
"statusMessage": "shellcheck"
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
|
@ -1,126 +0,0 @@
|
|||
name: Build + push backend image (Forgejo registry)
|
||||
|
||||
# Builds the BACKEND domain's image (backend/Dockerfile, context backend/) and
|
||||
# pushes it to Forgejo's built-in container registry, tagged by git SHA (every
|
||||
# push) and semver (version tags) -- build-once, deploy-everywhere.
|
||||
#
|
||||
# Monorepo port (reunification): the split repos each derived IMAGE_PATH from
|
||||
# github.repository, which now collides for every domain -- so each domain's
|
||||
# build-push names its image explicitly: this one is emi/thermograph/backend
|
||||
# (frontend-build-push.yml publishes emi/thermograph/frontend). infra's
|
||||
# compose/stack files reference the two paths independently
|
||||
# (BACKEND_IMAGE_PATH / FRONTEND_IMAGE_PATH -- their defaults were renamed to
|
||||
# match in the same cutover).
|
||||
#
|
||||
# Path-filtered: only a push that touches backend/** builds this image. A
|
||||
# frontend-only push builds nothing here; deploy.sh's persisted
|
||||
# .image-tags.env keeps the backend's live tag for the sibling's deploy.
|
||||
# EXCEPTION: version-tag pushes (v*.*.*) carry no paths context, so a semver
|
||||
# tag builds BOTH domain images -- intended, a release names the whole set.
|
||||
#
|
||||
# Runs on the `docker` label -- the LAN runner's job containers get the host's
|
||||
# docker.sock automounted in (forgejo-runner config: container.docker_host:
|
||||
# automount; see infra/deploy/forgejo/register-lan-runner.sh),
|
||||
# Docker-outside-of-Docker rather than DinD -- no privileged mode needed.
|
||||
# The job image itself (node:20-bookworm) has no docker CLI, hence the
|
||||
# "Install Docker CLI" step.
|
||||
#
|
||||
# The registry is deliberately mesh-only: git.thermograph.org's public DNS
|
||||
# resolves to beta's public IP, which Caddy's ACL rejects for /v2/ paths, so
|
||||
# any runner host that pushes/pulls needs a /etc/hosts entry pointing
|
||||
# git.thermograph.org at beta's WireGuard IP (10.10.0.2). The docker
|
||||
# build/push commands run against the automounted HOST daemon, so it's the
|
||||
# runner HOST's own resolver that has to route over the mesh, not anything
|
||||
# settable from inside the job container.
|
||||
#
|
||||
# REGISTRY is the vars.REGISTRY_HOST Actions variable (git.thermograph.org),
|
||||
# NOT github.server_url -- that context resolves to the bare mesh IP:port
|
||||
# (http://10.10.0.2:3080, no TLS), which docker CLI refuses ("server gave HTTP
|
||||
# response to HTTPS client"). git.thermograph.org gets real TLS via Caddy and
|
||||
# routes over the mesh once /etc/hosts is set, so it's the only correct value.
|
||||
# deploy.sh resolves the SHA tag at deploy time and `docker compose pull`s it.
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [dev, main, release]
|
||||
paths: ['backend/**']
|
||||
tags: ['v*.*.*']
|
||||
workflow_dispatch: {}
|
||||
|
||||
# The runner has capacity: 1 (one physical LAN box) -- a burst of pushes to the
|
||||
# same ref would otherwise queue whole builds back to back even though only the
|
||||
# last one still matters. Keying on domain + github.ref means each domain's
|
||||
# dev/main/release/tag builds get their own group, so this only cancels a
|
||||
# superseded build of the SAME domain on the SAME ref -- and a frontend build
|
||||
# never cancels a backend one.
|
||||
concurrency:
|
||||
group: backend-build-push-${{ github.ref }}
|
||||
cancel-in-progress: true
|
||||
|
||||
jobs:
|
||||
build-push:
|
||||
runs-on: docker
|
||||
env:
|
||||
REGISTRY: https://${{ vars.REGISTRY_HOST }}
|
||||
IMAGE_PATH: ${{ github.repository }}/backend
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
|
||||
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
|
||||
# often an unrelated domain's commit, and a depth-1 clone can't see
|
||||
# past it to find the real key.
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Install Docker CLI
|
||||
run: |
|
||||
apt-get update -qq
|
||||
apt-get install -y -qq docker.io
|
||||
|
||||
- name: Compute image ref + tags
|
||||
id: image
|
||||
run: |
|
||||
host="${REGISTRY#https://}"; host="${host#http://}"
|
||||
path="$(printf '%s' "$IMAGE_PATH" | tr '[:upper:]' '[:lower:]')"
|
||||
ref="$host/$path"
|
||||
# Key the tag to the last commit that touched backend/ -- the SAME key
|
||||
# every backend deploy workflow computes -- so build and deploy always
|
||||
# agree even when the branch tip is another domain's commit. (A tag
|
||||
# keyed to the push tip breaks the moment an infra-only commit lands:
|
||||
# deploys go looking for an image no build ever produced.)
|
||||
domain_sha="$(git log -1 --format=%H -- backend/)"
|
||||
sha_tag="$ref:sha-${domain_sha:0:12}"
|
||||
echo "host=$host" >> "$GITHUB_OUTPUT"
|
||||
echo "sha_tag=$sha_tag" >> "$GITHUB_OUTPUT"
|
||||
if [[ "$GITHUB_REF" == refs/tags/v* ]]; then
|
||||
echo "semver_tag=$ref:${GITHUB_REF#refs/tags/}" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
|
||||
- name: Log in to the Forgejo registry
|
||||
run: |
|
||||
# NOT secrets.GITHUB_TOKEN -- Forgejo's per-job auto-injected token
|
||||
# can never push to the container registry (a known Forgejo
|
||||
# limitation independent of any permission setting). Needs a
|
||||
# manually provisioned PAT with write:package scope, stored as the
|
||||
# REGISTRY_TOKEN secret.
|
||||
echo "${{ secrets.REGISTRY_TOKEN }}" | docker login "${{ steps.image.outputs.host }}" \
|
||||
--username emi --password-stdin
|
||||
|
||||
- name: Build
|
||||
run: |
|
||||
# One build, both tags -- the semver tag (tag pushes only) rides
|
||||
# along with the SHA tag instead of a second tag+push round trip.
|
||||
# Context is the DOMAIN dir, so backend/Dockerfile sees the same
|
||||
# tree it saw as a repo root in the split era.
|
||||
tag_args=(-t "${{ steps.image.outputs.sha_tag }}")
|
||||
if [[ -n "${{ steps.image.outputs.semver_tag }}" ]]; then
|
||||
tag_args+=(-t "${{ steps.image.outputs.semver_tag }}")
|
||||
fi
|
||||
docker build "${tag_args[@]}" backend/
|
||||
|
||||
- name: Push (SHA tag, every push)
|
||||
run: docker push "${{ steps.image.outputs.sha_tag }}"
|
||||
|
||||
- name: Push (semver tag, tag pushes only)
|
||||
if: startsWith(github.ref, 'refs/tags/v')
|
||||
run: docker push "${{ steps.image.outputs.semver_tag }}"
|
||||
|
|
@ -1,79 +0,0 @@
|
|||
name: Deploy backend to LAN dev server
|
||||
|
||||
# Fires on a push to `dev` touching backend/** -- a direct push, or the real
|
||||
# merge commit a native "auto merge when checks succeed" produces (see
|
||||
# pr-build.yml). A `build` job (the reusable build.yml sanity check) gates a
|
||||
# `deploy` job that rolls the LAN dev box.
|
||||
#
|
||||
# CUTOVER NOTE (monorepo reunification): the LAN box's ~/thermograph-dev is a
|
||||
# thermograph-infra checkout from the split era. This job calls
|
||||
# ~/thermograph-dev/infra/deploy/deploy-dev.sh -- the MONOREPO layout -- so it
|
||||
# is INERT until ~/thermograph-dev is re-pointed at the monorepo (a fresh
|
||||
# clone; deploy-dev.sh then self-updates via deploy.sh's git reset like the
|
||||
# beta/prod paths).
|
||||
#
|
||||
# SERVICE=backend + BACKEND_IMAGE_TAG is the same env-var contract the
|
||||
# beta/prod deploy workflows use against infra/deploy/deploy.sh, so a dev-only
|
||||
# rollout never touches the frontend's running container or tag.
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [dev]
|
||||
paths: ['backend/**']
|
||||
workflow_dispatch: {}
|
||||
|
||||
jobs:
|
||||
build:
|
||||
name: build
|
||||
# Fixed (unparameterized) group: if two backend pushes to dev land close
|
||||
# together, the newer push's build cancels the older's still-running one
|
||||
# on the single physical runner -- the older result would be thrown away
|
||||
# once needs:build gates its deploy anyway. Safe to cancel mid-run:
|
||||
# build.yml only runs a throwaway `docker build`, nothing persisted.
|
||||
concurrency:
|
||||
group: dev-push-build-backend
|
||||
cancel-in-progress: true
|
||||
uses: ./.forgejo/workflows/build.yml
|
||||
with:
|
||||
domain: backend
|
||||
|
||||
deploy:
|
||||
needs: build
|
||||
# NOT [self-hosted, thermograph-lan] -- GitHub implicitly tags every
|
||||
# self-hosted runner with "self-hosted"; Forgejo's runner has only the
|
||||
# labels it was registered with ("docker" + "thermograph-lan"). The array
|
||||
# form silently makes this job unschedulable on every push.
|
||||
runs-on: thermograph-lan
|
||||
permissions:
|
||||
contents: read
|
||||
concurrency:
|
||||
group: dev-lan-deploy-backend
|
||||
cancel-in-progress: false
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
|
||||
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
|
||||
# often an unrelated domain's commit, and a depth-1 clone can't see
|
||||
# past it to find the real key.
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Compute image tag
|
||||
id: tag
|
||||
# Keyed to the last commit that touched backend/ -- must match
|
||||
# backend-build-push.yml's tag key exactly (see its comment).
|
||||
run: |
|
||||
domain_sha="$(git log -1 --format=%H -- backend/)"
|
||||
echo "tag=sha-${domain_sha:0:12}" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Deploy to the LAN dev server
|
||||
# thermograph-lan is host-native (the runner job runs directly on the
|
||||
# LAN box, not over SSH like beta/prod), so plain env vars on the
|
||||
# command reach deploy-dev.sh directly -- no ssh-action needed here.
|
||||
# The lake creds ride the Actions S3 secrets: dev renders no vault
|
||||
# (deploy-dev.sh's no-op secrets path), and compose interpolates
|
||||
# ${THERMOGRAPH_LAKE_S3_*} straight from the deploy environment.
|
||||
env:
|
||||
THERMOGRAPH_LAKE_S3_ACCESS_KEY: ${{ secrets.S3_ACCESS_KEY }}
|
||||
THERMOGRAPH_LAKE_S3_SECRET_KEY: ${{ secrets.S3_SECRET_KEY }}
|
||||
run: SERVICE=backend BACKEND_IMAGE_TAG=${{ steps.tag.outputs.tag }} bash "$HOME/thermograph-dev/infra/deploy/deploy-dev.sh"
|
||||
|
|
@ -1,56 +0,0 @@
|
|||
name: Deploy backend to prod VPS
|
||||
|
||||
# On a push to `release` that touches backend/**, SSH to prod and roll ONLY
|
||||
# the backend service onto the image backend-build-push.yml published for this
|
||||
# commit. Same shape as backend-deploy.yml (the `main`->beta path) against a
|
||||
# completely separate secret set (PROD_SSH_HOST/USER/KEY/PORT) so a beta
|
||||
# credential leak can't reach prod. Frontend deploys itself independently via
|
||||
# frontend-deploy-prod.yml -- a backend change ships to prod without touching
|
||||
# frontend and vice versa, exactly as in the split era.
|
||||
#
|
||||
# The tag MUST match backend-build-push.yml's last-backend-commit key (12
|
||||
# hex) -- see backend-deploy.yml for why it's truncated here. deploy.sh
|
||||
# retries the pull because the build for this push may still be in flight.
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [release]
|
||||
paths: ['backend/**']
|
||||
workflow_dispatch: {}
|
||||
|
||||
concurrency:
|
||||
group: prod-deploy-backend
|
||||
cancel-in-progress: false
|
||||
|
||||
jobs:
|
||||
deploy:
|
||||
runs-on: docker
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
|
||||
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
|
||||
# often an unrelated domain's commit, and a depth-1 clone can't see
|
||||
# past it to find the real key.
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Compute image tag (last backend-touching commit)
|
||||
id: tag
|
||||
run: |
|
||||
domain_sha="$(git log -1 --format=%H -- backend/)"
|
||||
echo "tag=sha-${domain_sha:0:12}" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Deploy backend over SSH
|
||||
uses: https://github.com/appleboy/ssh-action@v1.2.0
|
||||
with:
|
||||
host: ${{ secrets.PROD_SSH_HOST }}
|
||||
username: ${{ secrets.PROD_SSH_USER }}
|
||||
key: ${{ secrets.PROD_SSH_KEY }}
|
||||
port: ${{ secrets.PROD_SSH_PORT }}
|
||||
# Forward the target service + the image tag to run. deploy.sh reads
|
||||
# BACKEND_IMAGE_TAG and rolls just `backend`.
|
||||
envs: SERVICE,BACKEND_IMAGE_TAG
|
||||
script: /opt/thermograph/infra/deploy/deploy.sh
|
||||
env:
|
||||
SERVICE: backend
|
||||
BACKEND_IMAGE_TAG: ${{ steps.tag.outputs.tag }}
|
||||
|
|
@ -1,65 +0,0 @@
|
|||
name: Deploy backend to beta VPS
|
||||
|
||||
# On a push to `main` that touches backend/**, SSH to beta and roll ONLY the
|
||||
# backend service onto the image backend-build-push.yml just published for
|
||||
# this commit. Frontend is deployed independently by frontend-deploy.yml --
|
||||
# the FE/BE CI-CD split survives reunification intact: a backend change ships
|
||||
# without touching frontend and vice versa. infra/deploy/deploy.sh persists
|
||||
# each service's live tag host-side, so a single-service roll never disturbs
|
||||
# the other. An infra-only push rolls nothing (infra-sync.yml syncs the
|
||||
# host checkouts instead); a mixed backend+frontend push fires both deploy
|
||||
# workflows, serialized host-side by deploy.sh's flock.
|
||||
#
|
||||
# The SSH_HOST/USER/KEY/PORT secrets point at beta. appleboy/ssh-action is
|
||||
# referenced by full GitHub URL because it isn't mirrored in Forgejo's default
|
||||
# action registry.
|
||||
#
|
||||
# The tag MUST match backend-build-push.yml's `sha-${GITHUB_SHA:0:12}` (12
|
||||
# hex), not the full 40-char ${{ github.sha }} -- default Actions expressions
|
||||
# have no substring function, so the compute step below truncates it.
|
||||
# deploy.sh also retries the pull for ~5 min because Forgejo `needs:` can't
|
||||
# gate across the separate build-push workflow, so this push's build may still
|
||||
# be in flight.
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
paths: ['backend/**']
|
||||
workflow_dispatch: {}
|
||||
|
||||
concurrency:
|
||||
group: beta-deploy-backend
|
||||
cancel-in-progress: false
|
||||
|
||||
jobs:
|
||||
deploy:
|
||||
runs-on: docker
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
|
||||
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
|
||||
# often an unrelated domain's commit, and a depth-1 clone can't see
|
||||
# past it to find the real key.
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Compute image tag (last backend-touching commit)
|
||||
id: tag
|
||||
run: |
|
||||
domain_sha="$(git log -1 --format=%H -- backend/)"
|
||||
echo "tag=sha-${domain_sha:0:12}" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Deploy backend over SSH
|
||||
uses: https://github.com/appleboy/ssh-action@v1.2.0
|
||||
with:
|
||||
host: ${{ secrets.SSH_HOST }}
|
||||
username: ${{ secrets.SSH_USER }}
|
||||
key: ${{ secrets.SSH_KEY }}
|
||||
port: ${{ secrets.SSH_PORT }}
|
||||
# Forward the target service + the image tag to run. deploy.sh reads
|
||||
# BACKEND_IMAGE_TAG and rolls just `backend`.
|
||||
envs: SERVICE,BACKEND_IMAGE_TAG
|
||||
script: /opt/thermograph/infra/deploy/deploy.sh
|
||||
env:
|
||||
SERVICE: backend
|
||||
BACKEND_IMAGE_TAG: ${{ steps.tag.outputs.tag }}
|
||||
144
.forgejo/workflows/build-push.yml
Normal file
144
.forgejo/workflows/build-push.yml
Normal file
|
|
@ -0,0 +1,144 @@
|
|||
name: Build + push images (Forgejo registry)
|
||||
|
||||
# backend-build-push.yml and frontend-build-push.yml, collapsed into one.
|
||||
# They were identical apart from the domain string: the paths filter, the
|
||||
# concurrency group, IMAGE_PATH, the tag key and the build context.
|
||||
#
|
||||
# Images stay separate and independently deployable -- emi/thermograph/backend
|
||||
# and emi/thermograph/frontend -- exactly as before. Build-once,
|
||||
# deploy-everywhere; nothing about the published artefacts changes here.
|
||||
#
|
||||
# SEMVER EXCEPTION, preserved: a v*.*.* tag push carries no paths context, so a
|
||||
# release tag builds BOTH domain images. That is intended -- a release names the
|
||||
# whole set, not whichever domain happened to move last.
|
||||
#
|
||||
# Runs on the `docker` label. The LAN runner's job containers get the host's
|
||||
# docker.sock automounted (Docker-outside-of-Docker, no privileged mode); the
|
||||
# job image (node:20-bookworm) ships no docker CLI, hence the install step.
|
||||
#
|
||||
# The registry is deliberately mesh-only: git.thermograph.org's public DNS
|
||||
# resolves to beta's public IP, whose Caddy ACL rejects /v2/ paths, so any
|
||||
# runner host that pushes or pulls needs an /etc/hosts entry.
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [dev, main, release]
|
||||
paths:
|
||||
- 'backend/**'
|
||||
- 'frontend/**'
|
||||
tags: ['v*.*.*']
|
||||
workflow_dispatch: {}
|
||||
|
||||
jobs:
|
||||
build-push:
|
||||
strategy:
|
||||
fail-fast: false
|
||||
# capacity: 1 physical runner. Serialising keeps a two-domain push from
|
||||
# queueing two whole image builds against each other.
|
||||
max-parallel: 1
|
||||
matrix:
|
||||
domain: [backend, frontend]
|
||||
|
||||
runs-on: docker
|
||||
|
||||
# Per-domain, per-ref: only a superseded build of the SAME domain on the
|
||||
# SAME ref is cancelled, and a frontend build never cancels a backend one.
|
||||
concurrency:
|
||||
group: build-push-${{ matrix.domain }}-${{ github.ref }}
|
||||
cancel-in-progress: true
|
||||
|
||||
env:
|
||||
REGISTRY: https://${{ vars.REGISTRY_HOST }}
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
|
||||
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
|
||||
# often an unrelated domain's commit, and a depth-1 clone can't see past
|
||||
# it to find the real key.
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Compute image ref + tags
|
||||
id: image
|
||||
run: |
|
||||
set -euo pipefail
|
||||
|
||||
host="${REGISTRY#https://}"; host="${host#http://}"
|
||||
path="$(printf '%s/%s' "${{ github.repository }}" "${{ matrix.domain }}" | tr '[:upper:]' '[:lower:]')"
|
||||
ref="$host/$path"
|
||||
|
||||
# Does this push touch this domain? The workflow-level paths filter
|
||||
# only says backend OR frontend moved; without this a backend-only
|
||||
# push would also rebuild and republish the frontend image.
|
||||
before="${{ github.event.before }}"
|
||||
if [[ "$GITHUB_REF" == refs/tags/v* ]]; then
|
||||
changed=true # semver tag: build the whole set, see header
|
||||
elif [ -z "$before" ] || [ "$before" = "0000000000000000000000000000000000000000" ] \
|
||||
|| ! git cat-file -e "$before^{commit}" 2>/dev/null; then
|
||||
changed=true # no usable base; build rather than silently skip
|
||||
elif git diff --name-only "$before" "${{ github.sha }}" | grep -q "^${{ matrix.domain }}/"; then
|
||||
changed=true
|
||||
else
|
||||
changed=false
|
||||
fi
|
||||
|
||||
# Key the tag to the last commit that touched this domain -- the SAME
|
||||
# key every deploy computes -- so build and deploy always agree even
|
||||
# when the branch tip is another domain's commit. (A tag keyed to the
|
||||
# push tip breaks the moment an infra-only commit lands: deploys go
|
||||
# looking for an image no build ever produced.)
|
||||
domain_sha="$(git log -1 --format=%H -- "${{ matrix.domain }}/")"
|
||||
|
||||
{
|
||||
echo "changed=$changed"
|
||||
echo "host=$host"
|
||||
echo "sha_tag=$ref:sha-${domain_sha:0:12}"
|
||||
} >> "$GITHUB_OUTPUT"
|
||||
|
||||
if [[ "$GITHUB_REF" == refs/tags/v* ]]; then
|
||||
echo "semver_tag=$ref:${GITHUB_REF#refs/tags/}" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
|
||||
echo "==> ${{ matrix.domain }} | changed=$changed | $ref:sha-${domain_sha:0:12}"
|
||||
|
||||
- name: Install Docker CLI
|
||||
if: steps.image.outputs.changed == 'true'
|
||||
run: |
|
||||
apt-get update -qq
|
||||
apt-get install -y -qq docker.io
|
||||
|
||||
- name: Log in to the Forgejo registry
|
||||
if: steps.image.outputs.changed == 'true'
|
||||
run: |
|
||||
# NOT secrets.GITHUB_TOKEN -- Forgejo's per-job auto-injected token can
|
||||
# never push to the container registry (a known Forgejo limitation,
|
||||
# independent of any permission setting). Needs a manually provisioned
|
||||
# PAT with write:package scope, stored as REGISTRY_TOKEN.
|
||||
echo "${{ secrets.REGISTRY_TOKEN }}" | docker login "${{ steps.image.outputs.host }}" \
|
||||
--username emi --password-stdin
|
||||
|
||||
- name: Build
|
||||
if: steps.image.outputs.changed == 'true'
|
||||
run: |
|
||||
# One build, both tags -- the semver tag (tag pushes only) rides along
|
||||
# with the SHA tag instead of a second tag+push round trip. Context is
|
||||
# the DOMAIN dir, so the Dockerfile sees the same tree it saw as a repo
|
||||
# root in the split era.
|
||||
tag_args=(-t "${{ steps.image.outputs.sha_tag }}")
|
||||
if [[ -n "${{ steps.image.outputs.semver_tag }}" ]]; then
|
||||
tag_args+=(-t "${{ steps.image.outputs.semver_tag }}")
|
||||
fi
|
||||
docker build "${tag_args[@]}" "${{ matrix.domain }}/"
|
||||
|
||||
- name: Push (SHA tag, every push)
|
||||
if: steps.image.outputs.changed == 'true'
|
||||
run: docker push "${{ steps.image.outputs.sha_tag }}"
|
||||
|
||||
- name: Push (semver tag, tag pushes only)
|
||||
if: steps.image.outputs.changed == 'true' && startsWith(github.ref, 'refs/tags/v')
|
||||
run: docker push "${{ steps.image.outputs.semver_tag }}"
|
||||
|
||||
- name: Skipped
|
||||
if: steps.image.outputs.changed != 'true'
|
||||
run: echo "${{ matrix.domain }} unchanged in this push — no image built."
|
||||
166
.forgejo/workflows/deploy.yml
Normal file
166
.forgejo/workflows/deploy.yml
Normal file
|
|
@ -0,0 +1,166 @@
|
|||
name: Deploy
|
||||
|
||||
# The six per-domain-per-environment deploy workflows, collapsed into one.
|
||||
#
|
||||
# backend-deploy{,-dev,-prod}.yml and frontend-deploy{,-dev,-prod}.yml were the
|
||||
# same file written six times: they differed only in a branch name, a paths
|
||||
# filter, a concurrency group, the service name, its *_IMAGE_TAG variable and a
|
||||
# secret prefix. The contract into infra/deploy/deploy.sh
|
||||
# (SERVICE + BACKEND_IMAGE_TAG/FRONTEND_IMAGE_TAG) was already fully
|
||||
# parameterised, so the duplication bought nothing and cost six files to keep in
|
||||
# step.
|
||||
#
|
||||
# The two *-deploy-dev.yml workflows are NOT ported. They were documented as
|
||||
# inert: they call the monorepo layout at ~/thermograph-dev on the LAN box,
|
||||
# which is still a split-era thermograph-infra checkout, so that path does not
|
||||
# exist there. LAN dev is a local `make dev-up` concern, not a CI environment.
|
||||
#
|
||||
# DELIBERATELY BORING EXPRESSIONS. No dynamic matrix (fromJSON), no
|
||||
# `cond && secrets.A || secrets.B` ternary. Those are GitHub idioms that a
|
||||
# Forgejo/act runner may evaluate differently, and the failure mode here is
|
||||
# "production does not deploy" or, worse, "deploys with an empty SSH host". The
|
||||
# two environments therefore get two explicit, mutually exclusive steps.
|
||||
#
|
||||
# What is preserved from the originals, all of it load-bearing:
|
||||
# - fetch-depth: 0, because the image tag is keyed to the LAST COMMIT THAT
|
||||
# TOUCHED THAT DOMAIN, not the branch tip. In a path-filtered monorepo the
|
||||
# tip is often an unrelated domain's commit and a depth-1 clone cannot see
|
||||
# past it.
|
||||
# - The 12-hex truncation, matching build-push exactly.
|
||||
# - Per-service, per-environment concurrency with cancel-in-progress: false --
|
||||
# a half-finished deploy must never be cancelled by a newer one.
|
||||
# - Separate PROD_SSH_* credentials, so a beta credential leak cannot reach
|
||||
# prod.
|
||||
# - appleboy/ssh-action by full URL; it is not mirrored in Forgejo's default
|
||||
# action registry.
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [main, release]
|
||||
paths:
|
||||
- 'backend/**'
|
||||
- 'frontend/**'
|
||||
workflow_dispatch: {}
|
||||
|
||||
jobs:
|
||||
deploy:
|
||||
strategy:
|
||||
fail-fast: false
|
||||
# One physical runner, and deploy.sh serialises host-side with flock
|
||||
# anyway. Rolling one service at a time keeps the logs readable and
|
||||
# matches what the six separate workflows effectively did.
|
||||
max-parallel: 1
|
||||
matrix:
|
||||
service: [backend, frontend]
|
||||
|
||||
runs-on: docker
|
||||
|
||||
# Keyed by ref AND service, reproducing the old per-file groups
|
||||
# (beta-deploy-backend, prod-deploy-frontend, ...).
|
||||
concurrency:
|
||||
group: deploy-${{ github.ref_name }}-${{ matrix.service }}
|
||||
cancel-in-progress: false
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Plan this leg
|
||||
id: plan
|
||||
run: |
|
||||
set -euo pipefail
|
||||
|
||||
# Branch selects the environment. main -> beta, release -> prod.
|
||||
case "${{ github.ref_name }}" in
|
||||
main) environment=beta ;;
|
||||
release) environment=prod ;;
|
||||
*) echo "::error::Deploy triggered on unexpected ref '${{ github.ref_name }}'"; exit 1 ;;
|
||||
esac
|
||||
|
||||
# Did THIS push touch THIS domain? The workflow-level paths filter only
|
||||
# tells us backend OR frontend moved; without this refinement a
|
||||
# backend-only push would also roll the frontend, losing the
|
||||
# independent-deploy property the FE/BE split exists for.
|
||||
before="${{ github.event.before }}"
|
||||
if [ -z "$before" ] || [ "$before" = "0000000000000000000000000000000000000000" ] \
|
||||
|| ! git cat-file -e "$before^{commit}" 2>/dev/null; then
|
||||
# No usable base (first push, force push, or workflow_dispatch).
|
||||
# Deploy rather than skip: rolling onto the tag already running is a
|
||||
# no-op for deploy.sh, whereas skipping silently strands a change.
|
||||
changed=true
|
||||
echo "no usable before-sha; defaulting to deploy"
|
||||
elif git diff --name-only "$before" "${{ github.sha }}" | grep -q "^${{ matrix.service }}/"; then
|
||||
changed=true
|
||||
else
|
||||
changed=false
|
||||
fi
|
||||
|
||||
# The tag is the last commit that touched this domain, truncated to 12
|
||||
# hex to match build-push.yml exactly. Actions expressions have no
|
||||
# substring function, which is why this is computed in shell.
|
||||
domain_sha="$(git log -1 --format=%H -- "${{ matrix.service }}/")"
|
||||
|
||||
tag="sha-${domain_sha:0:12}"
|
||||
|
||||
{
|
||||
echo "environment=$environment"
|
||||
echo "changed=$changed"
|
||||
echo "tag=$tag"
|
||||
} >> "$GITHUB_OUTPUT"
|
||||
|
||||
# Export the deploy.sh contract as real environment variables, chosen
|
||||
# in shell rather than with a `matrix.service == 'x' && a || b`
|
||||
# expression. That idiom is a GitHub convention a Forgejo/act runner
|
||||
# may evaluate differently, and the failure here would be silent: an
|
||||
# empty *_IMAGE_TAG makes deploy.sh fall back to the tag already
|
||||
# running, so the job goes green having deployed nothing.
|
||||
{
|
||||
echo "SERVICE=${{ matrix.service }}"
|
||||
if [ "${{ matrix.service }}" = "backend" ]; then
|
||||
echo "BACKEND_IMAGE_TAG=$tag"
|
||||
else
|
||||
echo "FRONTEND_IMAGE_TAG=$tag"
|
||||
fi
|
||||
} >> "$GITHUB_ENV"
|
||||
|
||||
echo "==> ${{ matrix.service }} -> $environment | changed=$changed | tag=$tag"
|
||||
|
||||
- name: Deploy to beta
|
||||
if: steps.plan.outputs.changed == 'true' && steps.plan.outputs.environment == 'beta'
|
||||
uses: https://github.com/appleboy/ssh-action@v1.2.0
|
||||
with:
|
||||
host: ${{ secrets.SSH_HOST }}
|
||||
username: ${{ secrets.SSH_USER }}
|
||||
key: ${{ secrets.SSH_KEY }}
|
||||
port: ${{ secrets.SSH_PORT }}
|
||||
envs: SERVICE,BACKEND_IMAGE_TAG,FRONTEND_IMAGE_TAG
|
||||
script: /opt/thermograph/infra/deploy/deploy.sh
|
||||
env:
|
||||
SERVICE: ${{ matrix.service }}
|
||||
# Only the matching one is read by deploy.sh for a single-service roll;
|
||||
# the other stays empty and the persisted .image-tags.env supplies the
|
||||
# sibling's live tag, so this roll never disturbs it.
|
||||
BACKEND_IMAGE_TAG: ${{ matrix.service == 'backend' && steps.plan.outputs.tag || '' }}
|
||||
FRONTEND_IMAGE_TAG: ${{ matrix.service == 'frontend' && steps.plan.outputs.tag || '' }}
|
||||
|
||||
- name: Deploy to prod
|
||||
if: steps.plan.outputs.changed == 'true' && steps.plan.outputs.environment == 'prod'
|
||||
uses: https://github.com/appleboy/ssh-action@v1.2.0
|
||||
with:
|
||||
# A completely separate secret set from beta's, deliberately: a beta
|
||||
# credential leak must not reach prod.
|
||||
host: ${{ secrets.PROD_SSH_HOST }}
|
||||
username: ${{ secrets.PROD_SSH_USER }}
|
||||
key: ${{ secrets.PROD_SSH_KEY }}
|
||||
port: ${{ secrets.PROD_SSH_PORT }}
|
||||
envs: SERVICE,BACKEND_IMAGE_TAG,FRONTEND_IMAGE_TAG
|
||||
script: /opt/thermograph/infra/deploy/deploy.sh
|
||||
env:
|
||||
SERVICE: ${{ matrix.service }}
|
||||
BACKEND_IMAGE_TAG: ${{ matrix.service == 'backend' && steps.plan.outputs.tag || '' }}
|
||||
FRONTEND_IMAGE_TAG: ${{ matrix.service == 'frontend' && steps.plan.outputs.tag || '' }}
|
||||
|
||||
- name: Skipped
|
||||
if: steps.plan.outputs.changed != 'true'
|
||||
run: echo "${{ matrix.service }} unchanged in this push — nothing to roll."
|
||||
|
|
@ -1,126 +0,0 @@
|
|||
name: Build + push frontend image (Forgejo registry)
|
||||
|
||||
# Builds the FRONTEND domain's image (frontend/Dockerfile, context frontend/) and
|
||||
# pushes it to Forgejo's built-in container registry, tagged by git SHA (every
|
||||
# push) and semver (version tags) -- build-once, deploy-everywhere.
|
||||
#
|
||||
# Monorepo port (reunification): the split repos each derived IMAGE_PATH from
|
||||
# github.repository, which now collides for every domain -- so each domain's
|
||||
# build-push names its image explicitly: this one is emi/thermograph/frontend
|
||||
# (backend-build-push.yml publishes emi/thermograph/backend). infra's
|
||||
# compose/stack files reference the two paths independently
|
||||
# (FRONTEND_IMAGE_PATH / BACKEND_IMAGE_PATH -- their defaults were renamed to
|
||||
# match in the same cutover).
|
||||
#
|
||||
# Path-filtered: only a push that touches frontend/** builds this image. A
|
||||
# backend-only push builds nothing here; deploy.sh's persisted
|
||||
# .image-tags.env keeps the frontend's live tag for the sibling's deploy.
|
||||
# EXCEPTION: version-tag pushes (v*.*.*) carry no paths context, so a semver
|
||||
# tag builds BOTH domain images -- intended, a release names the whole set.
|
||||
#
|
||||
# Runs on the `docker` label -- the LAN runner's job containers get the host's
|
||||
# docker.sock automounted in (forgejo-runner config: container.docker_host:
|
||||
# automount; see infra/deploy/forgejo/register-lan-runner.sh),
|
||||
# Docker-outside-of-Docker rather than DinD -- no privileged mode needed.
|
||||
# The job image itself (node:20-bookworm) has no docker CLI, hence the
|
||||
# "Install Docker CLI" step.
|
||||
#
|
||||
# The registry is deliberately mesh-only: git.thermograph.org's public DNS
|
||||
# resolves to beta's public IP, which Caddy's ACL rejects for /v2/ paths, so
|
||||
# any runner host that pushes/pulls needs a /etc/hosts entry pointing
|
||||
# git.thermograph.org at beta's WireGuard IP (10.10.0.2). The docker
|
||||
# build/push commands run against the automounted HOST daemon, so it's the
|
||||
# runner HOST's own resolver that has to route over the mesh, not anything
|
||||
# settable from inside the job container.
|
||||
#
|
||||
# REGISTRY is the vars.REGISTRY_HOST Actions variable (git.thermograph.org),
|
||||
# NOT github.server_url -- that context resolves to the bare mesh IP:port
|
||||
# (http://10.10.0.2:3080, no TLS), which docker CLI refuses ("server gave HTTP
|
||||
# response to HTTPS client"). git.thermograph.org gets real TLS via Caddy and
|
||||
# routes over the mesh once /etc/hosts is set, so it's the only correct value.
|
||||
# deploy.sh resolves the SHA tag at deploy time and `docker compose pull`s it.
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [dev, main, release]
|
||||
paths: ['frontend/**']
|
||||
tags: ['v*.*.*']
|
||||
workflow_dispatch: {}
|
||||
|
||||
# The runner has capacity: 1 (one physical LAN box) -- a burst of pushes to the
|
||||
# same ref would otherwise queue whole builds back to back even though only the
|
||||
# last one still matters. Keying on domain + github.ref means each domain's
|
||||
# dev/main/release/tag builds get their own group, so this only cancels a
|
||||
# superseded build of the SAME domain on the SAME ref -- and a backend build
|
||||
# never cancels a frontend one.
|
||||
concurrency:
|
||||
group: frontend-build-push-${{ github.ref }}
|
||||
cancel-in-progress: true
|
||||
|
||||
jobs:
|
||||
build-push:
|
||||
runs-on: docker
|
||||
env:
|
||||
REGISTRY: https://${{ vars.REGISTRY_HOST }}
|
||||
IMAGE_PATH: ${{ github.repository }}/frontend
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
|
||||
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
|
||||
# often an unrelated domain's commit, and a depth-1 clone can't see
|
||||
# past it to find the real key.
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Install Docker CLI
|
||||
run: |
|
||||
apt-get update -qq
|
||||
apt-get install -y -qq docker.io
|
||||
|
||||
- name: Compute image ref + tags
|
||||
id: image
|
||||
run: |
|
||||
host="${REGISTRY#https://}"; host="${host#http://}"
|
||||
path="$(printf '%s' "$IMAGE_PATH" | tr '[:upper:]' '[:lower:]')"
|
||||
ref="$host/$path"
|
||||
# Key the tag to the last commit that touched frontend/ -- the SAME key
|
||||
# every frontend deploy workflow computes -- so build and deploy always
|
||||
# agree even when the branch tip is another domain's commit. (A tag
|
||||
# keyed to the push tip breaks the moment an infra-only commit lands:
|
||||
# deploys go looking for an image no build ever produced.)
|
||||
domain_sha="$(git log -1 --format=%H -- frontend/)"
|
||||
sha_tag="$ref:sha-${domain_sha:0:12}"
|
||||
echo "host=$host" >> "$GITHUB_OUTPUT"
|
||||
echo "sha_tag=$sha_tag" >> "$GITHUB_OUTPUT"
|
||||
if [[ "$GITHUB_REF" == refs/tags/v* ]]; then
|
||||
echo "semver_tag=$ref:${GITHUB_REF#refs/tags/}" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
|
||||
- name: Log in to the Forgejo registry
|
||||
run: |
|
||||
# NOT secrets.GITHUB_TOKEN -- Forgejo's per-job auto-injected token
|
||||
# can never push to the container registry (a known Forgejo
|
||||
# limitation independent of any permission setting). Needs a
|
||||
# manually provisioned PAT with write:package scope, stored as the
|
||||
# REGISTRY_TOKEN secret.
|
||||
echo "${{ secrets.REGISTRY_TOKEN }}" | docker login "${{ steps.image.outputs.host }}" \
|
||||
--username emi --password-stdin
|
||||
|
||||
- name: Build
|
||||
run: |
|
||||
# One build, both tags -- the semver tag (tag pushes only) rides
|
||||
# along with the SHA tag instead of a second tag+push round trip.
|
||||
# Context is the DOMAIN dir, so frontend/Dockerfile sees the same
|
||||
# tree it saw as a repo root in the split era.
|
||||
tag_args=(-t "${{ steps.image.outputs.sha_tag }}")
|
||||
if [[ -n "${{ steps.image.outputs.semver_tag }}" ]]; then
|
||||
tag_args+=(-t "${{ steps.image.outputs.semver_tag }}")
|
||||
fi
|
||||
docker build "${tag_args[@]}" frontend/
|
||||
|
||||
- name: Push (SHA tag, every push)
|
||||
run: docker push "${{ steps.image.outputs.sha_tag }}"
|
||||
|
||||
- name: Push (semver tag, tag pushes only)
|
||||
if: startsWith(github.ref, 'refs/tags/v')
|
||||
run: docker push "${{ steps.image.outputs.semver_tag }}"
|
||||
|
|
@ -1,73 +0,0 @@
|
|||
name: Deploy frontend to LAN dev server
|
||||
|
||||
# Fires on a push to `dev` touching frontend/** -- a direct push, or the real
|
||||
# merge commit a native "auto merge when checks succeed" produces (see
|
||||
# pr-build.yml). A `build` job (the reusable build.yml sanity check) gates a
|
||||
# `deploy` job that rolls the LAN dev box.
|
||||
#
|
||||
# CUTOVER NOTE (monorepo reunification): the LAN box's ~/thermograph-dev is a
|
||||
# thermograph-infra checkout from the split era. This job calls
|
||||
# ~/thermograph-dev/infra/deploy/deploy-dev.sh -- the MONOREPO layout -- so it
|
||||
# is INERT until ~/thermograph-dev is re-pointed at the monorepo (a fresh
|
||||
# clone; deploy-dev.sh then self-updates via deploy.sh's git reset like the
|
||||
# beta/prod paths).
|
||||
#
|
||||
# SERVICE=frontend + FRONTEND_IMAGE_TAG is the same env-var contract the
|
||||
# beta/prod deploy workflows use against infra/deploy/deploy.sh, so a dev-only
|
||||
# rollout never touches the backend's running container or tag.
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [dev]
|
||||
paths: ['frontend/**']
|
||||
workflow_dispatch: {}
|
||||
|
||||
jobs:
|
||||
build:
|
||||
name: build
|
||||
# Fixed (unparameterized) group: if two frontend pushes to dev land close
|
||||
# together, the newer push's build cancels the older's still-running one
|
||||
# on the single physical runner -- the older result would be thrown away
|
||||
# once needs:build gates its deploy anyway. Safe to cancel mid-run:
|
||||
# build.yml only runs a throwaway `docker build`, nothing persisted.
|
||||
concurrency:
|
||||
group: dev-push-build-frontend
|
||||
cancel-in-progress: true
|
||||
uses: ./.forgejo/workflows/build.yml
|
||||
with:
|
||||
domain: frontend
|
||||
|
||||
deploy:
|
||||
needs: build
|
||||
# NOT [self-hosted, thermograph-lan] -- GitHub implicitly tags every
|
||||
# self-hosted runner with "self-hosted"; Forgejo's runner has only the
|
||||
# labels it was registered with ("docker" + "thermograph-lan"). The array
|
||||
# form silently makes this job unschedulable on every push.
|
||||
runs-on: thermograph-lan
|
||||
permissions:
|
||||
contents: read
|
||||
concurrency:
|
||||
group: dev-lan-deploy-frontend
|
||||
cancel-in-progress: false
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
|
||||
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
|
||||
# often an unrelated domain's commit, and a depth-1 clone can't see
|
||||
# past it to find the real key.
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Compute image tag
|
||||
id: tag
|
||||
# Keyed to the last commit that touched frontend/ -- must match
|
||||
# frontend-build-push.yml's tag key exactly (see its comment).
|
||||
run: |
|
||||
domain_sha="$(git log -1 --format=%H -- frontend/)"
|
||||
echo "tag=sha-${domain_sha:0:12}" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Deploy to the LAN dev server
|
||||
# thermograph-lan is host-native (the runner job runs directly on the
|
||||
# LAN box, not over SSH like beta/prod), so plain env vars on the
|
||||
# command reach deploy-dev.sh directly -- no ssh-action needed here.
|
||||
run: SERVICE=frontend FRONTEND_IMAGE_TAG=${{ steps.tag.outputs.tag }} bash "$HOME/thermograph-dev/infra/deploy/deploy-dev.sh"
|
||||
|
|
@ -1,61 +0,0 @@
|
|||
name: Deploy frontend to prod VPS
|
||||
|
||||
# On a push to `release` that touches frontend/**, SSH to prod and roll ONLY
|
||||
# the frontend service onto the image frontend-build-push.yml published for this
|
||||
# commit. Same shape as frontend-deploy.yml (the `main`->beta path) against a
|
||||
# completely separate secret set (PROD_SSH_HOST/USER/KEY/PORT) so a beta
|
||||
# credential leak can't reach prod. Backend deploys itself independently via
|
||||
# backend-deploy-prod.yml -- a frontend change ships to prod without touching
|
||||
# backend and vice versa, exactly as in the split era.
|
||||
#
|
||||
# The tag MUST match frontend-build-push.yml's last-frontend-commit key (12
|
||||
# hex) -- see frontend-deploy.yml for why it's truncated here. deploy.sh
|
||||
# retries the pull because the build for this push may still be in flight.
|
||||
|
||||
#
|
||||
# Frontend's own register() fetches the IndexNow key from backend at boot, so
|
||||
# deploy.sh waits on a healthy backend before declaring the frontend roll OK
|
||||
# (its compose depends_on already encodes that ordering).
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [release]
|
||||
paths: ['frontend/**']
|
||||
workflow_dispatch: {}
|
||||
|
||||
concurrency:
|
||||
group: prod-deploy-frontend
|
||||
cancel-in-progress: false
|
||||
|
||||
jobs:
|
||||
deploy:
|
||||
runs-on: docker
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
|
||||
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
|
||||
# often an unrelated domain's commit, and a depth-1 clone can't see
|
||||
# past it to find the real key.
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Compute image tag (last frontend-touching commit)
|
||||
id: tag
|
||||
run: |
|
||||
domain_sha="$(git log -1 --format=%H -- frontend/)"
|
||||
echo "tag=sha-${domain_sha:0:12}" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Deploy frontend over SSH
|
||||
uses: https://github.com/appleboy/ssh-action@v1.2.0
|
||||
with:
|
||||
host: ${{ secrets.PROD_SSH_HOST }}
|
||||
username: ${{ secrets.PROD_SSH_USER }}
|
||||
key: ${{ secrets.PROD_SSH_KEY }}
|
||||
port: ${{ secrets.PROD_SSH_PORT }}
|
||||
# Forward the target service + the image tag to run. deploy.sh reads
|
||||
# FRONTEND_IMAGE_TAG and rolls just `frontend`.
|
||||
envs: SERVICE,FRONTEND_IMAGE_TAG
|
||||
script: /opt/thermograph/infra/deploy/deploy.sh
|
||||
env:
|
||||
SERVICE: frontend
|
||||
FRONTEND_IMAGE_TAG: ${{ steps.tag.outputs.tag }}
|
||||
|
|
@ -1,70 +0,0 @@
|
|||
name: Deploy frontend to beta VPS
|
||||
|
||||
# On a push to `main` that touches frontend/**, SSH to beta and roll ONLY the
|
||||
# frontend service onto the image frontend-build-push.yml just published for
|
||||
# this commit. Backend is deployed independently by backend-deploy.yml --
|
||||
# the FE/BE CI-CD split survives reunification intact: a frontend change ships
|
||||
# without touching backend and vice versa. infra/deploy/deploy.sh persists
|
||||
# each service's live tag host-side, so a single-service roll never disturbs
|
||||
# the other. An infra-only push rolls nothing (infra-sync.yml syncs the
|
||||
# host checkouts instead); a mixed frontend+backend push fires both deploy
|
||||
# workflows, serialized host-side by deploy.sh's flock.
|
||||
#
|
||||
# The SSH_HOST/USER/KEY/PORT secrets point at beta. appleboy/ssh-action is
|
||||
# referenced by full GitHub URL because it isn't mirrored in Forgejo's default
|
||||
# action registry.
|
||||
#
|
||||
# The tag MUST match frontend-build-push.yml's last-frontend-commit key (12
|
||||
# hex), not the full 40-char ${{ github.sha }} -- default Actions expressions
|
||||
# have no substring function, so the compute step below truncates it.
|
||||
# deploy.sh also retries the pull for ~5 min because Forgejo `needs:` can't
|
||||
# gate across the separate build-push workflow, so this push's build may still
|
||||
# be in flight.
|
||||
|
||||
#
|
||||
# Frontend's own register() fetches the IndexNow key from backend at boot, so
|
||||
# deploy.sh waits on a healthy backend before declaring the frontend roll OK
|
||||
# (its compose depends_on already encodes that ordering).
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
paths: ['frontend/**']
|
||||
workflow_dispatch: {}
|
||||
|
||||
concurrency:
|
||||
group: beta-deploy-frontend
|
||||
cancel-in-progress: false
|
||||
|
||||
jobs:
|
||||
deploy:
|
||||
runs-on: docker
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
|
||||
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
|
||||
# often an unrelated domain's commit, and a depth-1 clone can't see
|
||||
# past it to find the real key.
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Compute image tag (last frontend-touching commit)
|
||||
id: tag
|
||||
run: |
|
||||
domain_sha="$(git log -1 --format=%H -- frontend/)"
|
||||
echo "tag=sha-${domain_sha:0:12}" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Deploy frontend over SSH
|
||||
uses: https://github.com/appleboy/ssh-action@v1.2.0
|
||||
with:
|
||||
host: ${{ secrets.SSH_HOST }}
|
||||
username: ${{ secrets.SSH_USER }}
|
||||
key: ${{ secrets.SSH_KEY }}
|
||||
port: ${{ secrets.SSH_PORT }}
|
||||
# Forward the target service + the image tag to run. deploy.sh reads
|
||||
# FRONTEND_IMAGE_TAG and rolls just `frontend`.
|
||||
envs: SERVICE,FRONTEND_IMAGE_TAG
|
||||
script: /opt/thermograph/infra/deploy/deploy.sh
|
||||
env:
|
||||
SERVICE: frontend
|
||||
FRONTEND_IMAGE_TAG: ${{ steps.tag.outputs.tag }}
|
||||
105
CLAUDE.md
105
CLAUDE.md
|
|
@ -1,27 +1,88 @@
|
|||
# thermograph monorepo — agent instructions
|
||||
|
||||
Reunified monorepo (2026-07-22): `backend/`, `frontend/`, `infra/`,
|
||||
`observability/` — each a former split repo subtree-merged with history.
|
||||
`thermograph-docs` remains a separate repo; cross-cutting decision docs and
|
||||
operator runbooks belong there, not here.
|
||||
One repo, four domains: `backend/`, `frontend/`, `infra/`, `observability/`.
|
||||
Reunified 2026-07-22 from four split repos, with history, via subtree merges.
|
||||
`thermograph-docs` is deliberately a separate repo — cross-cutting decision
|
||||
records and operator runbooks go there, not here.
|
||||
|
||||
- **Each domain keeps its own `CLAUDE.md`** — read the one for the domain you
|
||||
touch. They still use repo-era wording ("this repo", sibling-repo names);
|
||||
mentally map `thermograph-backend` → `backend/` etc., and fix wording as you
|
||||
touch files near it.
|
||||
- **CI lives only in root `.forgejo/workflows/`**, path-filtered per domain
|
||||
(see README). Never re-add workflows under a domain's own `.forgejo/` — they
|
||||
are inert there and become a trap.
|
||||
- Images are `emi/thermograph/backend` and `emi/thermograph/frontend`
|
||||
(`sha-<12hex>`). Deploy = `infra/deploy/deploy.sh` over SSH, per-service.
|
||||
**Rule for this file and every domain `CLAUDE.md`: only statements that would
|
||||
break CI if they became false, or that name a file that exists.** Background,
|
||||
history and rationale belong in `thermograph-docs`. These files are read before
|
||||
every change, so a stale one is a correctness bug, not a documentation bug.
|
||||
|
||||
## Branches and environments
|
||||
|
||||
| Branch | Deploys to | Workflow |
|
||||
|---|---|---|
|
||||
| feature branch | nothing | PR into `dev` |
|
||||
| `dev` | nothing | integration branch only — see below |
|
||||
| `main` | beta (beta.thermograph.org) | `deploy.yml` |
|
||||
| `release` | prod (thermograph.org) | `deploy.yml` |
|
||||
|
||||
`dev`, `main` and `release` are protected: **everything is a PR**, for humans and
|
||||
agents alike. Promotion is one PR per hop, `dev` → `main` → `release`.
|
||||
|
||||
**`dev` deploys nowhere.** It is purely the integration branch that feature PRs
|
||||
land on before promotion to `main`. The two LAN-dev deploy workflows were deleted
|
||||
rather than kept: they called a monorepo path that does not exist on the LAN box
|
||||
(`~/thermograph-dev` is still a split-era `thermograph-infra` checkout), so they
|
||||
had been inert since cutover. Run LAN dev locally with `infra/`'s `make dev-up`.
|
||||
|
||||
**One workflow deploys everything.** `deploy.yml` handles both services and both
|
||||
environments: the branch selects the environment (`main` → beta, `release` →
|
||||
prod), a matrix covers backend and frontend, and each leg checks whether this
|
||||
push actually touched its domain before rolling. `build-push.yml` is the same
|
||||
shape for images. They replaced six and two near-identical files respectively.
|
||||
|
||||
## The deploy contract
|
||||
|
||||
One entry point, two modes, one contract:
|
||||
|
||||
```
|
||||
SERVICE=backend|frontend|all BACKEND_IMAGE_TAG=sha-<12hex> FRONTEND_IMAGE_TAG=sha-<12hex> \
|
||||
/opt/thermograph/infra/deploy/deploy.sh
|
||||
```
|
||||
|
||||
`deploy.sh` resets the host checkout, renders secrets from the SOPS vault, then
|
||||
either rolls compose services or — if `/etc/thermograph/deploy-mode` contains
|
||||
`stack` — execs `infra/deploy/stack/deploy-stack.sh`.
|
||||
|
||||
- **prod runs Swarm.** Its stack is `infra/deploy/stack/thermograph-stack.yml`
|
||||
(db, web, worker, lake, daemon, frontend, autoscaler, autoscaler-lake).
|
||||
`deploy-stack.sh` also offers `STACK_TEST=1`: a full parallel rehearsal on
|
||||
throwaway volumes and ports that cannot touch live data.
|
||||
- **beta and LAN dev run compose**, from `infra/docker-compose.yml`
|
||||
(db, backend, lake, daemon, frontend).
|
||||
- Each service's live tag is persisted host-side, so a single-service roll never
|
||||
disturbs the sibling's running tag.
|
||||
|
||||
Images are `emi/thermograph/backend` and `emi/thermograph/frontend`, tagged
|
||||
`sha-<12hex>` on every push and by semver on `v*.*.*` tags. They build and deploy
|
||||
**independently** — a backend change ships without a frontend deploy and vice
|
||||
versa. That independence is the point of the FE/BE split and survived
|
||||
reunification.
|
||||
|
||||
## Rules that bind across domains
|
||||
|
||||
- **CI lives only in root `.forgejo/workflows/`**, path-filtered per domain.
|
||||
Never add workflows under a domain's own `.forgejo/` — they are inert there and
|
||||
become a trap.
|
||||
- **Secrets only via the SOPS vault** (`infra/deploy/secrets/`). Never hand-edit
|
||||
`/etc/thermograph.env` on a host — it is a rendered artifact. `secrets-guard`
|
||||
CI rejects plaintext. `seed-from-live.sh` reads production secrets and is
|
||||
explicitly **not** for an agent to run.
|
||||
- FE and BE ship out of lockstep, so the `/api/version` contract and
|
||||
`PAYLOAD_VER` discipline are load-bearing — see `backend/CLAUDE.md`.
|
||||
- Anything named "prefetch" must never spend the Open-Meteo quota.
|
||||
Nominatim ≤ 1 req/s.
|
||||
- The compose project name is pinned (`name: thermograph` in
|
||||
`infra/docker-compose.yml`); LAN dev overrides via
|
||||
`COMPOSE_PROJECT_NAME=thermograph-dev`. Don't remove either half.
|
||||
- Cross-domain rules that still bind: FE/BE ship independently → keep the
|
||||
`/api/version` contract and `PAYLOAD_VER` discipline (see
|
||||
`backend/CLAUDE.md`); anything named "prefetch" must never spend the
|
||||
Open-Meteo quota; Nominatim ≤ 1 req/s; secrets only via the SOPS vault
|
||||
(`infra/deploy/secrets/`).
|
||||
`infra/docker-compose.yml`); LAN dev overrides with
|
||||
`COMPOSE_PROJECT_NAME=thermograph-dev`. Don't remove either half — the pinned
|
||||
name is what makes the Swarm stack's external volume names line up.
|
||||
- `CUTOVER-NOTES.md` is the source of truth for what is and isn't live yet.
|
||||
- Commits/PRs: concise and technical; never mention AI/assistants/automated
|
||||
authorship.
|
||||
|
||||
## Commits & PRs
|
||||
|
||||
Describe only the substance of the change; concise and technical. Never mention
|
||||
AI, Claude, assistants or automated authorship anywhere — no trailers,
|
||||
co-authors, emoji, or "as requested" narration.
|
||||
|
|
|
|||
|
|
@ -13,6 +13,16 @@ subtree-merging each split repo's `origin/main` (full history, 380 commits):
|
|||
**CI (all in root `.forgejo/workflows/`; the subtree copies were deleted —
|
||||
Forgejo only reads root workflows, but dead copies are a trap):**
|
||||
|
||||
> **Superseded 2026-07-25.** The per-domain workflow names below describe the
|
||||
> cutover as it landed, and are kept for that history. They no longer exist:
|
||||
> the six `*-deploy*.yml` files were collapsed into a single `deploy.yml`
|
||||
> (branch selects environment, matrix covers the services) and the two
|
||||
> `*-build-push.yml` into a single `build-push.yml`. The two LAN-dev deploy
|
||||
> workflows were deleted outright rather than ported, having been inert since
|
||||
> cutover. Everything described below about *behaviour* — path filtering, the
|
||||
> domain-keyed 12-hex tag, per-domain concurrency, the `v*.*.*` both-images
|
||||
> exception, separate `SSH_*`/`PROD_SSH_*` secret sets — is unchanged.
|
||||
|
||||
- `backend-build-push.yml` / `frontend-build-push.yml` — the old per-repo
|
||||
build-push, path-filtered (`paths: ['backend/**']` etc.), image paths now
|
||||
**explicit**: `emi/thermograph/backend|frontend` (the old
|
||||
|
|
|
|||
|
|
@ -8,9 +8,9 @@ for: **per-domain images, per-domain deploys, and an async FE/BE contract**.
|
|||
|
||||
| Dir | What | CI |
|
||||
|---|---|---|
|
||||
| `backend/` | FastAPI graded-climate API, accounts, notifications (Discord bot, push, mail), data pipeline | `backend-build-push` → image `emi/thermograph/backend`; `backend-deploy[-prod\|-dev]` |
|
||||
| `frontend/` | Public client: static JS/CSS + SSR pages | `frontend-*` mirrors of the above; image `emi/thermograph/frontend` |
|
||||
| `infra/` | Compose, deploy scripts, terraform, SOPS secrets vault, ops cron | `infra-sync` (host checkout + secrets render), `secrets-guard`, `ops-cron` |
|
||||
| `backend/` | FastAPI graded-climate API, accounts, notifications (Discord bot, push, mail), data pipeline | `build-push` → image `emi/thermograph/backend`; `deploy` |
|
||||
| `frontend/` | Public client: static JS/CSS + SSR pages | same `build-push` / `deploy` workflows, matrixed by domain; image `emi/thermograph/frontend` |
|
||||
| `infra/` | Compose (beta, LAN dev) + the Swarm stack (prod), deploy scripts, terraform, SOPS secrets vault, ops cron | `infra-sync` (host checkout + secrets render), `secrets-guard`, `ops-cron` |
|
||||
| `observability/` | Loki + Grafana + Alloy stack | `observability-validate` |
|
||||
|
||||
`thermograph-docs` deliberately **stays its own repo** (ADRs + runbooks, no
|
||||
|
|
|
|||
|
|
@ -1,161 +1,66 @@
|
|||
# thermograph-backend — agent instructions
|
||||
# backend/ — agent instructions
|
||||
|
||||
## What this repo is
|
||||
The Thermograph **API service**: a FastAPI app serving the graded-climate API
|
||||
(`/api/v2/...`), user accounts, notifications, and the SSR-content JSON API the
|
||||
frontend renders pages from. It owns the climate data pipeline and the
|
||||
derived-payload cache. It serves **no HTML/CSS/JS** — that is `frontend/`.
|
||||
|
||||
The Thermograph **API service**: FastAPI app serving the graded-climate API
|
||||
(`/api/v2/...`), user accounts (`accounts/`, fastapi-users + Postgres via
|
||||
alembic), push/email/Discord notifications (`notifications/`), and the
|
||||
SSR-content JSON API (`api/content_routes.py` + `api/content_payloads.py`) that
|
||||
the separate `thermograph-frontend`'s server-rendered pages consume. It owns
|
||||
the climate data pipeline (`data/` — grid snapping, Open-Meteo fetch + parquet
|
||||
cache, percentile grading/scoring, places/cities) and the derived-payload
|
||||
cache/ETag store (`data/store.py`, `api/payloads.py`). It does **not** serve
|
||||
any HTML/CSS/JS itself — that moved to `thermograph-frontend` (repo-split
|
||||
Stage 7a); this process is a pure JSON API + DB-backed service, launched as
|
||||
`app:app` (a shim re-exporting `web.app:app`, kept stable for systemd/CI/`run.sh`
|
||||
callers that still expect that target).
|
||||
Read the root `CLAUDE.md` first for the branch model, deploy contract and image
|
||||
names.
|
||||
|
||||
## Split topology
|
||||
## Build & verify
|
||||
|
||||
Split out of the `emi/thermograph` monorepo (see
|
||||
`MONOREPO-CUTOVER-PLAN.md` at the workspace root during the transition). Sibling
|
||||
repos:
|
||||
- **`thermograph-frontend`** — the client the public sees: static
|
||||
JS/CSS/HTML + `frontend_ssr`'s server-rendered climate hub/city/month pages,
|
||||
calling this repo's `/api/v2/...` and `/content/...` JSON endpoints.
|
||||
- **`thermograph-infra`** — deploy/terraform/Swarm-Compose plumbing:
|
||||
`docker-compose.yml`, `deploy/deploy.sh`, secrets (SOPS-encrypted
|
||||
`deploy/secrets/*.yaml`), Terraform for prod/beta provisioning. This repo has
|
||||
no deploy scripts or terraform of its own — it only builds an image and hands
|
||||
infra a tag.
|
||||
- **`thermograph-docs`** — non-code planning layer: architecture decision
|
||||
records and operator runbooks. Nothing here should duplicate that; if you're
|
||||
writing a cross-cutting decision doc or a step-by-step operator runbook,
|
||||
it belongs there, not in this repo's README/CLAUDE.md.
|
||||
- **`thermograph-observability`** — monitoring/metrics stack, separate from
|
||||
this repo's own lightweight in-process `core/metrics.py`.
|
||||
- `make test` — builds a py3.12 venv and runs the hermetic pytest suite
|
||||
(`scripts/test.sh`; pass `ARGS=...`). Tests are hermetic per `tests/conftest.py`:
|
||||
no real Open-Meteo/Nominatim/GeoNames calls, throwaway SQLite, notifier thread
|
||||
disabled.
|
||||
- `make smoke` — boots the built image plus a throwaway db and asserts
|
||||
`/healthz` + `/api/version`.
|
||||
- CI runs this suite **inside the built image** (`build.yml`), so the exact
|
||||
interpreter and deps that ship are what gets tested.
|
||||
|
||||
## Branch, PR & deploy flow
|
||||
`app.py` at the root is a one-line re-export shim for `web/app.py`, kept stable
|
||||
for systemd/CI callers. Entry target is `app:app`.
|
||||
|
||||
- `build-push.yml` builds **this repo's own image**
|
||||
(`git.thermograph.org/emi/thermograph-backend/app`, tagged `sha-<12hex>` on
|
||||
every push to `dev`/`main`/`release` and additionally by semver on `v*.*.*`
|
||||
tags) — no more shared monorepo image. `build.yml` (push/PR to `main`) is a
|
||||
build-only sanity check (`docker build`), not a boot/health check — repeated
|
||||
attempts at a live boot+healthz check hit this runner's own quirks (bridge-IP
|
||||
reachability, non-unique `$$`, squatted host ports) without ever going
|
||||
reliably green; the image's actual boot is verified by hand (`docker run` +
|
||||
alembic + `/healthz` 200 + a real `/api/v2/place` 200) instead.
|
||||
- **`main` → beta**: `deploy.yml` SSHes to beta and runs
|
||||
`thermograph-infra`'s `/opt/thermograph/deploy/deploy.sh` with
|
||||
`SERVICE=backend` + `BACKEND_IMAGE_TAG=sha-<12hex>` (must match
|
||||
build-push.yml's 12-hex SHA tag exactly — Actions expressions have no
|
||||
substring function for the full 40-char SHA). `deploy.sh` rolls only the
|
||||
`backend` compose service (`--no-deps`) and persists the tag in
|
||||
`deploy/.image-tags.env` host-side, so a backend-only deploy never touches
|
||||
the frontend's running container or tag.
|
||||
- **`release` → prod**: `deploy-prod.yml` is the same shape against a
|
||||
completely separate secret set (`PROD_SSH_HOST/USER/KEY/PORT`) so a beta
|
||||
credential leak can't reach prod. Also `SERVICE=backend` +
|
||||
`BACKEND_IMAGE_TAG`.
|
||||
- The frontend deploys itself the same way, independently — that
|
||||
independence (a backend change ships without a frontend deploy and vice
|
||||
versa) is the entire point of the FE/BE CI-CD split.
|
||||
- No `dev`-branch LAN deploy path exists in this repo (the monorepo's
|
||||
`deploy-dev.yml`/`docker-compose.dev.yml` never made it across the split —
|
||||
known gap, see the cutover plan if resurrecting it).
|
||||
## Contracts that break silently if you change them
|
||||
|
||||
## Run / test locally
|
||||
|
||||
There is **no Makefile in this repo** — the monorepo's app-level `make`
|
||||
targets (`lan-run`, `test`, `venv`, ...) haven't been given a new home yet
|
||||
(tracked as a gap in the cutover plan). Until that lands, drive it directly:
|
||||
|
||||
```bash
|
||||
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt -r requirements-dev.txt
|
||||
.venv/bin/uvicorn app:app --host 0.0.0.0 --port 8137 # serve
|
||||
.venv/bin/python -m pytest tests # test
|
||||
```
|
||||
|
||||
The monorepo's Makefile preferred `uv venv --python 3.12 .venv` (falling back
|
||||
to plain `python3 -m venv` only if `uv` isn't installed) — prefer `uv` here too
|
||||
if it's on the box. There is no committed `.venv` in this repo; if one exists
|
||||
locally and looks broken (import errors, wrong Python), just `rm -rf .venv` and
|
||||
recreate rather than debugging it — it's gitignored, disposable state, not
|
||||
something the repo depends on being a specific way.
|
||||
|
||||
Tests are hermetic (`tests/conftest.py`): no real Open-Meteo/Nominatim/GeoNames
|
||||
calls, a throwaway SQLite for accounts + the derived-payload store, and the
|
||||
notifier thread disabled. `THERMOGRAPH_FRONTEND_BASE_INTERNAL` is set to an
|
||||
unreachable placeholder in conftest — `web/app.py` fails loud at import if it's
|
||||
unset in a real run (it's the internal URL back to `frontend_ssr` for the
|
||||
non-API HTML routes this process no longer serves itself).
|
||||
|
||||
## API version contract
|
||||
|
||||
- Everything real lives under **`/api/v2/...`**; `/api/` and `/api/v1/` stay
|
||||
mounted as aliases of the same handlers (they always were — not new).
|
||||
- **`GET /api/version`** (`web/app.py`) is the capability/negotiation endpoint
|
||||
a split-out frontend can call at boot: `{backend_version, min_frontend,
|
||||
payload_ver}`, driven by two module constants next to it —
|
||||
`API_CONTRACT_VERSION` (bump only on a breaking change to any `/api/v{N}`
|
||||
route or payload shape) and `MIN_SUPPORTED_FRONTEND` (oldest frontend
|
||||
contract version this backend still serves correctly). A genuine breaking
|
||||
change ships as a new `v3` router mounted alongside `v2` re-registering only
|
||||
the changed handlers, `v2` kept working, `API_CONTRACT_VERSION` bumped in the
|
||||
same PR.
|
||||
- **`PAYLOAD_VER`** (`api/payloads.py`, currently `"p2"`) is a *separate*
|
||||
concern from URL versioning — it's the cache/ETag invalidation token baked
|
||||
into every derived-store validity token (`history_token()`). Bump it
|
||||
whenever a response payload's shape changes (new metric, renamed/removed
|
||||
field) so one bump atomically orphans every pre-upgrade cached row instead
|
||||
of a stale row's `history_end` happening to still validate under the old
|
||||
shape.
|
||||
- Both endpoints (`/healthz`, `/api/version`) are deliberately I/O-free so they
|
||||
stay cheap under tight healthcheck intervals.
|
||||
|
||||
## Live cross-repo contracts (don't break silently)
|
||||
|
||||
- **`/cell` bundle + ETag/If-None-Match**: every graded payload (`grade`,
|
||||
`calendar`, `day`) is cached in SQLite by `(kind, cell_id, key)` keyed to a
|
||||
validity token that only advances when the underlying history actually
|
||||
changes (`_etag_for` in `web/app.py`, `history_token`/`PAYLOAD_VER` in
|
||||
`api/payloads.py`). The same token doubles as a weak ETag; a client sending
|
||||
`If-None-Match` gets an empty 304 without the payload being rebuilt or even
|
||||
loaded. `expose_headers=["ETag"]` in the CORS middleware matters as much as
|
||||
`allow_origins` — without it the frontend's `cache.js` reading
|
||||
`res.headers.get("ETag")` cross-origin silently gets `null` and never
|
||||
revalidates correctly. Don't change the ETag derivation or the CORS
|
||||
`expose_headers` list without checking the frontend's cache layer.
|
||||
- **Percentile/unit parity with `thermograph-frontend`'s `static/shared.js`**:
|
||||
the frontend's percentile-to-ordinal display logic explicitly mirrors this
|
||||
repo's `data/grading.py::pct_ordinal()` (floors an empirical percentile into
|
||||
1..99 — a rank against the sample can round to 0 or 100, and "100th
|
||||
percentile" is never shown). `TEMP_BANDS`/`RAIN_BANDS` tier boundaries in
|
||||
`data/grading.py` are the source of truth for the tier names/thresholds
|
||||
(Near Record / High / Above Normal / Normal / Below Normal / Low, and the
|
||||
five rain-intensity tiers) — changing them without a corresponding frontend
|
||||
update produces a client that draws differently-colored tiers than the
|
||||
labels the API returns.
|
||||
- **SSR content API** (`api/content_routes.py`): a separate ETag'd JSON surface
|
||||
(`/content/hub`, `/content/city/{slug}`, `/content/city/{slug}/month/{month}`,
|
||||
`/content/city/{slug}/records`, `/content/home`, `/content/sitemap`,
|
||||
`/content/indexnow-key`) that `frontend_ssr` calls in-process-free (no shared
|
||||
Python import) to render pages. Same caching pattern as the main API; keep
|
||||
payload shape changes coordinated with whatever consumes it on the frontend
|
||||
side.
|
||||
- **`GET /api/version`** → `{backend_version, min_frontend, payload_ver}`, driven
|
||||
by `API_CONTRACT_VERSION` and `MIN_SUPPORTED_FRONTEND` in `web/app.py`. A
|
||||
breaking change ships as a new `v3` router mounted **alongside** `v2`,
|
||||
re-registering only changed handlers, with the constant bumped in the same PR.
|
||||
`/api` and `/api/v1` stay mounted as aliases.
|
||||
- **`PAYLOAD_VER`** (`api/payloads.py`) is separate from URL versioning — it is
|
||||
the cache/ETag invalidation token. Bump it whenever a response payload's shape
|
||||
changes, so one bump atomically orphans every pre-upgrade cached row.
|
||||
- **ETag / `If-None-Match`** — every graded payload is cached against a validity
|
||||
token that doubles as a weak ETag. `expose_headers=["ETag"]` in the CORS
|
||||
middleware matters as much as `allow_origins`: without it the frontend's
|
||||
`cache.js` reads `null` cross-origin and never revalidates. Don't change the
|
||||
ETag derivation or that list without checking `frontend/static/cache.js`.
|
||||
- **`data/grading.py::pct_ordinal()`** is mirrored by `frontend`'s
|
||||
`shared.js::pctOrd()` — floor a percentile into `1..99`, never 0 or 100.
|
||||
`TEMP_BANDS`/`RAIN_BANDS` are the source of truth for tier names and
|
||||
thresholds; changing them without the frontend produces tiers drawn in colours
|
||||
that disagree with the labels the API returns.
|
||||
- **Fahrenheit country set** — `api/content_payloads.py`'s `F_COUNTRIES` must stay
|
||||
identical to `frontend`'s `format.py::F_COUNTRIES` and `static/units.js`'s
|
||||
`F_REGIONS`. There is a test asserting this.
|
||||
- **`/healthz` and `/api/version` are deliberately I/O-free** so they stay cheap
|
||||
under tight healthcheck intervals.
|
||||
|
||||
## Layout
|
||||
|
||||
`accounts/` (fastapi-users models/schemas/db, `alembic/` migrations),
|
||||
`api/` (versioned payload builders + SSR content routes), `core/` (metrics,
|
||||
audit logging, a singleton helper), `data/` (grid, climate fetch/cache, grading/
|
||||
scoring, places/cities, the derived-payload store), `notifications/` (push,
|
||||
email, Discord bot + interactions + linking, digest, scheduler), `web/app.py`
|
||||
(the actual FastAPI app; `app.py` at the repo root is a one-line re-export
|
||||
shim). `paths.py` resolves every filesystem location (climate parquet cache,
|
||||
accounts DB, logs, bundled `cities.json`/`cities_flavor.json`) from the repo
|
||||
root — don't reintroduce `__file__`-relative paths in a module, it breaks the
|
||||
moment a module moves. `deploy/entrypoint.sh` is the container entrypoint
|
||||
(alembic migrate-then-serve, with a Swarm-secrets-as-files shim); actual deploy
|
||||
orchestration lives in `thermograph-infra`, not here.
|
||||
`accounts/` (fastapi-users + alembic), `api/` (payload builders + SSR content
|
||||
routes), `core/` (metrics, audit, singleton helper), `data/` (grid, climate
|
||||
fetch/cache, grading/scoring, places/cities, derived-payload store),
|
||||
`notifications/` (push, email, Discord bot, digest, scheduler), `web/app.py`
|
||||
(the real app), `daemon/`.
|
||||
|
||||
`paths.py` resolves every filesystem location from the repo root — don't
|
||||
reintroduce `__file__`-relative paths in a module; it breaks the moment a module
|
||||
moves. `deploy/entrypoint.sh` is the container entrypoint (alembic migrate, then
|
||||
serve, with a Swarm-secrets-as-files shim).
|
||||
|
||||
## Commits & PRs
|
||||
|
||||
Concise and technical. Never mention AI, assistants or automated authorship.
|
||||
|
|
|
|||
|
|
@ -936,29 +936,27 @@ def _fetch_recent_forecast(cell: dict) -> pl.DataFrame:
|
|||
.sort("date"))
|
||||
|
||||
|
||||
def _om_local_today(payload: dict) -> "datetime.date | None":
|
||||
"""The calendar date it is *now* at the cell, from the UTC offset Open-Meteo
|
||||
reports for a ``timezone=auto`` request. None when the offset is absent, which
|
||||
tells the caller to skip the in-progress-day guard rather than guess a date."""
|
||||
offset = payload.get("utc_offset_seconds")
|
||||
if offset is None:
|
||||
return None
|
||||
return (datetime.datetime.now(datetime.timezone.utc)
|
||||
+ datetime.timedelta(seconds=int(offset))).date()
|
||||
|
||||
|
||||
def _fetch_recent_forecast_om(cell: dict) -> pl.DataFrame:
|
||||
"""Fallback recent+forecast bundle from the Open-Meteo forecast API (the former
|
||||
primary): recent past + forward days in one call.
|
||||
"""Primary recent+forecast bundle from the Open-Meteo forecast API: recent past
|
||||
and forward days in one call, covering today + 7 (see FORECAST_DAYS).
|
||||
|
||||
The cell's in-progress *local* day is dropped from the bundle. Open-Meteo's
|
||||
daily high/low for today aggregates only the hours elapsed so far, so grading it
|
||||
reads a still-unfolding day as complete and produces spurious extremes (a cool
|
||||
morning served as a record-low high, seen live at the 1st percentile). This is
|
||||
the same failure the MET Norway path guards against with its diurnal-coverage
|
||||
gate (see _metno_to_frame); the fallback needs the equivalent. Past days are
|
||||
complete and future days are whole-day forecasts, so only today is excluded —
|
||||
the day lands in the record once it is over."""
|
||||
The cell's in-progress *local* day is KEPT. An earlier revision dropped it on
|
||||
the assumption that Open-Meteo's daily high/low for today aggregates only the
|
||||
hours elapsed so far — true of the MET Norway series, which genuinely starts
|
||||
mid-day and is gated on diurnal coverage (see _metno_to_frame), but NOT of this
|
||||
endpoint: Open-Meteo backfills the rest of today from the model run, so today's
|
||||
daily value spans the whole local day exactly as tomorrow's does.
|
||||
|
||||
Verified against the live API at four cells spanning 01:00-19:00 local: the
|
||||
reported daily max equals the max over all 24 hourly values, and diverges from
|
||||
the elapsed-hours-only max wherever the day still has hours left (Los Angeles
|
||||
at 09:16 local reported 93.9F, the whole-day figure, where the hours elapsed
|
||||
topped out at 76.1F). Re-check it that way before reintroducing any exclusion.
|
||||
|
||||
Dropping today cost every freshly-synced cell its current day — a hole between
|
||||
a complete yesterday and a forecast tomorrow, which the UI renders as a missing
|
||||
column. A day that genuinely cannot be graded, because it came back with no
|
||||
high/low, is already dropped upstream by _to_frame."""
|
||||
params = {
|
||||
"latitude": cell["center_lat"],
|
||||
"longitude": cell["center_lon"],
|
||||
|
|
@ -971,12 +969,7 @@ def _fetch_recent_forecast_om(cell: dict) -> pl.DataFrame:
|
|||
"forecast_days": FORECAST_DAYS,
|
||||
}
|
||||
r = _request(FORECAST_URL, params, 60, phase="recent_forecast_fetch")
|
||||
payload = r.json()
|
||||
df = _to_frame(payload["daily"])
|
||||
local_today = _om_local_today(payload)
|
||||
if local_today is not None:
|
||||
df = df.filter(pl.col("date") != local_today)
|
||||
return df
|
||||
return _to_frame(r.json()["daily"])
|
||||
|
||||
|
||||
def _load_recent_forecast(cell: dict) -> pl.DataFrame:
|
||||
|
|
|
|||
|
|
@ -33,6 +33,7 @@ services:
|
|||
THERMOGRAPH_FRONTEND_BASE_INTERNAL: http://127.0.0.1:1
|
||||
THERMOGRAPH_BASE: "/"
|
||||
THERMOGRAPH_ENABLE_NOTIFIER: "0"
|
||||
THERMOGRAPH_ENABLE_HEARTBEAT: "0"
|
||||
PORT: "8137"
|
||||
WORKERS: "1"
|
||||
ports:
|
||||
|
|
|
|||
|
|
@ -28,6 +28,7 @@ _TMP = tempfile.mkdtemp(prefix="thermograph-tests-")
|
|||
# start the background notifier thread during tests. Set before db is imported.
|
||||
os.environ.setdefault("THERMOGRAPH_ACCOUNTS_DB", os.path.join(_TMP, "accounts.sqlite"))
|
||||
os.environ.setdefault("THERMOGRAPH_ENABLE_NOTIFIER", "0")
|
||||
os.environ.setdefault("THERMOGRAPH_ENABLE_HEARTBEAT", "0")
|
||||
# web/app.py (repo-split Stage 4) fails loud at import if this is unset -- tests
|
||||
# never actually hit a frontend-owned path through the live proxy (that's
|
||||
# frontend_ssr/tests' job), so an unreachable placeholder is fine here.
|
||||
|
|
|
|||
|
|
@ -215,20 +215,20 @@ def test_recent_forecast_fallback_merges_archive_past_and_metno_forward(monkeypa
|
|||
datetime.date(2026, 7, 16), datetime.date(2026, 7, 17)]
|
||||
|
||||
|
||||
def test_recent_forecast_om_drops_the_in_progress_local_day(monkeypatch):
|
||||
"""The Open-Meteo fallback must not grade the cell's in-progress local day: its
|
||||
daily high/low is only a partial aggregate of the hours elapsed so far (a cool
|
||||
morning would read as a record-low high). Past days and future forecast days
|
||||
survive; today — per the UTC offset Open-Meteo reports for a timezone=auto
|
||||
request — is dropped, matching the MET path's diurnal-coverage gate."""
|
||||
offset = 3 * 3600 # UTC+3, e.g. Europe/Vilnius (the reported Ringaudai incident)
|
||||
def test_recent_forecast_om_keeps_the_in_progress_local_day(monkeypatch):
|
||||
"""Today must survive the Open-Meteo bundle. Unlike the MET Norway series (which
|
||||
starts mid-day and is gated on diurnal coverage), this endpoint backfills the
|
||||
rest of today from the model run, so today's high/low is a whole-day value of
|
||||
the same kind as tomorrow's. Dropping it left a hole between a complete
|
||||
yesterday and a forecast tomorrow, which the UI renders as a missing column."""
|
||||
offset = 3 * 3600 # UTC+3, e.g. Europe/Vilnius (the reported Ringaudai cell)
|
||||
local_today = (datetime.datetime.now(datetime.timezone.utc)
|
||||
+ datetime.timedelta(seconds=offset)).date()
|
||||
days = [local_today + datetime.timedelta(days=n) for n in (-2, -1, 0, 1)]
|
||||
daily = {
|
||||
"time": [d.isoformat() for d in days],
|
||||
"temperature_2m_max": [80.0, 82.0, 61.0, 84.0], # today's 61 is the partial value
|
||||
"temperature_2m_min": [60.0, 61.0, 57.0, 62.0],
|
||||
"temperature_2m_max": [80.0, 82.0, 76.1, 84.0],
|
||||
"temperature_2m_min": [60.0, 61.0, 55.9, 62.0],
|
||||
"precipitation_sum": [0.0, 0.0, 0.0, 0.0],
|
||||
}
|
||||
payload = {"utc_offset_seconds": offset, "daily": daily}
|
||||
|
|
@ -237,60 +237,42 @@ def test_recent_forecast_om_drops_the_in_progress_local_day(monkeypatch):
|
|||
def json(self): return payload
|
||||
monkeypatch.setattr(climate, "_request", lambda *a, **k: Resp())
|
||||
|
||||
got = climate._fetch_recent_forecast_om({"center_lat": 54.9, "center_lon": 23.8})
|
||||
assert got["date"].to_list() == days # every day present, no gap at today
|
||||
row = got.filter(pl.col("date") == local_today)
|
||||
assert row["tmax"].item() == 76.1 # today's whole-day high, ungraded-down
|
||||
assert row["tmin"].item() == 55.9
|
||||
|
||||
|
||||
def test_recent_forecast_om_still_drops_a_today_with_no_high_low(monkeypatch):
|
||||
"""The one case today is excluded: it came back with no high/low and so cannot be
|
||||
graded at all. _to_frame enforces that for every day, today included."""
|
||||
offset = 3 * 3600
|
||||
local_today = (datetime.datetime.now(datetime.timezone.utc)
|
||||
+ datetime.timedelta(seconds=offset)).date()
|
||||
days = [local_today + datetime.timedelta(days=n) for n in (-1, 0, 1)]
|
||||
daily = {
|
||||
"time": [d.isoformat() for d in days],
|
||||
"temperature_2m_max": [82.0, None, 84.0],
|
||||
"temperature_2m_min": [61.0, None, 62.0],
|
||||
"precipitation_sum": [0.0, 0.0, 0.0],
|
||||
}
|
||||
payload = {"utc_offset_seconds": offset, "daily": daily}
|
||||
|
||||
class Resp:
|
||||
def json(self): return payload
|
||||
monkeypatch.setattr(climate, "_request", lambda *a, **k: Resp())
|
||||
|
||||
got = climate._fetch_recent_forecast_om(
|
||||
{"center_lat": 54.9, "center_lon": 23.8})["date"].to_list()
|
||||
assert local_today not in got # partial today dropped
|
||||
assert local_today - datetime.timedelta(days=1) in got # yesterday kept
|
||||
assert local_today + datetime.timedelta(days=1) in got # tomorrow's forecast kept
|
||||
assert len(got) == 3
|
||||
assert local_today not in got
|
||||
assert got == [local_today - datetime.timedelta(days=1),
|
||||
local_today + datetime.timedelta(days=1)]
|
||||
|
||||
|
||||
def test_recent_forecast_om_keeps_all_days_without_a_utc_offset(monkeypatch):
|
||||
"""No UTC offset reported -> skip the in-progress-day guard rather than guess a
|
||||
date, so the bundle passes through as before (only the usual null-day filter)."""
|
||||
payload = {"daily": _om_daily(3)} # no utc_offset_seconds
|
||||
|
||||
class Resp:
|
||||
def json(self): return payload
|
||||
monkeypatch.setattr(climate, "_request", lambda *a, **k: Resp())
|
||||
|
||||
df = climate._fetch_recent_forecast_om({"center_lat": 1.0, "center_lon": 2.0})
|
||||
assert df.height == climate._to_frame(_om_daily(3)).height
|
||||
|
||||
|
||||
def test_recent_forecast_om_drops_the_in_progress_local_day(monkeypatch):
|
||||
"""The Open-Meteo fallback must not grade the cell's in-progress local day: its
|
||||
daily high/low is only a partial aggregate of the hours elapsed so far (a cool
|
||||
morning would read as a record-low high). Past days and future forecast days
|
||||
survive; today — per the UTC offset Open-Meteo reports for a timezone=auto
|
||||
request — is dropped, matching the MET path's diurnal-coverage gate."""
|
||||
offset = 3 * 3600 # UTC+3, e.g. Europe/Vilnius (the reported Ringaudai incident)
|
||||
local_today = (datetime.datetime.now(datetime.timezone.utc)
|
||||
+ datetime.timedelta(seconds=offset)).date()
|
||||
days = [local_today + datetime.timedelta(days=n) for n in (-2, -1, 0, 1)]
|
||||
daily = {
|
||||
"time": [d.isoformat() for d in days],
|
||||
"temperature_2m_max": [80.0, 82.0, 61.0, 84.0], # today's 61 is the partial value
|
||||
"temperature_2m_min": [60.0, 61.0, 57.0, 62.0],
|
||||
"precipitation_sum": [0.0, 0.0, 0.0, 0.0],
|
||||
}
|
||||
payload = {"utc_offset_seconds": offset, "daily": daily}
|
||||
|
||||
class Resp:
|
||||
def json(self): return payload
|
||||
monkeypatch.setattr(climate, "_request", lambda *a, **k: Resp())
|
||||
|
||||
got = climate._fetch_recent_forecast_om(
|
||||
{"center_lat": 54.9, "center_lon": 23.8})["date"].to_list()
|
||||
assert local_today not in got # partial today dropped
|
||||
assert local_today - datetime.timedelta(days=1) in got # yesterday kept
|
||||
assert local_today + datetime.timedelta(days=1) in got # tomorrow's forecast kept
|
||||
assert len(got) == 3
|
||||
|
||||
|
||||
def test_recent_forecast_om_keeps_all_days_without_a_utc_offset(monkeypatch):
|
||||
"""No UTC offset reported -> skip the in-progress-day guard rather than guess a
|
||||
date, so the bundle passes through as before (only the usual null-day filter)."""
|
||||
"""No UTC offset reported changes nothing: the bundle passes through with only
|
||||
the usual null-day filter."""
|
||||
payload = {"daily": _om_daily(3)} # no utc_offset_seconds
|
||||
|
||||
class Resp:
|
||||
|
|
|
|||
91
backend/tests/web/test_heartbeat.py
Normal file
91
backend/tests/web/test_heartbeat.py
Normal file
|
|
@ -0,0 +1,91 @@
|
|||
"""Process-level liveness heartbeat for web/worker (app.py _heartbeat_loop /
|
||||
_start_heartbeat) — the fix for ProdWebContainerSilent / ProdWorkerContainerSilent
|
||||
false-firing on container-stdout silence that isn't a real liveness signal for
|
||||
either role."""
|
||||
import json
|
||||
import time
|
||||
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from core import audit
|
||||
from web import app as appmod
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _reset_heartbeat_thread():
|
||||
# _HEARTBEAT_THREAD is a module-level singleton guard (mirrors notify.py's own
|
||||
# _STOP/_THREAD pattern) -- reset it around every test so one test starting the
|
||||
# thread doesn't leave _start_heartbeat() a no-op for the next.
|
||||
appmod._HEARTBEAT_STOP.clear()
|
||||
appmod._HEARTBEAT_THREAD = None
|
||||
yield
|
||||
appmod._HEARTBEAT_STOP.set()
|
||||
appmod._HEARTBEAT_THREAD = None
|
||||
|
||||
|
||||
@pytest.mark.parametrize("role", ["web", "worker", "all"])
|
||||
def test_heartbeat_loop_tags_the_beat_with_role(monkeypatch, role):
|
||||
monkeypatch.setattr(appmod, "ROLE", role)
|
||||
calls = []
|
||||
monkeypatch.setattr(audit, "log_heartbeat", lambda daemon, interval_s: calls.append((daemon, interval_s)))
|
||||
monkeypatch.setattr(appmod.metrics, "record_heartbeat", lambda daemon, interval_s: None)
|
||||
|
||||
# Setting the stop event first means .wait() returns True immediately, so the
|
||||
# loop beats exactly once (the pre-loop beat) and exits without blocking.
|
||||
appmod._HEARTBEAT_STOP.set()
|
||||
appmod._heartbeat_loop()
|
||||
|
||||
assert calls == [(role, appmod.HEARTBEAT_INTERVAL)]
|
||||
|
||||
|
||||
def test_start_heartbeat_respects_the_disable_flag(monkeypatch):
|
||||
monkeypatch.setenv("THERMOGRAPH_ENABLE_HEARTBEAT", "0")
|
||||
started = []
|
||||
monkeypatch.setattr(appmod.threading, "Thread", lambda **kw: started.append(kw))
|
||||
|
||||
appmod._start_heartbeat()
|
||||
|
||||
assert appmod._HEARTBEAT_THREAD is None
|
||||
assert started == []
|
||||
|
||||
|
||||
def test_start_heartbeat_is_idempotent(monkeypatch):
|
||||
monkeypatch.delenv("THERMOGRAPH_ENABLE_HEARTBEAT", raising=False)
|
||||
starts = []
|
||||
|
||||
class _FakeThread:
|
||||
def __init__(self, **kw):
|
||||
starts.append(kw)
|
||||
|
||||
def start(self):
|
||||
pass
|
||||
|
||||
monkeypatch.setattr(appmod.threading, "Thread", _FakeThread)
|
||||
|
||||
appmod._start_heartbeat()
|
||||
appmod._start_heartbeat()
|
||||
|
||||
# A second call is a no-op once a thread handle already exists -- mirrors
|
||||
# notify.start()'s own idempotency, so a double-entry into the lifespan
|
||||
# (or a stray manual call) can't spawn a second beat loop.
|
||||
assert len(starts) == 1
|
||||
|
||||
|
||||
def test_app_boot_emits_a_heartbeat_within_one_interval(monkeypatch, tmp_path):
|
||||
monkeypatch.setattr(appmod, "ROLE", "web")
|
||||
monkeypatch.setattr(appmod, "HEARTBEAT_INTERVAL", 0.01)
|
||||
monkeypatch.setattr(audit, "HEARTBEAT_DIR", str(tmp_path))
|
||||
monkeypatch.delenv("THERMOGRAPH_ENABLE_HEARTBEAT", raising=False)
|
||||
|
||||
with TestClient(appmod.app):
|
||||
# The pre-loop beat fires synchronously on thread start; a short sleep is
|
||||
# just slack for the thread to actually get scheduled. The loop itself
|
||||
# keeps running on HEARTBEAT_INTERVAL until lifespan shutdown sets the
|
||||
# stop event, so there's nothing to join here.
|
||||
time.sleep(0.2)
|
||||
|
||||
files = list(tmp_path.glob("heartbeat-*.jsonl"))
|
||||
assert files, "expected a heartbeat-*.jsonl file to be written during app lifespan"
|
||||
lines = [json.loads(line) for line in files[0].read_text().splitlines() if line]
|
||||
assert any(rec["daemon"] == "web" and rec["tag"] == "heartbeat" for rec in lines)
|
||||
|
|
@ -154,6 +154,39 @@ def _should_run_notifier() -> bool:
|
|||
return ROLE in ("worker", "all") and singleton.claim_leader()
|
||||
|
||||
|
||||
_HEARTBEAT_STOP = threading.Event()
|
||||
_HEARTBEAT_THREAD: threading.Thread | None = None
|
||||
HEARTBEAT_INTERVAL = float(os.environ.get("THERMOGRAPH_HEARTBEAT_INTERVAL", "300")) # 5 min
|
||||
|
||||
|
||||
def _heartbeat_loop() -> None:
|
||||
metrics.record_heartbeat(ROLE, HEARTBEAT_INTERVAL)
|
||||
audit.log_heartbeat(ROLE, HEARTBEAT_INTERVAL)
|
||||
while not _HEARTBEAT_STOP.wait(HEARTBEAT_INTERVAL):
|
||||
metrics.record_heartbeat(ROLE, HEARTBEAT_INTERVAL)
|
||||
audit.log_heartbeat(ROLE, HEARTBEAT_INTERVAL)
|
||||
|
||||
|
||||
def _start_heartbeat() -> None:
|
||||
"""Process-level liveness beat, tagged daemon=ROLE, independent of whether this
|
||||
process also runs the notifier. web's and worker's own stdout is near-silent by
|
||||
design (real logging goes to the app-json file source; worker's Docker/stdout
|
||||
stream is dropped entirely by Alloy), so container-log-silence alerting can't
|
||||
tell a live process from a dead one for either role — this heartbeat is the
|
||||
signal observability actually alerts on instead. Deliberately unguarded across
|
||||
every uvicorn worker process and every autoscaled replica: a beat is a single
|
||||
local file append (no upstream quota, no DB, no lock), so duplicate beats are
|
||||
harmless, and only the process actually being dead stops them — see
|
||||
_should_run_notifier's leader election for the contrasting case where
|
||||
duplication would be a real cost."""
|
||||
global _HEARTBEAT_THREAD
|
||||
if os.environ.get("THERMOGRAPH_ENABLE_HEARTBEAT", "1") == "0" or _HEARTBEAT_THREAD is not None:
|
||||
return
|
||||
_HEARTBEAT_STOP.clear()
|
||||
_HEARTBEAT_THREAD = threading.Thread(target=_heartbeat_loop, name=f"heartbeat-{ROLE}", daemon=True)
|
||||
_HEARTBEAT_THREAD.start()
|
||||
|
||||
|
||||
@contextlib.asynccontextmanager
|
||||
async def _lifespan(app):
|
||||
# Warm the local place-name index (/suggest's typo tolerance) in the
|
||||
|
|
@ -167,6 +200,7 @@ async def _lifespan(app):
|
|||
await db.create_db_and_tables()
|
||||
places.start_loading()
|
||||
threading.Thread(target=_neighbor_warmer, name="neighbor-warmer", daemon=True).start()
|
||||
_start_heartbeat()
|
||||
# Background subscription evaluator. Its periodic sweep fetches upstream on a timer
|
||||
# (not per request), so with multiple workers — or multiple *hosts* under Swarm —
|
||||
# it must run in exactly one: elect a leader (see singleton.claim_leader, which
|
||||
|
|
@ -189,6 +223,7 @@ async def _lifespan(app):
|
|||
yield
|
||||
audit.log_activity("app.stop", {"role": ROLE, "pid": os.getpid()})
|
||||
notify.stop()
|
||||
_HEARTBEAT_STOP.set()
|
||||
await _frontend_client.aclose()
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -1,181 +1,72 @@
|
|||
# Thermograph frontend — agent instructions
|
||||
# frontend/ — agent instructions
|
||||
|
||||
This repo is the **SSR + static-asset service** split out of the `emi/thermograph`
|
||||
monorepo (repo-split Stage 7). It has no climate data, no DB, and does no
|
||||
polars/compute work — everything comes from the backend's content API over
|
||||
HTTP. It is one of four sibling repos in `thermograph-repos/`:
|
||||
The **SSR + static-asset service**: server-rendered crawlable pages, the
|
||||
interactive tool's SPA shells, and every static asset. No climate data, no DB,
|
||||
no compute — everything comes from `backend/`'s content API over HTTP.
|
||||
|
||||
- **`thermograph-backend`** — FastAPI API + grading/scoring/grid compute + DB.
|
||||
This repo's only dependency.
|
||||
- **`thermograph-frontend`** (this repo) — SSR content pages + the interactive
|
||||
tool's SPA shells + every static asset.
|
||||
- **`thermograph-infra`** — Terraform, `docker-compose*.yml`, `deploy/` (incl.
|
||||
`deploy.sh`, the SOPS secrets vault), Caddy config. No app code.
|
||||
- **`thermograph-docs`** — architecture decision records + operator runbooks.
|
||||
No code.
|
||||
Read the root `CLAUDE.md` first for the branch model, deploy contract and image
|
||||
names.
|
||||
|
||||
See `MONOREPO-CUTOVER-PLAN.md` (one level up, in `thermograph-repos/`) for the
|
||||
full split status and the remaining gap list before the monorepo can be
|
||||
archived.
|
||||
## Build & verify
|
||||
|
||||
## What this repo is
|
||||
- `make test` — whole suite (`scripts/test.sh`).
|
||||
- `make test-unit` — hermetic unit tier: SSR rendering fed committed fixtures,
|
||||
no Docker. **This is the tier CI runs.**
|
||||
- `make test-integration` — pulls and runs the real backend image and tests the
|
||||
live contract. Local only; CI does not run it.
|
||||
- `make backend-up` / `make backend-down` — a local backend container for dev.
|
||||
- `make capture-fixtures` — refresh `tests/fixtures/*.json` from a live backend.
|
||||
|
||||
- **`content.py`** — server-rendered, crawlable pages (climate hub, per-city,
|
||||
month, records, glossary, about, privacy) + `robots.txt` + `sitemap.xml`.
|
||||
Every route fetches its data from the backend's content API
|
||||
(`api_client.py`) instead of computing in-process; `content_payloads.py` on
|
||||
the backend owns `page_title`/`canonical_path`/`breadcrumb`/`jsonld`, not
|
||||
this repo.
|
||||
- **`static/*.js`** — the interactive tool: `app.js` (map/search/graded
|
||||
results + inline SVG chart), `calendar.js`/`day.js`/`score.js`/`compare.js`
|
||||
(SPA shells served by `app.py`'s `_page()`), `account.js` (auth), `cache.js`
|
||||
(IndexedDB response cache + `/cell` bundle prefetch), `shared.js` (format
|
||||
helpers shared across views).
|
||||
- **`static/style.css`** — the single hand-written stylesheet; all design
|
||||
tokens live here (see `DESIGN.md`).
|
||||
- **`templates/*.html.j2`** — Jinja templates `content.py` renders.
|
||||
- **`content/*.yaml`** — structured SSR copy (glossary, static-page SEO meta),
|
||||
loaded by `content_loader.py`. Committed here as a starter copy extracted
|
||||
alongside the split; real cross-repo copy vendoring (a pinned
|
||||
`thermograph-copy` checkout at build time) is deferred, unbuilt follow-up —
|
||||
don't assume it exists.
|
||||
## Running it
|
||||
|
||||
## Branch/PR + deploy flow
|
||||
|
||||
- Work happens on feature branches; PR into `main`. (The monorepo's
|
||||
`dev`→`main`→`release` three-stage promotion and its LAN `deploy-dev.yml`
|
||||
path do **not** exist in this repo yet — a documented gap, see the cutover
|
||||
plan §4. Don't assume a `dev` branch here.)
|
||||
- **`.forgejo/workflows/build-push.yml`** — on push to `dev`/`main`/`release`
|
||||
or a `v*.*.*` tag: builds THIS repo's own `Dockerfile` and pushes
|
||||
`git.thermograph.org/emi/thermograph-frontend/app`, tagged `sha-<12 hex>`
|
||||
(every push) and the semver tag (tag pushes only). Backend publishes its own
|
||||
separate image the same way — the two are no longer one shared
|
||||
`emi/thermograph/app` image.
|
||||
- **`.forgejo/workflows/build.yml`** — push/PR to `main`: proves the Dockerfile
|
||||
builds. It is a **build check only**, not a boot/health check — booting the
|
||||
real app crashes at import without a reachable backend (see API-version
|
||||
section below), so a standalone boot check would need to check out
|
||||
`thermograph-backend` too. Not yet built; flagged, not silently skipped.
|
||||
- **`.forgejo/workflows/deploy.yml`** — push to `main`: SSH to beta, run
|
||||
`SERVICE=frontend FRONTEND_IMAGE_TAG=sha-<12 hex> /opt/thermograph/deploy/deploy.sh`
|
||||
(that script lives in `thermograph-infra`). Rolls **only** the frontend
|
||||
container (`--no-deps`); backend is untouched and deployed independently by
|
||||
its own repo's workflow. `deploy.sh` retries the image pull for ~5 min
|
||||
(Forgejo has no cross-workflow `needs:`, so this deploy can race ahead of
|
||||
`build-push.yml`) and waits for a healthy backend before declaring the roll
|
||||
OK (frontend's boot fetches the IndexNow key from backend — see below).
|
||||
- **`.forgejo/workflows/deploy-prod.yml`** — push to `release`: same shape,
|
||||
targets prod via its own `PROD_SSH_*` secret set (fully separate from beta's,
|
||||
so a beta credential leak can't touch prod). Nothing else deploys to prod;
|
||||
there is no release-triggered Terraform apply from this repo.
|
||||
- The `SERVICE=frontend` + `FRONTEND_IMAGE_TAG=sha-<12 hex>` pair is the entire
|
||||
contract into `thermograph-infra/deploy/deploy.sh`: it persists each
|
||||
service's live tag in `deploy/.image-tags.env` (host-side, untracked) so a
|
||||
frontend-only roll never disturbs backend's currently-running tag, and
|
||||
vice versa.
|
||||
|
||||
## How to run / test
|
||||
|
||||
There is no `run.sh`/`Makefile` in this repo yet (the monorepo's `make
|
||||
lan-run`/`make run`/`make stop`/`venv` targets have no home here — see the
|
||||
cutover plan's gap list). Run directly:
|
||||
`THERMOGRAPH_API_BASE_INTERNAL` is **required** — `api_client.py` raises at
|
||||
import if unset.
|
||||
|
||||
```bash
|
||||
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
|
||||
THERMOGRAPH_API_BASE_INTERNAL=http://127.0.0.1:8137 \
|
||||
THERMOGRAPH_BASE=/thermograph \
|
||||
THERMOGRAPH_API_BASE_INTERNAL=http://127.0.0.1:8137 THERMOGRAPH_BASE=/thermograph \
|
||||
.venv/bin/uvicorn app:app --host 0.0.0.0 --port 8080
|
||||
```
|
||||
|
||||
`THERMOGRAPH_API_BASE_INTERNAL` is **required** — `api_client.py` raises
|
||||
`RuntimeError` at import if it's unset. Other env vars: `THERMOGRAPH_BASE`
|
||||
(default `/thermograph`; the Dockerfile sets `/` for the deployed clean-root
|
||||
topology), `THERMOGRAPH_API_VERSION` (default `v2`, see below),
|
||||
`THERMOGRAPH_API_BASE_PUBLIC` (browser-facing backend origin for asset URLs
|
||||
when frontend and backend are cross-origin; empty = same-origin, today's
|
||||
default), `THERMOGRAPH_SSR_CACHE_TTL` (default 600s, the content-API response
|
||||
cache in `api_client.py`), `THERMOGRAPH_GOOGLE_VERIFY`/`THERMOGRAPH_BING_VERIFY`
|
||||
(search-console `<meta>` tags).
|
||||
Other env: `THERMOGRAPH_BASE` (default `/thermograph`; the Dockerfile sets `/`),
|
||||
`THERMOGRAPH_API_VERSION` (default `v2`), `THERMOGRAPH_API_BASE_PUBLIC`
|
||||
(browser-facing backend origin when cross-origin; empty = same-origin, today's
|
||||
default), `THERMOGRAPH_SSR_CACHE_TTL` (default 600s).
|
||||
|
||||
**Known gap — `tests/conftest.py` does not run standalone in this repo.** It
|
||||
still assumes the pre-split layout: a sibling `backend/` checkout with its own
|
||||
`tests/conftest.py` (`make_history`/`make_recent`) and an `api.content_payloads`
|
||||
module, neither of which exists here. It was extracted verbatim as reference
|
||||
material for the real fix (a genuine cross-repo contract-test job, or a
|
||||
self-contained set of fixtures owned in this repo) — not yet built. Do **not**
|
||||
assume `pytest tests` passes here until that lands; `.forgejo/workflows/build.yml`
|
||||
only proves the image builds, for the same reason. There is also no
|
||||
`requirements-dev.txt` in this repo yet (pytest/httpx test deps aren't pinned).
|
||||
## Layout
|
||||
|
||||
## API-version pinning contract
|
||||
- **`content.py`** — SSR routes (climate hub, city, month, records, glossary,
|
||||
about, privacy) + `robots.txt` + `sitemap.xml`. Each fetches from the backend
|
||||
via `api_client.py`; the backend's `content_payloads.py` owns `page_title` /
|
||||
`canonical_path` / `breadcrumb` / `jsonld`, not this domain.
|
||||
- **`static/*.js`** — `app.js` (map/search/results + inline SVG chart),
|
||||
`calendar.js`/`day.js`/`score.js`/`compare.js` (SPA shells), `account.js`
|
||||
(auth), `cache.js` (IndexedDB cache + `/cell` bundle prefetch), `shared.js`.
|
||||
- **`static/style.css`** — the single hand-written stylesheet; design tokens live
|
||||
here, documented in `DESIGN.md`.
|
||||
- **`templates/*.html.j2`**, **`content/*.yaml`** (structured SSR copy, loaded by
|
||||
`content_loader.py`).
|
||||
|
||||
Every backend call in this repo goes through a single pinned constant instead
|
||||
of scattered `api/v2/...` literals:
|
||||
## Contracts with the backend
|
||||
|
||||
- **Python (SSR):** `api_client.py`'s module-level `API_VERSION` (default
|
||||
`"v2"`, overridable via `THERMOGRAPH_API_VERSION`). Every path builder
|
||||
(`hub()`, `sitemap()`, `indexnow_key()`, `home()`, `city()`, `city_month()`,
|
||||
`city_records()`) reads it.
|
||||
- **JS (interactive tool):** `static/account.js`'s exported `API_VERSION`
|
||||
constant and the `uv(path)` helper — every fetch in `cache.js`, `account.js`
|
||||
itself, etc. is built by wrapping the path in `uv(...)` rather than
|
||||
hardcoding a version.
|
||||
|
||||
**Bump only in lockstep with a verified backend `/api/version` check** — the
|
||||
backend exposes `GET {BASE}/api/version` → `{backend_version, min_frontend,
|
||||
payload_ver}` (`web/app.py`'s `API_CONTRACT_VERSION`/`MIN_SUPPORTED_FRONTEND`).
|
||||
Before bumping this repo's `API_VERSION`, confirm the target backend's
|
||||
`backend_version` actually supports it and its `min_frontend` doesn't already
|
||||
exclude the version you're moving *away* from.
|
||||
|
||||
**A v3 cutover would work like this:** backend mounts a new `v3 = APIRouter()`
|
||||
alongside `v2`, re-registering only the changed handlers (v2 keeps serving
|
||||
old clients); backend bumps `API_CONTRACT_VERSION` to `"3"` in that same PR.
|
||||
Only once that's deployed and verified does this repo bump `API_VERSION` to
|
||||
`"v3"` in both `api_client.py` and `account.js` — one PR, both pins together,
|
||||
never one without the other. `v2` (and the `/api`, `/api/v1` aliases) stay
|
||||
mounted and working until no client depends on them.
|
||||
|
||||
The frontend's own boot is **resilient to the backend being down**: the
|
||||
`indexnow_key()` fetch in `content.py`'s `register()` is tried once eagerly
|
||||
(so the common case still serves `/<key>.txt` as a plain static route), but a
|
||||
failure is caught, logged, and falls back to a lazy per-request lookup
|
||||
instead of crashing boot — asynchronous frontend/backend deploys depend on
|
||||
this (a briefly-unreachable backend at frontend boot must be survivable, not
|
||||
fatal).
|
||||
|
||||
## Cross-repo contracts this repo depends on
|
||||
|
||||
These are **live contracts with the backend** — a change on either side
|
||||
without the other breaks something silently, not loudly:
|
||||
|
||||
- **The `/cell` bundle + ETag/`If-None-Match`** — `static/cache.js` fetches
|
||||
`/api/v2/cell` once per view-set and slices it for calendar/day/score/compare
|
||||
instead of one request per view; conditional refetch is via `If-None-Match`
|
||||
against the stored `ETag`, so an unchanged payload costs an empty 304. The
|
||||
URL map (which slice serves which view) must stay in sync with
|
||||
`thermograph-backend`'s route/payload shape.
|
||||
- **`shared.js`'s `pctOrd()` must mirror `thermograph-backend`'s
|
||||
`data/grading.py`'s `pct_ordinal()`** byte-for-byte in behavior: floor a
|
||||
percentile into `1..99` (never round to 100/0 — "100th percentile" reads as
|
||||
measurement error, not "as extreme as it has ever been"). Every percentile
|
||||
shown anywhere (Day page, calendar tooltip, chart, city pages, homepage
|
||||
strip) goes through one of these two functions; if they diverge, the same
|
||||
reading says two different things on two different surfaces.
|
||||
- **Unit/region logic** (`format.py`'s `F_COUNTRIES` / `static/units.js`'s
|
||||
`F_REGIONS` / backend's `api/content_payloads.py`'s `F_COUNTRIES`) — the set
|
||||
of Fahrenheit-using country codes must stay identical across all three;
|
||||
backend has a test asserting this.
|
||||
|
||||
## Design & visual verification
|
||||
|
||||
Design tokens + conventions are documented in `DESIGN.md` (source of truth:
|
||||
`static/style.css`). To see a change rendered, see `DESIGN.md`'s `make shots`
|
||||
section (`tools/shoot.py`).
|
||||
- **API version is pinned in exactly two places** — `api_client.py`'s
|
||||
`API_VERSION` (Python) and `static/account.js`'s exported `API_VERSION` plus the
|
||||
`uv(path)` helper (JS). Never hardcode `api/v2/...` anywhere else. Bump both in
|
||||
one PR, and only after the target backend's `/api/version` confirms it serves
|
||||
that version and its `min_frontend` doesn't exclude the one you're leaving.
|
||||
- **`shared.js::pctOrd()` must mirror `backend`'s `data/grading.py::pct_ordinal()`**
|
||||
in behaviour — floor into `1..99`, never 0 or 100. Every percentile on every
|
||||
surface goes through one of the two; if they diverge, the same reading says two
|
||||
different things in two places.
|
||||
- **`/cell` bundle + ETag** — `cache.js` fetches `/api/v2/cell` once per view-set
|
||||
and slices it, revalidating with `If-None-Match`. The slice→view map must track
|
||||
the backend's payload shape.
|
||||
- **Fahrenheit country set** — `format.py::F_COUNTRIES` and `static/units.js`'s
|
||||
`F_REGIONS` must stay identical to the backend's `F_COUNTRIES`.
|
||||
- **Boot must survive an unreachable backend.** `content.py`'s `register()` tries
|
||||
the IndexNow-key fetch once, catches failure, and falls back to a lazy
|
||||
per-request lookup. Asynchronous FE/BE deploys depend on this — don't make it
|
||||
fatal.
|
||||
|
||||
## Commits & PRs
|
||||
|
||||
Describe only the substance of the change; concise, technical. Never mention
|
||||
AI/Claude/assistants or automated authorship anywhere — no trailers,
|
||||
co-authors, emoji, or "as requested"/"per the agent" narration.
|
||||
Concise and technical. Never mention AI, assistants or automated authorship.
|
||||
|
|
|
|||
4
infra/.gitignore
vendored
4
infra/.gitignore
vendored
|
|
@ -27,3 +27,7 @@ __pycache__/
|
|||
deploy/.image-tags.env
|
||||
deploy/.deploy.lock
|
||||
deploy/.stack-image-tags.env
|
||||
|
||||
# LAN dev's mirror of the rendered secrets, for snap-confined Docker (see
|
||||
# deploy-dev.sh / render-secrets.sh) -- plaintext, re-rendered every deploy.
|
||||
deploy/dev-secrets.env
|
||||
|
|
|
|||
|
|
@ -1,42 +1,58 @@
|
|||
# thermograph-infra — agent instructions
|
||||
# infra/ — agent instructions
|
||||
|
||||
How and where the already-built Thermograph app images run. This repo owns
|
||||
Terraform, the SOPS+age secrets vault, compose files, deploy scripts, host
|
||||
provisioning, and the ops cron (DB backup + IndexNow). The app repos
|
||||
(`thermograph-backend`, `thermograph-frontend`) own building and testing the
|
||||
images; this repo never checks out app source. The old monorepo
|
||||
(`emi/thermograph`) is archived — do not point anything at it.
|
||||
How and where the already-built app images run: Terraform, the SOPS+age secrets
|
||||
vault, compose and Swarm stack files, deploy scripts, host provisioning, mail,
|
||||
and the ops cron (DB backup + IndexNow). Read the root `CLAUDE.md` first — it
|
||||
owns the branch model, the deploy contract, and which environment runs which
|
||||
orchestrator.
|
||||
|
||||
## The four machines
|
||||
## The machines
|
||||
|
||||
Same as ACCESS.md: **dev machine** (operator's box, LAN dev server, CI runner),
|
||||
Per `ACCESS.md`: **dev** (operator's box — LAN dev server and CI runner),
|
||||
**prod** (`169.58.46.181`, thermograph.org, `agent` user, passwordless sudo),
|
||||
**beta** (`75.119.132.91`, beta.thermograph.org + Forgejo, `agent` user). Hosts'
|
||||
`/opt/thermograph` (and LAN's `~/thermograph-dev`) are checkouts of THIS repo.
|
||||
**beta** (`75.119.132.91`, beta.thermograph.org + Forgejo + Grafana/Loki,
|
||||
`agent` user). All four are on a WireGuard mesh. Hosts' `/opt/thermograph` is a
|
||||
checkout of **this monorepo**, not an infra-only repo.
|
||||
|
||||
## Deploys
|
||||
## Deploy paths
|
||||
|
||||
- `deploy/deploy.sh` — the one deploy path (prod/beta): resets this repo's
|
||||
checkout, renders secrets from the SOPS vault, pulls the per-service image
|
||||
tags (`BACKEND_IMAGE_TAG`/`FRONTEND_IMAGE_TAG`, persisted in untracked
|
||||
`deploy/.image-tags.env`), rolls the target service, health-checks via
|
||||
container healthchecks. Invoked over SSH by the app repos' deploy workflows.
|
||||
- `deploy/deploy-dev.sh` — thin LAN-dev wrapper around deploy.sh (dev compose
|
||||
overlay, `~/thermograph-dev`, branch `dev`).
|
||||
- Branch model: see README — infra `main` is live on prod+beta, `dev` on LAN;
|
||||
`release` is currently unused. App code is env-staged via image tags; infra
|
||||
is not.
|
||||
- **`deploy/deploy.sh`** — the single entry point for beta and prod. Resets the
|
||||
host checkout to `BRANCH` (default `main`), renders secrets, then routes:
|
||||
if `/etc/thermograph/deploy-mode` says `stack` it execs
|
||||
`deploy/stack/deploy-stack.sh`, otherwise it rolls compose services. Per-service
|
||||
tags persist in untracked `deploy/.image-tags.env` (compose) and
|
||||
`deploy/.stack-image-tags.env` (stack) so the two never mix.
|
||||
- **`deploy/stack/deploy-stack.sh`** — the Swarm path, live on prod. `backend`
|
||||
rolls web **and** worker; `frontend` rolls frontend; `all` runs a full
|
||||
`docker stack deploy`. Start-first, health-gated, auto-rollback.
|
||||
`STACK_TEST=1` rehearses the whole thing under stack name `thermograph-test`
|
||||
on throwaway volumes and ports `18137`/`18080`.
|
||||
- **`deploy/deploy-dev.sh`** — thin LAN-dev wrapper (dev compose overlay,
|
||||
`~/thermograph-dev`). Its CI trigger is inert; see root `CLAUDE.md`.
|
||||
- **`Makefile`** — compose orchestration only: `up`, `down`, `db-up`, `dev-up`,
|
||||
`om-up`, `om-backfill`.
|
||||
|
||||
Volumes are the reason the compose project name is pinned: compose creates
|
||||
`thermograph_pgdata`/`_appdata`/`_applogs`, and the Swarm stack declares those
|
||||
same names as `external: true` at the same mount paths.
|
||||
|
||||
## Rules
|
||||
|
||||
- **Never run `terraform apply` casually** — no tfstate exists anywhere (never
|
||||
persisted), so an apply would attempt full re-provisioning of live hosts.
|
||||
Terraform here is executable documentation until state is bootstrapped.
|
||||
- Secrets: only via the SOPS vault (`deploy/secrets/*.yaml`, `sops edit` +
|
||||
commit + deploy). Never hand-edit `/etc/thermograph.env` on a SOPS-enabled
|
||||
host — it's a rendered artifact. `secrets-guard` CI rejects plaintext.
|
||||
- The ops cron (`.forgejo/workflows/ops-cron.yml`) is THE prod backup — it uses
|
||||
this repo's `PROD_SSH_*` Actions secrets. If you touch it, verify a dump
|
||||
actually lands in `agent@prod:~/thermograph-backups/`.
|
||||
- Commits/PRs: concise and technical; never mention AI/assistants/automated
|
||||
authorship.
|
||||
- **Never run `terraform apply` casually.** No tfstate is persisted anywhere, so
|
||||
an apply would attempt full re-provisioning of live hosts. Terraform here is
|
||||
executable documentation until state is bootstrapped.
|
||||
- **Secrets: SOPS vault only** (`deploy/secrets/*.yaml`, `sops edit` → commit →
|
||||
deploy). Never hand-edit `/etc/thermograph.env` — it is a rendered artifact.
|
||||
`secrets-guard` CI rejects plaintext. `deploy/secrets/seed-from-live.sh` reads
|
||||
production secrets and is **not** for an agent to run.
|
||||
- **The ops cron (`.forgejo/workflows/ops-cron.yml`, at the repo root) is THE
|
||||
prod backup.** It uses `PROD_SSH_*`, not `SSH_*` — an earlier revision reused
|
||||
`SSH_*` and silently dumped beta while prod had no backup at all. If you touch
|
||||
it, verify a dump actually lands in `agent@prod:~/thermograph-backups/`.
|
||||
- Shell here runs as root over SSH against live hosts with no test suite in
|
||||
front of it. `shell-lint` CI (pinned shellcheck) is the only guard — keep the
|
||||
tree at zero findings.
|
||||
|
||||
## Commits & PRs
|
||||
|
||||
Concise and technical. Never mention AI, assistants or automated authorship.
|
||||
|
|
|
|||
|
|
@ -1,46 +1,49 @@
|
|||
# thermograph-infra
|
||||
# infra/
|
||||
|
||||
Infrastructure for [Thermograph](https://thermograph.org): Terraform host
|
||||
provisioning, the SOPS+age secrets vault, Docker Swarm/WireGuard networking,
|
||||
Forgejo, Caddy, and the deploy scripts that run the already-built app image on
|
||||
each host. Extracted from the app monorepo (`emi/thermograph`) — the app repo
|
||||
owns building and testing the app; this repo owns running it.
|
||||
provisioning, the SOPS+age secrets vault, WireGuard/Swarm networking, Forgejo,
|
||||
Caddy, mail, and the deploy scripts that run the already-built app images on each
|
||||
host. This is a domain of the `emi/thermograph` monorepo — hosts' `/opt/thermograph`
|
||||
is a checkout of the whole monorepo, and `infra/` never builds app source; it only
|
||||
runs published images.
|
||||
|
||||
- **`terraform/`** — provisions/configures hosts (SSH-driven by default; an
|
||||
optional GCP-creating module is scaffolded, no live resources yet) and
|
||||
triggers each deploy. See `terraform/README.md`.
|
||||
- **`deploy/secrets/`** — the git-native SOPS+age secrets vault (every app
|
||||
secret, encrypted at rest, rendered at deploy time). See
|
||||
`deploy/secrets/README.md`.
|
||||
- **`deploy/swarm/`, `deploy/forgejo/`** — the 3-node WireGuard/Swarm cluster
|
||||
that hosts Forgejo (git + CI + registry); does not run the app itself. See
|
||||
`ACCESS.md` and the READMEs under each directory.
|
||||
- **`deploy/deploy.sh`** — pulls the pinned app image (`IMAGE_TAG`) and rolls
|
||||
the compose stack; invoked by Terraform and by the app repo's
|
||||
`.forgejo/workflows/deploy.yml` over SSH.
|
||||
- **`docker-compose*.yml`, `docker-stack.yml`** — how the app image runs
|
||||
(compose in production today; `docker-stack.yml` is a design record for a
|
||||
possible future Swarm-based app deploy, not currently live).
|
||||
optional GCP-creating module is scaffolded, no live resources yet). See
|
||||
`terraform/README.md`. No tfstate is persisted anywhere — treat `apply` as
|
||||
executable documentation, not a routine operation.
|
||||
- **`deploy/secrets/`** — the git-native SOPS+age secrets vault (every app secret,
|
||||
encrypted at rest, rendered at deploy time). See `deploy/secrets/README.md`.
|
||||
- **`deploy/swarm/`, `deploy/forgejo/`** — the WireGuard/Swarm cluster hosting
|
||||
Forgejo (git + CI + registry). See `ACCESS.md`.
|
||||
- **`deploy/deploy.sh`** — the single deploy entry point for beta and prod.
|
||||
Takes `SERVICE=backend|frontend|all` plus `BACKEND_IMAGE_TAG`/`FRONTEND_IMAGE_TAG`,
|
||||
resets the host checkout, renders secrets, and routes to the right orchestrator.
|
||||
- **`deploy/stack/`** — the **Swarm** path, live on **prod**:
|
||||
`thermograph-stack.yml` (db, web, worker, lake, daemon, frontend, autoscaler,
|
||||
autoscaler-lake), `deploy-stack.sh`, `autoscale.sh`, and the LB. Rolling updates
|
||||
are start-first, health-gated, with auto-rollback. `STACK_TEST=1` rehearses the
|
||||
whole stack on throwaway volumes and ports.
|
||||
- **`docker-compose*.yml`** — the **compose** path, live on **beta** and LAN dev
|
||||
(db, backend, lake, daemon, frontend). `docker-compose.dev.yml` is the LAN
|
||||
overlay; `docker-compose.openmeteo.yml` is the self-hosted Open-Meteo overlay.
|
||||
|
||||
The app's own source, `Dockerfile`, and build/test CI stay in the app repo —
|
||||
this repo never checks out app source; hosts only pull tagged images from the
|
||||
registry. See `ACCESS.md` for host access and the Swarm/Forgejo topology, and
|
||||
`terraform/README.md` for the day-to-day `plan`/`apply` workflow.
|
||||
Which path a host takes is decided by `/etc/thermograph/deploy-mode`: the string
|
||||
`stack` makes `deploy.sh` exec `deploy/stack/deploy-stack.sh`; anything else is
|
||||
compose. The workflows never need to know which mode a host runs.
|
||||
|
||||
## Branches & how changes reach each environment
|
||||
|
||||
- **`main`** — what **prod and beta** run: their `/opt/thermograph` checkouts
|
||||
`git reset --hard origin/main` at the start of every deploy (`deploy/deploy.sh`).
|
||||
A merge to `main` reaches those hosts on the next app deploy (or a by-hand
|
||||
`deploy.sh` run); there is no separate infra deploy trigger.
|
||||
- **`dev`** — what **LAN dev** runs: `~/thermograph-dev` resets to it via
|
||||
`deploy/deploy-dev.sh`. Keep it fast-forwarded to `main` (infra changes are not
|
||||
environment-staged today; the branches exist so LAN dev *can* trail or lead
|
||||
when needed).
|
||||
- **`release`** — currently consumed by nothing (prod tracks `main`, not
|
||||
`release`). It exists to mirror the app repos' dev→main→release promotion
|
||||
shape if per-environment infra staging is ever wanted; until then, treat
|
||||
`main` as live-everywhere.
|
||||
- **`main`** — what **prod and beta** run. `infra-sync.yml` fires on a push to
|
||||
`main` touching `infra/**`, fast-forwards each host's `/opt/thermograph` checkout
|
||||
and re-renders `/etc/thermograph.env` from the vault. It deliberately does
|
||||
**not** roll any service: image tags are the app domains' axis, not infra's. A
|
||||
compose or stack change that must recreate containers takes effect on the next
|
||||
app deploy, or a by-hand `SERVICE=all … deploy/deploy.sh`.
|
||||
- **`dev`** — what LAN dev would run via `deploy/deploy-dev.sh`. The CI trigger
|
||||
for this is currently inert (the LAN box still holds a split-era checkout); use
|
||||
`make dev-up` locally.
|
||||
- **`release`** — consumed by app deploys only. Both hosts track infra via `main`;
|
||||
prod's *app images* are staged by `release`, but its checkout follows `main`.
|
||||
|
||||
Note the asymmetry with the app repos: app code IS environment-staged
|
||||
(dev→main→release maps to LAN→beta→prod via image tags), infra is not.
|
||||
Note the asymmetry with the app domains: app code IS environment-staged
|
||||
(`dev`→`main`→`release` maps to LAN→beta→prod via image tags); infra is not.
|
||||
|
|
|
|||
|
|
@ -72,6 +72,15 @@ export COMPOSE_PROJECT_NAME="${COMPOSE_PROJECT_NAME:-thermograph-dev}"
|
|||
# genuinely needs, and nothing else.
|
||||
export THERMOGRAPH_SECRETS_SKIP_COMMON=1
|
||||
|
||||
# (1b) Docker on this box is the snap package, confined to $HOME (plus a short
|
||||
# allowlist) -- it silently can't see /etc/thermograph.env at all, so
|
||||
# env_file: /etc/thermograph.env resolves to nothing for every container no
|
||||
# matter how correctly it's rendered. Mirror the render to a path under
|
||||
# $APP_DIR too; docker-compose.dev.yml points env_file at this copy instead of
|
||||
# /etc/thermograph.env. Untracked (see infra/.gitignore); re-rendered on every
|
||||
# deploy, never committed.
|
||||
export THERMOGRAPH_SECRETS_ENV_FILE_MIRROR="$APP_DIR/infra/deploy/dev-secrets.env"
|
||||
|
||||
# (2) The age key is read from wherever it is READABLE, not necessarily
|
||||
# /etc/thermograph/age.key. render-secrets.sh falls back to `sudo cat` for a
|
||||
# root-owned 0400 key, and this box has no passwordless sudo -- under the Forgejo
|
||||
|
|
|
|||
|
|
@ -146,7 +146,7 @@ bug hit during the frontend Go rewrite). `ci-runner` adds `docker-ce-cli` +
|
|||
Current tag: `git.thermograph.org/emi/thermograph/ci-runner:v2` (`v1` is
|
||||
broken — do not register any runner against it). Rebuild/push (requires a PAT
|
||||
with `write:package` scope — the embedded git-remote token lacks it, same
|
||||
requirement documented in `backend-build-push.yml`):
|
||||
requirement documented in `build-push.yml`):
|
||||
|
||||
```bash
|
||||
docker build -t git.thermograph.org/emi/thermograph/ci-runner:vN deploy/forgejo/ci-runner
|
||||
|
|
|
|||
|
|
@ -82,6 +82,11 @@ After=network-online.target
|
|||
|
||||
[Service]
|
||||
WorkingDirectory=${RUNNER_DIR}
|
||||
# CI job steps shell out to user-installed tools (sops, age, ...) that live in
|
||||
# ~/.local/bin, not on systemd --user's default PATH. Without this, any
|
||||
# workflow that needs one (e.g. render-secrets.sh once a host is
|
||||
# SOPS-configured) fails with "not installed" even though it plainly is.
|
||||
Environment=PATH=%h/.local/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/usr/games:/usr/local/games:/snap/bin
|
||||
ExecStart=${RUNNER_DIR}/forgejo-runner daemon -c config.yaml
|
||||
Restart=on-failure
|
||||
RestartSec=5
|
||||
|
|
|
|||
|
|
@ -127,6 +127,23 @@ render_thermograph_secrets() {
|
|||
echo "!! cannot write /etc/thermograph.env (need file write access or passwordless sudo)" >&2
|
||||
rc=1
|
||||
fi
|
||||
|
||||
# Optional second copy under $HOME, for a host whose docker CLI cannot read
|
||||
# /etc/thermograph.env at all -- the LAN dev box's snap-packaged Docker is
|
||||
# confined to $HOME (plus a short allowlist) and silently treats anything
|
||||
# under /etc as nonexistent, so env_file: /etc/thermograph.env resolves to
|
||||
# nothing for every container even though the file is right there and the
|
||||
# shell that rendered it can read it fine. deploy-dev.sh sets this;
|
||||
# prod/beta (apt-installed Docker, no snap confinement) leave it unset and
|
||||
# this is a no-op for them.
|
||||
if [ "$rc" = 0 ] && [ -n "${THERMOGRAPH_SECRETS_ENV_FILE_MIRROR:-}" ]; then
|
||||
if ! mkdir -p "$(dirname "$THERMOGRAPH_SECRETS_ENV_FILE_MIRROR")" \
|
||||
|| ! install -m 0600 "$tmp" "$THERMOGRAPH_SECRETS_ENV_FILE_MIRROR"; then
|
||||
echo "!! cannot write mirror at $THERMOGRAPH_SECRETS_ENV_FILE_MIRROR" >&2
|
||||
rc=1
|
||||
fi
|
||||
fi
|
||||
|
||||
rm -f "$tmp"
|
||||
return "$rc"
|
||||
}
|
||||
|
|
|
|||
|
|
@ -27,12 +27,23 @@
|
|||
# prod's backend=4 / frontend=2 / db=2 allocation. `!reset` drops the base
|
||||
# value (both the top-level `cpus:` and the Swarm-style `deploy.resources`
|
||||
# block the base file carries for parity).
|
||||
# 4. backend/daemon/lake get a second env_file entry pointing at a copy of the
|
||||
# render under $APP_DIR (deploy-dev.sh sets THERMOGRAPH_SECRETS_ENV_FILE_MIRROR
|
||||
# to produce it). This box's Docker is the snap package, confined to $HOME --
|
||||
# it can't see /etc/thermograph.env at all (not a permissions error, it just
|
||||
# doesn't exist as far as snap-confined Docker is concerned), so the base
|
||||
# file's env_file entry silently loads nothing here. compose appends env_file
|
||||
# lists across overlays (last-wins on duplicate keys), so this is additive:
|
||||
# prod/beta never set the mirror var, so they only ever get the base entry.
|
||||
services:
|
||||
backend:
|
||||
ports: !override
|
||||
- "8137:8137"
|
||||
cpus: !reset null
|
||||
deploy: !reset null
|
||||
env_file:
|
||||
- path: ./deploy/dev-secrets.env
|
||||
required: false
|
||||
frontend:
|
||||
ports: !reset null
|
||||
cpus: !reset null
|
||||
|
|
@ -42,6 +53,9 @@ services:
|
|||
daemon:
|
||||
cpus: !reset null
|
||||
deploy: !reset null
|
||||
env_file:
|
||||
- path: ./deploy/dev-secrets.env
|
||||
required: false
|
||||
db:
|
||||
cpus: !reset null
|
||||
deploy: !reset null
|
||||
|
|
@ -53,3 +67,6 @@ services:
|
|||
# the default DB_MEMORY (raise DB_MEMORY to give dev more); shm_size stays (it's
|
||||
# required shared memory for parallel query, not a limit).
|
||||
mem_limit: !reset null
|
||||
env_file:
|
||||
- path: ./deploy/dev-secrets.env
|
||||
required: false
|
||||
|
|
|
|||
|
|
@ -46,8 +46,10 @@ services:
|
|||
# to an exact minor (e.g. 2.17.2-pg18) before any host of this stack could ever
|
||||
# replicate with another — a floating tag risks two hosts landing on different
|
||||
# extension minors, which blocks a physical replica and risks compressed-chunk
|
||||
# corruption on restore. docker-stack.yml (the Swarm interim stack) REQUIRES an
|
||||
# exact pin for exactly this reason; use the SAME tag on both once you set one.
|
||||
# corruption on restore. The Swarm path (deploy/stack/) enforces this a
|
||||
# different way: deploy-stack.sh resolves TIMESCALEDB_IMAGE to the digest of
|
||||
# whatever db is ALREADY running -- including this compose stack's
|
||||
# thermograph-db-1 -- so the image under an existing volume can never drift.
|
||||
image: timescale/timescaledb:${TIMESCALEDB_TAG:-latest-pg18}
|
||||
environment:
|
||||
POSTGRES_USER: thermograph
|
||||
|
|
@ -256,7 +258,7 @@ services:
|
|||
restart: unless-stopped
|
||||
|
||||
frontend:
|
||||
# Frontend's OWN image, published by frontend-build-push.yml from
|
||||
# Frontend's OWN image, published by build-push.yml (frontend leg) from
|
||||
# frontend/Dockerfile (which starts the thermograph-frontend Go binary
|
||||
# directly -- no THERMOGRAPH_SERVICE_ROLE process-picking). Independent
|
||||
# FRONTEND_IMAGE_TAG so a frontend deploy never disturbs the backend's tag.
|
||||
|
|
|
|||
|
|
@ -1,173 +0,0 @@
|
|||
# Docker Swarm stack for the hop-1 interim cutover — see
|
||||
# thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md (which this file implements) and
|
||||
# thermograph-docs/architecture/repo-topology-and-infrastructure.md §6/§9.
|
||||
#
|
||||
# Distinct from docker-compose.yml (today's plain-compose deploy, unaffected by
|
||||
# this file): Swarm pulls a pre-built image from the registry (IMAGE_TAG) rather
|
||||
# than building in place, and needs Swarm-specific keys (deploy:, networks:,
|
||||
# secrets:, no host-published ports on the overlay-fronted services).
|
||||
#
|
||||
# docker stack deploy -c docker-stack.yml thermograph
|
||||
#
|
||||
# `docker stack deploy` only interpolates from the invoking shell's environment,
|
||||
# not a .env file — export the required vars first (or `set -a; . ./stack.env;
|
||||
# set +a; docker stack deploy ...`). Required shell vars: IMAGE_TAG (the tag CI
|
||||
# pushed to the registry), TIMESCALEDB_TAG (an EXACT pinned minor — see the db
|
||||
# service below, hazard #7). POSTGRES_PASSWORD is NOT a shell var here: unlike
|
||||
# docker-compose.yml (which interpolates it into THERMOGRAPH_DATABASE_URL at
|
||||
# compose time), this file reads it as the `postgres_password` Swarm secret —
|
||||
# POSTGRES_PASSWORD_FILE for db (native support in the postgres/timescaledb
|
||||
# image), and the chunk-3 entrypoint shim builds THERMOGRAPH_DATABASE_URL for
|
||||
# app/worker from the same secret file at container start. All `postgres_password`
|
||||
# / `thermograph_*` secrets below must already exist in the Swarm (`docker secret
|
||||
# create <name> -`) before the first deploy — Track B step 7 in
|
||||
# thermograph-docs/runbooks/implementation-handoff.md.
|
||||
#
|
||||
# Migrations are NOT run inline by app/worker here (RUN_MIGRATIONS=0) — the
|
||||
# runbook's Stage F brings the schema to head via a one-shot task BEFORE scaling
|
||||
# app up, so multiple replicas never race Alembic and nothing runs DDL against a
|
||||
# still-read-only standby mid-cutover (hazard #9):
|
||||
#
|
||||
# docker run --rm --network thermograph_internal \
|
||||
# -e THERMOGRAPH_DATABASE_URL=... <image> migrate
|
||||
|
||||
services:
|
||||
db:
|
||||
# Pin the EXACT TimescaleDB minor — never latest-pg18 here. A floating tag can
|
||||
# give two hosts different extension minors, which blocks a physical replica
|
||||
# (a newer .so over an older catalog won't start) and risks compressed-chunk
|
||||
# corruption on timescaledb_post_restore() (hazard #7). Get the exact X.Y.Z
|
||||
# from the source database: SELECT extversion FROM pg_extension WHERE
|
||||
# extname='timescaledb' — use the SAME tag docker-compose.yml's TIMESCALEDB_TAG
|
||||
# is pinned to, on every host that could ever replicate with this one.
|
||||
image: timescale/timescaledb:${TIMESCALEDB_TAG:?set TIMESCALEDB_TAG to an exact pinned minor, e.g. 2.17.2-pg18 -- not latest-pg18}
|
||||
environment:
|
||||
POSTGRES_USER: thermograph
|
||||
POSTGRES_PASSWORD_FILE: /run/secrets/postgres_password
|
||||
POSTGRES_DB: thermograph
|
||||
DB_MEMORY: ${DB_MEMORY:-8g}
|
||||
volumes:
|
||||
- pgdata:/var/lib/postgresql
|
||||
- ./deploy/db/init:/docker-entrypoint-initdb.d
|
||||
networks:
|
||||
- internal
|
||||
secrets:
|
||||
- postgres_password
|
||||
deploy:
|
||||
# Pin to the labelled DB node so replication/IO never crosses the slow WG
|
||||
# uplink (hazard #14): `docker node update --label-add db=true <node>` once,
|
||||
# on whichever node holds pgdata (Track B).
|
||||
placement:
|
||||
constraints: ["node.labels.db == true"]
|
||||
resources:
|
||||
limits:
|
||||
cpus: "${DB_CPUS:-2}"
|
||||
memory: ${DB_MEMORY:-8g}
|
||||
restart_policy:
|
||||
condition: on-failure
|
||||
# No published port: reachable only as db:5432 on the `internal` overlay —
|
||||
# Postgres must never be public (design doc §6).
|
||||
|
||||
# Stateless, freely-replicable web tier. ROLE=web means _should_run_notifier()
|
||||
# is False unconditionally (web/app.py) — it never starts the notifier or the
|
||||
# worker scheduler even if it would otherwise win leader election.
|
||||
app:
|
||||
image: ${IMAGE_TAG:?set IMAGE_TAG to the image CI pushed, e.g. forge.example/thermograph/app:sha-abc123}
|
||||
environment:
|
||||
THERMOGRAPH_BASE: /
|
||||
PORT: 8137
|
||||
WORKERS: ${WORKERS:-4}
|
||||
THERMOGRAPH_ROLE: web
|
||||
RUN_MIGRATIONS: "0"
|
||||
volumes:
|
||||
- appdata:/app/data
|
||||
- applogs:/app/logs
|
||||
networks:
|
||||
- internal
|
||||
secrets:
|
||||
- postgres_password
|
||||
- thermograph_auth_secret
|
||||
- thermograph_vapid_private_key
|
||||
- thermograph_vapid_public_key
|
||||
deploy:
|
||||
replicas: 1 # single replica this hop -- homepage.json now lives in
|
||||
# Postgres (chunk 4), but the notifier/scheduler split (chunk
|
||||
# 2/5) is what actually lets this go multi-replica in Phase 2
|
||||
resources:
|
||||
limits:
|
||||
cpus: "${APP_CPUS:-4}"
|
||||
restart_policy:
|
||||
condition: on-failure
|
||||
update_config:
|
||||
order: start-first # new task must pass the image's HEALTHCHECK (now
|
||||
# GET /healthz) before the old one is stopped
|
||||
# No `ports:` — 127.0.0.1:8137:8137 (docker-compose.yml) has no Swarm
|
||||
# equivalent: Swarm's routing mesh publishes on 0.0.0.0, which would expose
|
||||
# the plaintext app un-fronted (hazard #6). Only the host's Caddy reaches
|
||||
# `app`, over the `internal` overlay network — see the Caddyfile templates'
|
||||
# health-gated reverse_proxy.
|
||||
|
||||
# Owns the notifier + worker scheduler (chunks 1/2/5). Exactly one replica —
|
||||
# the leader-election guard is belt-and-suspenders, not a substitute for it.
|
||||
worker:
|
||||
image: ${IMAGE_TAG:?set IMAGE_TAG to the image CI pushed, e.g. forge.example/thermograph/app:sha-abc123}
|
||||
environment:
|
||||
THERMOGRAPH_BASE: /
|
||||
WORKERS: "1"
|
||||
THERMOGRAPH_ROLE: worker
|
||||
# Cluster-wide advisory lock, not the flock: the flock only arbitrates
|
||||
# workers on ONE host, so under Swarm it guards nothing across replicas.
|
||||
THERMOGRAPH_SINGLETON_PG: "1"
|
||||
RUN_MIGRATIONS: "0"
|
||||
volumes:
|
||||
- appdata:/app/data
|
||||
- applogs:/app/logs
|
||||
networks:
|
||||
- internal
|
||||
secrets:
|
||||
- postgres_password
|
||||
- thermograph_auth_secret
|
||||
- thermograph_vapid_private_key
|
||||
- thermograph_vapid_public_key
|
||||
- thermograph_discord_webhook
|
||||
deploy:
|
||||
replicas: 1
|
||||
resources:
|
||||
limits:
|
||||
cpus: "${WORKER_CPUS:-1}"
|
||||
restart_policy:
|
||||
condition: on-failure
|
||||
# No `ports:` — the worker serves no public traffic; GET /healthz on its own
|
||||
# container port is for the image's own HEALTHCHECK (Swarm task health) only.
|
||||
|
||||
networks:
|
||||
internal:
|
||||
driver: overlay
|
||||
driver_opts:
|
||||
# VXLAN-over-WireGuard double encapsulation needs a lower MTU than the
|
||||
# ~1450 overlay default, or large payloads (a /cell bundle, the homepage
|
||||
# feed) silently stall while small packets (health checks) keep passing
|
||||
# (hazard #13). Validate with a real large-payload transfer, not just a
|
||||
# ping, once the mesh is up (Track B).
|
||||
com.docker.network.driver.mtu: "1370"
|
||||
|
||||
volumes:
|
||||
pgdata: {}
|
||||
appdata: {}
|
||||
applogs: {}
|
||||
|
||||
# Declared here, provisioned externally (docker secret create <name> - < file) —
|
||||
# never by this stack, and never committed. See Stage 0 of the cutover runbook
|
||||
# for which values are continuity-critical (AUTH_SECRET, VAPID, POSTGRES_PASSWORD
|
||||
# must be the EXISTING live values, not freshly generated ones).
|
||||
secrets:
|
||||
postgres_password:
|
||||
external: true
|
||||
thermograph_auth_secret:
|
||||
external: true
|
||||
thermograph_vapid_private_key:
|
||||
external: true
|
||||
thermograph_vapid_public_key:
|
||||
external: true
|
||||
thermograph_discord_webhook:
|
||||
external: true
|
||||
|
|
@ -1,55 +1,52 @@
|
|||
# thermograph-observability — agent instructions
|
||||
# observability/ — agent instructions
|
||||
|
||||
The **logging/metrics stack** for the Thermograph fleet: Loki + Grafana on beta,
|
||||
with a Grafana Alloy agent on every node (prod, beta, dev) shipping container and
|
||||
app logs over the WireGuard mesh. Grafana is fronted by beta's Caddy at
|
||||
`dashboard.thermograph.org` (Google SSO, pre-provisioned users only).
|
||||
The **logging stack** for the fleet: Loki + Grafana on beta, with a Grafana Alloy
|
||||
agent on every node (prod, beta, dev) shipping container and app logs over the
|
||||
WireGuard mesh. Grafana is fronted by beta's Caddy at
|
||||
**`dashboard.thermograph.org`** (Google SSO, pre-provisioned users only) — use
|
||||
that hostname everywhere, never `grafana.thermograph.org`.
|
||||
|
||||
This is **operational infra config, not application code** — there is no build.
|
||||
It deploys by hand: `docker compose up -d` on beta for Loki+Grafana, and the
|
||||
Alloy agent (`alloy/docker-compose.agent.yml`) on each node. See the README for
|
||||
the full per-node procedure.
|
||||
This is operational config, not application code: there is **no build** and no
|
||||
deploy automation. It ships by hand — `docker compose up -d` on beta for
|
||||
Loki+Grafana, and `alloy/docker-compose.agent.yml` on each node. Read the root
|
||||
`CLAUDE.md` first; the branch model there applies here too.
|
||||
|
||||
## Layout
|
||||
|
||||
- `docker-compose.yml` — the Loki + Grafana stack (runs on beta).
|
||||
- `loki/config.yml` — Loki config (mesh-only, filesystem storage).
|
||||
- `grafana/provisioning/` — datasource (Loki) + dashboard provider, auto-loaded
|
||||
at startup. `grafana/dashboards/*.json` — the dashboards themselves.
|
||||
- `grafana/provisioning/alerting/` — the alert rules, the Discord contact point
|
||||
and the notification policy, also auto-loaded at startup. Same rule as the
|
||||
dashboards: **the repo is the only durable path**, UI edits get overwritten.
|
||||
Alerts go to Discord `#ops-alerts`, never email — beta's Grafana relays SMTP
|
||||
through prod's Postfix, so email dies exactly when prod does. Every rule is
|
||||
LogQL (there is no Prometheus anywhere in the fleet). Thresholds were derived
|
||||
from real Loki data and the working is in the comments beside each rule —
|
||||
re-derive before changing a number rather than guessing.
|
||||
- `alloy/config.alloy` + `alloy/docker-compose.agent.yml` — the per-node log
|
||||
shipper. The node name comes from `ALLOY_NODE` (set per host).
|
||||
- `caddy-grafana.conf` — the Caddy vhost for `dashboard.thermograph.org` (lives
|
||||
in beta's Caddy config, kept here for reference).
|
||||
- `.env.example` — copy to `.env` on beta before `docker compose up`. `.env` is
|
||||
gitignored; never commit real OAuth secrets or the admin password.
|
||||
- `docker-compose.yml` — the Loki + Grafana stack (beta).
|
||||
- `loki/config.yml` — mesh-only, filesystem storage.
|
||||
- `grafana/provisioning/` — datasource + dashboard provider, auto-loaded at
|
||||
startup. `grafana/dashboards/*.json` — the dashboards.
|
||||
- `grafana/provisioning/alerting/` — alert rules, the Discord contact point, and
|
||||
the notification policy.
|
||||
- `alloy/config.alloy` + `alloy/docker-compose.agent.yml` — the per-node shipper.
|
||||
Node name comes from `ALLOY_NODE`, set per host.
|
||||
- `caddy-grafana.conf` — the beta Caddy vhost, kept here for reference.
|
||||
- `.env.example` — copy to `.env` on beta. `.env` is gitignored; never commit
|
||||
OAuth secrets or the admin password.
|
||||
|
||||
## Conventions
|
||||
## Rules
|
||||
|
||||
- **The repo is the only durable path.** Dashboards and alerting are provisioned
|
||||
from this directory at startup; edits made through Grafana's UI or API are
|
||||
overwritten on the next provision. Changes go through a PR.
|
||||
- **Every artifact that ships to a node is CI-validated**
|
||||
(`.forgejo/workflows/observability-validate.yml`): both compose files parse, all
|
||||
dashboard JSON is valid, and the Loki/provisioning YAML parses. Keep new
|
||||
dashboards as valid JSON and new config as valid YAML or CI fails. The alerting
|
||||
config gets a stricter third step — every rule's `condition` must name a refId
|
||||
that exists, every policy must route to a receiver that exists, and a literal
|
||||
Discord webhook URL in the repo is a hard failure. (A rule pointing at a missing
|
||||
refId is valid YAML, provisions cleanly, and then never fires; that is precisely
|
||||
the silent-no-op this whole domain exists to prevent.) The Alloy config is still
|
||||
not CI-validated (needs the `alloy` binary).
|
||||
- The public hostname is **`dashboard.thermograph.org`** everywhere — not
|
||||
`grafana.thermograph.org`. Match it in any new comment/config.
|
||||
- Secrets (OAuth client id/secret, admin password) live only in the host `.env`,
|
||||
never in the repo.
|
||||
(`.forgejo/workflows/observability-validate.yml`, at the repo root): both
|
||||
compose files parse, all dashboard JSON is valid, the Loki and provisioning
|
||||
YAML parse, and the Alloy config is checked with the pinned `alloy` binary
|
||||
(v1.9.1, matching what the fleet runs).
|
||||
- **Alerting gets a stricter third step.** Every rule's `condition` must name a
|
||||
refId that exists, every policy must route to a receiver that exists, and a
|
||||
literal Discord webhook URL in the repo is a hard failure. A rule pointing at a
|
||||
missing refId is valid YAML, provisions cleanly, and then never fires — that
|
||||
silent no-op is what this check exists to prevent.
|
||||
- **Alerts go to Discord `#ops-alerts`, never email.** Beta's Grafana relays SMTP
|
||||
through prod's Postfix, so email dies exactly when prod does.
|
||||
- **Every rule is LogQL** — there is no Prometheus anywhere in the fleet.
|
||||
Thresholds were derived from real Loki data and the working is in the comments
|
||||
beside each rule; re-derive before changing a number rather than guessing.
|
||||
- Secrets (OAuth client id/secret, admin password) live only in the host `.env`.
|
||||
|
||||
## Branching
|
||||
## Commits & PRs
|
||||
|
||||
Single `main` branch, no protection. Open a PR against `main`; the validator
|
||||
gates it. Deploying the change to beta / the nodes is a separate manual step
|
||||
(this repo has no deploy automation — a known gap).
|
||||
Concise and technical. Never mention AI, assistants or automated authorship.
|
||||
|
|
|
|||
Loading…
Reference in a new issue