promote: main → release (#93)
All checks were successful
secrets-guard / encrypted (push) Successful in 18s
shell-lint / shellcheck (push) Successful in 16s
Build + push images (Forgejo registry) / build-push (backend) (push) Successful in 1m27s
Build + push images (Forgejo registry) / build-push (frontend) (push) Successful in 1m25s
Deploy / deploy (backend) (push) Successful in 1m31s
Deploy / deploy (frontend) (push) Successful in 1m45s

Reviewed-on: https://git.thermograph.org/emi/thermograph/pulls/93
This commit is contained in:
emi 2026-07-25 17:24:10 +00:00
commit 90f699ceaf
37 changed files with 1243 additions and 1372 deletions

84
.claude/hooks/README.md Normal file
View file

@ -0,0 +1,84 @@
# .claude/hooks — enforcement, not advice
`CLAUDE.md` files can only ask. These hooks are the part that enforces, and they
travel with the repo so a session started anywhere in this checkout gets them.
They are wired up in `.claude/settings.json`.
| Hook | Event | What it does |
|---|---|---|
| `prod-guard.sh` | PreToolUse | Live-host commands: reads run unprompted, mutations ask first |
| `secrets-guard.sh` | PreToolUse | Denies any direct write to `infra/deploy/secrets/*.yaml` |
| `lint-after-edit.sh` | PostToolUse | shellchecks an edited `*.sh` and feeds findings back immediately |
## Why prod-guard exists
`agent` has passwordless sudo on prod and beta, and the operator's global settings
allow `Bash(ssh prod:*)` outright. Nothing about `ssh prod 'docker service rm …'`
goes through a pull request, so CI cannot see it — before this hook, any session in
any directory could delete production with no prompt.
**It classifies by allowlist, not blocklist.** Only commands positively recognised as
read-only are allowed through; everything else asks. A blocklist of dangerous verbs
is wrong by construction — the first destructive verb nobody thought of sails
straight through. Being wrong in the "ask" direction costs a keystroke; being wrong
the other way costs production.
Beta is guarded as strictly as prod: it serves beta.thermograph.org *and* hosts
Forgejo, so a destructive command there takes out git, CI and the registry at once.
The LAN dev box is deliberately unguarded.
## Changing the classifier
`is_readonly()` in `prod-guard.sh` splits a command on `&&`, `||`, `;` and `|`, then
checks each segment's leading word. Redirection to a file, command substitution,
`sed -i`, and mutating `docker`/`git`/`systemctl` subcommands all count as writes.
If you add a command to the allowlist, **re-run the test matrix**. There is a
non-obvious failure mode this already hit once: the loop is fed by
`printf '%s\n' | sed`, and dropping that trailing newline makes `read` return
non-zero on the only line, so the loop body never runs and *every* command
classifies as read-only — a silent, total bypass that still looks like it works.
Test by piping a payload straight in:
```bash
echo '{"tool_name":"Bash","tool_input":{"command":"ssh prod '\''rm -rf /'\''"}}' \
| .claude/hooks/prod-guard.sh
# expect: permissionDecision "ask"
```
Exit 0 with no output means "no opinion" and the normal permission flow proceeds.
That is also what happens if `jq` is missing or the script errors, so a bug here
fails toward the prompt rather than toward silent execution.
## lint-after-edit needs shellcheck on PATH
It exits quietly when shellcheck is absent — a missing local tool must not break
editing, and `shell-lint` CI is still the backstop. v0.11.0 (the version CI pins) is
installed at `~/.local/bin/shellcheck`; if you move machines, reinstall it:
```bash
VER=v0.11.0
curl -fsSL "https://github.com/koalaman/shellcheck/releases/download/${VER}/shellcheck-${VER}.linux.x86_64.tar.xz" \
| tar -xJ --strip-components=1 -C ~/.local/bin "shellcheck-${VER}/shellcheck"
```
Keep it matched to `shell-lint.yml`'s pin. A drifted shellcheck grows *new* warnings
and fails CI on an unrelated push, which is why that workflow pins the version and
its sha256 rather than using `apt-get install`.
## The global copy
`~/.claude/settings.json` runs `prod-guard.sh` from `~/.claude/hooks/` as well, so a
session started **outside** this repo is covered too — otherwise the guard would be
trivially bypassed by working from `~/Code`. That copy is a mirror; this one is
canonical. If you change the classifier here, copy it over:
```bash
cp .claude/hooks/prod-guard.sh ~/.claude/hooks/prod-guard.sh
```
The blanket `Bash(ssh prod:*)` / `Bash(ssh beta:*)` allow rules were removed from
global settings at the same time, so the hook is the only thing granting live-host
access. If it is missing or broken, nothing allows the command and you get a prompt.

View file

@ -0,0 +1,34 @@
#!/usr/bin/env bash
# PostToolUse: shellcheck a shell script the moment it is edited.
#
# The scripts under infra/deploy/ run as root over SSH against live hosts with no
# test suite in front of them — a quoting bug or an unset-variable typo is found on
# the host, or not at all. shell-lint.yml already guards this, but only after a
# push: a five-minute round trip to learn about a missing quote, by which point the
# context that produced it is gone.
#
# Findings are fed back to the model rather than blocking the edit. The file is
# already written; the useful move is to fix it now, in the same turn.
#
# Matches shell-lint.yml's pinned shellcheck when one is on PATH. If shellcheck is
# not installed this exits quietly — CI is still the backstop, and a missing local
# tool must not break editing.
set -uo pipefail
payload=$(cat)
command -v jq >/dev/null 2>&1 || exit 0
command -v shellcheck >/dev/null 2>&1 || exit 0
path=$(printf '%s' "$payload" | jq -r '.tool_response.filePath // .tool_input.file_path // ""')
[ -n "$path" ] || exit 0
case "$path" in *.sh) ;; *) exit 0 ;; esac
[ -f "$path" ] || exit 0
if out=$(shellcheck -f gcc "$path" 2>&1); then
exit 0
fi
# Keep the feedback small: findings only, capped.
out=$(printf '%s' "$out" | head -30)
jq -n --arg p "$path" --arg o "$out" \
'{decision:"block", reason:("shellcheck findings in \($p) — shell-lint CI will fail on these, so fix them now:\n\($o)")}'

156
.claude/hooks/prod-guard.sh Executable file
View file

@ -0,0 +1,156 @@
#!/usr/bin/env bash
# PreToolUse guard for commands that reach a live host.
#
# Policy: reads stay friction-free, mutations ask first.
#
# The estate's widest privilege is that `agent` has passwordless sudo on prod and
# beta, and the global settings allow `ssh prod:*` outright — so before this hook
# existed, any session in any directory could roll, delete or restart production
# with no prompt. CI cannot help here: nothing about `ssh prod 'docker service rm'`
# goes through a pull request.
#
# Classification is an ALLOWLIST of read-only commands, not a blocklist of
# dangerous ones. A blocklist is wrong by construction: the first destructive verb
# nobody thought of sails straight through. Anything this script does not
# positively recognise as a read becomes an "ask", which is the safe direction to
# be wrong in.
#
# Exit 0 with no output = "no opinion", and the normal permission flow proceeds.
# That is also what happens if this script errors or jq is missing, so a bug here
# fails toward the prompt rather than toward silent execution.
#
# Beta is guarded as strictly as prod: it serves beta.thermograph.org AND hosts
# Forgejo, so a destructive command there can take out the git server, CI and the
# container registry at once.
set -uo pipefail
payload=$(cat)
command -v jq >/dev/null 2>&1 || exit 0
tool=$(printf '%s' "$payload" | jq -r '.tool_name // ""')
# An ssh target naming a live host, matched as a whole token after any `user@`.
# Backslashes are doubled because awk -v processes escape sequences in the value
# before the regex ever sees it; a single \. arrives as a bare dot (and warns).
readonly HOST_TOKEN_RE='^(prod|beta|169\\.58\\.46\\.181|75\\.119\\.132\\.91|10\\.10\\.0\\.[12])$'
emit() { # $1=allow|ask|deny $2=reason
jq -n --arg d "$1" --arg r "$2" \
'{hookSpecificOutput:{hookEventName:"PreToolUse",permissionDecision:$d,permissionDecisionReason:$r}}'
exit 0
}
# --- is this shell command read-only? ----------------------------------------
# Returns 0 only when every segment leads with a recognised read-only command.
is_readonly() {
local cmd="$1" seg w1 w2 stripped
# Discard the harmless redirections that appear in ordinary read commands,
# then treat any surviving redirect as a write.
stripped=${cmd//2>&1/}
stripped=${stripped//&>\/dev\/null/}
stripped=${stripped//2>\/dev\/null/}
stripped=${stripped//>\/dev\/null/}
stripped=${stripped//> \/dev\/null/}
case "$stripped" in *'>'*) return 1 ;; esac
# Command substitution can hide anything; refuse to reason about it.
# shellcheck disable=SC2016 # single quotes are the point: these are literal
# `$(` and backtick characters being searched for inside the command text, not
# expansions to perform. Expanding them here would defeat the check entirely.
case "$cmd" in *'$('*|*'`'*) return 1 ;; esac
while IFS= read -r seg; do
seg="${seg#"${seg%%[![:space:]]*}"}"
[ -z "$seg" ] && continue
w1=$(printf '%s' "$seg" | awk '{print $1}')
w2=$(printf '%s' "$seg" | awk '{print $2}')
case "$w1" in
cat|ls|head|tail|wc|grep|egrep|fgrep|rg|find|stat|df|du|free|uptime|whoami) ;;
id|hostname|date|ps|env|printenv|pwd|which|uname|lsblk|ip|ss|netstat) ;;
pg_isready|echo|true|sort|uniq|cut|tr|awk|jq|nproc|readlink|realpath|test) ;;
sed) case "$seg" in *' -i'*) return 1 ;; esac ;;
curl) case "$seg" in *' -X'*|*' -d'*|*'--data'*|*' -T'*|*' -F'*|*'--upload'*) return 1 ;; esac ;;
journalctl) ;;
systemctl)
case "$w2" in status|show|is-active|is-enabled|is-failed|list-units|cat) ;; *) return 1 ;; esac ;;
docker)
case "$w2" in
ps|logs|inspect|stats|images|version|info|top|port|events|history|diff) ;;
service) case "$seg" in *'service ls'*|*'service ps'*|*'service inspect'*|*'service logs'*) ;; *) return 1 ;; esac ;;
stack) case "$seg" in *'stack ls'*|*'stack ps'*|*'stack services'*) ;; *) return 1 ;; esac ;;
compose) case "$seg" in *'compose ps'*|*'compose logs'*|*'compose config'*|*'compose top'*|*'compose images'*) ;; *) return 1 ;; esac ;;
node) case "$seg" in *'node ls'*|*'node inspect'*) ;; *) return 1 ;; esac ;;
volume) case "$seg" in *'volume ls'*|*'volume inspect'*) ;; *) return 1 ;; esac ;;
*) return 1 ;;
esac ;;
git)
case "$w2" in status|log|diff|show|branch|remote|rev-parse|describe|blame|ls-files|config) ;; *) return 1 ;; esac ;;
*) return 1 ;;
esac
# NOTE: the trailing newline matters. Without it `read` returns non-zero on the
# final (only) line, the loop body never runs, and every command classifies as
# read-only — a silent false negative that allows anything.
done < <(printf '%s\n' "$cmd" | sed 's/&&/\n/g; s/||/\n/g; s/;/\n/g; s/|/\n/g')
return 0
}
# --- find an ssh target and return what would run on it -----------------------
# Prints "<host>\t<remote command>" when a live host is addressed, nothing
# otherwise. The remote command is empty for a bare login shell.
ssh_target_of() {
printf '%s' "$1" | awk -v re="$HOST_TOKEN_RE" '
{
for (i = 1; i <= NF; i++) {
t = $i; sub(/^[^@]*@/, "", t)
if (t ~ re) {
out = ""
for (j = i + 1; j <= NF; j++) out = out (out == "" ? "" : " ") $j
printf "%s\t%s\n", t, out
exit
}
}
}'
}
case "$tool" in
Bash)
cmd=$(printf '%s' "$payload" | jq -r '.tool_input.command // ""')
case "$cmd" in *ssh*) ;; *) exit 0 ;; esac
found=$(ssh_target_of "$cmd")
[ -n "$found" ] || exit 0 # no live host addressed
host=${found%%$'\t'*}
remote=${found#*$'\t'}
remote=$(printf '%s' "$remote" | sed -E "s/^'(.*)'$/\1/; s/^\"(.*)\"$/\1/")
[ -z "$remote" ] && exit 0 # bare login shell
if is_readonly "$remote"; then
emit allow "read-only command on ${host}"
fi
emit ask "This mutates ${host}, a live host where the agent user has passwordless sudo. Command: ${remote}"
;;
mcp__centralis__run_on_host)
env=$(printf '%s' "$payload" | jq -r '.tool_input.env // ""')
cmd=$(printf '%s' "$payload" | jq -r '.tool_input.command // ""')
case "$env" in prod|beta) ;; *) exit 0 ;; esac
if is_readonly "$cmd"; then
emit allow "read-only command on ${env}"
fi
emit ask "run_on_host would mutate ${env}. Command: ${cmd}"
;;
mcp__centralis__sql_query)
env=$(printf '%s' "$payload" | jq -r '.tool_input.env // ""')
write=$(printf '%s' "$payload" | jq -r '.tool_input.write // false')
[ "$write" = "true" ] || exit 0
case "$env" in prod|beta) ;; *) exit 0 ;; esac
sql=$(printf '%s' "$payload" | jq -r '.tool_input.sql // ""' | tr '\n' ' ' | cut -c1-300)
emit ask "Write query against the ${env} database. sql_query escalates to the app role and has no confirmation of its own. SQL: ${sql}"
;;
mcp__centralis__rollback_to|mcp__centralis__promote|mcp__centralis__secrets_rotate)
emit ask "${tool#mcp__centralis__} changes production state directly."
;;
*) exit 0 ;;
esac

34
.claude/hooks/secrets-guard.sh Executable file
View file

@ -0,0 +1,34 @@
#!/usr/bin/env bash
# PreToolUse guard: never write infra/deploy/secrets/*.yaml directly.
#
# Every file matching that glob is SOPS-encrypted, and secrets-guard.yml fails the
# build if one is not. This is the same rule enforced one step earlier: a direct
# Write or Edit produces plaintext, and by the time CI catches it the secret has
# already been written to disk and probably committed.
#
# The supported path is `sops edit <file>`, which decrypts to a temp file, opens an
# editor and re-encrypts on save — the plaintext never touches the working tree.
#
# Denies rather than asks: there is no legitimate reason to hand-write one of these,
# so a prompt would only be an invitation to click through.
set -uo pipefail
payload=$(cat)
command -v jq >/dev/null 2>&1 || exit 0
path=$(printf '%s' "$payload" | jq -r '.tool_input.file_path // ""')
[ -n "$path" ] || exit 0
case "$path" in
*/infra/deploy/secrets/*.yaml)
jq -n --arg p "$path" '{
hookSpecificOutput: {
hookEventName: "PreToolUse",
permissionDecision: "deny",
permissionDecisionReason: ("Refusing to write \($p) directly — every file under infra/deploy/secrets/ is SOPS-encrypted and a direct write would land plaintext. Use `sops edit \($p)` instead, which re-encrypts on save. (secrets-guard.yml would fail the build on this anyway, but only after the secret was already on disk.)")
}
}'
exit 0
;;
*) exit 0 ;;
esac

42
.claude/settings.json Normal file
View file

@ -0,0 +1,42 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"hooks": {
"PreToolUse": [
{
"matcher": "Bash|mcp__centralis__run_on_host|mcp__centralis__sql_query|mcp__centralis__rollback_to|mcp__centralis__promote|mcp__centralis__secrets_rotate",
"hooks": [
{
"type": "command",
"command": "\"${CLAUDE_PROJECT_DIR:-.}/.claude/hooks/prod-guard.sh\"",
"timeout": 10,
"statusMessage": "Checking live-host access"
}
]
},
{
"matcher": "Write|Edit",
"hooks": [
{
"type": "command",
"command": "\"${CLAUDE_PROJECT_DIR:-.}/.claude/hooks/secrets-guard.sh\"",
"timeout": 10,
"statusMessage": "Checking SOPS vault"
}
]
}
],
"PostToolUse": [
{
"matcher": "Write|Edit",
"hooks": [
{
"type": "command",
"command": "\"${CLAUDE_PROJECT_DIR:-.}/.claude/hooks/lint-after-edit.sh\"",
"timeout": 30,
"statusMessage": "shellcheck"
}
]
}
]
}
}

View file

@ -1,126 +0,0 @@
name: Build + push backend image (Forgejo registry)
# Builds the BACKEND domain's image (backend/Dockerfile, context backend/) and
# pushes it to Forgejo's built-in container registry, tagged by git SHA (every
# push) and semver (version tags) -- build-once, deploy-everywhere.
#
# Monorepo port (reunification): the split repos each derived IMAGE_PATH from
# github.repository, which now collides for every domain -- so each domain's
# build-push names its image explicitly: this one is emi/thermograph/backend
# (frontend-build-push.yml publishes emi/thermograph/frontend). infra's
# compose/stack files reference the two paths independently
# (BACKEND_IMAGE_PATH / FRONTEND_IMAGE_PATH -- their defaults were renamed to
# match in the same cutover).
#
# Path-filtered: only a push that touches backend/** builds this image. A
# frontend-only push builds nothing here; deploy.sh's persisted
# .image-tags.env keeps the backend's live tag for the sibling's deploy.
# EXCEPTION: version-tag pushes (v*.*.*) carry no paths context, so a semver
# tag builds BOTH domain images -- intended, a release names the whole set.
#
# Runs on the `docker` label -- the LAN runner's job containers get the host's
# docker.sock automounted in (forgejo-runner config: container.docker_host:
# automount; see infra/deploy/forgejo/register-lan-runner.sh),
# Docker-outside-of-Docker rather than DinD -- no privileged mode needed.
# The job image itself (node:20-bookworm) has no docker CLI, hence the
# "Install Docker CLI" step.
#
# The registry is deliberately mesh-only: git.thermograph.org's public DNS
# resolves to beta's public IP, which Caddy's ACL rejects for /v2/ paths, so
# any runner host that pushes/pulls needs a /etc/hosts entry pointing
# git.thermograph.org at beta's WireGuard IP (10.10.0.2). The docker
# build/push commands run against the automounted HOST daemon, so it's the
# runner HOST's own resolver that has to route over the mesh, not anything
# settable from inside the job container.
#
# REGISTRY is the vars.REGISTRY_HOST Actions variable (git.thermograph.org),
# NOT github.server_url -- that context resolves to the bare mesh IP:port
# (http://10.10.0.2:3080, no TLS), which docker CLI refuses ("server gave HTTP
# response to HTTPS client"). git.thermograph.org gets real TLS via Caddy and
# routes over the mesh once /etc/hosts is set, so it's the only correct value.
# deploy.sh resolves the SHA tag at deploy time and `docker compose pull`s it.
on:
push:
branches: [dev, main, release]
paths: ['backend/**']
tags: ['v*.*.*']
workflow_dispatch: {}
# The runner has capacity: 1 (one physical LAN box) -- a burst of pushes to the
# same ref would otherwise queue whole builds back to back even though only the
# last one still matters. Keying on domain + github.ref means each domain's
# dev/main/release/tag builds get their own group, so this only cancels a
# superseded build of the SAME domain on the SAME ref -- and a frontend build
# never cancels a backend one.
concurrency:
group: backend-build-push-${{ github.ref }}
cancel-in-progress: true
jobs:
build-push:
runs-on: docker
env:
REGISTRY: https://${{ vars.REGISTRY_HOST }}
IMAGE_PATH: ${{ github.repository }}/backend
steps:
- uses: actions/checkout@v4
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
# often an unrelated domain's commit, and a depth-1 clone can't see
# past it to find the real key.
with:
fetch-depth: 0
- name: Install Docker CLI
run: |
apt-get update -qq
apt-get install -y -qq docker.io
- name: Compute image ref + tags
id: image
run: |
host="${REGISTRY#https://}"; host="${host#http://}"
path="$(printf '%s' "$IMAGE_PATH" | tr '[:upper:]' '[:lower:]')"
ref="$host/$path"
# Key the tag to the last commit that touched backend/ -- the SAME key
# every backend deploy workflow computes -- so build and deploy always
# agree even when the branch tip is another domain's commit. (A tag
# keyed to the push tip breaks the moment an infra-only commit lands:
# deploys go looking for an image no build ever produced.)
domain_sha="$(git log -1 --format=%H -- backend/)"
sha_tag="$ref:sha-${domain_sha:0:12}"
echo "host=$host" >> "$GITHUB_OUTPUT"
echo "sha_tag=$sha_tag" >> "$GITHUB_OUTPUT"
if [[ "$GITHUB_REF" == refs/tags/v* ]]; then
echo "semver_tag=$ref:${GITHUB_REF#refs/tags/}" >> "$GITHUB_OUTPUT"
fi
- name: Log in to the Forgejo registry
run: |
# NOT secrets.GITHUB_TOKEN -- Forgejo's per-job auto-injected token
# can never push to the container registry (a known Forgejo
# limitation independent of any permission setting). Needs a
# manually provisioned PAT with write:package scope, stored as the
# REGISTRY_TOKEN secret.
echo "${{ secrets.REGISTRY_TOKEN }}" | docker login "${{ steps.image.outputs.host }}" \
--username emi --password-stdin
- name: Build
run: |
# One build, both tags -- the semver tag (tag pushes only) rides
# along with the SHA tag instead of a second tag+push round trip.
# Context is the DOMAIN dir, so backend/Dockerfile sees the same
# tree it saw as a repo root in the split era.
tag_args=(-t "${{ steps.image.outputs.sha_tag }}")
if [[ -n "${{ steps.image.outputs.semver_tag }}" ]]; then
tag_args+=(-t "${{ steps.image.outputs.semver_tag }}")
fi
docker build "${tag_args[@]}" backend/
- name: Push (SHA tag, every push)
run: docker push "${{ steps.image.outputs.sha_tag }}"
- name: Push (semver tag, tag pushes only)
if: startsWith(github.ref, 'refs/tags/v')
run: docker push "${{ steps.image.outputs.semver_tag }}"

View file

@ -1,79 +0,0 @@
name: Deploy backend to LAN dev server
# Fires on a push to `dev` touching backend/** -- a direct push, or the real
# merge commit a native "auto merge when checks succeed" produces (see
# pr-build.yml). A `build` job (the reusable build.yml sanity check) gates a
# `deploy` job that rolls the LAN dev box.
#
# CUTOVER NOTE (monorepo reunification): the LAN box's ~/thermograph-dev is a
# thermograph-infra checkout from the split era. This job calls
# ~/thermograph-dev/infra/deploy/deploy-dev.sh -- the MONOREPO layout -- so it
# is INERT until ~/thermograph-dev is re-pointed at the monorepo (a fresh
# clone; deploy-dev.sh then self-updates via deploy.sh's git reset like the
# beta/prod paths).
#
# SERVICE=backend + BACKEND_IMAGE_TAG is the same env-var contract the
# beta/prod deploy workflows use against infra/deploy/deploy.sh, so a dev-only
# rollout never touches the frontend's running container or tag.
on:
push:
branches: [dev]
paths: ['backend/**']
workflow_dispatch: {}
jobs:
build:
name: build
# Fixed (unparameterized) group: if two backend pushes to dev land close
# together, the newer push's build cancels the older's still-running one
# on the single physical runner -- the older result would be thrown away
# once needs:build gates its deploy anyway. Safe to cancel mid-run:
# build.yml only runs a throwaway `docker build`, nothing persisted.
concurrency:
group: dev-push-build-backend
cancel-in-progress: true
uses: ./.forgejo/workflows/build.yml
with:
domain: backend
deploy:
needs: build
# NOT [self-hosted, thermograph-lan] -- GitHub implicitly tags every
# self-hosted runner with "self-hosted"; Forgejo's runner has only the
# labels it was registered with ("docker" + "thermograph-lan"). The array
# form silently makes this job unschedulable on every push.
runs-on: thermograph-lan
permissions:
contents: read
concurrency:
group: dev-lan-deploy-backend
cancel-in-progress: false
steps:
- uses: actions/checkout@v4
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
# often an unrelated domain's commit, and a depth-1 clone can't see
# past it to find the real key.
with:
fetch-depth: 0
- name: Compute image tag
id: tag
# Keyed to the last commit that touched backend/ -- must match
# backend-build-push.yml's tag key exactly (see its comment).
run: |
domain_sha="$(git log -1 --format=%H -- backend/)"
echo "tag=sha-${domain_sha:0:12}" >> "$GITHUB_OUTPUT"
- name: Deploy to the LAN dev server
# thermograph-lan is host-native (the runner job runs directly on the
# LAN box, not over SSH like beta/prod), so plain env vars on the
# command reach deploy-dev.sh directly -- no ssh-action needed here.
# The lake creds ride the Actions S3 secrets: dev renders no vault
# (deploy-dev.sh's no-op secrets path), and compose interpolates
# ${THERMOGRAPH_LAKE_S3_*} straight from the deploy environment.
env:
THERMOGRAPH_LAKE_S3_ACCESS_KEY: ${{ secrets.S3_ACCESS_KEY }}
THERMOGRAPH_LAKE_S3_SECRET_KEY: ${{ secrets.S3_SECRET_KEY }}
run: SERVICE=backend BACKEND_IMAGE_TAG=${{ steps.tag.outputs.tag }} bash "$HOME/thermograph-dev/infra/deploy/deploy-dev.sh"

View file

@ -1,56 +0,0 @@
name: Deploy backend to prod VPS
# On a push to `release` that touches backend/**, SSH to prod and roll ONLY
# the backend service onto the image backend-build-push.yml published for this
# commit. Same shape as backend-deploy.yml (the `main`->beta path) against a
# completely separate secret set (PROD_SSH_HOST/USER/KEY/PORT) so a beta
# credential leak can't reach prod. Frontend deploys itself independently via
# frontend-deploy-prod.yml -- a backend change ships to prod without touching
# frontend and vice versa, exactly as in the split era.
#
# The tag MUST match backend-build-push.yml's last-backend-commit key (12
# hex) -- see backend-deploy.yml for why it's truncated here. deploy.sh
# retries the pull because the build for this push may still be in flight.
on:
push:
branches: [release]
paths: ['backend/**']
workflow_dispatch: {}
concurrency:
group: prod-deploy-backend
cancel-in-progress: false
jobs:
deploy:
runs-on: docker
steps:
- uses: actions/checkout@v4
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
# often an unrelated domain's commit, and a depth-1 clone can't see
# past it to find the real key.
with:
fetch-depth: 0
- name: Compute image tag (last backend-touching commit)
id: tag
run: |
domain_sha="$(git log -1 --format=%H -- backend/)"
echo "tag=sha-${domain_sha:0:12}" >> "$GITHUB_OUTPUT"
- name: Deploy backend over SSH
uses: https://github.com/appleboy/ssh-action@v1.2.0
with:
host: ${{ secrets.PROD_SSH_HOST }}
username: ${{ secrets.PROD_SSH_USER }}
key: ${{ secrets.PROD_SSH_KEY }}
port: ${{ secrets.PROD_SSH_PORT }}
# Forward the target service + the image tag to run. deploy.sh reads
# BACKEND_IMAGE_TAG and rolls just `backend`.
envs: SERVICE,BACKEND_IMAGE_TAG
script: /opt/thermograph/infra/deploy/deploy.sh
env:
SERVICE: backend
BACKEND_IMAGE_TAG: ${{ steps.tag.outputs.tag }}

View file

@ -1,65 +0,0 @@
name: Deploy backend to beta VPS
# On a push to `main` that touches backend/**, SSH to beta and roll ONLY the
# backend service onto the image backend-build-push.yml just published for
# this commit. Frontend is deployed independently by frontend-deploy.yml --
# the FE/BE CI-CD split survives reunification intact: a backend change ships
# without touching frontend and vice versa. infra/deploy/deploy.sh persists
# each service's live tag host-side, so a single-service roll never disturbs
# the other. An infra-only push rolls nothing (infra-sync.yml syncs the
# host checkouts instead); a mixed backend+frontend push fires both deploy
# workflows, serialized host-side by deploy.sh's flock.
#
# The SSH_HOST/USER/KEY/PORT secrets point at beta. appleboy/ssh-action is
# referenced by full GitHub URL because it isn't mirrored in Forgejo's default
# action registry.
#
# The tag MUST match backend-build-push.yml's `sha-${GITHUB_SHA:0:12}` (12
# hex), not the full 40-char ${{ github.sha }} -- default Actions expressions
# have no substring function, so the compute step below truncates it.
# deploy.sh also retries the pull for ~5 min because Forgejo `needs:` can't
# gate across the separate build-push workflow, so this push's build may still
# be in flight.
on:
push:
branches: [main]
paths: ['backend/**']
workflow_dispatch: {}
concurrency:
group: beta-deploy-backend
cancel-in-progress: false
jobs:
deploy:
runs-on: docker
steps:
- uses: actions/checkout@v4
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
# often an unrelated domain's commit, and a depth-1 clone can't see
# past it to find the real key.
with:
fetch-depth: 0
- name: Compute image tag (last backend-touching commit)
id: tag
run: |
domain_sha="$(git log -1 --format=%H -- backend/)"
echo "tag=sha-${domain_sha:0:12}" >> "$GITHUB_OUTPUT"
- name: Deploy backend over SSH
uses: https://github.com/appleboy/ssh-action@v1.2.0
with:
host: ${{ secrets.SSH_HOST }}
username: ${{ secrets.SSH_USER }}
key: ${{ secrets.SSH_KEY }}
port: ${{ secrets.SSH_PORT }}
# Forward the target service + the image tag to run. deploy.sh reads
# BACKEND_IMAGE_TAG and rolls just `backend`.
envs: SERVICE,BACKEND_IMAGE_TAG
script: /opt/thermograph/infra/deploy/deploy.sh
env:
SERVICE: backend
BACKEND_IMAGE_TAG: ${{ steps.tag.outputs.tag }}

View file

@ -0,0 +1,144 @@
name: Build + push images (Forgejo registry)
# backend-build-push.yml and frontend-build-push.yml, collapsed into one.
# They were identical apart from the domain string: the paths filter, the
# concurrency group, IMAGE_PATH, the tag key and the build context.
#
# Images stay separate and independently deployable -- emi/thermograph/backend
# and emi/thermograph/frontend -- exactly as before. Build-once,
# deploy-everywhere; nothing about the published artefacts changes here.
#
# SEMVER EXCEPTION, preserved: a v*.*.* tag push carries no paths context, so a
# release tag builds BOTH domain images. That is intended -- a release names the
# whole set, not whichever domain happened to move last.
#
# Runs on the `docker` label. The LAN runner's job containers get the host's
# docker.sock automounted (Docker-outside-of-Docker, no privileged mode); the
# job image (node:20-bookworm) ships no docker CLI, hence the install step.
#
# The registry is deliberately mesh-only: git.thermograph.org's public DNS
# resolves to beta's public IP, whose Caddy ACL rejects /v2/ paths, so any
# runner host that pushes or pulls needs an /etc/hosts entry.
on:
push:
branches: [dev, main, release]
paths:
- 'backend/**'
- 'frontend/**'
tags: ['v*.*.*']
workflow_dispatch: {}
jobs:
build-push:
strategy:
fail-fast: false
# capacity: 1 physical runner. Serialising keeps a two-domain push from
# queueing two whole image builds against each other.
max-parallel: 1
matrix:
domain: [backend, frontend]
runs-on: docker
# Per-domain, per-ref: only a superseded build of the SAME domain on the
# SAME ref is cancelled, and a frontend build never cancels a backend one.
concurrency:
group: build-push-${{ matrix.domain }}-${{ github.ref }}
cancel-in-progress: true
env:
REGISTRY: https://${{ vars.REGISTRY_HOST }}
steps:
- uses: actions/checkout@v4
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
# often an unrelated domain's commit, and a depth-1 clone can't see past
# it to find the real key.
with:
fetch-depth: 0
- name: Compute image ref + tags
id: image
run: |
set -euo pipefail
host="${REGISTRY#https://}"; host="${host#http://}"
path="$(printf '%s/%s' "${{ github.repository }}" "${{ matrix.domain }}" | tr '[:upper:]' '[:lower:]')"
ref="$host/$path"
# Does this push touch this domain? The workflow-level paths filter
# only says backend OR frontend moved; without this a backend-only
# push would also rebuild and republish the frontend image.
before="${{ github.event.before }}"
if [[ "$GITHUB_REF" == refs/tags/v* ]]; then
changed=true # semver tag: build the whole set, see header
elif [ -z "$before" ] || [ "$before" = "0000000000000000000000000000000000000000" ] \
|| ! git cat-file -e "$before^{commit}" 2>/dev/null; then
changed=true # no usable base; build rather than silently skip
elif git diff --name-only "$before" "${{ github.sha }}" | grep -q "^${{ matrix.domain }}/"; then
changed=true
else
changed=false
fi
# Key the tag to the last commit that touched this domain -- the SAME
# key every deploy computes -- so build and deploy always agree even
# when the branch tip is another domain's commit. (A tag keyed to the
# push tip breaks the moment an infra-only commit lands: deploys go
# looking for an image no build ever produced.)
domain_sha="$(git log -1 --format=%H -- "${{ matrix.domain }}/")"
{
echo "changed=$changed"
echo "host=$host"
echo "sha_tag=$ref:sha-${domain_sha:0:12}"
} >> "$GITHUB_OUTPUT"
if [[ "$GITHUB_REF" == refs/tags/v* ]]; then
echo "semver_tag=$ref:${GITHUB_REF#refs/tags/}" >> "$GITHUB_OUTPUT"
fi
echo "==> ${{ matrix.domain }} | changed=$changed | $ref:sha-${domain_sha:0:12}"
- name: Install Docker CLI
if: steps.image.outputs.changed == 'true'
run: |
apt-get update -qq
apt-get install -y -qq docker.io
- name: Log in to the Forgejo registry
if: steps.image.outputs.changed == 'true'
run: |
# NOT secrets.GITHUB_TOKEN -- Forgejo's per-job auto-injected token can
# never push to the container registry (a known Forgejo limitation,
# independent of any permission setting). Needs a manually provisioned
# PAT with write:package scope, stored as REGISTRY_TOKEN.
echo "${{ secrets.REGISTRY_TOKEN }}" | docker login "${{ steps.image.outputs.host }}" \
--username emi --password-stdin
- name: Build
if: steps.image.outputs.changed == 'true'
run: |
# One build, both tags -- the semver tag (tag pushes only) rides along
# with the SHA tag instead of a second tag+push round trip. Context is
# the DOMAIN dir, so the Dockerfile sees the same tree it saw as a repo
# root in the split era.
tag_args=(-t "${{ steps.image.outputs.sha_tag }}")
if [[ -n "${{ steps.image.outputs.semver_tag }}" ]]; then
tag_args+=(-t "${{ steps.image.outputs.semver_tag }}")
fi
docker build "${tag_args[@]}" "${{ matrix.domain }}/"
- name: Push (SHA tag, every push)
if: steps.image.outputs.changed == 'true'
run: docker push "${{ steps.image.outputs.sha_tag }}"
- name: Push (semver tag, tag pushes only)
if: steps.image.outputs.changed == 'true' && startsWith(github.ref, 'refs/tags/v')
run: docker push "${{ steps.image.outputs.semver_tag }}"
- name: Skipped
if: steps.image.outputs.changed != 'true'
run: echo "${{ matrix.domain }} unchanged in this push — no image built."

View file

@ -0,0 +1,166 @@
name: Deploy
# The six per-domain-per-environment deploy workflows, collapsed into one.
#
# backend-deploy{,-dev,-prod}.yml and frontend-deploy{,-dev,-prod}.yml were the
# same file written six times: they differed only in a branch name, a paths
# filter, a concurrency group, the service name, its *_IMAGE_TAG variable and a
# secret prefix. The contract into infra/deploy/deploy.sh
# (SERVICE + BACKEND_IMAGE_TAG/FRONTEND_IMAGE_TAG) was already fully
# parameterised, so the duplication bought nothing and cost six files to keep in
# step.
#
# The two *-deploy-dev.yml workflows are NOT ported. They were documented as
# inert: they call the monorepo layout at ~/thermograph-dev on the LAN box,
# which is still a split-era thermograph-infra checkout, so that path does not
# exist there. LAN dev is a local `make dev-up` concern, not a CI environment.
#
# DELIBERATELY BORING EXPRESSIONS. No dynamic matrix (fromJSON), no
# `cond && secrets.A || secrets.B` ternary. Those are GitHub idioms that a
# Forgejo/act runner may evaluate differently, and the failure mode here is
# "production does not deploy" or, worse, "deploys with an empty SSH host". The
# two environments therefore get two explicit, mutually exclusive steps.
#
# What is preserved from the originals, all of it load-bearing:
# - fetch-depth: 0, because the image tag is keyed to the LAST COMMIT THAT
# TOUCHED THAT DOMAIN, not the branch tip. In a path-filtered monorepo the
# tip is often an unrelated domain's commit and a depth-1 clone cannot see
# past it.
# - The 12-hex truncation, matching build-push exactly.
# - Per-service, per-environment concurrency with cancel-in-progress: false --
# a half-finished deploy must never be cancelled by a newer one.
# - Separate PROD_SSH_* credentials, so a beta credential leak cannot reach
# prod.
# - appleboy/ssh-action by full URL; it is not mirrored in Forgejo's default
# action registry.
on:
push:
branches: [main, release]
paths:
- 'backend/**'
- 'frontend/**'
workflow_dispatch: {}
jobs:
deploy:
strategy:
fail-fast: false
# One physical runner, and deploy.sh serialises host-side with flock
# anyway. Rolling one service at a time keeps the logs readable and
# matches what the six separate workflows effectively did.
max-parallel: 1
matrix:
service: [backend, frontend]
runs-on: docker
# Keyed by ref AND service, reproducing the old per-file groups
# (beta-deploy-backend, prod-deploy-frontend, ...).
concurrency:
group: deploy-${{ github.ref_name }}-${{ matrix.service }}
cancel-in-progress: false
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Plan this leg
id: plan
run: |
set -euo pipefail
# Branch selects the environment. main -> beta, release -> prod.
case "${{ github.ref_name }}" in
main) environment=beta ;;
release) environment=prod ;;
*) echo "::error::Deploy triggered on unexpected ref '${{ github.ref_name }}'"; exit 1 ;;
esac
# Did THIS push touch THIS domain? The workflow-level paths filter only
# tells us backend OR frontend moved; without this refinement a
# backend-only push would also roll the frontend, losing the
# independent-deploy property the FE/BE split exists for.
before="${{ github.event.before }}"
if [ -z "$before" ] || [ "$before" = "0000000000000000000000000000000000000000" ] \
|| ! git cat-file -e "$before^{commit}" 2>/dev/null; then
# No usable base (first push, force push, or workflow_dispatch).
# Deploy rather than skip: rolling onto the tag already running is a
# no-op for deploy.sh, whereas skipping silently strands a change.
changed=true
echo "no usable before-sha; defaulting to deploy"
elif git diff --name-only "$before" "${{ github.sha }}" | grep -q "^${{ matrix.service }}/"; then
changed=true
else
changed=false
fi
# The tag is the last commit that touched this domain, truncated to 12
# hex to match build-push.yml exactly. Actions expressions have no
# substring function, which is why this is computed in shell.
domain_sha="$(git log -1 --format=%H -- "${{ matrix.service }}/")"
tag="sha-${domain_sha:0:12}"
{
echo "environment=$environment"
echo "changed=$changed"
echo "tag=$tag"
} >> "$GITHUB_OUTPUT"
# Export the deploy.sh contract as real environment variables, chosen
# in shell rather than with a `matrix.service == 'x' && a || b`
# expression. That idiom is a GitHub convention a Forgejo/act runner
# may evaluate differently, and the failure here would be silent: an
# empty *_IMAGE_TAG makes deploy.sh fall back to the tag already
# running, so the job goes green having deployed nothing.
{
echo "SERVICE=${{ matrix.service }}"
if [ "${{ matrix.service }}" = "backend" ]; then
echo "BACKEND_IMAGE_TAG=$tag"
else
echo "FRONTEND_IMAGE_TAG=$tag"
fi
} >> "$GITHUB_ENV"
echo "==> ${{ matrix.service }} -> $environment | changed=$changed | tag=$tag"
- name: Deploy to beta
if: steps.plan.outputs.changed == 'true' && steps.plan.outputs.environment == 'beta'
uses: https://github.com/appleboy/ssh-action@v1.2.0
with:
host: ${{ secrets.SSH_HOST }}
username: ${{ secrets.SSH_USER }}
key: ${{ secrets.SSH_KEY }}
port: ${{ secrets.SSH_PORT }}
envs: SERVICE,BACKEND_IMAGE_TAG,FRONTEND_IMAGE_TAG
script: /opt/thermograph/infra/deploy/deploy.sh
env:
SERVICE: ${{ matrix.service }}
# Only the matching one is read by deploy.sh for a single-service roll;
# the other stays empty and the persisted .image-tags.env supplies the
# sibling's live tag, so this roll never disturbs it.
BACKEND_IMAGE_TAG: ${{ matrix.service == 'backend' && steps.plan.outputs.tag || '' }}
FRONTEND_IMAGE_TAG: ${{ matrix.service == 'frontend' && steps.plan.outputs.tag || '' }}
- name: Deploy to prod
if: steps.plan.outputs.changed == 'true' && steps.plan.outputs.environment == 'prod'
uses: https://github.com/appleboy/ssh-action@v1.2.0
with:
# A completely separate secret set from beta's, deliberately: a beta
# credential leak must not reach prod.
host: ${{ secrets.PROD_SSH_HOST }}
username: ${{ secrets.PROD_SSH_USER }}
key: ${{ secrets.PROD_SSH_KEY }}
port: ${{ secrets.PROD_SSH_PORT }}
envs: SERVICE,BACKEND_IMAGE_TAG,FRONTEND_IMAGE_TAG
script: /opt/thermograph/infra/deploy/deploy.sh
env:
SERVICE: ${{ matrix.service }}
BACKEND_IMAGE_TAG: ${{ matrix.service == 'backend' && steps.plan.outputs.tag || '' }}
FRONTEND_IMAGE_TAG: ${{ matrix.service == 'frontend' && steps.plan.outputs.tag || '' }}
- name: Skipped
if: steps.plan.outputs.changed != 'true'
run: echo "${{ matrix.service }} unchanged in this push — nothing to roll."

View file

@ -1,126 +0,0 @@
name: Build + push frontend image (Forgejo registry)
# Builds the FRONTEND domain's image (frontend/Dockerfile, context frontend/) and
# pushes it to Forgejo's built-in container registry, tagged by git SHA (every
# push) and semver (version tags) -- build-once, deploy-everywhere.
#
# Monorepo port (reunification): the split repos each derived IMAGE_PATH from
# github.repository, which now collides for every domain -- so each domain's
# build-push names its image explicitly: this one is emi/thermograph/frontend
# (backend-build-push.yml publishes emi/thermograph/backend). infra's
# compose/stack files reference the two paths independently
# (FRONTEND_IMAGE_PATH / BACKEND_IMAGE_PATH -- their defaults were renamed to
# match in the same cutover).
#
# Path-filtered: only a push that touches frontend/** builds this image. A
# backend-only push builds nothing here; deploy.sh's persisted
# .image-tags.env keeps the frontend's live tag for the sibling's deploy.
# EXCEPTION: version-tag pushes (v*.*.*) carry no paths context, so a semver
# tag builds BOTH domain images -- intended, a release names the whole set.
#
# Runs on the `docker` label -- the LAN runner's job containers get the host's
# docker.sock automounted in (forgejo-runner config: container.docker_host:
# automount; see infra/deploy/forgejo/register-lan-runner.sh),
# Docker-outside-of-Docker rather than DinD -- no privileged mode needed.
# The job image itself (node:20-bookworm) has no docker CLI, hence the
# "Install Docker CLI" step.
#
# The registry is deliberately mesh-only: git.thermograph.org's public DNS
# resolves to beta's public IP, which Caddy's ACL rejects for /v2/ paths, so
# any runner host that pushes/pulls needs a /etc/hosts entry pointing
# git.thermograph.org at beta's WireGuard IP (10.10.0.2). The docker
# build/push commands run against the automounted HOST daemon, so it's the
# runner HOST's own resolver that has to route over the mesh, not anything
# settable from inside the job container.
#
# REGISTRY is the vars.REGISTRY_HOST Actions variable (git.thermograph.org),
# NOT github.server_url -- that context resolves to the bare mesh IP:port
# (http://10.10.0.2:3080, no TLS), which docker CLI refuses ("server gave HTTP
# response to HTTPS client"). git.thermograph.org gets real TLS via Caddy and
# routes over the mesh once /etc/hosts is set, so it's the only correct value.
# deploy.sh resolves the SHA tag at deploy time and `docker compose pull`s it.
on:
push:
branches: [dev, main, release]
paths: ['frontend/**']
tags: ['v*.*.*']
workflow_dispatch: {}
# The runner has capacity: 1 (one physical LAN box) -- a burst of pushes to the
# same ref would otherwise queue whole builds back to back even though only the
# last one still matters. Keying on domain + github.ref means each domain's
# dev/main/release/tag builds get their own group, so this only cancels a
# superseded build of the SAME domain on the SAME ref -- and a backend build
# never cancels a frontend one.
concurrency:
group: frontend-build-push-${{ github.ref }}
cancel-in-progress: true
jobs:
build-push:
runs-on: docker
env:
REGISTRY: https://${{ vars.REGISTRY_HOST }}
IMAGE_PATH: ${{ github.repository }}/frontend
steps:
- uses: actions/checkout@v4
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
# often an unrelated domain's commit, and a depth-1 clone can't see
# past it to find the real key.
with:
fetch-depth: 0
- name: Install Docker CLI
run: |
apt-get update -qq
apt-get install -y -qq docker.io
- name: Compute image ref + tags
id: image
run: |
host="${REGISTRY#https://}"; host="${host#http://}"
path="$(printf '%s' "$IMAGE_PATH" | tr '[:upper:]' '[:lower:]')"
ref="$host/$path"
# Key the tag to the last commit that touched frontend/ -- the SAME key
# every frontend deploy workflow computes -- so build and deploy always
# agree even when the branch tip is another domain's commit. (A tag
# keyed to the push tip breaks the moment an infra-only commit lands:
# deploys go looking for an image no build ever produced.)
domain_sha="$(git log -1 --format=%H -- frontend/)"
sha_tag="$ref:sha-${domain_sha:0:12}"
echo "host=$host" >> "$GITHUB_OUTPUT"
echo "sha_tag=$sha_tag" >> "$GITHUB_OUTPUT"
if [[ "$GITHUB_REF" == refs/tags/v* ]]; then
echo "semver_tag=$ref:${GITHUB_REF#refs/tags/}" >> "$GITHUB_OUTPUT"
fi
- name: Log in to the Forgejo registry
run: |
# NOT secrets.GITHUB_TOKEN -- Forgejo's per-job auto-injected token
# can never push to the container registry (a known Forgejo
# limitation independent of any permission setting). Needs a
# manually provisioned PAT with write:package scope, stored as the
# REGISTRY_TOKEN secret.
echo "${{ secrets.REGISTRY_TOKEN }}" | docker login "${{ steps.image.outputs.host }}" \
--username emi --password-stdin
- name: Build
run: |
# One build, both tags -- the semver tag (tag pushes only) rides
# along with the SHA tag instead of a second tag+push round trip.
# Context is the DOMAIN dir, so frontend/Dockerfile sees the same
# tree it saw as a repo root in the split era.
tag_args=(-t "${{ steps.image.outputs.sha_tag }}")
if [[ -n "${{ steps.image.outputs.semver_tag }}" ]]; then
tag_args+=(-t "${{ steps.image.outputs.semver_tag }}")
fi
docker build "${tag_args[@]}" frontend/
- name: Push (SHA tag, every push)
run: docker push "${{ steps.image.outputs.sha_tag }}"
- name: Push (semver tag, tag pushes only)
if: startsWith(github.ref, 'refs/tags/v')
run: docker push "${{ steps.image.outputs.semver_tag }}"

View file

@ -1,73 +0,0 @@
name: Deploy frontend to LAN dev server
# Fires on a push to `dev` touching frontend/** -- a direct push, or the real
# merge commit a native "auto merge when checks succeed" produces (see
# pr-build.yml). A `build` job (the reusable build.yml sanity check) gates a
# `deploy` job that rolls the LAN dev box.
#
# CUTOVER NOTE (monorepo reunification): the LAN box's ~/thermograph-dev is a
# thermograph-infra checkout from the split era. This job calls
# ~/thermograph-dev/infra/deploy/deploy-dev.sh -- the MONOREPO layout -- so it
# is INERT until ~/thermograph-dev is re-pointed at the monorepo (a fresh
# clone; deploy-dev.sh then self-updates via deploy.sh's git reset like the
# beta/prod paths).
#
# SERVICE=frontend + FRONTEND_IMAGE_TAG is the same env-var contract the
# beta/prod deploy workflows use against infra/deploy/deploy.sh, so a dev-only
# rollout never touches the backend's running container or tag.
on:
push:
branches: [dev]
paths: ['frontend/**']
workflow_dispatch: {}
jobs:
build:
name: build
# Fixed (unparameterized) group: if two frontend pushes to dev land close
# together, the newer push's build cancels the older's still-running one
# on the single physical runner -- the older result would be thrown away
# once needs:build gates its deploy anyway. Safe to cancel mid-run:
# build.yml only runs a throwaway `docker build`, nothing persisted.
concurrency:
group: dev-push-build-frontend
cancel-in-progress: true
uses: ./.forgejo/workflows/build.yml
with:
domain: frontend
deploy:
needs: build
# NOT [self-hosted, thermograph-lan] -- GitHub implicitly tags every
# self-hosted runner with "self-hosted"; Forgejo's runner has only the
# labels it was registered with ("docker" + "thermograph-lan"). The array
# form silently makes this job unschedulable on every push.
runs-on: thermograph-lan
permissions:
contents: read
concurrency:
group: dev-lan-deploy-frontend
cancel-in-progress: false
steps:
- uses: actions/checkout@v4
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
# often an unrelated domain's commit, and a depth-1 clone can't see
# past it to find the real key.
with:
fetch-depth: 0
- name: Compute image tag
id: tag
# Keyed to the last commit that touched frontend/ -- must match
# frontend-build-push.yml's tag key exactly (see its comment).
run: |
domain_sha="$(git log -1 --format=%H -- frontend/)"
echo "tag=sha-${domain_sha:0:12}" >> "$GITHUB_OUTPUT"
- name: Deploy to the LAN dev server
# thermograph-lan is host-native (the runner job runs directly on the
# LAN box, not over SSH like beta/prod), so plain env vars on the
# command reach deploy-dev.sh directly -- no ssh-action needed here.
run: SERVICE=frontend FRONTEND_IMAGE_TAG=${{ steps.tag.outputs.tag }} bash "$HOME/thermograph-dev/infra/deploy/deploy-dev.sh"

View file

@ -1,61 +0,0 @@
name: Deploy frontend to prod VPS
# On a push to `release` that touches frontend/**, SSH to prod and roll ONLY
# the frontend service onto the image frontend-build-push.yml published for this
# commit. Same shape as frontend-deploy.yml (the `main`->beta path) against a
# completely separate secret set (PROD_SSH_HOST/USER/KEY/PORT) so a beta
# credential leak can't reach prod. Backend deploys itself independently via
# backend-deploy-prod.yml -- a frontend change ships to prod without touching
# backend and vice versa, exactly as in the split era.
#
# The tag MUST match frontend-build-push.yml's last-frontend-commit key (12
# hex) -- see frontend-deploy.yml for why it's truncated here. deploy.sh
# retries the pull because the build for this push may still be in flight.
#
# Frontend's own register() fetches the IndexNow key from backend at boot, so
# deploy.sh waits on a healthy backend before declaring the frontend roll OK
# (its compose depends_on already encodes that ordering).
on:
push:
branches: [release]
paths: ['frontend/**']
workflow_dispatch: {}
concurrency:
group: prod-deploy-frontend
cancel-in-progress: false
jobs:
deploy:
runs-on: docker
steps:
- uses: actions/checkout@v4
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
# often an unrelated domain's commit, and a depth-1 clone can't see
# past it to find the real key.
with:
fetch-depth: 0
- name: Compute image tag (last frontend-touching commit)
id: tag
run: |
domain_sha="$(git log -1 --format=%H -- frontend/)"
echo "tag=sha-${domain_sha:0:12}" >> "$GITHUB_OUTPUT"
- name: Deploy frontend over SSH
uses: https://github.com/appleboy/ssh-action@v1.2.0
with:
host: ${{ secrets.PROD_SSH_HOST }}
username: ${{ secrets.PROD_SSH_USER }}
key: ${{ secrets.PROD_SSH_KEY }}
port: ${{ secrets.PROD_SSH_PORT }}
# Forward the target service + the image tag to run. deploy.sh reads
# FRONTEND_IMAGE_TAG and rolls just `frontend`.
envs: SERVICE,FRONTEND_IMAGE_TAG
script: /opt/thermograph/infra/deploy/deploy.sh
env:
SERVICE: frontend
FRONTEND_IMAGE_TAG: ${{ steps.tag.outputs.tag }}

View file

@ -1,70 +0,0 @@
name: Deploy frontend to beta VPS
# On a push to `main` that touches frontend/**, SSH to beta and roll ONLY the
# frontend service onto the image frontend-build-push.yml just published for
# this commit. Backend is deployed independently by backend-deploy.yml --
# the FE/BE CI-CD split survives reunification intact: a frontend change ships
# without touching backend and vice versa. infra/deploy/deploy.sh persists
# each service's live tag host-side, so a single-service roll never disturbs
# the other. An infra-only push rolls nothing (infra-sync.yml syncs the
# host checkouts instead); a mixed frontend+backend push fires both deploy
# workflows, serialized host-side by deploy.sh's flock.
#
# The SSH_HOST/USER/KEY/PORT secrets point at beta. appleboy/ssh-action is
# referenced by full GitHub URL because it isn't mirrored in Forgejo's default
# action registry.
#
# The tag MUST match frontend-build-push.yml's last-frontend-commit key (12
# hex), not the full 40-char ${{ github.sha }} -- default Actions expressions
# have no substring function, so the compute step below truncates it.
# deploy.sh also retries the pull for ~5 min because Forgejo `needs:` can't
# gate across the separate build-push workflow, so this push's build may still
# be in flight.
#
# Frontend's own register() fetches the IndexNow key from backend at boot, so
# deploy.sh waits on a healthy backend before declaring the frontend roll OK
# (its compose depends_on already encodes that ordering).
on:
push:
branches: [main]
paths: ['frontend/**']
workflow_dispatch: {}
concurrency:
group: beta-deploy-frontend
cancel-in-progress: false
jobs:
deploy:
runs-on: docker
steps:
- uses: actions/checkout@v4
# fetch-depth 0: the tag is keyed to the LAST COMMIT THAT TOUCHED THIS
# DOMAIN, not the branch tip -- in a path-filtered monorepo the tip is
# often an unrelated domain's commit, and a depth-1 clone can't see
# past it to find the real key.
with:
fetch-depth: 0
- name: Compute image tag (last frontend-touching commit)
id: tag
run: |
domain_sha="$(git log -1 --format=%H -- frontend/)"
echo "tag=sha-${domain_sha:0:12}" >> "$GITHUB_OUTPUT"
- name: Deploy frontend over SSH
uses: https://github.com/appleboy/ssh-action@v1.2.0
with:
host: ${{ secrets.SSH_HOST }}
username: ${{ secrets.SSH_USER }}
key: ${{ secrets.SSH_KEY }}
port: ${{ secrets.SSH_PORT }}
# Forward the target service + the image tag to run. deploy.sh reads
# FRONTEND_IMAGE_TAG and rolls just `frontend`.
envs: SERVICE,FRONTEND_IMAGE_TAG
script: /opt/thermograph/infra/deploy/deploy.sh
env:
SERVICE: frontend
FRONTEND_IMAGE_TAG: ${{ steps.tag.outputs.tag }}

105
CLAUDE.md
View file

@ -1,27 +1,88 @@
# thermograph monorepo — agent instructions
Reunified monorepo (2026-07-22): `backend/`, `frontend/`, `infra/`,
`observability/` — each a former split repo subtree-merged with history.
`thermograph-docs` remains a separate repo; cross-cutting decision docs and
operator runbooks belong there, not here.
One repo, four domains: `backend/`, `frontend/`, `infra/`, `observability/`.
Reunified 2026-07-22 from four split repos, with history, via subtree merges.
`thermograph-docs` is deliberately a separate repo — cross-cutting decision
records and operator runbooks go there, not here.
- **Each domain keeps its own `CLAUDE.md`** — read the one for the domain you
touch. They still use repo-era wording ("this repo", sibling-repo names);
mentally map `thermograph-backend``backend/` etc., and fix wording as you
touch files near it.
- **CI lives only in root `.forgejo/workflows/`**, path-filtered per domain
(see README). Never re-add workflows under a domain's own `.forgejo/` — they
are inert there and become a trap.
- Images are `emi/thermograph/backend` and `emi/thermograph/frontend`
(`sha-<12hex>`). Deploy = `infra/deploy/deploy.sh` over SSH, per-service.
**Rule for this file and every domain `CLAUDE.md`: only statements that would
break CI if they became false, or that name a file that exists.** Background,
history and rationale belong in `thermograph-docs`. These files are read before
every change, so a stale one is a correctness bug, not a documentation bug.
## Branches and environments
| Branch | Deploys to | Workflow |
|---|---|---|
| feature branch | nothing | PR into `dev` |
| `dev` | nothing | integration branch only — see below |
| `main` | beta (beta.thermograph.org) | `deploy.yml` |
| `release` | prod (thermograph.org) | `deploy.yml` |
`dev`, `main` and `release` are protected: **everything is a PR**, for humans and
agents alike. Promotion is one PR per hop, `dev``main``release`.
**`dev` deploys nowhere.** It is purely the integration branch that feature PRs
land on before promotion to `main`. The two LAN-dev deploy workflows were deleted
rather than kept: they called a monorepo path that does not exist on the LAN box
(`~/thermograph-dev` is still a split-era `thermograph-infra` checkout), so they
had been inert since cutover. Run LAN dev locally with `infra/`'s `make dev-up`.
**One workflow deploys everything.** `deploy.yml` handles both services and both
environments: the branch selects the environment (`main` → beta, `release`
prod), a matrix covers backend and frontend, and each leg checks whether this
push actually touched its domain before rolling. `build-push.yml` is the same
shape for images. They replaced six and two near-identical files respectively.
## The deploy contract
One entry point, two modes, one contract:
```
SERVICE=backend|frontend|all BACKEND_IMAGE_TAG=sha-<12hex> FRONTEND_IMAGE_TAG=sha-<12hex> \
/opt/thermograph/infra/deploy/deploy.sh
```
`deploy.sh` resets the host checkout, renders secrets from the SOPS vault, then
either rolls compose services or — if `/etc/thermograph/deploy-mode` contains
`stack` — execs `infra/deploy/stack/deploy-stack.sh`.
- **prod runs Swarm.** Its stack is `infra/deploy/stack/thermograph-stack.yml`
(db, web, worker, lake, daemon, frontend, autoscaler, autoscaler-lake).
`deploy-stack.sh` also offers `STACK_TEST=1`: a full parallel rehearsal on
throwaway volumes and ports that cannot touch live data.
- **beta and LAN dev run compose**, from `infra/docker-compose.yml`
(db, backend, lake, daemon, frontend).
- Each service's live tag is persisted host-side, so a single-service roll never
disturbs the sibling's running tag.
Images are `emi/thermograph/backend` and `emi/thermograph/frontend`, tagged
`sha-<12hex>` on every push and by semver on `v*.*.*` tags. They build and deploy
**independently** — a backend change ships without a frontend deploy and vice
versa. That independence is the point of the FE/BE split and survived
reunification.
## Rules that bind across domains
- **CI lives only in root `.forgejo/workflows/`**, path-filtered per domain.
Never add workflows under a domain's own `.forgejo/` — they are inert there and
become a trap.
- **Secrets only via the SOPS vault** (`infra/deploy/secrets/`). Never hand-edit
`/etc/thermograph.env` on a host — it is a rendered artifact. `secrets-guard`
CI rejects plaintext. `seed-from-live.sh` reads production secrets and is
explicitly **not** for an agent to run.
- FE and BE ship out of lockstep, so the `/api/version` contract and
`PAYLOAD_VER` discipline are load-bearing — see `backend/CLAUDE.md`.
- Anything named "prefetch" must never spend the Open-Meteo quota.
Nominatim ≤ 1 req/s.
- The compose project name is pinned (`name: thermograph` in
`infra/docker-compose.yml`); LAN dev overrides via
`COMPOSE_PROJECT_NAME=thermograph-dev`. Don't remove either half.
- Cross-domain rules that still bind: FE/BE ship independently → keep the
`/api/version` contract and `PAYLOAD_VER` discipline (see
`backend/CLAUDE.md`); anything named "prefetch" must never spend the
Open-Meteo quota; Nominatim ≤ 1 req/s; secrets only via the SOPS vault
(`infra/deploy/secrets/`).
`infra/docker-compose.yml`); LAN dev overrides with
`COMPOSE_PROJECT_NAME=thermograph-dev`. Don't remove either half — the pinned
name is what makes the Swarm stack's external volume names line up.
- `CUTOVER-NOTES.md` is the source of truth for what is and isn't live yet.
- Commits/PRs: concise and technical; never mention AI/assistants/automated
authorship.
## Commits & PRs
Describe only the substance of the change; concise and technical. Never mention
AI, Claude, assistants or automated authorship anywhere — no trailers,
co-authors, emoji, or "as requested" narration.

View file

@ -13,6 +13,16 @@ subtree-merging each split repo's `origin/main` (full history, 380 commits):
**CI (all in root `.forgejo/workflows/`; the subtree copies were deleted —
Forgejo only reads root workflows, but dead copies are a trap):**
> **Superseded 2026-07-25.** The per-domain workflow names below describe the
> cutover as it landed, and are kept for that history. They no longer exist:
> the six `*-deploy*.yml` files were collapsed into a single `deploy.yml`
> (branch selects environment, matrix covers the services) and the two
> `*-build-push.yml` into a single `build-push.yml`. The two LAN-dev deploy
> workflows were deleted outright rather than ported, having been inert since
> cutover. Everything described below about *behaviour* — path filtering, the
> domain-keyed 12-hex tag, per-domain concurrency, the `v*.*.*` both-images
> exception, separate `SSH_*`/`PROD_SSH_*` secret sets — is unchanged.
- `backend-build-push.yml` / `frontend-build-push.yml` — the old per-repo
build-push, path-filtered (`paths: ['backend/**']` etc.), image paths now
**explicit**: `emi/thermograph/backend|frontend` (the old

View file

@ -8,9 +8,9 @@ for: **per-domain images, per-domain deploys, and an async FE/BE contract**.
| Dir | What | CI |
|---|---|---|
| `backend/` | FastAPI graded-climate API, accounts, notifications (Discord bot, push, mail), data pipeline | `backend-build-push` → image `emi/thermograph/backend`; `backend-deploy[-prod\|-dev]` |
| `frontend/` | Public client: static JS/CSS + SSR pages | `frontend-*` mirrors of the above; image `emi/thermograph/frontend` |
| `infra/` | Compose, deploy scripts, terraform, SOPS secrets vault, ops cron | `infra-sync` (host checkout + secrets render), `secrets-guard`, `ops-cron` |
| `backend/` | FastAPI graded-climate API, accounts, notifications (Discord bot, push, mail), data pipeline | `build-push` → image `emi/thermograph/backend`; `deploy` |
| `frontend/` | Public client: static JS/CSS + SSR pages | same `build-push` / `deploy` workflows, matrixed by domain; image `emi/thermograph/frontend` |
| `infra/` | Compose (beta, LAN dev) + the Swarm stack (prod), deploy scripts, terraform, SOPS secrets vault, ops cron | `infra-sync` (host checkout + secrets render), `secrets-guard`, `ops-cron` |
| `observability/` | Loki + Grafana + Alloy stack | `observability-validate` |
`thermograph-docs` deliberately **stays its own repo** (ADRs + runbooks, no

View file

@ -1,161 +1,66 @@
# thermograph-backend — agent instructions
# backend/ — agent instructions
## What this repo is
The Thermograph **API service**: a FastAPI app serving the graded-climate API
(`/api/v2/...`), user accounts, notifications, and the SSR-content JSON API the
frontend renders pages from. It owns the climate data pipeline and the
derived-payload cache. It serves **no HTML/CSS/JS** — that is `frontend/`.
The Thermograph **API service**: FastAPI app serving the graded-climate API
(`/api/v2/...`), user accounts (`accounts/`, fastapi-users + Postgres via
alembic), push/email/Discord notifications (`notifications/`), and the
SSR-content JSON API (`api/content_routes.py` + `api/content_payloads.py`) that
the separate `thermograph-frontend`'s server-rendered pages consume. It owns
the climate data pipeline (`data/` — grid snapping, Open-Meteo fetch + parquet
cache, percentile grading/scoring, places/cities) and the derived-payload
cache/ETag store (`data/store.py`, `api/payloads.py`). It does **not** serve
any HTML/CSS/JS itself — that moved to `thermograph-frontend` (repo-split
Stage 7a); this process is a pure JSON API + DB-backed service, launched as
`app:app` (a shim re-exporting `web.app:app`, kept stable for systemd/CI/`run.sh`
callers that still expect that target).
Read the root `CLAUDE.md` first for the branch model, deploy contract and image
names.
## Split topology
## Build & verify
Split out of the `emi/thermograph` monorepo (see
`MONOREPO-CUTOVER-PLAN.md` at the workspace root during the transition). Sibling
repos:
- **`thermograph-frontend`** — the client the public sees: static
JS/CSS/HTML + `frontend_ssr`'s server-rendered climate hub/city/month pages,
calling this repo's `/api/v2/...` and `/content/...` JSON endpoints.
- **`thermograph-infra`** — deploy/terraform/Swarm-Compose plumbing:
`docker-compose.yml`, `deploy/deploy.sh`, secrets (SOPS-encrypted
`deploy/secrets/*.yaml`), Terraform for prod/beta provisioning. This repo has
no deploy scripts or terraform of its own — it only builds an image and hands
infra a tag.
- **`thermograph-docs`** — non-code planning layer: architecture decision
records and operator runbooks. Nothing here should duplicate that; if you're
writing a cross-cutting decision doc or a step-by-step operator runbook,
it belongs there, not in this repo's README/CLAUDE.md.
- **`thermograph-observability`** — monitoring/metrics stack, separate from
this repo's own lightweight in-process `core/metrics.py`.
- `make test` — builds a py3.12 venv and runs the hermetic pytest suite
(`scripts/test.sh`; pass `ARGS=...`). Tests are hermetic per `tests/conftest.py`:
no real Open-Meteo/Nominatim/GeoNames calls, throwaway SQLite, notifier thread
disabled.
- `make smoke` — boots the built image plus a throwaway db and asserts
`/healthz` + `/api/version`.
- CI runs this suite **inside the built image** (`build.yml`), so the exact
interpreter and deps that ship are what gets tested.
## Branch, PR & deploy flow
`app.py` at the root is a one-line re-export shim for `web/app.py`, kept stable
for systemd/CI callers. Entry target is `app:app`.
- `build-push.yml` builds **this repo's own image**
(`git.thermograph.org/emi/thermograph-backend/app`, tagged `sha-<12hex>` on
every push to `dev`/`main`/`release` and additionally by semver on `v*.*.*`
tags) — no more shared monorepo image. `build.yml` (push/PR to `main`) is a
build-only sanity check (`docker build`), not a boot/health check — repeated
attempts at a live boot+healthz check hit this runner's own quirks (bridge-IP
reachability, non-unique `$$`, squatted host ports) without ever going
reliably green; the image's actual boot is verified by hand (`docker run` +
alembic + `/healthz` 200 + a real `/api/v2/place` 200) instead.
- **`main` → beta**: `deploy.yml` SSHes to beta and runs
`thermograph-infra`'s `/opt/thermograph/deploy/deploy.sh` with
`SERVICE=backend` + `BACKEND_IMAGE_TAG=sha-<12hex>` (must match
build-push.yml's 12-hex SHA tag exactly — Actions expressions have no
substring function for the full 40-char SHA). `deploy.sh` rolls only the
`backend` compose service (`--no-deps`) and persists the tag in
`deploy/.image-tags.env` host-side, so a backend-only deploy never touches
the frontend's running container or tag.
- **`release` → prod**: `deploy-prod.yml` is the same shape against a
completely separate secret set (`PROD_SSH_HOST/USER/KEY/PORT`) so a beta
credential leak can't reach prod. Also `SERVICE=backend` +
`BACKEND_IMAGE_TAG`.
- The frontend deploys itself the same way, independently — that
independence (a backend change ships without a frontend deploy and vice
versa) is the entire point of the FE/BE CI-CD split.
- No `dev`-branch LAN deploy path exists in this repo (the monorepo's
`deploy-dev.yml`/`docker-compose.dev.yml` never made it across the split —
known gap, see the cutover plan if resurrecting it).
## Contracts that break silently if you change them
## Run / test locally
There is **no Makefile in this repo** — the monorepo's app-level `make`
targets (`lan-run`, `test`, `venv`, ...) haven't been given a new home yet
(tracked as a gap in the cutover plan). Until that lands, drive it directly:
```bash
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt -r requirements-dev.txt
.venv/bin/uvicorn app:app --host 0.0.0.0 --port 8137 # serve
.venv/bin/python -m pytest tests # test
```
The monorepo's Makefile preferred `uv venv --python 3.12 .venv` (falling back
to plain `python3 -m venv` only if `uv` isn't installed) — prefer `uv` here too
if it's on the box. There is no committed `.venv` in this repo; if one exists
locally and looks broken (import errors, wrong Python), just `rm -rf .venv` and
recreate rather than debugging it — it's gitignored, disposable state, not
something the repo depends on being a specific way.
Tests are hermetic (`tests/conftest.py`): no real Open-Meteo/Nominatim/GeoNames
calls, a throwaway SQLite for accounts + the derived-payload store, and the
notifier thread disabled. `THERMOGRAPH_FRONTEND_BASE_INTERNAL` is set to an
unreachable placeholder in conftest — `web/app.py` fails loud at import if it's
unset in a real run (it's the internal URL back to `frontend_ssr` for the
non-API HTML routes this process no longer serves itself).
## API version contract
- Everything real lives under **`/api/v2/...`**; `/api/` and `/api/v1/` stay
mounted as aliases of the same handlers (they always were — not new).
- **`GET /api/version`** (`web/app.py`) is the capability/negotiation endpoint
a split-out frontend can call at boot: `{backend_version, min_frontend,
payload_ver}`, driven by two module constants next to it —
`API_CONTRACT_VERSION` (bump only on a breaking change to any `/api/v{N}`
route or payload shape) and `MIN_SUPPORTED_FRONTEND` (oldest frontend
contract version this backend still serves correctly). A genuine breaking
change ships as a new `v3` router mounted alongside `v2` re-registering only
the changed handlers, `v2` kept working, `API_CONTRACT_VERSION` bumped in the
same PR.
- **`PAYLOAD_VER`** (`api/payloads.py`, currently `"p2"`) is a *separate*
concern from URL versioning — it's the cache/ETag invalidation token baked
into every derived-store validity token (`history_token()`). Bump it
whenever a response payload's shape changes (new metric, renamed/removed
field) so one bump atomically orphans every pre-upgrade cached row instead
of a stale row's `history_end` happening to still validate under the old
shape.
- Both endpoints (`/healthz`, `/api/version`) are deliberately I/O-free so they
stay cheap under tight healthcheck intervals.
## Live cross-repo contracts (don't break silently)
- **`/cell` bundle + ETag/If-None-Match**: every graded payload (`grade`,
`calendar`, `day`) is cached in SQLite by `(kind, cell_id, key)` keyed to a
validity token that only advances when the underlying history actually
changes (`_etag_for` in `web/app.py`, `history_token`/`PAYLOAD_VER` in
`api/payloads.py`). The same token doubles as a weak ETag; a client sending
`If-None-Match` gets an empty 304 without the payload being rebuilt or even
loaded. `expose_headers=["ETag"]` in the CORS middleware matters as much as
`allow_origins` — without it the frontend's `cache.js` reading
`res.headers.get("ETag")` cross-origin silently gets `null` and never
revalidates correctly. Don't change the ETag derivation or the CORS
`expose_headers` list without checking the frontend's cache layer.
- **Percentile/unit parity with `thermograph-frontend`'s `static/shared.js`**:
the frontend's percentile-to-ordinal display logic explicitly mirrors this
repo's `data/grading.py::pct_ordinal()` (floors an empirical percentile into
1..99 — a rank against the sample can round to 0 or 100, and "100th
percentile" is never shown). `TEMP_BANDS`/`RAIN_BANDS` tier boundaries in
`data/grading.py` are the source of truth for the tier names/thresholds
(Near Record / High / Above Normal / Normal / Below Normal / Low, and the
five rain-intensity tiers) — changing them without a corresponding frontend
update produces a client that draws differently-colored tiers than the
labels the API returns.
- **SSR content API** (`api/content_routes.py`): a separate ETag'd JSON surface
(`/content/hub`, `/content/city/{slug}`, `/content/city/{slug}/month/{month}`,
`/content/city/{slug}/records`, `/content/home`, `/content/sitemap`,
`/content/indexnow-key`) that `frontend_ssr` calls in-process-free (no shared
Python import) to render pages. Same caching pattern as the main API; keep
payload shape changes coordinated with whatever consumes it on the frontend
side.
- **`GET /api/version`** → `{backend_version, min_frontend, payload_ver}`, driven
by `API_CONTRACT_VERSION` and `MIN_SUPPORTED_FRONTEND` in `web/app.py`. A
breaking change ships as a new `v3` router mounted **alongside** `v2`,
re-registering only changed handlers, with the constant bumped in the same PR.
`/api` and `/api/v1` stay mounted as aliases.
- **`PAYLOAD_VER`** (`api/payloads.py`) is separate from URL versioning — it is
the cache/ETag invalidation token. Bump it whenever a response payload's shape
changes, so one bump atomically orphans every pre-upgrade cached row.
- **ETag / `If-None-Match`** — every graded payload is cached against a validity
token that doubles as a weak ETag. `expose_headers=["ETag"]` in the CORS
middleware matters as much as `allow_origins`: without it the frontend's
`cache.js` reads `null` cross-origin and never revalidates. Don't change the
ETag derivation or that list without checking `frontend/static/cache.js`.
- **`data/grading.py::pct_ordinal()`** is mirrored by `frontend`'s
`shared.js::pctOrd()` — floor a percentile into `1..99`, never 0 or 100.
`TEMP_BANDS`/`RAIN_BANDS` are the source of truth for tier names and
thresholds; changing them without the frontend produces tiers drawn in colours
that disagree with the labels the API returns.
- **Fahrenheit country set**`api/content_payloads.py`'s `F_COUNTRIES` must stay
identical to `frontend`'s `format.py::F_COUNTRIES` and `static/units.js`'s
`F_REGIONS`. There is a test asserting this.
- **`/healthz` and `/api/version` are deliberately I/O-free** so they stay cheap
under tight healthcheck intervals.
## Layout
`accounts/` (fastapi-users models/schemas/db, `alembic/` migrations),
`api/` (versioned payload builders + SSR content routes), `core/` (metrics,
audit logging, a singleton helper), `data/` (grid, climate fetch/cache, grading/
scoring, places/cities, the derived-payload store), `notifications/` (push,
email, Discord bot + interactions + linking, digest, scheduler), `web/app.py`
(the actual FastAPI app; `app.py` at the repo root is a one-line re-export
shim). `paths.py` resolves every filesystem location (climate parquet cache,
accounts DB, logs, bundled `cities.json`/`cities_flavor.json`) from the repo
root — don't reintroduce `__file__`-relative paths in a module, it breaks the
moment a module moves. `deploy/entrypoint.sh` is the container entrypoint
(alembic migrate-then-serve, with a Swarm-secrets-as-files shim); actual deploy
orchestration lives in `thermograph-infra`, not here.
`accounts/` (fastapi-users + alembic), `api/` (payload builders + SSR content
routes), `core/` (metrics, audit, singleton helper), `data/` (grid, climate
fetch/cache, grading/scoring, places/cities, derived-payload store),
`notifications/` (push, email, Discord bot, digest, scheduler), `web/app.py`
(the real app), `daemon/`.
`paths.py` resolves every filesystem location from the repo root — don't
reintroduce `__file__`-relative paths in a module; it breaks the moment a module
moves. `deploy/entrypoint.sh` is the container entrypoint (alembic migrate, then
serve, with a Swarm-secrets-as-files shim).
## Commits & PRs
Concise and technical. Never mention AI, assistants or automated authorship.

View file

@ -936,29 +936,27 @@ def _fetch_recent_forecast(cell: dict) -> pl.DataFrame:
.sort("date"))
def _om_local_today(payload: dict) -> "datetime.date | None":
"""The calendar date it is *now* at the cell, from the UTC offset Open-Meteo
reports for a ``timezone=auto`` request. None when the offset is absent, which
tells the caller to skip the in-progress-day guard rather than guess a date."""
offset = payload.get("utc_offset_seconds")
if offset is None:
return None
return (datetime.datetime.now(datetime.timezone.utc)
+ datetime.timedelta(seconds=int(offset))).date()
def _fetch_recent_forecast_om(cell: dict) -> pl.DataFrame:
"""Fallback recent+forecast bundle from the Open-Meteo forecast API (the former
primary): recent past + forward days in one call.
"""Primary recent+forecast bundle from the Open-Meteo forecast API: recent past
and forward days in one call, covering today + 7 (see FORECAST_DAYS).
The cell's in-progress *local* day is dropped from the bundle. Open-Meteo's
daily high/low for today aggregates only the hours elapsed so far, so grading it
reads a still-unfolding day as complete and produces spurious extremes (a cool
morning served as a record-low high, seen live at the 1st percentile). This is
the same failure the MET Norway path guards against with its diurnal-coverage
gate (see _metno_to_frame); the fallback needs the equivalent. Past days are
complete and future days are whole-day forecasts, so only today is excluded
the day lands in the record once it is over."""
The cell's in-progress *local* day is KEPT. An earlier revision dropped it on
the assumption that Open-Meteo's daily high/low for today aggregates only the
hours elapsed so far true of the MET Norway series, which genuinely starts
mid-day and is gated on diurnal coverage (see _metno_to_frame), but NOT of this
endpoint: Open-Meteo backfills the rest of today from the model run, so today's
daily value spans the whole local day exactly as tomorrow's does.
Verified against the live API at four cells spanning 01:00-19:00 local: the
reported daily max equals the max over all 24 hourly values, and diverges from
the elapsed-hours-only max wherever the day still has hours left (Los Angeles
at 09:16 local reported 93.9F, the whole-day figure, where the hours elapsed
topped out at 76.1F). Re-check it that way before reintroducing any exclusion.
Dropping today cost every freshly-synced cell its current day a hole between
a complete yesterday and a forecast tomorrow, which the UI renders as a missing
column. A day that genuinely cannot be graded, because it came back with no
high/low, is already dropped upstream by _to_frame."""
params = {
"latitude": cell["center_lat"],
"longitude": cell["center_lon"],
@ -971,12 +969,7 @@ def _fetch_recent_forecast_om(cell: dict) -> pl.DataFrame:
"forecast_days": FORECAST_DAYS,
}
r = _request(FORECAST_URL, params, 60, phase="recent_forecast_fetch")
payload = r.json()
df = _to_frame(payload["daily"])
local_today = _om_local_today(payload)
if local_today is not None:
df = df.filter(pl.col("date") != local_today)
return df
return _to_frame(r.json()["daily"])
def _load_recent_forecast(cell: dict) -> pl.DataFrame:

View file

@ -33,6 +33,7 @@ services:
THERMOGRAPH_FRONTEND_BASE_INTERNAL: http://127.0.0.1:1
THERMOGRAPH_BASE: "/"
THERMOGRAPH_ENABLE_NOTIFIER: "0"
THERMOGRAPH_ENABLE_HEARTBEAT: "0"
PORT: "8137"
WORKERS: "1"
ports:

View file

@ -28,6 +28,7 @@ _TMP = tempfile.mkdtemp(prefix="thermograph-tests-")
# start the background notifier thread during tests. Set before db is imported.
os.environ.setdefault("THERMOGRAPH_ACCOUNTS_DB", os.path.join(_TMP, "accounts.sqlite"))
os.environ.setdefault("THERMOGRAPH_ENABLE_NOTIFIER", "0")
os.environ.setdefault("THERMOGRAPH_ENABLE_HEARTBEAT", "0")
# web/app.py (repo-split Stage 4) fails loud at import if this is unset -- tests
# never actually hit a frontend-owned path through the live proxy (that's
# frontend_ssr/tests' job), so an unreachable placeholder is fine here.

View file

@ -215,20 +215,20 @@ def test_recent_forecast_fallback_merges_archive_past_and_metno_forward(monkeypa
datetime.date(2026, 7, 16), datetime.date(2026, 7, 17)]
def test_recent_forecast_om_drops_the_in_progress_local_day(monkeypatch):
"""The Open-Meteo fallback must not grade the cell's in-progress local day: its
daily high/low is only a partial aggregate of the hours elapsed so far (a cool
morning would read as a record-low high). Past days and future forecast days
survive; today per the UTC offset Open-Meteo reports for a timezone=auto
request is dropped, matching the MET path's diurnal-coverage gate."""
offset = 3 * 3600 # UTC+3, e.g. Europe/Vilnius (the reported Ringaudai incident)
def test_recent_forecast_om_keeps_the_in_progress_local_day(monkeypatch):
"""Today must survive the Open-Meteo bundle. Unlike the MET Norway series (which
starts mid-day and is gated on diurnal coverage), this endpoint backfills the
rest of today from the model run, so today's high/low is a whole-day value of
the same kind as tomorrow's. Dropping it left a hole between a complete
yesterday and a forecast tomorrow, which the UI renders as a missing column."""
offset = 3 * 3600 # UTC+3, e.g. Europe/Vilnius (the reported Ringaudai cell)
local_today = (datetime.datetime.now(datetime.timezone.utc)
+ datetime.timedelta(seconds=offset)).date()
days = [local_today + datetime.timedelta(days=n) for n in (-2, -1, 0, 1)]
daily = {
"time": [d.isoformat() for d in days],
"temperature_2m_max": [80.0, 82.0, 61.0, 84.0], # today's 61 is the partial value
"temperature_2m_min": [60.0, 61.0, 57.0, 62.0],
"temperature_2m_max": [80.0, 82.0, 76.1, 84.0],
"temperature_2m_min": [60.0, 61.0, 55.9, 62.0],
"precipitation_sum": [0.0, 0.0, 0.0, 0.0],
}
payload = {"utc_offset_seconds": offset, "daily": daily}
@ -237,60 +237,42 @@ def test_recent_forecast_om_drops_the_in_progress_local_day(monkeypatch):
def json(self): return payload
monkeypatch.setattr(climate, "_request", lambda *a, **k: Resp())
got = climate._fetch_recent_forecast_om({"center_lat": 54.9, "center_lon": 23.8})
assert got["date"].to_list() == days # every day present, no gap at today
row = got.filter(pl.col("date") == local_today)
assert row["tmax"].item() == 76.1 # today's whole-day high, ungraded-down
assert row["tmin"].item() == 55.9
def test_recent_forecast_om_still_drops_a_today_with_no_high_low(monkeypatch):
"""The one case today is excluded: it came back with no high/low and so cannot be
graded at all. _to_frame enforces that for every day, today included."""
offset = 3 * 3600
local_today = (datetime.datetime.now(datetime.timezone.utc)
+ datetime.timedelta(seconds=offset)).date()
days = [local_today + datetime.timedelta(days=n) for n in (-1, 0, 1)]
daily = {
"time": [d.isoformat() for d in days],
"temperature_2m_max": [82.0, None, 84.0],
"temperature_2m_min": [61.0, None, 62.0],
"precipitation_sum": [0.0, 0.0, 0.0],
}
payload = {"utc_offset_seconds": offset, "daily": daily}
class Resp:
def json(self): return payload
monkeypatch.setattr(climate, "_request", lambda *a, **k: Resp())
got = climate._fetch_recent_forecast_om(
{"center_lat": 54.9, "center_lon": 23.8})["date"].to_list()
assert local_today not in got # partial today dropped
assert local_today - datetime.timedelta(days=1) in got # yesterday kept
assert local_today + datetime.timedelta(days=1) in got # tomorrow's forecast kept
assert len(got) == 3
assert local_today not in got
assert got == [local_today - datetime.timedelta(days=1),
local_today + datetime.timedelta(days=1)]
def test_recent_forecast_om_keeps_all_days_without_a_utc_offset(monkeypatch):
"""No UTC offset reported -> skip the in-progress-day guard rather than guess a
date, so the bundle passes through as before (only the usual null-day filter)."""
payload = {"daily": _om_daily(3)} # no utc_offset_seconds
class Resp:
def json(self): return payload
monkeypatch.setattr(climate, "_request", lambda *a, **k: Resp())
df = climate._fetch_recent_forecast_om({"center_lat": 1.0, "center_lon": 2.0})
assert df.height == climate._to_frame(_om_daily(3)).height
def test_recent_forecast_om_drops_the_in_progress_local_day(monkeypatch):
"""The Open-Meteo fallback must not grade the cell's in-progress local day: its
daily high/low is only a partial aggregate of the hours elapsed so far (a cool
morning would read as a record-low high). Past days and future forecast days
survive; today per the UTC offset Open-Meteo reports for a timezone=auto
request is dropped, matching the MET path's diurnal-coverage gate."""
offset = 3 * 3600 # UTC+3, e.g. Europe/Vilnius (the reported Ringaudai incident)
local_today = (datetime.datetime.now(datetime.timezone.utc)
+ datetime.timedelta(seconds=offset)).date()
days = [local_today + datetime.timedelta(days=n) for n in (-2, -1, 0, 1)]
daily = {
"time": [d.isoformat() for d in days],
"temperature_2m_max": [80.0, 82.0, 61.0, 84.0], # today's 61 is the partial value
"temperature_2m_min": [60.0, 61.0, 57.0, 62.0],
"precipitation_sum": [0.0, 0.0, 0.0, 0.0],
}
payload = {"utc_offset_seconds": offset, "daily": daily}
class Resp:
def json(self): return payload
monkeypatch.setattr(climate, "_request", lambda *a, **k: Resp())
got = climate._fetch_recent_forecast_om(
{"center_lat": 54.9, "center_lon": 23.8})["date"].to_list()
assert local_today not in got # partial today dropped
assert local_today - datetime.timedelta(days=1) in got # yesterday kept
assert local_today + datetime.timedelta(days=1) in got # tomorrow's forecast kept
assert len(got) == 3
def test_recent_forecast_om_keeps_all_days_without_a_utc_offset(monkeypatch):
"""No UTC offset reported -> skip the in-progress-day guard rather than guess a
date, so the bundle passes through as before (only the usual null-day filter)."""
"""No UTC offset reported changes nothing: the bundle passes through with only
the usual null-day filter."""
payload = {"daily": _om_daily(3)} # no utc_offset_seconds
class Resp:

View file

@ -0,0 +1,91 @@
"""Process-level liveness heartbeat for web/worker (app.py _heartbeat_loop /
_start_heartbeat) the fix for ProdWebContainerSilent / ProdWorkerContainerSilent
false-firing on container-stdout silence that isn't a real liveness signal for
either role."""
import json
import time
import pytest
from fastapi.testclient import TestClient
from core import audit
from web import app as appmod
@pytest.fixture(autouse=True)
def _reset_heartbeat_thread():
# _HEARTBEAT_THREAD is a module-level singleton guard (mirrors notify.py's own
# _STOP/_THREAD pattern) -- reset it around every test so one test starting the
# thread doesn't leave _start_heartbeat() a no-op for the next.
appmod._HEARTBEAT_STOP.clear()
appmod._HEARTBEAT_THREAD = None
yield
appmod._HEARTBEAT_STOP.set()
appmod._HEARTBEAT_THREAD = None
@pytest.mark.parametrize("role", ["web", "worker", "all"])
def test_heartbeat_loop_tags_the_beat_with_role(monkeypatch, role):
monkeypatch.setattr(appmod, "ROLE", role)
calls = []
monkeypatch.setattr(audit, "log_heartbeat", lambda daemon, interval_s: calls.append((daemon, interval_s)))
monkeypatch.setattr(appmod.metrics, "record_heartbeat", lambda daemon, interval_s: None)
# Setting the stop event first means .wait() returns True immediately, so the
# loop beats exactly once (the pre-loop beat) and exits without blocking.
appmod._HEARTBEAT_STOP.set()
appmod._heartbeat_loop()
assert calls == [(role, appmod.HEARTBEAT_INTERVAL)]
def test_start_heartbeat_respects_the_disable_flag(monkeypatch):
monkeypatch.setenv("THERMOGRAPH_ENABLE_HEARTBEAT", "0")
started = []
monkeypatch.setattr(appmod.threading, "Thread", lambda **kw: started.append(kw))
appmod._start_heartbeat()
assert appmod._HEARTBEAT_THREAD is None
assert started == []
def test_start_heartbeat_is_idempotent(monkeypatch):
monkeypatch.delenv("THERMOGRAPH_ENABLE_HEARTBEAT", raising=False)
starts = []
class _FakeThread:
def __init__(self, **kw):
starts.append(kw)
def start(self):
pass
monkeypatch.setattr(appmod.threading, "Thread", _FakeThread)
appmod._start_heartbeat()
appmod._start_heartbeat()
# A second call is a no-op once a thread handle already exists -- mirrors
# notify.start()'s own idempotency, so a double-entry into the lifespan
# (or a stray manual call) can't spawn a second beat loop.
assert len(starts) == 1
def test_app_boot_emits_a_heartbeat_within_one_interval(monkeypatch, tmp_path):
monkeypatch.setattr(appmod, "ROLE", "web")
monkeypatch.setattr(appmod, "HEARTBEAT_INTERVAL", 0.01)
monkeypatch.setattr(audit, "HEARTBEAT_DIR", str(tmp_path))
monkeypatch.delenv("THERMOGRAPH_ENABLE_HEARTBEAT", raising=False)
with TestClient(appmod.app):
# The pre-loop beat fires synchronously on thread start; a short sleep is
# just slack for the thread to actually get scheduled. The loop itself
# keeps running on HEARTBEAT_INTERVAL until lifespan shutdown sets the
# stop event, so there's nothing to join here.
time.sleep(0.2)
files = list(tmp_path.glob("heartbeat-*.jsonl"))
assert files, "expected a heartbeat-*.jsonl file to be written during app lifespan"
lines = [json.loads(line) for line in files[0].read_text().splitlines() if line]
assert any(rec["daemon"] == "web" and rec["tag"] == "heartbeat" for rec in lines)

View file

@ -154,6 +154,39 @@ def _should_run_notifier() -> bool:
return ROLE in ("worker", "all") and singleton.claim_leader()
_HEARTBEAT_STOP = threading.Event()
_HEARTBEAT_THREAD: threading.Thread | None = None
HEARTBEAT_INTERVAL = float(os.environ.get("THERMOGRAPH_HEARTBEAT_INTERVAL", "300")) # 5 min
def _heartbeat_loop() -> None:
metrics.record_heartbeat(ROLE, HEARTBEAT_INTERVAL)
audit.log_heartbeat(ROLE, HEARTBEAT_INTERVAL)
while not _HEARTBEAT_STOP.wait(HEARTBEAT_INTERVAL):
metrics.record_heartbeat(ROLE, HEARTBEAT_INTERVAL)
audit.log_heartbeat(ROLE, HEARTBEAT_INTERVAL)
def _start_heartbeat() -> None:
"""Process-level liveness beat, tagged daemon=ROLE, independent of whether this
process also runs the notifier. web's and worker's own stdout is near-silent by
design (real logging goes to the app-json file source; worker's Docker/stdout
stream is dropped entirely by Alloy), so container-log-silence alerting can't
tell a live process from a dead one for either role this heartbeat is the
signal observability actually alerts on instead. Deliberately unguarded across
every uvicorn worker process and every autoscaled replica: a beat is a single
local file append (no upstream quota, no DB, no lock), so duplicate beats are
harmless, and only the process actually being dead stops them see
_should_run_notifier's leader election for the contrasting case where
duplication would be a real cost."""
global _HEARTBEAT_THREAD
if os.environ.get("THERMOGRAPH_ENABLE_HEARTBEAT", "1") == "0" or _HEARTBEAT_THREAD is not None:
return
_HEARTBEAT_STOP.clear()
_HEARTBEAT_THREAD = threading.Thread(target=_heartbeat_loop, name=f"heartbeat-{ROLE}", daemon=True)
_HEARTBEAT_THREAD.start()
@contextlib.asynccontextmanager
async def _lifespan(app):
# Warm the local place-name index (/suggest's typo tolerance) in the
@ -167,6 +200,7 @@ async def _lifespan(app):
await db.create_db_and_tables()
places.start_loading()
threading.Thread(target=_neighbor_warmer, name="neighbor-warmer", daemon=True).start()
_start_heartbeat()
# Background subscription evaluator. Its periodic sweep fetches upstream on a timer
# (not per request), so with multiple workers — or multiple *hosts* under Swarm —
# it must run in exactly one: elect a leader (see singleton.claim_leader, which
@ -189,6 +223,7 @@ async def _lifespan(app):
yield
audit.log_activity("app.stop", {"role": ROLE, "pid": os.getpid()})
notify.stop()
_HEARTBEAT_STOP.set()
await _frontend_client.aclose()

View file

@ -1,181 +1,72 @@
# Thermograph frontend — agent instructions
# frontend/ — agent instructions
This repo is the **SSR + static-asset service** split out of the `emi/thermograph`
monorepo (repo-split Stage 7). It has no climate data, no DB, and does no
polars/compute work — everything comes from the backend's content API over
HTTP. It is one of four sibling repos in `thermograph-repos/`:
The **SSR + static-asset service**: server-rendered crawlable pages, the
interactive tool's SPA shells, and every static asset. No climate data, no DB,
no compute — everything comes from `backend/`'s content API over HTTP.
- **`thermograph-backend`** — FastAPI API + grading/scoring/grid compute + DB.
This repo's only dependency.
- **`thermograph-frontend`** (this repo) — SSR content pages + the interactive
tool's SPA shells + every static asset.
- **`thermograph-infra`** — Terraform, `docker-compose*.yml`, `deploy/` (incl.
`deploy.sh`, the SOPS secrets vault), Caddy config. No app code.
- **`thermograph-docs`** — architecture decision records + operator runbooks.
No code.
Read the root `CLAUDE.md` first for the branch model, deploy contract and image
names.
See `MONOREPO-CUTOVER-PLAN.md` (one level up, in `thermograph-repos/`) for the
full split status and the remaining gap list before the monorepo can be
archived.
## Build & verify
## What this repo is
- `make test` — whole suite (`scripts/test.sh`).
- `make test-unit` — hermetic unit tier: SSR rendering fed committed fixtures,
no Docker. **This is the tier CI runs.**
- `make test-integration` — pulls and runs the real backend image and tests the
live contract. Local only; CI does not run it.
- `make backend-up` / `make backend-down` — a local backend container for dev.
- `make capture-fixtures` — refresh `tests/fixtures/*.json` from a live backend.
- **`content.py`** — server-rendered, crawlable pages (climate hub, per-city,
month, records, glossary, about, privacy) + `robots.txt` + `sitemap.xml`.
Every route fetches its data from the backend's content API
(`api_client.py`) instead of computing in-process; `content_payloads.py` on
the backend owns `page_title`/`canonical_path`/`breadcrumb`/`jsonld`, not
this repo.
- **`static/*.js`** — the interactive tool: `app.js` (map/search/graded
results + inline SVG chart), `calendar.js`/`day.js`/`score.js`/`compare.js`
(SPA shells served by `app.py`'s `_page()`), `account.js` (auth), `cache.js`
(IndexedDB response cache + `/cell` bundle prefetch), `shared.js` (format
helpers shared across views).
- **`static/style.css`** — the single hand-written stylesheet; all design
tokens live here (see `DESIGN.md`).
- **`templates/*.html.j2`** — Jinja templates `content.py` renders.
- **`content/*.yaml`** — structured SSR copy (glossary, static-page SEO meta),
loaded by `content_loader.py`. Committed here as a starter copy extracted
alongside the split; real cross-repo copy vendoring (a pinned
`thermograph-copy` checkout at build time) is deferred, unbuilt follow-up —
don't assume it exists.
## Running it
## Branch/PR + deploy flow
- Work happens on feature branches; PR into `main`. (The monorepo's
`dev`→`main`→`release` three-stage promotion and its LAN `deploy-dev.yml`
path do **not** exist in this repo yet — a documented gap, see the cutover
plan §4. Don't assume a `dev` branch here.)
- **`.forgejo/workflows/build-push.yml`** — on push to `dev`/`main`/`release`
or a `v*.*.*` tag: builds THIS repo's own `Dockerfile` and pushes
`git.thermograph.org/emi/thermograph-frontend/app`, tagged `sha-<12 hex>`
(every push) and the semver tag (tag pushes only). Backend publishes its own
separate image the same way — the two are no longer one shared
`emi/thermograph/app` image.
- **`.forgejo/workflows/build.yml`** — push/PR to `main`: proves the Dockerfile
builds. It is a **build check only**, not a boot/health check — booting the
real app crashes at import without a reachable backend (see API-version
section below), so a standalone boot check would need to check out
`thermograph-backend` too. Not yet built; flagged, not silently skipped.
- **`.forgejo/workflows/deploy.yml`** — push to `main`: SSH to beta, run
`SERVICE=frontend FRONTEND_IMAGE_TAG=sha-<12 hex> /opt/thermograph/deploy/deploy.sh`
(that script lives in `thermograph-infra`). Rolls **only** the frontend
container (`--no-deps`); backend is untouched and deployed independently by
its own repo's workflow. `deploy.sh` retries the image pull for ~5 min
(Forgejo has no cross-workflow `needs:`, so this deploy can race ahead of
`build-push.yml`) and waits for a healthy backend before declaring the roll
OK (frontend's boot fetches the IndexNow key from backend — see below).
- **`.forgejo/workflows/deploy-prod.yml`** — push to `release`: same shape,
targets prod via its own `PROD_SSH_*` secret set (fully separate from beta's,
so a beta credential leak can't touch prod). Nothing else deploys to prod;
there is no release-triggered Terraform apply from this repo.
- The `SERVICE=frontend` + `FRONTEND_IMAGE_TAG=sha-<12 hex>` pair is the entire
contract into `thermograph-infra/deploy/deploy.sh`: it persists each
service's live tag in `deploy/.image-tags.env` (host-side, untracked) so a
frontend-only roll never disturbs backend's currently-running tag, and
vice versa.
## How to run / test
There is no `run.sh`/`Makefile` in this repo yet (the monorepo's `make
lan-run`/`make run`/`make stop`/`venv` targets have no home here — see the
cutover plan's gap list). Run directly:
`THERMOGRAPH_API_BASE_INTERNAL` is **required**`api_client.py` raises at
import if unset.
```bash
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
THERMOGRAPH_API_BASE_INTERNAL=http://127.0.0.1:8137 \
THERMOGRAPH_BASE=/thermograph \
THERMOGRAPH_API_BASE_INTERNAL=http://127.0.0.1:8137 THERMOGRAPH_BASE=/thermograph \
.venv/bin/uvicorn app:app --host 0.0.0.0 --port 8080
```
`THERMOGRAPH_API_BASE_INTERNAL` is **required**`api_client.py` raises
`RuntimeError` at import if it's unset. Other env vars: `THERMOGRAPH_BASE`
(default `/thermograph`; the Dockerfile sets `/` for the deployed clean-root
topology), `THERMOGRAPH_API_VERSION` (default `v2`, see below),
`THERMOGRAPH_API_BASE_PUBLIC` (browser-facing backend origin for asset URLs
when frontend and backend are cross-origin; empty = same-origin, today's
default), `THERMOGRAPH_SSR_CACHE_TTL` (default 600s, the content-API response
cache in `api_client.py`), `THERMOGRAPH_GOOGLE_VERIFY`/`THERMOGRAPH_BING_VERIFY`
(search-console `<meta>` tags).
Other env: `THERMOGRAPH_BASE` (default `/thermograph`; the Dockerfile sets `/`),
`THERMOGRAPH_API_VERSION` (default `v2`), `THERMOGRAPH_API_BASE_PUBLIC`
(browser-facing backend origin when cross-origin; empty = same-origin, today's
default), `THERMOGRAPH_SSR_CACHE_TTL` (default 600s).
**Known gap — `tests/conftest.py` does not run standalone in this repo.** It
still assumes the pre-split layout: a sibling `backend/` checkout with its own
`tests/conftest.py` (`make_history`/`make_recent`) and an `api.content_payloads`
module, neither of which exists here. It was extracted verbatim as reference
material for the real fix (a genuine cross-repo contract-test job, or a
self-contained set of fixtures owned in this repo) — not yet built. Do **not**
assume `pytest tests` passes here until that lands; `.forgejo/workflows/build.yml`
only proves the image builds, for the same reason. There is also no
`requirements-dev.txt` in this repo yet (pytest/httpx test deps aren't pinned).
## Layout
## API-version pinning contract
- **`content.py`** — SSR routes (climate hub, city, month, records, glossary,
about, privacy) + `robots.txt` + `sitemap.xml`. Each fetches from the backend
via `api_client.py`; the backend's `content_payloads.py` owns `page_title` /
`canonical_path` / `breadcrumb` / `jsonld`, not this domain.
- **`static/*.js`** — `app.js` (map/search/results + inline SVG chart),
`calendar.js`/`day.js`/`score.js`/`compare.js` (SPA shells), `account.js`
(auth), `cache.js` (IndexedDB cache + `/cell` bundle prefetch), `shared.js`.
- **`static/style.css`** — the single hand-written stylesheet; design tokens live
here, documented in `DESIGN.md`.
- **`templates/*.html.j2`**, **`content/*.yaml`** (structured SSR copy, loaded by
`content_loader.py`).
Every backend call in this repo goes through a single pinned constant instead
of scattered `api/v2/...` literals:
## Contracts with the backend
- **Python (SSR):** `api_client.py`'s module-level `API_VERSION` (default
`"v2"`, overridable via `THERMOGRAPH_API_VERSION`). Every path builder
(`hub()`, `sitemap()`, `indexnow_key()`, `home()`, `city()`, `city_month()`,
`city_records()`) reads it.
- **JS (interactive tool):** `static/account.js`'s exported `API_VERSION`
constant and the `uv(path)` helper — every fetch in `cache.js`, `account.js`
itself, etc. is built by wrapping the path in `uv(...)` rather than
hardcoding a version.
**Bump only in lockstep with a verified backend `/api/version` check** — the
backend exposes `GET {BASE}/api/version` → `{backend_version, min_frontend,
payload_ver}` (`web/app.py`'s `API_CONTRACT_VERSION`/`MIN_SUPPORTED_FRONTEND`).
Before bumping this repo's `API_VERSION`, confirm the target backend's
`backend_version` actually supports it and its `min_frontend` doesn't already
exclude the version you're moving *away* from.
**A v3 cutover would work like this:** backend mounts a new `v3 = APIRouter()`
alongside `v2`, re-registering only the changed handlers (v2 keeps serving
old clients); backend bumps `API_CONTRACT_VERSION` to `"3"` in that same PR.
Only once that's deployed and verified does this repo bump `API_VERSION` to
`"v3"` in both `api_client.py` and `account.js` — one PR, both pins together,
never one without the other. `v2` (and the `/api`, `/api/v1` aliases) stay
mounted and working until no client depends on them.
The frontend's own boot is **resilient to the backend being down**: the
`indexnow_key()` fetch in `content.py`'s `register()` is tried once eagerly
(so the common case still serves `/<key>.txt` as a plain static route), but a
failure is caught, logged, and falls back to a lazy per-request lookup
instead of crashing boot — asynchronous frontend/backend deploys depend on
this (a briefly-unreachable backend at frontend boot must be survivable, not
fatal).
## Cross-repo contracts this repo depends on
These are **live contracts with the backend** — a change on either side
without the other breaks something silently, not loudly:
- **The `/cell` bundle + ETag/`If-None-Match`**`static/cache.js` fetches
`/api/v2/cell` once per view-set and slices it for calendar/day/score/compare
instead of one request per view; conditional refetch is via `If-None-Match`
against the stored `ETag`, so an unchanged payload costs an empty 304. The
URL map (which slice serves which view) must stay in sync with
`thermograph-backend`'s route/payload shape.
- **`shared.js`'s `pctOrd()` must mirror `thermograph-backend`'s
`data/grading.py`'s `pct_ordinal()`** byte-for-byte in behavior: floor a
percentile into `1..99` (never round to 100/0 — "100th percentile" reads as
measurement error, not "as extreme as it has ever been"). Every percentile
shown anywhere (Day page, calendar tooltip, chart, city pages, homepage
strip) goes through one of these two functions; if they diverge, the same
reading says two different things on two different surfaces.
- **Unit/region logic** (`format.py`'s `F_COUNTRIES` / `static/units.js`'s
`F_REGIONS` / backend's `api/content_payloads.py`'s `F_COUNTRIES`) — the set
of Fahrenheit-using country codes must stay identical across all three;
backend has a test asserting this.
## Design & visual verification
Design tokens + conventions are documented in `DESIGN.md` (source of truth:
`static/style.css`). To see a change rendered, see `DESIGN.md`'s `make shots`
section (`tools/shoot.py`).
- **API version is pinned in exactly two places**`api_client.py`'s
`API_VERSION` (Python) and `static/account.js`'s exported `API_VERSION` plus the
`uv(path)` helper (JS). Never hardcode `api/v2/...` anywhere else. Bump both in
one PR, and only after the target backend's `/api/version` confirms it serves
that version and its `min_frontend` doesn't exclude the one you're leaving.
- **`shared.js::pctOrd()` must mirror `backend`'s `data/grading.py::pct_ordinal()`**
in behaviour — floor into `1..99`, never 0 or 100. Every percentile on every
surface goes through one of the two; if they diverge, the same reading says two
different things in two places.
- **`/cell` bundle + ETag** — `cache.js` fetches `/api/v2/cell` once per view-set
and slices it, revalidating with `If-None-Match`. The slice→view map must track
the backend's payload shape.
- **Fahrenheit country set**`format.py::F_COUNTRIES` and `static/units.js`'s
`F_REGIONS` must stay identical to the backend's `F_COUNTRIES`.
- **Boot must survive an unreachable backend.** `content.py`'s `register()` tries
the IndexNow-key fetch once, catches failure, and falls back to a lazy
per-request lookup. Asynchronous FE/BE deploys depend on this — don't make it
fatal.
## Commits & PRs
Describe only the substance of the change; concise, technical. Never mention
AI/Claude/assistants or automated authorship anywhere — no trailers,
co-authors, emoji, or "as requested"/"per the agent" narration.
Concise and technical. Never mention AI, assistants or automated authorship.

4
infra/.gitignore vendored
View file

@ -27,3 +27,7 @@ __pycache__/
deploy/.image-tags.env
deploy/.deploy.lock
deploy/.stack-image-tags.env
# LAN dev's mirror of the rendered secrets, for snap-confined Docker (see
# deploy-dev.sh / render-secrets.sh) -- plaintext, re-rendered every deploy.
deploy/dev-secrets.env

View file

@ -1,42 +1,58 @@
# thermograph-infra — agent instructions
# infra/ — agent instructions
How and where the already-built Thermograph app images run. This repo owns
Terraform, the SOPS+age secrets vault, compose files, deploy scripts, host
provisioning, and the ops cron (DB backup + IndexNow). The app repos
(`thermograph-backend`, `thermograph-frontend`) own building and testing the
images; this repo never checks out app source. The old monorepo
(`emi/thermograph`) is archived — do not point anything at it.
How and where the already-built app images run: Terraform, the SOPS+age secrets
vault, compose and Swarm stack files, deploy scripts, host provisioning, mail,
and the ops cron (DB backup + IndexNow). Read the root `CLAUDE.md` first — it
owns the branch model, the deploy contract, and which environment runs which
orchestrator.
## The four machines
## The machines
Same as ACCESS.md: **dev machine** (operator's box, LAN dev server, CI runner),
Per `ACCESS.md`: **dev** (operator's box — LAN dev server and CI runner),
**prod** (`169.58.46.181`, thermograph.org, `agent` user, passwordless sudo),
**beta** (`75.119.132.91`, beta.thermograph.org + Forgejo, `agent` user). Hosts'
`/opt/thermograph` (and LAN's `~/thermograph-dev`) are checkouts of THIS repo.
**beta** (`75.119.132.91`, beta.thermograph.org + Forgejo + Grafana/Loki,
`agent` user). All four are on a WireGuard mesh. Hosts' `/opt/thermograph` is a
checkout of **this monorepo**, not an infra-only repo.
## Deploys
## Deploy paths
- `deploy/deploy.sh` — the one deploy path (prod/beta): resets this repo's
checkout, renders secrets from the SOPS vault, pulls the per-service image
tags (`BACKEND_IMAGE_TAG`/`FRONTEND_IMAGE_TAG`, persisted in untracked
`deploy/.image-tags.env`), rolls the target service, health-checks via
container healthchecks. Invoked over SSH by the app repos' deploy workflows.
- `deploy/deploy-dev.sh` — thin LAN-dev wrapper around deploy.sh (dev compose
overlay, `~/thermograph-dev`, branch `dev`).
- Branch model: see README — infra `main` is live on prod+beta, `dev` on LAN;
`release` is currently unused. App code is env-staged via image tags; infra
is not.
- **`deploy/deploy.sh`** — the single entry point for beta and prod. Resets the
host checkout to `BRANCH` (default `main`), renders secrets, then routes:
if `/etc/thermograph/deploy-mode` says `stack` it execs
`deploy/stack/deploy-stack.sh`, otherwise it rolls compose services. Per-service
tags persist in untracked `deploy/.image-tags.env` (compose) and
`deploy/.stack-image-tags.env` (stack) so the two never mix.
- **`deploy/stack/deploy-stack.sh`** — the Swarm path, live on prod. `backend`
rolls web **and** worker; `frontend` rolls frontend; `all` runs a full
`docker stack deploy`. Start-first, health-gated, auto-rollback.
`STACK_TEST=1` rehearses the whole thing under stack name `thermograph-test`
on throwaway volumes and ports `18137`/`18080`.
- **`deploy/deploy-dev.sh`** — thin LAN-dev wrapper (dev compose overlay,
`~/thermograph-dev`). Its CI trigger is inert; see root `CLAUDE.md`.
- **`Makefile`** — compose orchestration only: `up`, `down`, `db-up`, `dev-up`,
`om-up`, `om-backfill`.
Volumes are the reason the compose project name is pinned: compose creates
`thermograph_pgdata`/`_appdata`/`_applogs`, and the Swarm stack declares those
same names as `external: true` at the same mount paths.
## Rules
- **Never run `terraform apply` casually** — no tfstate exists anywhere (never
persisted), so an apply would attempt full re-provisioning of live hosts.
Terraform here is executable documentation until state is bootstrapped.
- Secrets: only via the SOPS vault (`deploy/secrets/*.yaml`, `sops edit` +
commit + deploy). Never hand-edit `/etc/thermograph.env` on a SOPS-enabled
host — it's a rendered artifact. `secrets-guard` CI rejects plaintext.
- The ops cron (`.forgejo/workflows/ops-cron.yml`) is THE prod backup — it uses
this repo's `PROD_SSH_*` Actions secrets. If you touch it, verify a dump
actually lands in `agent@prod:~/thermograph-backups/`.
- Commits/PRs: concise and technical; never mention AI/assistants/automated
authorship.
- **Never run `terraform apply` casually.** No tfstate is persisted anywhere, so
an apply would attempt full re-provisioning of live hosts. Terraform here is
executable documentation until state is bootstrapped.
- **Secrets: SOPS vault only** (`deploy/secrets/*.yaml`, `sops edit` → commit →
deploy). Never hand-edit `/etc/thermograph.env` — it is a rendered artifact.
`secrets-guard` CI rejects plaintext. `deploy/secrets/seed-from-live.sh` reads
production secrets and is **not** for an agent to run.
- **The ops cron (`.forgejo/workflows/ops-cron.yml`, at the repo root) is THE
prod backup.** It uses `PROD_SSH_*`, not `SSH_*` — an earlier revision reused
`SSH_*` and silently dumped beta while prod had no backup at all. If you touch
it, verify a dump actually lands in `agent@prod:~/thermograph-backups/`.
- Shell here runs as root over SSH against live hosts with no test suite in
front of it. `shell-lint` CI (pinned shellcheck) is the only guard — keep the
tree at zero findings.
## Commits & PRs
Concise and technical. Never mention AI, assistants or automated authorship.

View file

@ -1,46 +1,49 @@
# thermograph-infra
# infra/
Infrastructure for [Thermograph](https://thermograph.org): Terraform host
provisioning, the SOPS+age secrets vault, Docker Swarm/WireGuard networking,
Forgejo, Caddy, and the deploy scripts that run the already-built app image on
each host. Extracted from the app monorepo (`emi/thermograph`) — the app repo
owns building and testing the app; this repo owns running it.
provisioning, the SOPS+age secrets vault, WireGuard/Swarm networking, Forgejo,
Caddy, mail, and the deploy scripts that run the already-built app images on each
host. This is a domain of the `emi/thermograph` monorepo — hosts' `/opt/thermograph`
is a checkout of the whole monorepo, and `infra/` never builds app source; it only
runs published images.
- **`terraform/`** — provisions/configures hosts (SSH-driven by default; an
optional GCP-creating module is scaffolded, no live resources yet) and
triggers each deploy. See `terraform/README.md`.
- **`deploy/secrets/`** — the git-native SOPS+age secrets vault (every app
secret, encrypted at rest, rendered at deploy time). See
`deploy/secrets/README.md`.
- **`deploy/swarm/`, `deploy/forgejo/`** — the 3-node WireGuard/Swarm cluster
that hosts Forgejo (git + CI + registry); does not run the app itself. See
`ACCESS.md` and the READMEs under each directory.
- **`deploy/deploy.sh`** — pulls the pinned app image (`IMAGE_TAG`) and rolls
the compose stack; invoked by Terraform and by the app repo's
`.forgejo/workflows/deploy.yml` over SSH.
- **`docker-compose*.yml`, `docker-stack.yml`** — how the app image runs
(compose in production today; `docker-stack.yml` is a design record for a
possible future Swarm-based app deploy, not currently live).
optional GCP-creating module is scaffolded, no live resources yet). See
`terraform/README.md`. No tfstate is persisted anywhere — treat `apply` as
executable documentation, not a routine operation.
- **`deploy/secrets/`** — the git-native SOPS+age secrets vault (every app secret,
encrypted at rest, rendered at deploy time). See `deploy/secrets/README.md`.
- **`deploy/swarm/`, `deploy/forgejo/`** — the WireGuard/Swarm cluster hosting
Forgejo (git + CI + registry). See `ACCESS.md`.
- **`deploy/deploy.sh`** — the single deploy entry point for beta and prod.
Takes `SERVICE=backend|frontend|all` plus `BACKEND_IMAGE_TAG`/`FRONTEND_IMAGE_TAG`,
resets the host checkout, renders secrets, and routes to the right orchestrator.
- **`deploy/stack/`** — the **Swarm** path, live on **prod**:
`thermograph-stack.yml` (db, web, worker, lake, daemon, frontend, autoscaler,
autoscaler-lake), `deploy-stack.sh`, `autoscale.sh`, and the LB. Rolling updates
are start-first, health-gated, with auto-rollback. `STACK_TEST=1` rehearses the
whole stack on throwaway volumes and ports.
- **`docker-compose*.yml`** — the **compose** path, live on **beta** and LAN dev
(db, backend, lake, daemon, frontend). `docker-compose.dev.yml` is the LAN
overlay; `docker-compose.openmeteo.yml` is the self-hosted Open-Meteo overlay.
The app's own source, `Dockerfile`, and build/test CI stay in the app repo —
this repo never checks out app source; hosts only pull tagged images from the
registry. See `ACCESS.md` for host access and the Swarm/Forgejo topology, and
`terraform/README.md` for the day-to-day `plan`/`apply` workflow.
Which path a host takes is decided by `/etc/thermograph/deploy-mode`: the string
`stack` makes `deploy.sh` exec `deploy/stack/deploy-stack.sh`; anything else is
compose. The workflows never need to know which mode a host runs.
## Branches & how changes reach each environment
- **`main`** — what **prod and beta** run: their `/opt/thermograph` checkouts
`git reset --hard origin/main` at the start of every deploy (`deploy/deploy.sh`).
A merge to `main` reaches those hosts on the next app deploy (or a by-hand
`deploy.sh` run); there is no separate infra deploy trigger.
- **`dev`** — what **LAN dev** runs: `~/thermograph-dev` resets to it via
`deploy/deploy-dev.sh`. Keep it fast-forwarded to `main` (infra changes are not
environment-staged today; the branches exist so LAN dev *can* trail or lead
when needed).
- **`release`** — currently consumed by nothing (prod tracks `main`, not
`release`). It exists to mirror the app repos' dev→main→release promotion
shape if per-environment infra staging is ever wanted; until then, treat
`main` as live-everywhere.
- **`main`** — what **prod and beta** run. `infra-sync.yml` fires on a push to
`main` touching `infra/**`, fast-forwards each host's `/opt/thermograph` checkout
and re-renders `/etc/thermograph.env` from the vault. It deliberately does
**not** roll any service: image tags are the app domains' axis, not infra's. A
compose or stack change that must recreate containers takes effect on the next
app deploy, or a by-hand `SERVICE=all … deploy/deploy.sh`.
- **`dev`** — what LAN dev would run via `deploy/deploy-dev.sh`. The CI trigger
for this is currently inert (the LAN box still holds a split-era checkout); use
`make dev-up` locally.
- **`release`** — consumed by app deploys only. Both hosts track infra via `main`;
prod's *app images* are staged by `release`, but its checkout follows `main`.
Note the asymmetry with the app repos: app code IS environment-staged
(dev→main→release maps to LAN→beta→prod via image tags), infra is not.
Note the asymmetry with the app domains: app code IS environment-staged
(`dev`→`main`→`release` maps to LAN→beta→prod via image tags); infra is not.

View file

@ -72,6 +72,15 @@ export COMPOSE_PROJECT_NAME="${COMPOSE_PROJECT_NAME:-thermograph-dev}"
# genuinely needs, and nothing else.
export THERMOGRAPH_SECRETS_SKIP_COMMON=1
# (1b) Docker on this box is the snap package, confined to $HOME (plus a short
# allowlist) -- it silently can't see /etc/thermograph.env at all, so
# env_file: /etc/thermograph.env resolves to nothing for every container no
# matter how correctly it's rendered. Mirror the render to a path under
# $APP_DIR too; docker-compose.dev.yml points env_file at this copy instead of
# /etc/thermograph.env. Untracked (see infra/.gitignore); re-rendered on every
# deploy, never committed.
export THERMOGRAPH_SECRETS_ENV_FILE_MIRROR="$APP_DIR/infra/deploy/dev-secrets.env"
# (2) The age key is read from wherever it is READABLE, not necessarily
# /etc/thermograph/age.key. render-secrets.sh falls back to `sudo cat` for a
# root-owned 0400 key, and this box has no passwordless sudo -- under the Forgejo

View file

@ -146,7 +146,7 @@ bug hit during the frontend Go rewrite). `ci-runner` adds `docker-ce-cli` +
Current tag: `git.thermograph.org/emi/thermograph/ci-runner:v2` (`v1` is
broken — do not register any runner against it). Rebuild/push (requires a PAT
with `write:package` scope — the embedded git-remote token lacks it, same
requirement documented in `backend-build-push.yml`):
requirement documented in `build-push.yml`):
```bash
docker build -t git.thermograph.org/emi/thermograph/ci-runner:vN deploy/forgejo/ci-runner

View file

@ -82,6 +82,11 @@ After=network-online.target
[Service]
WorkingDirectory=${RUNNER_DIR}
# CI job steps shell out to user-installed tools (sops, age, ...) that live in
# ~/.local/bin, not on systemd --user's default PATH. Without this, any
# workflow that needs one (e.g. render-secrets.sh once a host is
# SOPS-configured) fails with "not installed" even though it plainly is.
Environment=PATH=%h/.local/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/usr/games:/usr/local/games:/snap/bin
ExecStart=${RUNNER_DIR}/forgejo-runner daemon -c config.yaml
Restart=on-failure
RestartSec=5

View file

@ -127,6 +127,23 @@ render_thermograph_secrets() {
echo "!! cannot write /etc/thermograph.env (need file write access or passwordless sudo)" >&2
rc=1
fi
# Optional second copy under $HOME, for a host whose docker CLI cannot read
# /etc/thermograph.env at all -- the LAN dev box's snap-packaged Docker is
# confined to $HOME (plus a short allowlist) and silently treats anything
# under /etc as nonexistent, so env_file: /etc/thermograph.env resolves to
# nothing for every container even though the file is right there and the
# shell that rendered it can read it fine. deploy-dev.sh sets this;
# prod/beta (apt-installed Docker, no snap confinement) leave it unset and
# this is a no-op for them.
if [ "$rc" = 0 ] && [ -n "${THERMOGRAPH_SECRETS_ENV_FILE_MIRROR:-}" ]; then
if ! mkdir -p "$(dirname "$THERMOGRAPH_SECRETS_ENV_FILE_MIRROR")" \
|| ! install -m 0600 "$tmp" "$THERMOGRAPH_SECRETS_ENV_FILE_MIRROR"; then
echo "!! cannot write mirror at $THERMOGRAPH_SECRETS_ENV_FILE_MIRROR" >&2
rc=1
fi
fi
rm -f "$tmp"
return "$rc"
}

View file

@ -27,12 +27,23 @@
# prod's backend=4 / frontend=2 / db=2 allocation. `!reset` drops the base
# value (both the top-level `cpus:` and the Swarm-style `deploy.resources`
# block the base file carries for parity).
# 4. backend/daemon/lake get a second env_file entry pointing at a copy of the
# render under $APP_DIR (deploy-dev.sh sets THERMOGRAPH_SECRETS_ENV_FILE_MIRROR
# to produce it). This box's Docker is the snap package, confined to $HOME --
# it can't see /etc/thermograph.env at all (not a permissions error, it just
# doesn't exist as far as snap-confined Docker is concerned), so the base
# file's env_file entry silently loads nothing here. compose appends env_file
# lists across overlays (last-wins on duplicate keys), so this is additive:
# prod/beta never set the mirror var, so they only ever get the base entry.
services:
backend:
ports: !override
- "8137:8137"
cpus: !reset null
deploy: !reset null
env_file:
- path: ./deploy/dev-secrets.env
required: false
frontend:
ports: !reset null
cpus: !reset null
@ -42,6 +53,9 @@ services:
daemon:
cpus: !reset null
deploy: !reset null
env_file:
- path: ./deploy/dev-secrets.env
required: false
db:
cpus: !reset null
deploy: !reset null
@ -53,3 +67,6 @@ services:
# the default DB_MEMORY (raise DB_MEMORY to give dev more); shm_size stays (it's
# required shared memory for parallel query, not a limit).
mem_limit: !reset null
env_file:
- path: ./deploy/dev-secrets.env
required: false

View file

@ -46,8 +46,10 @@ services:
# to an exact minor (e.g. 2.17.2-pg18) before any host of this stack could ever
# replicate with another — a floating tag risks two hosts landing on different
# extension minors, which blocks a physical replica and risks compressed-chunk
# corruption on restore. docker-stack.yml (the Swarm interim stack) REQUIRES an
# exact pin for exactly this reason; use the SAME tag on both once you set one.
# corruption on restore. The Swarm path (deploy/stack/) enforces this a
# different way: deploy-stack.sh resolves TIMESCALEDB_IMAGE to the digest of
# whatever db is ALREADY running -- including this compose stack's
# thermograph-db-1 -- so the image under an existing volume can never drift.
image: timescale/timescaledb:${TIMESCALEDB_TAG:-latest-pg18}
environment:
POSTGRES_USER: thermograph
@ -256,7 +258,7 @@ services:
restart: unless-stopped
frontend:
# Frontend's OWN image, published by frontend-build-push.yml from
# Frontend's OWN image, published by build-push.yml (frontend leg) from
# frontend/Dockerfile (which starts the thermograph-frontend Go binary
# directly -- no THERMOGRAPH_SERVICE_ROLE process-picking). Independent
# FRONTEND_IMAGE_TAG so a frontend deploy never disturbs the backend's tag.

View file

@ -1,173 +0,0 @@
# Docker Swarm stack for the hop-1 interim cutover — see
# thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md (which this file implements) and
# thermograph-docs/architecture/repo-topology-and-infrastructure.md §6/§9.
#
# Distinct from docker-compose.yml (today's plain-compose deploy, unaffected by
# this file): Swarm pulls a pre-built image from the registry (IMAGE_TAG) rather
# than building in place, and needs Swarm-specific keys (deploy:, networks:,
# secrets:, no host-published ports on the overlay-fronted services).
#
# docker stack deploy -c docker-stack.yml thermograph
#
# `docker stack deploy` only interpolates from the invoking shell's environment,
# not a .env file — export the required vars first (or `set -a; . ./stack.env;
# set +a; docker stack deploy ...`). Required shell vars: IMAGE_TAG (the tag CI
# pushed to the registry), TIMESCALEDB_TAG (an EXACT pinned minor — see the db
# service below, hazard #7). POSTGRES_PASSWORD is NOT a shell var here: unlike
# docker-compose.yml (which interpolates it into THERMOGRAPH_DATABASE_URL at
# compose time), this file reads it as the `postgres_password` Swarm secret —
# POSTGRES_PASSWORD_FILE for db (native support in the postgres/timescaledb
# image), and the chunk-3 entrypoint shim builds THERMOGRAPH_DATABASE_URL for
# app/worker from the same secret file at container start. All `postgres_password`
# / `thermograph_*` secrets below must already exist in the Swarm (`docker secret
# create <name> -`) before the first deploy — Track B step 7 in
# thermograph-docs/runbooks/implementation-handoff.md.
#
# Migrations are NOT run inline by app/worker here (RUN_MIGRATIONS=0) — the
# runbook's Stage F brings the schema to head via a one-shot task BEFORE scaling
# app up, so multiple replicas never race Alembic and nothing runs DDL against a
# still-read-only standby mid-cutover (hazard #9):
#
# docker run --rm --network thermograph_internal \
# -e THERMOGRAPH_DATABASE_URL=... <image> migrate
services:
db:
# Pin the EXACT TimescaleDB minor — never latest-pg18 here. A floating tag can
# give two hosts different extension minors, which blocks a physical replica
# (a newer .so over an older catalog won't start) and risks compressed-chunk
# corruption on timescaledb_post_restore() (hazard #7). Get the exact X.Y.Z
# from the source database: SELECT extversion FROM pg_extension WHERE
# extname='timescaledb' — use the SAME tag docker-compose.yml's TIMESCALEDB_TAG
# is pinned to, on every host that could ever replicate with this one.
image: timescale/timescaledb:${TIMESCALEDB_TAG:?set TIMESCALEDB_TAG to an exact pinned minor, e.g. 2.17.2-pg18 -- not latest-pg18}
environment:
POSTGRES_USER: thermograph
POSTGRES_PASSWORD_FILE: /run/secrets/postgres_password
POSTGRES_DB: thermograph
DB_MEMORY: ${DB_MEMORY:-8g}
volumes:
- pgdata:/var/lib/postgresql
- ./deploy/db/init:/docker-entrypoint-initdb.d
networks:
- internal
secrets:
- postgres_password
deploy:
# Pin to the labelled DB node so replication/IO never crosses the slow WG
# uplink (hazard #14): `docker node update --label-add db=true <node>` once,
# on whichever node holds pgdata (Track B).
placement:
constraints: ["node.labels.db == true"]
resources:
limits:
cpus: "${DB_CPUS:-2}"
memory: ${DB_MEMORY:-8g}
restart_policy:
condition: on-failure
# No published port: reachable only as db:5432 on the `internal` overlay —
# Postgres must never be public (design doc §6).
# Stateless, freely-replicable web tier. ROLE=web means _should_run_notifier()
# is False unconditionally (web/app.py) — it never starts the notifier or the
# worker scheduler even if it would otherwise win leader election.
app:
image: ${IMAGE_TAG:?set IMAGE_TAG to the image CI pushed, e.g. forge.example/thermograph/app:sha-abc123}
environment:
THERMOGRAPH_BASE: /
PORT: 8137
WORKERS: ${WORKERS:-4}
THERMOGRAPH_ROLE: web
RUN_MIGRATIONS: "0"
volumes:
- appdata:/app/data
- applogs:/app/logs
networks:
- internal
secrets:
- postgres_password
- thermograph_auth_secret
- thermograph_vapid_private_key
- thermograph_vapid_public_key
deploy:
replicas: 1 # single replica this hop -- homepage.json now lives in
# Postgres (chunk 4), but the notifier/scheduler split (chunk
# 2/5) is what actually lets this go multi-replica in Phase 2
resources:
limits:
cpus: "${APP_CPUS:-4}"
restart_policy:
condition: on-failure
update_config:
order: start-first # new task must pass the image's HEALTHCHECK (now
# GET /healthz) before the old one is stopped
# No `ports:` — 127.0.0.1:8137:8137 (docker-compose.yml) has no Swarm
# equivalent: Swarm's routing mesh publishes on 0.0.0.0, which would expose
# the plaintext app un-fronted (hazard #6). Only the host's Caddy reaches
# `app`, over the `internal` overlay network — see the Caddyfile templates'
# health-gated reverse_proxy.
# Owns the notifier + worker scheduler (chunks 1/2/5). Exactly one replica —
# the leader-election guard is belt-and-suspenders, not a substitute for it.
worker:
image: ${IMAGE_TAG:?set IMAGE_TAG to the image CI pushed, e.g. forge.example/thermograph/app:sha-abc123}
environment:
THERMOGRAPH_BASE: /
WORKERS: "1"
THERMOGRAPH_ROLE: worker
# Cluster-wide advisory lock, not the flock: the flock only arbitrates
# workers on ONE host, so under Swarm it guards nothing across replicas.
THERMOGRAPH_SINGLETON_PG: "1"
RUN_MIGRATIONS: "0"
volumes:
- appdata:/app/data
- applogs:/app/logs
networks:
- internal
secrets:
- postgres_password
- thermograph_auth_secret
- thermograph_vapid_private_key
- thermograph_vapid_public_key
- thermograph_discord_webhook
deploy:
replicas: 1
resources:
limits:
cpus: "${WORKER_CPUS:-1}"
restart_policy:
condition: on-failure
# No `ports:` — the worker serves no public traffic; GET /healthz on its own
# container port is for the image's own HEALTHCHECK (Swarm task health) only.
networks:
internal:
driver: overlay
driver_opts:
# VXLAN-over-WireGuard double encapsulation needs a lower MTU than the
# ~1450 overlay default, or large payloads (a /cell bundle, the homepage
# feed) silently stall while small packets (health checks) keep passing
# (hazard #13). Validate with a real large-payload transfer, not just a
# ping, once the mesh is up (Track B).
com.docker.network.driver.mtu: "1370"
volumes:
pgdata: {}
appdata: {}
applogs: {}
# Declared here, provisioned externally (docker secret create <name> - < file) —
# never by this stack, and never committed. See Stage 0 of the cutover runbook
# for which values are continuity-critical (AUTH_SECRET, VAPID, POSTGRES_PASSWORD
# must be the EXISTING live values, not freshly generated ones).
secrets:
postgres_password:
external: true
thermograph_auth_secret:
external: true
thermograph_vapid_private_key:
external: true
thermograph_vapid_public_key:
external: true
thermograph_discord_webhook:
external: true

View file

@ -1,55 +1,52 @@
# thermograph-observability — agent instructions
# observability/ — agent instructions
The **logging/metrics stack** for the Thermograph fleet: Loki + Grafana on beta,
with a Grafana Alloy agent on every node (prod, beta, dev) shipping container and
app logs over the WireGuard mesh. Grafana is fronted by beta's Caddy at
`dashboard.thermograph.org` (Google SSO, pre-provisioned users only).
The **logging stack** for the fleet: Loki + Grafana on beta, with a Grafana Alloy
agent on every node (prod, beta, dev) shipping container and app logs over the
WireGuard mesh. Grafana is fronted by beta's Caddy at
**`dashboard.thermograph.org`** (Google SSO, pre-provisioned users only) — use
that hostname everywhere, never `grafana.thermograph.org`.
This is **operational infra config, not application code** — there is no build.
It deploys by hand: `docker compose up -d` on beta for Loki+Grafana, and the
Alloy agent (`alloy/docker-compose.agent.yml`) on each node. See the README for
the full per-node procedure.
This is operational config, not application code: there is **no build** and no
deploy automation. It ships by hand — `docker compose up -d` on beta for
Loki+Grafana, and `alloy/docker-compose.agent.yml` on each node. Read the root
`CLAUDE.md` first; the branch model there applies here too.
## Layout
- `docker-compose.yml` — the Loki + Grafana stack (runs on beta).
- `loki/config.yml` — Loki config (mesh-only, filesystem storage).
- `grafana/provisioning/` — datasource (Loki) + dashboard provider, auto-loaded
at startup. `grafana/dashboards/*.json` — the dashboards themselves.
- `grafana/provisioning/alerting/` — the alert rules, the Discord contact point
and the notification policy, also auto-loaded at startup. Same rule as the
dashboards: **the repo is the only durable path**, UI edits get overwritten.
Alerts go to Discord `#ops-alerts`, never email — beta's Grafana relays SMTP
through prod's Postfix, so email dies exactly when prod does. Every rule is
LogQL (there is no Prometheus anywhere in the fleet). Thresholds were derived
from real Loki data and the working is in the comments beside each rule —
re-derive before changing a number rather than guessing.
- `alloy/config.alloy` + `alloy/docker-compose.agent.yml` — the per-node log
shipper. The node name comes from `ALLOY_NODE` (set per host).
- `caddy-grafana.conf` — the Caddy vhost for `dashboard.thermograph.org` (lives
in beta's Caddy config, kept here for reference).
- `.env.example` — copy to `.env` on beta before `docker compose up`. `.env` is
gitignored; never commit real OAuth secrets or the admin password.
- `docker-compose.yml` — the Loki + Grafana stack (beta).
- `loki/config.yml` — mesh-only, filesystem storage.
- `grafana/provisioning/` — datasource + dashboard provider, auto-loaded at
startup. `grafana/dashboards/*.json` — the dashboards.
- `grafana/provisioning/alerting/` — alert rules, the Discord contact point, and
the notification policy.
- `alloy/config.alloy` + `alloy/docker-compose.agent.yml` — the per-node shipper.
Node name comes from `ALLOY_NODE`, set per host.
- `caddy-grafana.conf` — the beta Caddy vhost, kept here for reference.
- `.env.example` — copy to `.env` on beta. `.env` is gitignored; never commit
OAuth secrets or the admin password.
## Conventions
## Rules
- **The repo is the only durable path.** Dashboards and alerting are provisioned
from this directory at startup; edits made through Grafana's UI or API are
overwritten on the next provision. Changes go through a PR.
- **Every artifact that ships to a node is CI-validated**
(`.forgejo/workflows/observability-validate.yml`): both compose files parse, all
dashboard JSON is valid, and the Loki/provisioning YAML parses. Keep new
dashboards as valid JSON and new config as valid YAML or CI fails. The alerting
config gets a stricter third step — every rule's `condition` must name a refId
that exists, every policy must route to a receiver that exists, and a literal
Discord webhook URL in the repo is a hard failure. (A rule pointing at a missing
refId is valid YAML, provisions cleanly, and then never fires; that is precisely
the silent-no-op this whole domain exists to prevent.) The Alloy config is still
not CI-validated (needs the `alloy` binary).
- The public hostname is **`dashboard.thermograph.org`** everywhere — not
`grafana.thermograph.org`. Match it in any new comment/config.
- Secrets (OAuth client id/secret, admin password) live only in the host `.env`,
never in the repo.
(`.forgejo/workflows/observability-validate.yml`, at the repo root): both
compose files parse, all dashboard JSON is valid, the Loki and provisioning
YAML parse, and the Alloy config is checked with the pinned `alloy` binary
(v1.9.1, matching what the fleet runs).
- **Alerting gets a stricter third step.** Every rule's `condition` must name a
refId that exists, every policy must route to a receiver that exists, and a
literal Discord webhook URL in the repo is a hard failure. A rule pointing at a
missing refId is valid YAML, provisions cleanly, and then never fires — that
silent no-op is what this check exists to prevent.
- **Alerts go to Discord `#ops-alerts`, never email.** Beta's Grafana relays SMTP
through prod's Postfix, so email dies exactly when prod does.
- **Every rule is LogQL** — there is no Prometheus anywhere in the fleet.
Thresholds were derived from real Loki data and the working is in the comments
beside each rule; re-derive before changing a number rather than guessing.
- Secrets (OAuth client id/secret, admin password) live only in the host `.env`.
## Branching
## Commits & PRs
Single `main` branch, no protection. Open a PR against `main`; the validator
gates it. Deploying the change to beta / the nodes is a separate manual step
(this repo has no deploy automation — a known gap).
Concise and technical. Never mention AI, assistants or automated authorship.