thermograph/CLAUDE.md
Emi Griffith af21d8e477
All checks were successful
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 10s
PR build (required check) / changes (pull_request) Successful in 16s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Successful in 20s
PR build (required check) / build-backend (pull_request) Successful in 45s
PR build (required check) / gate (pull_request) Successful in 5s
docs: correct file references and the dev reachability claim
Audited the five CLAUDE.md files and all twenty-one README.md files against the
tree, machine-checking every in-repo path they name and verifying the testable
claims against the live hosts.

The one that matters is in the root file: dev was documented as reachable on
the mesh at 10.10.0.2:8137. It is not, and never was from anywhere but vps1 —
infra/docker-compose.yml binds the port to 127.0.0.1, and the address answers
from neither vps2 nor vps1 itself. Anyone following it gets a connection
refused with nothing to explain it.

The rest are stale paths, several from the reunification:

  * assetlinks.json moved under frontend/static/ in the subtree merge; the TWA
    README kept the pre-merge path in both places it names it. Following it
    would put the file where nothing serves it and Android app-link
    verification would fail silently.
  * push.py and notify.py now live in backend/notifications/.
  * INFRA.md and deploy/stack/README have never existed in this repo, in any
    branch.
  * the Caddyfile is at deploy/stack/lb/Caddyfile.
  * three bare relative paths that resolve for a reader but not from the
    directory the file sits in: units.js is the frontend's, deploy.sh is
    infra's, entrypoint.sh is the backend's.

Also records why mesh clients must pin the ROOT_URL host and not only the image
host: the registry's bearer-token realm follows ROOT_URL, so pinning
git.thermograph.org alone still sends the token request out the public route,
where the /v2/* matcher returns 403 and docker falls back to anonymous. That
surfaces as `unauthorized: reqPackageAccess`, indistinguishable from a bad
credential.

Verified true and left alone: the four-domain layout, both .claude runbooks,
the absence of any domain-level .forgejo directory, the pinned compose project
name, the deploy contract, prod's eight stack services, beta's five prefixed
ones with no db of its own, dev's five, and every documented make target.
2026-08-01 11:49:28 -07:00

6 KiB

thermograph monorepo — agent instructions

One repo, four domains: backend/, frontend/, infra/, observability/. Reunified 2026-07-22 from four split repos, with history, via subtree merges. thermograph-docs is deliberately a separate repo — cross-cutting decision records and operator runbooks go there, not here.

Rule for this file and every domain CLAUDE.md: only statements that would break CI if they became false, or that name a file that exists. Background, history and rationale belong in thermograph-docs. These files are read before every change, so a stale one is a correctness bug, not a documentation bug.

Branches and environments

Branch Deploys to Workflow
feature branch nothing PR into dev
dev dev (vps1, mesh-only, own Postgres) deploy.yml
main beta (beta.thermograph.org, vps2) deploy.yml
release prod (thermograph.org, vps2) deploy.yml

dev, main and release are protected: everything is a PR, for humans and agents alike. Promotion is one PR per hop, devmainrelease.

.claude/BRANCHING.md is the orchestrator's runbook for that flow — the serialised merge queue, promotions, hotfix down-merges — and .claude/ownership.md is the module map saying which tasks may run concurrently. Read both before dispatching parallel work.

dev is a first-class hosted environment, not just an integration branch. It runs on vps1 — the same box as Forgejo and the monitoring stack — with its own Postgres container, reachable only from vps1 itself (127.0.0.1:8137infra/docker-compose.yml binds the port to loopback, so the mesh does not reach it either; use an SSH tunnel): no public DNS record, no Caddy site, no TLS. It is deployed by CI like beta and prod. The desktop hosts no Thermograph environment at all; make dev-up there is a laptop convenience for running the stack locally, not a deployment target.

One workflow deploys everything. deploy.yml handles both services and all three environments: the branch selects the environment (dev → vps1, main → beta on vps2, release → prod on vps2), a matrix covers backend and frontend, and each leg checks whether this push actually touched its domain before rolling. build-push.yml is the same shape for images. THERMOGRAPH_ENV is the input that tells a leg which environment it's deploying — load-bearing on vps2, which runs beta and prod side by side and has no other way to tell them apart. See infra/deploy/env-topology.sh for the single source of truth on where each environment's checkout, branch, stack and ports live.

The deploy contract

One entry point, two modes, one contract:

SERVICE=backend|frontend|all  BACKEND_IMAGE_TAG=sha-<12hex>  FRONTEND_IMAGE_TAG=sha-<12hex> \
  /opt/thermograph/infra/deploy/deploy.sh

deploy.sh resets the host checkout, renders secrets from the SOPS vault, then either rolls compose services or — if infra/deploy/env-topology.sh says the environment's deploy mode is stack — execs infra/deploy/stack/deploy-stack.sh. The old host-wide /etc/thermograph/deploy-mode marker survives only as a fallback for a by-hand run with no environment resolvable any other way; it cannot describe vps2 alone, since vps2 runs beta and prod side by side.

  • prod and beta both run Swarm, as two separate stacks co-resident on vps2. Prod's is infra/deploy/stack/thermograph-stack.yml (db, web, worker, lake, daemon, frontend, autoscaler, autoscaler-lake). Beta's is the separate infra/deploy/stack/thermograph-beta-stack.yml — its services are prefixed (beta-web, beta-worker, …) because Swarm registers a service's short name as a DNS alias on every network it joins, and beta shares prod's data network to reach the database. Beta has no db service of its own: one TimescaleDB instance serves both, on separate databases and separate NOSUPERUSER roles. deploy-stack.sh also offers STACK_TEST=1: a full parallel rehearsal on throwaway volumes and ports that cannot touch live data.
  • dev runs compose, from infra/docker-compose.yml (db, backend, lake, daemon, frontend), on its own host (vps1) with its own Postgres container — the one environment not sharing a database with anything else.
  • Each service's live tag is persisted host-side, so a single-service roll never disturbs the sibling's running tag.

Images are jinemi/thermograph/backend and jinemi/thermograph/frontend, tagged sha-<12hex> on every push and by semver on v*.*.* tags. They build and deploy independently — a backend change ships without a frontend deploy and vice versa. That independence is the point of the FE/BE split and survived reunification.

Rules that bind across domains

  • CI lives only in root .forgejo/workflows/, path-filtered per domain. Never add workflows under a domain's own .forgejo/ — they are inert there and become a trap.
  • Secrets only via the SOPS vault (infra/deploy/secrets/). Never hand-edit /etc/thermograph.env on a host — it is a rendered artifact. secrets-guard CI rejects plaintext. seed-from-live.sh reads production secrets and is explicitly not for an agent to run.
  • FE and BE ship out of lockstep, so the /api/version contract and PAYLOAD_VER discipline are load-bearing — see backend/CLAUDE.md.
  • Anything named "prefetch" must never spend the Open-Meteo quota. Nominatim ≤ 1 req/s.
  • The compose project name is pinned (name: thermograph in infra/docker-compose.yml); dev (and a local make dev-up) overrides with COMPOSE_PROJECT_NAME=thermograph-dev. Don't remove either half — the pinned name is what makes the Swarm stack's external volume names line up.
  • CUTOVER-NOTES.md is the source of truth for what is and isn't live yet.

Commits & PRs

Describe only the substance of the change; concise and technical. Never mention AI, Claude, assistants or automated authorship anywhere — no trailers, co-authors, emoji, or "as requested" narration.