Audited the five CLAUDE.md files and all twenty-one README.md files against the
tree, machine-checking every in-repo path they name and verifying the testable
claims against the live hosts.
The one that matters is in the root file: dev was documented as reachable on
the mesh at 10.10.0.2:8137. It is not, and never was from anywhere but vps1 —
infra/docker-compose.yml binds the port to 127.0.0.1, and the address answers
from neither vps2 nor vps1 itself. Anyone following it gets a connection
refused with nothing to explain it.
The rest are stale paths, several from the reunification:
* assetlinks.json moved under frontend/static/ in the subtree merge; the TWA
README kept the pre-merge path in both places it names it. Following it
would put the file where nothing serves it and Android app-link
verification would fail silently.
* push.py and notify.py now live in backend/notifications/.
* INFRA.md and deploy/stack/README have never existed in this repo, in any
branch.
* the Caddyfile is at deploy/stack/lb/Caddyfile.
* three bare relative paths that resolve for a reader but not from the
directory the file sits in: units.js is the frontend's, deploy.sh is
infra's, entrypoint.sh is the backend's.
Also records why mesh clients must pin the ROOT_URL host and not only the image
host: the registry's bearer-token realm follows ROOT_URL, so pinning
git.thermograph.org alone still sends the token request out the public route,
where the /v2/* matcher returns 403 and docker falls back to anonymous. That
surfaces as `unauthorized: reqPackageAccess`, indistinguishable from a bad
credential.
Verified true and left alone: the four-domain layout, both .claude runbooks,
the absence of any domain-level .forgejo directory, the pinned compose project
name, the deploy contract, prod's eight stack services, beta's five prefixed
ones with no db of its own, dev's five, and every documented make target.
6 KiB
thermograph monorepo — agent instructions
One repo, four domains: backend/, frontend/, infra/, observability/.
Reunified 2026-07-22 from four split repos, with history, via subtree merges.
thermograph-docs is deliberately a separate repo — cross-cutting decision
records and operator runbooks go there, not here.
Rule for this file and every domain CLAUDE.md: only statements that would
break CI if they became false, or that name a file that exists. Background,
history and rationale belong in thermograph-docs. These files are read before
every change, so a stale one is a correctness bug, not a documentation bug.
Branches and environments
| Branch | Deploys to | Workflow |
|---|---|---|
| feature branch | nothing | PR into dev |
dev |
dev (vps1, mesh-only, own Postgres) | deploy.yml |
main |
beta (beta.thermograph.org, vps2) | deploy.yml |
release |
prod (thermograph.org, vps2) | deploy.yml |
dev, main and release are protected: everything is a PR, for humans and
agents alike. Promotion is one PR per hop, dev → main → release.
.claude/BRANCHING.md is the orchestrator's runbook for that flow — the
serialised merge queue, promotions, hotfix down-merges — and
.claude/ownership.md is the module map saying which tasks may run
concurrently. Read both before dispatching parallel work.
dev is a first-class hosted environment, not just an integration branch.
It runs on vps1 — the same box as Forgejo and the monitoring stack — with its
own Postgres container, reachable only from vps1 itself (127.0.0.1:8137 —
infra/docker-compose.yml binds the port to loopback, so the mesh does not
reach it either; use an SSH tunnel): no public DNS record, no Caddy site, no
TLS. It is deployed
by CI like beta and prod. The desktop hosts no Thermograph environment at all;
make dev-up there is a laptop convenience for running the stack locally, not
a deployment target.
One workflow deploys everything. deploy.yml handles both services and all
three environments: the branch selects the environment (dev → vps1, main →
beta on vps2, release → prod on vps2), a matrix covers backend and frontend,
and each leg checks whether this push actually touched its domain before
rolling. build-push.yml is the same shape for images. THERMOGRAPH_ENV is the
input that tells a leg which environment it's deploying — load-bearing on vps2,
which runs beta and prod side by side and has no other way to tell them apart.
See infra/deploy/env-topology.sh for the single source of truth on where each
environment's checkout, branch, stack and ports live.
The deploy contract
One entry point, two modes, one contract:
SERVICE=backend|frontend|all BACKEND_IMAGE_TAG=sha-<12hex> FRONTEND_IMAGE_TAG=sha-<12hex> \
/opt/thermograph/infra/deploy/deploy.sh
deploy.sh resets the host checkout, renders secrets from the SOPS vault, then
either rolls compose services or — if infra/deploy/env-topology.sh says the
environment's deploy mode is stack — execs infra/deploy/stack/deploy-stack.sh.
The old host-wide /etc/thermograph/deploy-mode marker survives only as a
fallback for a by-hand run with no environment resolvable any other way; it
cannot describe vps2 alone, since vps2 runs beta and prod side by side.
- prod and beta both run Swarm, as two separate stacks co-resident on vps2.
Prod's is
infra/deploy/stack/thermograph-stack.yml(db, web, worker, lake, daemon, frontend, autoscaler, autoscaler-lake). Beta's is the separateinfra/deploy/stack/thermograph-beta-stack.yml— its services are prefixed (beta-web,beta-worker, …) because Swarm registers a service's short name as a DNS alias on every network it joins, and beta shares prod'sdatanetwork to reach the database. Beta has nodbservice of its own: one TimescaleDB instance serves both, on separate databases and separate NOSUPERUSER roles.deploy-stack.shalso offersSTACK_TEST=1: a full parallel rehearsal on throwaway volumes and ports that cannot touch live data. - dev runs compose, from
infra/docker-compose.yml(db, backend, lake, daemon, frontend), on its own host (vps1) with its own Postgres container — the one environment not sharing a database with anything else. - Each service's live tag is persisted host-side, so a single-service roll never disturbs the sibling's running tag.
Images are jinemi/thermograph/backend and jinemi/thermograph/frontend, tagged
sha-<12hex> on every push and by semver on v*.*.* tags. They build and deploy
independently — a backend change ships without a frontend deploy and vice
versa. That independence is the point of the FE/BE split and survived
reunification.
Rules that bind across domains
- CI lives only in root
.forgejo/workflows/, path-filtered per domain. Never add workflows under a domain's own.forgejo/— they are inert there and become a trap. - Secrets only via the SOPS vault (
infra/deploy/secrets/). Never hand-edit/etc/thermograph.envon a host — it is a rendered artifact.secrets-guardCI rejects plaintext.seed-from-live.shreads production secrets and is explicitly not for an agent to run. - FE and BE ship out of lockstep, so the
/api/versioncontract andPAYLOAD_VERdiscipline are load-bearing — seebackend/CLAUDE.md. - Anything named "prefetch" must never spend the Open-Meteo quota. Nominatim ≤ 1 req/s.
- The compose project name is pinned (
name: thermographininfra/docker-compose.yml); dev (and a localmake dev-up) overrides withCOMPOSE_PROJECT_NAME=thermograph-dev. Don't remove either half — the pinned name is what makes the Swarm stack's external volume names line up. CUTOVER-NOTES.mdis the source of truth for what is and isn't live yet.
Commits & PRs
Describe only the substance of the change; concise and technical. Never mention AI, Claude, assistants or automated authorship anywhere — no trailers, co-authors, emoji, or "as requested" narration.