.claude/BRANCHING.md — how the trunk-based flow is actually run: the serialised merge queue (Forgejo has no native one, so whoever orchestrates is it), the two promotions and who decides each, and the hotfix down-merge that has to happen in the same session or it becomes a regression scheduled for the next promotion. Two things it states because getting them wrong is expensive here. CI does not gate PRs into `release` — pr-build.yml fires only for dev and main — so a required check on that branch would make every production promotion unmergeable with a 405 naming no cause; closing that gap means changing the workflow first and turning the requirement on second. And the chain cannot fast-forward: every promotion leaves a merge commit on the target the source never receives, so promotions are merge commits until someone deliberately reconciles the branches. .claude/ownership.md — which tasks may run concurrently, seamed along the four domains, plus the shared spine that has to be serialised: the workflows (CI lives only in the root .forgejo/), env-topology.sh, the deploy entry points and the CLAUDE.md files. Names the recurring collisions too — a schema change spans alembic and whatever reads the column, and if the payload shape moves it is PAYLOAD_VER and the frontend as well, which is one task and not two.
5.8 KiB
thermograph monorepo — agent instructions
One repo, four domains: backend/, frontend/, infra/, observability/.
Reunified 2026-07-22 from four split repos, with history, via subtree merges.
thermograph-docs is deliberately a separate repo — cross-cutting decision
records and operator runbooks go there, not here.
Rule for this file and every domain CLAUDE.md: only statements that would
break CI if they became false, or that name a file that exists. Background,
history and rationale belong in thermograph-docs. These files are read before
every change, so a stale one is a correctness bug, not a documentation bug.
Branches and environments
| Branch | Deploys to | Workflow |
|---|---|---|
| feature branch | nothing | PR into dev |
dev |
dev (vps1, mesh-only, own Postgres) | deploy.yml |
main |
beta (beta.thermograph.org, vps2) | deploy.yml |
release |
prod (thermograph.org, vps2) | deploy.yml |
dev, main and release are protected: everything is a PR, for humans and
agents alike. Promotion is one PR per hop, dev → main → release.
.claude/BRANCHING.md is the orchestrator's runbook for that flow — the
serialised merge queue, promotions, hotfix down-merges — and
.claude/ownership.md is the module map saying which tasks may run
concurrently. Read both before dispatching parallel work.
dev is a first-class hosted environment, not just an integration branch.
It runs on vps1 — the same box as Forgejo and the monitoring stack — with its
own Postgres container, reachable only on the WireGuard mesh
(10.10.0.2:8137): no public DNS record, no Caddy site, no TLS. It is deployed
by CI like beta and prod. The desktop hosts no Thermograph environment at all;
make dev-up there is a laptop convenience for running the stack locally, not
a deployment target.
One workflow deploys everything. deploy.yml handles both services and all
three environments: the branch selects the environment (dev → vps1, main →
beta on vps2, release → prod on vps2), a matrix covers backend and frontend,
and each leg checks whether this push actually touched its domain before
rolling. build-push.yml is the same shape for images. THERMOGRAPH_ENV is the
input that tells a leg which environment it's deploying — load-bearing on vps2,
which runs beta and prod side by side and has no other way to tell them apart.
See infra/deploy/env-topology.sh for the single source of truth on where each
environment's checkout, branch, stack and ports live.
The deploy contract
One entry point, two modes, one contract:
SERVICE=backend|frontend|all BACKEND_IMAGE_TAG=sha-<12hex> FRONTEND_IMAGE_TAG=sha-<12hex> \
/opt/thermograph/infra/deploy/deploy.sh
deploy.sh resets the host checkout, renders secrets from the SOPS vault, then
either rolls compose services or — if infra/deploy/env-topology.sh says the
environment's deploy mode is stack — execs infra/deploy/stack/deploy-stack.sh.
The old host-wide /etc/thermograph/deploy-mode marker survives only as a
fallback for a by-hand run with no environment resolvable any other way; it
cannot describe vps2 alone, since vps2 runs beta and prod side by side.
- prod and beta both run Swarm, as two separate stacks co-resident on vps2.
Prod's is
infra/deploy/stack/thermograph-stack.yml(db, web, worker, lake, daemon, frontend, autoscaler, autoscaler-lake). Beta's is the separateinfra/deploy/stack/thermograph-beta-stack.yml— its services are prefixed (beta-web,beta-worker, …) because Swarm registers a service's short name as a DNS alias on every network it joins, and beta shares prod'sdatanetwork to reach the database. Beta has nodbservice of its own: one TimescaleDB instance serves both, on separate databases and separate NOSUPERUSER roles.deploy-stack.shalso offersSTACK_TEST=1: a full parallel rehearsal on throwaway volumes and ports that cannot touch live data. - dev runs compose, from
infra/docker-compose.yml(db, backend, lake, daemon, frontend), on its own host (vps1) with its own Postgres container — the one environment not sharing a database with anything else. - Each service's live tag is persisted host-side, so a single-service roll never disturbs the sibling's running tag.
Images are emi/thermograph/backend and emi/thermograph/frontend, tagged
sha-<12hex> on every push and by semver on v*.*.* tags. They build and deploy
independently — a backend change ships without a frontend deploy and vice
versa. That independence is the point of the FE/BE split and survived
reunification.
Rules that bind across domains
- CI lives only in root
.forgejo/workflows/, path-filtered per domain. Never add workflows under a domain's own.forgejo/— they are inert there and become a trap. - Secrets only via the SOPS vault (
infra/deploy/secrets/). Never hand-edit/etc/thermograph.envon a host — it is a rendered artifact.secrets-guardCI rejects plaintext.seed-from-live.shreads production secrets and is explicitly not for an agent to run. - FE and BE ship out of lockstep, so the
/api/versioncontract andPAYLOAD_VERdiscipline are load-bearing — seebackend/CLAUDE.md. - Anything named "prefetch" must never spend the Open-Meteo quota. Nominatim ≤ 1 req/s.
- The compose project name is pinned (
name: thermographininfra/docker-compose.yml); dev (and a localmake dev-up) overrides withCOMPOSE_PROJECT_NAME=thermograph-dev. Don't remove either half — the pinned name is what makes the Swarm stack's external volume names line up. CUTOVER-NOTES.mdis the source of truth for what is and isn't live yet.
Commits & PRs
Describe only the substance of the change; concise and technical. Never mention AI, Claude, assistants or automated authorship anywhere — no trailers, co-authors, emoji, or "as requested" narration.