From 890cb4a764461b5127b79286dc21f2a64607b8da Mon Sep 17 00:00:00 2001 From: Emi Griffith Date: Sun, 26 Jul 2026 12:53:33 -0700 Subject: [PATCH] docs: the orchestrator runbook and the module map MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit .claude/BRANCHING.md — how the trunk-based flow is actually run: the serialised merge queue (Forgejo has no native one, so whoever orchestrates is it), the two promotions and who decides each, and the hotfix down-merge that has to happen in the same session or it becomes a regression scheduled for the next promotion. Two things it states because getting them wrong is expensive here. CI does not gate PRs into `release` — pr-build.yml fires only for dev and main — so a required check on that branch would make every production promotion unmergeable with a 405 naming no cause; closing that gap means changing the workflow first and turning the requirement on second. And the chain cannot fast-forward: every promotion leaves a merge commit on the target the source never receives, so promotions are merge commits until someone deliberately reconciles the branches. .claude/ownership.md — which tasks may run concurrently, seamed along the four domains, plus the shared spine that has to be serialised: the workflows (CI lives only in the root .forgejo/), env-topology.sh, the deploy entry points and the CLAUDE.md files. Names the recurring collisions too — a schema change spans alembic and whatever reads the column, and if the payload shape moves it is PAYLOAD_VER and the frontend as well, which is one task and not two. --- .claude/BRANCHING.md | 190 +++++++++++++++++++++++++++++++++++++++++++ .claude/ownership.md | 60 ++++++++++++++ CLAUDE.md | 5 ++ 3 files changed, 255 insertions(+) create mode 100644 .claude/BRANCHING.md create mode 100644 .claude/ownership.md diff --git a/.claude/BRANCHING.md b/.claude/BRANCHING.md new file mode 100644 index 0000000..f4b7366 --- /dev/null +++ b/.claude/BRANCHING.md @@ -0,0 +1,190 @@ +# Orchestrator instructions — trunk-based flow + +`feat/*` → `dev` → `main` → `release`. One PR per hop, merges serialised by the +orchestrator. Forgejo has no native merge queue, so you are the merge queue. + +Use the Centralis `forge_*` tools for every remote operation — PRs, protection, +merges, tags. Use local git only for rebases you have to resolve by hand. +Never call the Forgejo API with curl and a pasted token; the server holds it. + +The same file exists in `centralis` with the same rules. Where the two repos +differ it is noted inline. + +--- + +## The branches + +Each branch has its own environment here, and `infra/deploy/env-topology.sh` is +the single source of truth for where each one lives: + +| Branch | Deploys to | Receives | +|---|---|---| +| `feat/*`, `fix/*` | nothing | your work | +| `dev` | dev — vps1, mesh-only, own Postgres | feature PRs | +| `main` | beta — beta.thermograph.org, vps2 | promotions from `dev` | +| `release` | prod — thermograph.org, vps2 | promotions from `main`, and hotfixes | + +`deploy.yml` handles all three: the branch selects the environment, a matrix +covers backend and frontend, and each leg checks whether the push actually +touched its domain before rolling. So **FE and BE ship out of lockstep** — a +backend change lands without a frontend deploy. That is the point of the split, +and it is why `/api/version` and `PAYLOAD_VER` discipline are load-bearing: +whatever you promote, the other half may already be running something older. + +`dev` is the default branch, so PRs target it without anyone remembering to +retarget. All three are protected: **everything is a PR**, for humans and agents +alike. + +## Who may merge what + +This is not the same question as how a change travels, and it is deliberately +asymmetric. `MERGE_AUTHORITY` in centralis' +`src/lib/promotion.ts` is the whole policy; `forge_pr_merge` enforces it. + +- **Into `dev` — free.** Merge feature PRs yourself, and do not leave them open. + A PR opened against `dev` and abandoned is an unfinished task wearing the + costume of a delivered one. +- **`dev` → `main` — yours, batched.** Beta is the test environment, so + withholding this merge withholds the testing. The judgement being asked of you + is *when a batch is coherent enough to be worth testing*, not whether you are + allowed to test it. +- **`main` → `release` — the owner's.** That merge deploys to prod. Call + `promote(from="main", to="release")`, which prepares the PR and returns what + would ship, the CI evidence and a recommendation. Then put it to the owner — + including "I would wait", if that is what you think. No flag overrides this, + and routing around it is not an option. + +`confirm_protected_base` is unrelated to the above: it is the hotfix path for a +*non-promotion* PR aimed straight at `main` or `release`, and it is audited +loudly. + +--- + +## Phase 1 — repo setup + +Idempotent. Re-run it whenever you are unsure; every step reports "already +correct" rather than erroring. + +1. **The three branches exist.** + `forge_branches(create={branch:"main", from:"dev"})` and the same for + `release`. If a branch exists but points elsewhere, the tool says so — that + is a divergence to reconcile deliberately, not to overwrite. +2. **`dev` is the default branch.** + `forge_repo_settings(apply_flow_defaults=true)` — sets the default branch and + permits every merge style the chain uses, including `fast-forward-only`, plus + rebase-updates (without which the merge queue's core primitive 403s). +3. **Protection.** `forge_branch_protection(branch="dev", apply=true)`, then + `main`, then `release`. + + `dev` and `main` get the required check; **`release` does not**, and that is + deliberate. `pr-build.yml` fires only for PRs into `dev` and `main`, so a PR + into `release` reports no status at all — requiring a context nothing + produces would make every production promotion permanently unmergeable, with + a 405 that names no cause. This estate has already lost an afternoon to + exactly that, from a single mistyped check name. If you want it closed + properly, do it in this order: add `release` to `pr-build.yml`'s + `on.pull_request.branches`, confirm on a real PR that the context appears, + *then* pass `require_checks_on_release: true`. The other order blocks prod. +4. **Stale branches.** `forge_branches()` lists every branch with its age, how + much unmerged work it carries, and a verdict: + - `merged` — nothing on it the trunk lacks. Delete freely. + - `stale` — unmerged work, untouched 14+ days. **File a `stale-branch: ` + issue summarising its diff first**, then delete. Never discard work + silently. + - `ageing` — 2+ days with work still on it. Report it; do not delete. +5. **Module map.** `.claude/ownership.md` — read it before partitioning work. + +## Phase 2 — the ongoing loop + +### Dispatching subagents + +Partition by module using `.claude/ownership.md`. Never give two concurrent +agents tasks touching the same files; if that is unavoidable, sequence them. + +Inject this into every subagent task: + +``` +Branch off latest `dev` as feat/. Touch only these files/modules: . +One logical change, small diff. Rebase onto dev before opening your PR. +PR targets `dev`. The description must list: what changed, how it was tested, +files touched. Do not merge your own PR — the orchestrator handles all merges. +``` + +Keep a branch alive **less than a day**. A task bigger than that is several +stacked tasks. + +### The merge queue + +You merge **one PR into `dev` at a time**. For each: + +1. `forge_prs(base="dev")` — the queue, with each PR's head sha and triage flag. +2. `forge_pr_update(pr=N)` — replay it onto the current tip. If the head sha + does not move, it was already current and no CI will re-fire; do not sit + waiting for a run that was never scheduled. +3. `forge_pr_await(pr=N, until="checks_complete")` — checks on the **rebased** + head. The run you saw before the update was for a merge base that no longer + exists. +4. `forge_pr_merge(pr=N, method="rebase", delete_branch=true)`. +5. Only then move to the next PR. + +`block_on_outdated_branch` is set on `dev`, so Forgejo enforces step 2 rather +than trusting you to remember it: after one merge, the next PR is refused until +it has been replayed. Treat that 405 as the queue working, not as a fault. + +If a rebase conflicts: resolve it yourself if it is trivial and mechanical +(under ~20 lines), otherwise hand it back to the owning agent with the conflict +context. Do not resolve a large conflict on someone else's behalf. + +### Promotion: `dev` → `main` + +Promote when `dev` is green and a coherent batch is complete — not on a timer, +not mid-feature. `promote(from="dev", to="main")` opens the PR with the +divergence, the conflict prediction and the changelog written into its body. +Merge it with `method="merge"`. + +Unready work belongs behind a feature flag, or not on `dev` yet. **Never promote +around it with cherry-picks.** + +`promote` refuses a hop whose two tips have identical trees — the commit counts +look like real work while the content is already on the target under different +SHAs, and merging it would fire the target's deploys for nothing. + +### Promotion: `main` → `release` + +Prepare with `promote`, then ask. On the owner's yes, merge with `method="merge"` +and tag it: `forge_releases(create={tag:"v", target:"release", ...})`. + +**On fast-forward.** The written policy is that `release` moves only by +fast-forward. It cannot today: every promotion so far has been a merge commit +*into* the target, which the source never receives, so the branches are mutually +divergent by construction and Forgejo will refuse a `fast-forward-only` merge — +correctly. The style is implemented and accepted by `forge_pr_merge`; adopting +it needs a one-time reconciliation of the protected branches, which is the +owner's decision. Until that happens, promotions are merge commits. Do not +attempt to "fix" this with a rebase or a force-push. + +### Hotfix + +1. `forge_branches(create={branch:"hotfix/", from:"release"})`. +2. Fix, PR into `release`, checks green, merge with + `confirm_protected_base: true`, tag `v`. +3. **Immediately down-merge** `release` → `main` and `main` → `dev`, as merges, + not cherry-picks, in the same session. A hotfix that exists only on `release` + is a regression scheduled for the next promotion. + +Note that CI does not gate PRs into `release` — `pr-build.yml` fires only for +`dev` and `main`. A hotfix's green evidence therefore has to come from somewhere +else; say plainly where it came from rather than implying a check passed. + +A hotfix that touches only one domain deploys only that domain. Check the other +half's `/api/version` before assuming the fix is live end to end. + +## Guardrails + +- Never force-push `dev`, `main` or `release`. +- Never merge a red PR, hotfix included. Fix CI, or fix the tests in the same PR. +- Never merge two green PRs back to back without re-checking the second against + the new tip. +- If two down-merges or promotions conflict beyond the trivial, stop and report + rather than resolving autonomously. +- Weekly: `forge_branches()` and report anything `ageing` or `stale`. diff --git a/.claude/ownership.md b/.claude/ownership.md new file mode 100644 index 0000000..dcf8239 --- /dev/null +++ b/.claude/ownership.md @@ -0,0 +1,60 @@ +# Module map — what can be worked on concurrently + +Used for partitioning work across concurrent subagents. The rule is one owner +per file at a time: **never** give two concurrent agents tasks that touch the +same module. If two tasks genuinely need the same file, sequence them instead of +racing them — a rebase conflict costs more than the parallelism saved. + +Read this before dispatching. It answers one question: *can these two tasks run +at the same time?* + +The monorepo's four domains are the natural seam, and they are a real one: FE +and BE build and deploy independently, and `deploy.yml` checks per-leg whether a +push actually touched that domain. Two agents in two different domains will not +collide in CI either. + +## Safe to assign concurrently + +| Module | Files | Notes | +|---|---|---| +| Backend — API | `backend/api/**`, `backend/app.py` | Route handlers and serialisation. | +| Backend — core | `backend/core/**` | Grading, percentiles, climatology. The domain logic. | +| Backend — accounts | `backend/accounts/**` | Auth, sessions, OAuth links. | +| Backend — daemon | `backend/daemon/**` | Scheduled work. Anything named "prefetch" must not spend the Open-Meteo quota. | +| Backend — data & lake | `backend/data/**`, `backend/gen_era5_lake.py`, `backend/gen_cities.py` | | +| Backend — migrations | `backend/alembic/**`, `backend/alembic.ini` | One owner, always. Two agents generating revisions produce two heads. | +| Frontend — server | `frontend/server/**`, `frontend/app.py`, `frontend/api_client.py` | Go SSR and the Python app. | +| Frontend — static | `frontend/static/**` | | +| Frontend — templates | `frontend/templates/**`, `frontend/content/**`, `frontend/content.py` | | +| Infra — deploy | `infra/deploy/**` (excluding `secrets/`), `infra/docker-compose*.yml` | | +| Infra — secrets | `infra/deploy/secrets/**` | SOPS vault. One owner, and never alongside `infra/deploy`. | +| Infra — ops | `infra/ops/**`, `infra/lake-iceberg/**` | | +| Infra — terraform | `infra/terraform/**` | | +| Observability | `observability/**` | Grafana dashboards are provisioned from this JSON; a UI edit is overwritten. | +| Docs | `docs/**` | | +| Agent tooling | `.claude/hooks/**` | | + +## Serialise — shared spine + +Touched by most changes. A change here is its own task with nothing else running +against it. + +| File | Why | +|---|---| +| `.forgejo/workflows/**` | CI lives **only** here, path-filtered per domain. Every domain's changes route through the same files, so two agents editing workflows conflict even when their domains do not. | +| `CLAUDE.md` and every domain `CLAUDE.md` | Read before every change; a stale one is a correctness bug. Small edits, landed fast. | +| `CUTOVER-NOTES.md` | The source of truth for what is and is not live. | +| `infra/deploy/env-topology.sh` | The single source of truth for where each environment's checkout, branch, stack and ports live. | +| `infra/deploy/deploy.sh`, `infra/deploy/stack/deploy-stack.sh` | The deploy contract's entry points. | +| `backend/CLAUDE.md` + the `/api/version` contract | FE and BE ship out of lockstep, so `PAYLOAD_VER` discipline is load-bearing. A change to the payload shape is one task spanning both domains — not two concurrent ones. | + +## The recurring collisions + +- **A schema change is not a backend-only task.** It is `backend/alembic` plus + whatever reads the column, and if the payload shape moves it is `PAYLOAD_VER` + and the frontend too. Scope it as one task across both domains rather than + two agents discovering each other mid-flight. +- **Adding CI for a domain edits the shared workflows.** Sequence it against any + other workflow work. +- **`infra/deploy/secrets/` and `infra/deploy/` are one owner between them.** + The rendered artifact and the thing that renders it move together. diff --git a/CLAUDE.md b/CLAUDE.md index 3b640aa..ddf58a1 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -22,6 +22,11 @@ every change, so a stale one is a correctness bug, not a documentation bug. `dev`, `main` and `release` are protected: **everything is a PR**, for humans and agents alike. Promotion is one PR per hop, `dev` → `main` → `release`. +`.claude/BRANCHING.md` is the orchestrator's runbook for that flow — the +serialised merge queue, promotions, hotfix down-merges — and +`.claude/ownership.md` is the module map saying which tasks may run +concurrently. Read both before dispatching parallel work. + **`dev` is a first-class hosted environment, not just an integration branch.** It runs on vps1 — the same box as Forgejo and the monitoring stack — with its own Postgres container, reachable only on the WireGuard mesh