render-secrets-openbao.sh resolved the AppRole file to a single host-wide default,
/etc/thermograph/openbao-approle — which is prod's. On vps2 that made beta's render
authenticate as tg-prod, and tg-host-prod.hcl denies thermograph/data/env/beta, so
beta could never render from OpenBao. It surfaces as `cannot read
thermograph/env/beta`: the policy working correctly against the wrong identity.
env-topology.sh exists precisely because vps2 runs two environments and one
host-wide marker cannot answer "which environment is this". The AppRole path is the
same question and had been given the same host-wide answer. It is now derived per
environment alongside host, checkout, branch, ports and DB role, and the renderer
resolves THERMOGRAPH_BAO_APPROLE, then TG_BAO_APPROLE, then the conventional path —
the same precedence render-secrets.sh:44 already uses for the backend selector.
Only beta differs. prod and dev keep the conventional path; dev is alone on vps1 so
there is no collision there.
Beta parity passes when the path is supplied by hand (24 keys = 8 + 16 shared);
this makes it automatic. Derivation confirmed by sourcing env-topology.sh for all
three environments.
bootstrap.sh could not run to completion on OpenBao 2.6.1:
- Release asset names were wrong. `bao_<ver>_linux_amd64.tar.gz` and
`bao_<ver>_SHA256SUMS` both 404; the assets are `openbao_<ver>_...` and
`checksums.txt`. The tarball is now saved under its real name and the checksum
grep anchored to end-of-line, because checksums.txt also lists a .sbom.json and
`sha256sum -c` resolves each line by the filename inside it.
- `openssl rand -base64 32` appends a newline and the static seal reads the key
file raw, so the service failed with `Error configuring seal "static": unknown
encoding for AES-256 key`.
- The audit stanza relied on the block label being the device type. 2.6.1 requires
explicit `type` and `path` with device settings under `options`. Worth noting the
intermediate state: with type and path but a bare `file_path`, the server starts
and silently ignores the log location, which is the worse failure given a wedged
audit device stops OpenBao answering at all.
- `disable_mlock` is unsupported in 2.6.1 and warned on every start.
- Nothing opened port 8200 and ufw defaults to deny(incoming), so vps1 could not
reach the vault. dev renders on vps1, so this would have surfaced as a timeout
during cutover rather than here.
bootstrap-policies.sh installed beta's credentials as `deploy:deploy`, but vps2 has
no `deploy` user and both environments deploy as `agent` (env-topology.sh sets
TG_SSH_TARGET=agent@ for both; deploy.yml uses one VPS2_SSH_USER). install_creds
also printed a success line after a failed chown: the call site wrapped it in
`|| echo`, which suppresses `set -e` for everything inside the function, so it fell
through to chmod and reported an ownership it never applied. It now removes the
half-written file and returns non-zero.
Corrects the isolation rationale accordingly — separate credentials buy audit
attribution and a policy boundary against mistakes, not deploy-user isolation.
Documents the 2.6.0 root-token change (HCSEC-2026-08): `generate-root` now targets
an authenticated endpoint, so recovery keys cannot mint a replacement token and
revoking the last root token before a second admin identity exists is a one-way
door. Also corrects the runbook's `verify-parity.sh --all`, which cannot work from
a single host given the AppRole CIDR bindings.
forgejo_db has been at 0/1 replicas since 2026-07-29 01:24 UTC. Postgres hit
an invalid data-directory lock file ("could not open file postmaster.pid ...
performing immediate shutdown because data directory lock file is invalid")
and exited 0. With restart_policy.condition=on-failure, Swarm read the zero
status as successful completion, marked the task Complete, and never
rescheduled it.
The forgejo service itself stayed Up and kept serving its homepage, so the
outage presented as every repository page, the whole API and all CI returning
500 with "dial tcp: lookup db on 127.0.0.11:53: no such host" — including the
auth path, which is why API calls reported "user does not exist [uid: 0]"
rather than a database error.
on-failure cannot distinguish "finished successfully" from "shut itself down
and should be restarted", and Postgres exits 0 on several such paths, so it is
the wrong policy for an always-on stateful service.
This is the durable fix; it does not restart the currently stopped task.
forgejo_db has been at 0/1 replicas since 2026-07-29 01:24 UTC. Postgres hit
an invalid data-directory lock file ("could not open file postmaster.pid ...
performing immediate shutdown because data directory lock file is invalid")
and exited 0. With restart_policy.condition=on-failure, Swarm read the zero
status as successful completion, marked the task Complete, and never
rescheduled it.
The forgejo service itself stayed Up and kept serving its homepage, so the
outage presented as every repository page, the whole API and all CI returning
500 with "dial tcp: lookup db on 127.0.0.11:53: no such host" — including the
auth path, which is why API calls reported "user does not exist [uid: 0]"
rather than a database error.
on-failure cannot distinguish "finished successfully" from "shut itself down
and should be restarted", and Postgres exits 0 on several such paths, so it is
the wrong policy for an always-on stateful service.
This is the durable fix; it does not restart the currently stopped task.
Phase 1 of moving the secret store to OpenBao. Nothing is cut over:
TG_SECRETS_BACKEND defaults to sops for all three environments, so SOPS
remains authoritative until an environment is flipped in a reviewed PR.
The reason for the migration is one specific failure class. infra/CLAUDE.md
requires that vps1 never hold prod credentials, but that rule is currently
enforced by a caller-side shell flag (THERMOGRAPH_SECRETS_SKIP_COMMON) at
three call sites. A fourth call site, or one by-hand render, writes both S3
keypairs, the VAPID private key, REGISTRY_TOKEN and the IndexNow/metrics
tokens onto the Forgejo CI-runner box and exits 0. policies/tg-host-dev.hcl
makes that a 403 instead: dev's identity cannot read the common path at all.
Shape:
- vps2 only, native systemd (not a container: a Swarm service is circular,
and a plain docker run is one `prune -a` from gone), raft storage (2.6.0
deprecated the file backend), static auto-unseal so a reboot cannot block
deploys, mesh-only self-signed TLS (not ACME — that would put DNS in the
boot chain of the secret store).
- AppRole per environment, CIDR-bound, 5-minute tokens. Each environment's
secret-id is owned by its own deploy user, which is a boundary that does
not exist today since prod and beta both reach the same root-owned age key.
- Render still produces the same raw unquoted dotenv: Centralis compares
/etc/thermograph.env textually and the Swarm entrypoint parses it by line.
The backend dispatch in render-secrets.sh returns early rather than
refactoring the shared write path, so the SOPS path stays byte-identical.
Verified by rendering dev (12 keys) and prod (32 keys) through it, matching
the live hosts. The duplicated write path is noted for consolidation once
prod has been on OpenBao for a release cycle.
The OpenBao renderer rejects values containing a newline or a shell
metacharacter. /etc/thermograph.env is sourced by bash in six places, so
such a value is a command-execution path on the deploy host; the SOPS path
has always had this hazard and avoids it only because every current value
happens to be alphanumeric.
Also records that the age keypair is NOT retired by this migration: it is
the backup encryption key for every off-box dump (ops-cron.yml:142,197,
30-day retention), so deleting it with the vault files would silently
destroy a month of database recoverability.
verify-parity.sh is the cutover gate — it renders both backends and diffs
them, reporting key names and counts only, never values.
Protection an admin walks through is documentation rather than a control, so
apply_to_admins goes on — but that removes the direct-push route for everyone,
including in an emergency, and the way back in should be written down before it
is needed rather than improvised.
There is no force flag; the supported route is lift the rule, act, restore it,
with both halves audited. Names the failure mode that actually justifies it: a
renamed status context makes the required check unsatisfiable, and with
apply_to_admins on that blocks every merge into dev and main with no exemption
for anyone.
.claude/BRANCHING.md — how the trunk-based flow is actually run: the serialised
merge queue (Forgejo has no native one, so whoever orchestrates is it), the two
promotions and who decides each, and the hotfix down-merge that has to happen in
the same session or it becomes a regression scheduled for the next promotion.
Two things it states because getting them wrong is expensive here. CI does not
gate PRs into `release` — pr-build.yml fires only for dev and main — so a
required check on that branch would make every production promotion unmergeable
with a 405 naming no cause; closing that gap means changing the workflow first
and turning the requirement on second. And the chain cannot fast-forward: every
promotion leaves a merge commit on the target the source never receives, so
promotions are merge commits until someone deliberately reconciles the branches.
.claude/ownership.md — which tasks may run concurrently, seamed along the four
domains, plus the shared spine that has to be serialised: the workflows (CI
lives only in the root .forgejo/), env-topology.sh, the deploy entry points and
the CLAUDE.md files. Names the recurring collisions too — a schema change spans
alembic and whatever reads the column, and if the payload shape moves it is
PAYLOAD_VER and the frontend as well, which is one task and not two.
Twelve documents under docs/onboarding/ covering orientation, local setup,
the repo map, per-domain deep dives, the cross-service contracts, CI and the
release flow, infra and secrets, observability, task recipes, and a list of
which docs in this tree are currently stale.
Every command in the setup guide was run against this checkout: the backend
suite (429 passed, 8 skipped), the frontend Go suite, a venv boot of the
backend, and a build+boot of the Go frontend.
Two findings recorded along the way:
- static/units.js's F_REGIONS is guarded by no test, despite three source
comments claiming "a test asserts all three stay identical". The Go test
only cross-checks the Go copy against the backend's Python. All four copies
are currently identical.
- backend/ and frontend/docker-compose.test.yml still default to the retired
emi/thermograph-backend/app image path, and the frontend harness pins the
split-era v0.0.2-split-ci tag.
Claude-Session: https://claude.ai/code/session_01AfXqHrxCJLs2D7hpQkiUiJ