thermograph/infra/openbao/README.md
Emi Griffith 6bd07d33cb openbao: add a dormant OpenBao backend alongside SOPS
Phase 1 of moving the secret store to OpenBao. Nothing is cut over:
TG_SECRETS_BACKEND defaults to sops for all three environments, so SOPS
remains authoritative until an environment is flipped in a reviewed PR.

The reason for the migration is one specific failure class. infra/CLAUDE.md
requires that vps1 never hold prod credentials, but that rule is currently
enforced by a caller-side shell flag (THERMOGRAPH_SECRETS_SKIP_COMMON) at
three call sites. A fourth call site, or one by-hand render, writes both S3
keypairs, the VAPID private key, REGISTRY_TOKEN and the IndexNow/metrics
tokens onto the Forgejo CI-runner box and exits 0. policies/tg-host-dev.hcl
makes that a 403 instead: dev's identity cannot read the common path at all.

Shape:
- vps2 only, native systemd (not a container: a Swarm service is circular,
  and a plain docker run is one `prune -a` from gone), raft storage (2.6.0
  deprecated the file backend), static auto-unseal so a reboot cannot block
  deploys, mesh-only self-signed TLS (not ACME — that would put DNS in the
  boot chain of the secret store).
- AppRole per environment, CIDR-bound, 5-minute tokens. Each environment's
  secret-id is owned by its own deploy user, which is a boundary that does
  not exist today since prod and beta both reach the same root-owned age key.
- Render still produces the same raw unquoted dotenv: Centralis compares
  /etc/thermograph.env textually and the Swarm entrypoint parses it by line.

The backend dispatch in render-secrets.sh returns early rather than
refactoring the shared write path, so the SOPS path stays byte-identical.
Verified by rendering dev (12 keys) and prod (32 keys) through it, matching
the live hosts. The duplicated write path is noted for consolidation once
prod has been on OpenBao for a release cycle.

The OpenBao renderer rejects values containing a newline or a shell
metacharacter. /etc/thermograph.env is sourced by bash in six places, so
such a value is a command-execution path on the deploy host; the SOPS path
has always had this hazard and avoids it only because every current value
happens to be alphanumeric.

Also records that the age keypair is NOT retired by this migration: it is
the backup encryption key for every off-box dump (ops-cron.yml:142,197,
30-day retention), so deleting it with the vault files would silently
destroy a month of database recoverability.

verify-parity.sh is the cutover gate — it renders both backends and diffs
them, reporting key names and counts only, never values.
2026-07-29 21:35:20 -07:00

14 KiB

OpenBao — the secret store, phase 1

Status: dormant. Every file here is committed but nothing is cut over. SOPS is still authoritative for dev, beta and prod, because TG_SECRETS_BACKEND defaults to sops for all three in infra/deploy/env-topology.sh. Flipping an environment is a one-line reviewed change, and it should not happen until verify-parity.sh passes.

Verified against OpenBao v2.6.1 (2026-07-22). The binary is bao.


Why do this at all

Not for encryption, and not for audit — though the audit log is a real gain. The decisive reason is one specific failure class that a file-based vault cannot close.

infra/CLAUDE.md states the rule: "vps1 must never hold prod credentials. It's the box that runs Forgejo, its CI runner, and dev's unreviewed branch — the opposite of an isolation boundary."

Today that rule is enforced by a shell flagTHERMOGRAPH_SECRETS_SKIP_COMMON=1 — set by the caller, at three separate call sites (env-topology.sh:196deploy.sh:67, deploy-dev.sh:86, infra-sync.yml:93-95). The renderer itself does not know that env=dev means never-common. A fourth call site, or one by-hand render_thermograph_secrets /opt/thermograph-dev/infra dev, writes both S3 keypairs (one read-write on the bucket holding prod's database backups), the VAPID private key that signs Web Push to real subscribers, REGISTRY_TOKEN, and the IndexNow and metrics tokens onto the CI-runner box — and exits 0.

Under policies/tg-host-dev.hcl that is not possible to get wrong. dev's identity cannot read thermograph/data/common. A misconfigured caller gets a 403 instead of a silent success. The failure class is gone, not guarded. That is what this migration buys, and it is worth the cost of running a stateful service.

The estate has already hardened this once by hand (commit 7583258, "infra-sync: refuse to render dev's vault onto a host that is not dev"). That is the shape of a problem that wants an ACL, not another guard.

What it does not buy — stated plainly

  • Rotation of provider-issued credentials. Roughly half the credential count is Discord, Contabo S3, the Forgejo registry token, VAPID, IndexNow. OpenBao cannot rotate any of them; they stay static KV entries. Contabo Object Storage is S3-compatible but exposes no IAM API, so the AWS secrets engine cannot issue against it either.
  • Dynamic database credentials. Inapplicable here for two independent reasons: the apps read THERMOGRAPH_DATABASE_URL once at import and hold a pool, so a TTL'd credential kills the pool at expiry; and there is no reload path anywhere in the codebase — no SIGHUP handler, no config re-read.
  • Better review than git. This is a real regression. Today a secret change is a PR with a diff, backstopped by secrets-guard CI. Under OpenBao it is an out-of-band API call with no diff and no review. Mitigated by committing a key-name manifest so CI can still assert no key vanished, plus the audit log and bao kv metadata as the change record — but not fully.
  • Fewer moving parts. SOPS needs no server. This adds a stateful service to keep alive, upgrade, back up and TLS-rotate, for ~40 secrets that change rarely.

Shape of the deployment

Decision Choice Why
Host vps2 only vps1 must never hold prod credentials, and with a static seal the box also holds the key that decrypts everything. Cost: vps1 needs vps2 up to deploy dev — acceptable, since when vps2 is down prod is down anyway.
Process native systemd, not a container A Swarm service is circular (deploy-stack.sh renders secrets to deploy the stack). A plain docker run makes the vault depend on the daemon deploys restart, and is one prune -a from gone. Native = one dependency, the disk.
Storage raft, single node 2.6.0 deprecated the file backend for removal in 2.7.0. Raft also has the only backup primitive (operator raft snapshot save). A two-node raft would be worse than one: quorum of two means losing either node loses writes.
Seal static auto-unseal See below.
TLS self-signed, 10y, on disk Must not come from Caddy/ACME — that would put DNS and the public internet in the boot chain of the secret store.
Auth AppRole per environment, CIDR-bound 5-minute tokens; secret-ids that only work from the right mesh address.
Consumption keep rendering a dotenv file Preserves every existing seam and keeps OpenBao a deploy-time dependency, never a runtime one.

The unseal decision, which is the crux

Static seal, key at /etc/openbao/unseal.key (0400), plus an off-box copy.

Shamir (-key-shares=5 -key-threshold=3) means the vault is sealed after every process restart — reboot, package upgrade, OOM kill — and a human must type three shares. render-secrets.sh is on the deploy path for all three environments, so a sealed vault blocks every deploy and every hotfix. With one operator the N-of-M threshold is theatre: one person holds all the shares, so it buys nothing against the real threat while costing an availability property the estate has today for free. A Contabo reboot currently needs no human at all.

OpenBao's docs hedge static seal — "carefully evaluate" — but that caveat is aimed at multi-tenant enterprises. The argument here is that static seal is not a downgrade from the status quo: /etc/thermograph/age.key is already one 0400 file per host that decrypts everything forever with no audit trail. A static seal key is the identical trust model. What changes is everything layered above it.

The one new single point of failure: lose that key and the raft snapshots are unrestorable. It must exist in ≥2 places, one not on vps2. Hard operator obligation.

A transit seal against a second OpenBao on vps1 is genuinely better — vps2's disk would no longer contain the unseal key — but the vps1 unsealer needs unsealing too, so the circularity moves rather than dissolves, and it adds "vps1 must be up before vps2's vault unseals" as a new failure mode. Revisit once the primary migration is boring.


Layout

thermograph/data/common          the 16 shared beta+prod values   (= common.yaml)
thermograph/data/env/prod        prod's 16                        (= prod.yaml)
thermograph/data/env/beta        beta's 8                         (= beta.yaml)
thermograph/data/env/dev         dev's 12  — inherits NOTHING     (= dev.yaml)
thermograph/data/centralis/prod  Centralis' 9              (= centralis.prod.yaml)
thermograph/data/ops/backup      S3 creds for ops-cron + the backup age recipient
thermograph/data/legacy/age      the age PRIVATE key — see "the age key survives"

Inheritance lives in the render order (common then env/<name>, last-wins); isolation lives in the policy. Today both live in one shell variable.

common is a top-level path rather than env/common on purpose: it makes the dev deny rule expressible as exactly one path, with no wildcard over env/* able to accidentally re-include it.

~20 of the 61 keys are not secrets (PORT, WORKERS, APP_CPUS, TIMESCALEDB_TAG, THERMOGRAPH_BASE_URL, the Discord channel IDs, the VAPID public key, …). They stay in the vault for the cutover, because byte-identical parity is the entire safety property and splitting them would destroy it. Moving pure config back into the repo is a clean follow-on that shrinks the vault-outage blast radius.


The age key survives this migration

This is the trap that would cause silent data loss.

ops-cron.yml:142 and :197 stream every off-box Postgres and Forgejo dump through pg_dump | age -r <recipient> | rclone rcat, with 30-day S3 retention. Restore uses /etc/thermograph/age.key. So the age keypair is not just the SOPS transport — it is the backup encryption key.

Deleting it along with the SOPS vault files would destroy up to 30 days of database recoverability, and nothing would notice until someone attempted a restore.

So: keep the age private key at thermograph/data/legacy/age and in the password manager, and keep it on disk until either the newest age-encrypted object in S3 has aged out, or the backup path is re-pointed at a dedicated backup recipient. Retiring SOPS and retiring age are two different projects.


Runbook

Stand it up (once, by the operator, on vps2)

sudo bash infra/openbao/bootstrap.sh          # binary, TLS, seal key, unit, init
# custody the recovery keys + /etc/openbao/unseal.key OFF the box, then:
export BAO_ADDR=https://127.0.0.1:8200 BAO_CACERT=/etc/openbao/tls/bao.crt
export BAO_TOKEN=<root token>
sudo -E bash infra/openbao/bootstrap-policies.sh   # mount, policies, approles
bao token revoke -self

Seed and prove (from your own machine — it needs your age key)

infra/openbao/seed-from-sops.sh --dry-run     # key names + counts, writes nothing
infra/openbao/seed-from-sops.sh --all
infra/openbao/verify-parity.sh --all          # MUST pass before any flip

seed-from-sops.sh reads every production secret in plaintext. Per infra/CLAUDE.md the equivalent seed-from-live.sh is explicitly not for an agent to run; this inherits that rule.

Cut an environment over

One PR per hop, following the estate's own promotion model.

   dev)
-      TG_SECRETS_BACKEND=sops
+      TG_SECRETS_BACKEND=openbao

Order: dev → beta → prod, with verify-parity.sh green throughout. Recommended gate before prod: 7 consecutive green parity runs on all three environments, run nightly from ops-cron.yml over SSH (which gives continuous evidence without granting CI any vault access).

TG_SECRETS_BACKEND=sops remains a working two-way door for the whole period. Keep the SOPS path for a full release cycle after prod flips — it is the best mitigation available for anything unforeseen.

Rollback

Revert the one-line PR and redeploy. The SOPS path is untouched by this work — verified by rendering dev (12 keys) and prod (32 keys) through it after the dispatch was added, matching the live hosts exactly.


Failure modes

Failure Effect Mitigation
Vault down/sealed at deploy Render returns 1, deploy aborts, running stack unaffected — temp-then-install means /etc/thermograph.env is never truncated Restart=always + static seal so a reboot self-heals; alert on /v1/sys/health; TG_SECRETS_BACKEND=sops is a working revert
Audit device wedges ⚠️ OpenBao stops answering requests entirely when no enabled audit device can record them — a full /var on vps2 blocks every deploy Two devices (file + syslog); logrotate signals SIGHUP or bao writes to an unlinked inode; disk alert on vps2
Static seal key lost Raft snapshots unrestorable ≥2 copies, one off-box. The one new SPOF
Raft corruption / disk loss Total vault loss Nightly operator raft snapshot save, age-encrypted to S3 beside the DB dumps. Verify a restore end-to-end before any consumer depends on it
Cold boot of vps2 Nothing needs the vault Swarm restarts from persisted specs; stack.env / thermograph.env / centralis.env all persist. This property exists because we render a file, and would be destroyed by app-native reads
Secret-id leak Useless off the mesh secret_id_bound_cidrs; revoke by accessor without knowing the value
beta credential reaches prod Cross-env compromise Separate AppRoles, separate policies, secret-ids owned by the separate deploy users (agent vs deploy) — a boundary that does not exist today, since both currently reach the same root-owned age key. Honest limit: root on vps2 defeats it, exactly as it defeats the single age key now
Upgrade to 2.7.0 Removes the file backend and built-in cloud KMS seals Already avoided by choosing raft + static. Pin the version

Deliberately out of scope

For ~40 credentials and one operator the real risk is over-engineering. Not adopted: HA/multi-node raft, namespaces, dynamic database credentials, bao agent, cert auth, PKI, identity groups, transit seal (revisit later), OIDC for the operator.

OIDC for CI is out of scope for a reason worth recording. Forgejo Actions does support OIDC — shipped in Forgejo v15.0, needing Runner > v12.5.0, enabled via enable-openid-connect: true rather than GitHub's permissions: id-token: write. But this estate runs codeberg.org/forgejo/forgejo:9-rootless (infra/deploy/forgejo/docker-stack.yml:59). That is a six-major-version upgrade, and coupling it to this migration would mean two large independent migrations at once with no way to tell which one broke. Revisit after Forgejo reaches v15; then role_type=jwt with bound_claims removes the last CI-side bearer credential.

CI keeps its SSH keys and REGISTRY_TOKEN as Forgejo secrets. CI gets no vault access at all — the SSH-in architecture already means the host renders, never the runner, which is the least-privilege topology and should be preserved rather than "upgraded".

Secret zero, honestly

Nothing eliminates it; you choose where it sits and how little of it there is. After this phase the chain terminates at 9 Forgejo Actions secrets + 1 static seal key + 4 AppRole secret-ids, versus today's 13 CI secrets + 1 omnipotent age key per host.


Verification gaps to close on the box

Flagged rather than guessed, because these could not be checked from a workstation:

  1. Whether a raft snapshot is restorable under a different seal key. Docs don't say. Test in phase 0, before anything depends on it.
  2. Exact -wrap-ttl flag spelling on bao write — documented in the response-wrapping concept page but absent from the commands index. Check bao write -h.
  3. Debian/Ubuntu package repo URL. The install docs claim .deb packages exist but give no repository; bootstrap.sh uses the GitHub release tarball with checksum verification, which is fine and also fixes a standing weakness — provision-secrets.sh installs sops and age with a bare sudo curl and no verification at all.
  4. Vault-era naming persists inside OpenBao config (vault { } stanza, X-Vault-Wrap-TTL header). Expect it; don't "fix" it.