promote: main → release (jinemi registry namespace, Caddy trusted_proxies, CI login fix) #159

Merged
admin_emi merged 26 commits from main into release 2026-08-01 19:38:50 +00:00
Owner

Catches prod up with 26 commits. No application source changes, no migrations — the only file under backend/ or frontend/ that is not documentation or test-compose is backend/scripts/smoke.sh, whose retired image path was corrected.

What prod actually gains

  • Registry namespace. Image paths move to jinemi/thermograph/*. Prod currently resolves emi/ through a user_redirect left from renaming admin_emi — a name that has to be freed. Until this lands, prod is the last environment holding that redirect open.
  • Caddy trusted_proxies. The LB is the second Caddy hop and was overwriting X-Forwarded-Proto with http, so every OAuth redirect URI the backend built came out non-HTTPS. Discord tolerates that; Google rejects it outright, so Google sign-in cannot work in prod without this. The equivalent is already live on beta.
  • CI login fix. Removes the empty ${{ }} that silently dropped the registry-login step, which made every push anonymous.
  • Plus the doc-accuracy pass, Forgejo/dev.jinemi.com work, openbao parity gate and secrets-rendering changes.

Pre-flight

  • Both tags prod will roll to are already published: thermograph/backend:sha-af21d8e47776 and thermograph/frontend:sha-df409f88b3fd. deploy.yml does not wait on build-push, so this was checked rather than assumed.
  • Baseline for rollback: all eight services 1/1 on emi/thermograph/*:sha-deb039ee249f, thermograph.org returning 200 and /api/version payload_ver: p2. That tag exists in both namespaces, so a roll back is available either way.
  • deploy-stack.sh applies deploy/stack/lb/Caddyfile, so the LB change goes live with this deploy.
Catches prod up with 26 commits. No application source changes, no migrations — the only file under `backend/` or `frontend/` that is not documentation or test-compose is `backend/scripts/smoke.sh`, whose retired image path was corrected. **What prod actually gains** - **Registry namespace.** Image paths move to `jinemi/thermograph/*`. Prod currently resolves `emi/` through a `user_redirect` left from renaming `admin_emi` — a name that has to be freed. Until this lands, prod is the last environment holding that redirect open. - **Caddy `trusted_proxies`.** The LB is the second Caddy hop and was overwriting `X-Forwarded-Proto` with `http`, so every OAuth redirect URI the backend built came out non-HTTPS. Discord tolerates that; Google rejects it outright, so Google sign-in cannot work in prod without this. The equivalent is already live on beta. - **CI login fix.** Removes the empty `${{ }}` that silently dropped the registry-login step, which made every push anonymous. - Plus the doc-accuracy pass, Forgejo/`dev.jinemi.com` work, openbao parity gate and secrets-rendering changes. **Pre-flight** - Both tags prod will roll to are already published: `thermograph/backend:sha-af21d8e47776` and `thermograph/frontend:sha-df409f88b3fd`. `deploy.yml` does not wait on `build-push`, so this was checked rather than assumed. - Baseline for rollback: all eight services 1/1 on `emi/thermograph/*:sha-deb039ee249f`, `thermograph.org` returning 200 and `/api/version` `payload_ver: p2`. That tag exists in both namespaces, so a roll back is available either way. - `deploy-stack.sh` applies `deploy/stack/lb/Caddyfile`, so the LB change goes live with this deploy.
emig added 26 commits 2026-08-01 19:01:37 +00:00
ci: add an always-on Actions runner on vps2
All checks were successful
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / changes (pull_request) Successful in 13s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 7s
secrets-guard / encrypted (push) Successful in 7s
shell-lint / shellcheck (push) Successful in 7s
001e6b1365
The estate had exactly one registered runner, on the desktop. It went offline
at 2026-07-31 16:31Z; for the next 21 hours no PR could satisfy a required
check, no deploy could run, and the 03:00Z ops-cron -- the only backup for both
application databases and for Forgejo -- did not fire. Forgejo queued that
scheduled run rather than dropping it, so it completed on reconnect and nothing
was lost. A longer outage would have meant real gaps.

Three files claimed an "always-on Swarm-hosted runner" existed and that the
estate therefore no longer depended on the desktop. It did not exist: an early
revision of docker-stack.yml ran one as a Docker-in-Docker sidecar and it was
removed. That claim is why a single point of failure sat unnoticed. Corrected
in docker-stack.yml, forgejo/README.md and register-lan-runner.sh.

The new runner is a plain restart:always container, not a Swarm service: a
Swarm-scheduled runner cannot redeploy the Swarm that schedules it, so CI would
be gone exactly when the cluster is what is broken. It runs from
/opt/forgejo-runner rather than in place, because the checkout is reset on
every prod deploy and one `git clean -fdx` there would destroy the
registration.

vps2 runs prod, so the socket mount is bounded rather than assumed benign:
capacity 1, --cpus=2/--memory=4g on job containers, valid_volumes empty so no
job can bind-mount /etc/thermograph.env, and no thermograph network joined.
This defends against accident, not against a hostile workflow author -- stated
plainly in the compose header rather than implied.

The desktop runner stays registered as extra capacity. Nothing may assume it
is up.
openbao: make the parity gate actually work, and run it nightly
All checks were successful
PR build (required check) / changes (pull_request) Successful in 7s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 6s
secrets-guard / encrypted (push) Successful in 5s
shell-lint / shellcheck (push) Successful in 7s
7b14b9062c
Three fixes, all on the path between "the migration looks ready" and "the
migration is ready".

verify-parity.sh never sourced env-topology.sh, so TG_BAO_APPROLE was unset and
render-secrets-openbao.sh fell through to the bare default -- prod's credential
-- for every environment. On vps2 that made `--env beta` authenticate as
tg-prod and take a 403 from tg-host-prod.hcl on thermograph/data/env/beta.
`--all` failed the same way, taking dev with it. deploy.sh and
deploy-stack.sh always sourced it; only the verifier did not, which is the
worst place for the omission: the tool whose job is to notice divergence was
itself diverging from the path it verifies. It now calls thermograph_topology
per environment and takes TG_SKIP_COMMON from there rather than re-deriving it.

infra-sync.yml renders all three env files on every push touching infra/**, and
none of its three jobs sourced env-topology.sh either. TG_SECRETS_BACKEND was
therefore unset and render-secrets.sh:44 defaulted to sops. Flipping the
selector would have changed deploy.sh and deploy-stack.sh but not this
workflow, which would have kept re-rendering from SOPS -- silent while parity
holds, and a hard failure the moment the SOPS files are retired. The one-line
cutover the README describes was never sufficient on its own.

ops-cron.yml gains a secrets-parity job: prod and beta from vps2, dev from
vps1, nightly, failing loudly on a mismatch. README.md called this the gate for
a cutover; nothing implemented it. A gate that exists only in prose gets
satisfied by assertion. It grants CI no vault access -- the host renders, CI
only asks it to.

First measured run, immediately after the verifier fix:

  dev   12 keys  PASS, byte-identical
  beta  24 keys  FAIL, 3 values differ
  prod  32 keys  FAIL, 3 values differ

The same three keys on both: THERMOGRAPH_S3_SECRET_KEY,
THERMOGRAPH_LAKE_S3_SECRET_KEY, THERMOGRAPH_VAPID_CONTACT. All three live in
common.yaml, seeded 2026-07-31 02:30Z -- before the Contabo rotation landed.
dev passes because dev never layers common. One `seed-from-sops.sh --env
common` fixes beta and prod together, and the 7-day clock starts after that.
Both hostnames go on a single Caddy site block. That is load-bearing rather
than cosmetic: the /v2/* mesh-only matcher that keeps the OCI registry API
off the public internet (hazard #15) is scoped per block, so a separate
block for the new name would serve the same Forgejo with the registry open.

git.thermograph.org stays canonical — Forgejo has one ROOT_URL and builds
every absolute URL from it, so clone URLs, the OAuth callback and post-login
redirects keep naming the .org host. Flipping FORGEJO_DOMAIN additionally
requires the Google OAuth redirect URI, the registry host baked into image
names and runner labels, and the mesh /etc/hosts pins; README documents
those prerequisites.

Config verified with `caddy validate` and by asserting against the adapted
JSON that the /v2 403 guard matches both hostnames.
infra-sync: render /etc/centralis.env from the vault
All checks were successful
PR build (required check) / changes (pull_request) Successful in 7s
secrets-guard / encrypted (pull_request) Successful in 6s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
shell-lint / shellcheck (push) Successful in 8s
secrets-guard / encrypted (push) Successful in 7s
b62eb60121
render_centralis_secrets has existed since the OpenBao work and had no caller
on any deploy path -- only verify-centralis-render.sh, a check. A renderer
nothing calls is not a migration, it is a file. /etc/centralis.env has stayed
hand-maintained the whole time, with the vault holding a copy nothing consumed.

That gap is why turning on Google sign-in for Centralis had no safe home: the
only way to set CENTRALIS_GOOGLE_CLIENT_ID was to hand-edit the live file,
which is precisely what render-secrets.sh's header warns against.

THE GUARD IS THE POINT. On 2026-07-24 this estate lost Centralis' token
registry by writing a file missing a key the live one had. The renderer proves
its own output round-trips through bash, but it cannot know about a key that
exists only in the live file. So this job renders to a scratch path first,
compares KEY SETS (names only, both sides sourced in `env -i` so the
comparison is over what bash actually ends up with), and refuses to write --
failing loudly -- if the live file has anything the vault does not. Rendering
turns from a way to lose a hand-edit into the thing that catches one.

Verified the guard against a live file carrying a hand-added
CENTRALIS_GOOGLE_CLIENT_ID: correctly identified as a key that would be
deleted, write refused. Also verified the empty/unreadable-live-file path,
which needs `|| true` to survive pipefail.

It recreates the container where the other three jobs deliberately roll
nothing, and the difference is real rather than an exception: Centralis'
compose file interpolates /etc/centralis.env at PARSE time, so rendering
without recreating changes nothing -- the silent no-op this estate keeps
getting bitten by. `docker restart` would not do it either. One container, no
replicas, a two-second recreate. It also skips the recreate entirely when the
rendered file is byte-identical to the live one, so a no-op push does not
bounce a healthy container.
gitignore: ignore .claude/settings.local.json
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 7s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
5b5f28560d
It was untracked but not ignored, so a single `git add -A` would have
committed one operator's local permission allowlist — including which vault
secrets they may decrypt — as shared repo config. `.claude/settings.json`
stays tracked as the reviewed team config.
openbao: make verify-parity --all host-aware
All checks were successful
secrets-guard / encrypted (pull_request) Successful in 4s
PR build (required check) / changes (pull_request) Successful in 7s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 9s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
shell-lint / shellcheck (push) Successful in 6s
Sync infra to hosts / sync-dev (push) Successful in 5s
Sync infra to hosts / sync-centralis (push) Has been skipped
secrets-guard / encrypted (push) Successful in 4s
75968980f9
--all expanded to (dev beta prod) and then tried to check all three from
wherever it was run. No host carries all three: dev is on vps1, beta and prod
are on vps2. So --all could not succeed anywhere, and the way it failed was
misleading rather than obvious -- on vps1, prod resolves TG_BAO_APPROLE to
/etc/thermograph/openbao-approle, which on THAT box holds dev's credential, so
prod's check authenticated as tg-dev and took a 403 from tg-host-prod.hcl.
That reads like a policy bug. It is "wrong box".

--all now means "every environment that lives on THIS host". Residency is
decided by the environment's own address being bound here (TG_SSH_HOST against
the local interfaces), and skips are announced rather than silent.

Residency is deliberately NOT decided by TG_APP_DIR existing. That was the
first attempt and it was wrong: vps1 still carries a stale /opt/thermograph
from before the estate split (d5357d0, #106), so prod looked resident there and
got checked anyway. Verified against both live hosts. A leftover directory is
exactly the kind of thing that outlives the arrangement it belonged to.

Zero environments checked is now a FAILURE. Previously a run that skipped
everything would exit 0, and a green run is precisely what the 7-day cutover
gate counts -- "nothing was wrong" and "nothing was examined" must not look
alike. The success line reports how many were checked and says plainly that a
green run on one host says nothing about the other.

Live-verified on both boxes:
  vps1 --all -> dev PASS, beta and prod skipped, exit 0
  vps2 --all -> dev skipped, beta and prod checked (both currently failing on
                the three known stale common.yaml keys)
lb: trust the host Caddy's X-Forwarded-Proto
All checks were successful
secrets-guard / encrypted (pull_request) Successful in 4s
PR build (required check) / changes (pull_request) Successful in 6s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 1s
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-centralis (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 6s
secrets-guard / encrypted (push) Successful in 4s
shell-lint / shellcheck (push) Successful in 6s
3993b7552f
Every OAuth redirect URI the backend builds is currently http://, whatever
scheme the user actually arrived with.

There are two Caddy hops in front of the app: the host Caddy terminates TLS
and sets X-Forwarded-Proto: https, then proxies to 127.0.0.1:8137 over plain
HTTP; the stack LB is the second hop. Caddy preserves an incoming
X-Forwarded-* header only when the immediate peer is a trusted proxy, and
trusted_proxies was never set here -- so the LB overwrote the header with the
scheme of the connection it had just received, which is http.

accounts/oauth.py:_redirect_uri builds the callback from x-forwarded-proto,
and that URI must match what is registered in the provider's console.

Measured on vps2, same request to each hop:

  direct to web, XFP=https      -> https://thermograph.org/api/v2/...
  through the LB, XFP=https     -> http://thermograph.org/api/v2/...
  through the PATCHED LB        -> https://thermograph.org/api/v2/...
  through the PATCHED LB, no XFP-> http://thermograph.org/api/v2/...

The last line is the point of using trusted_proxies rather than hardcoding
`header_up X-Forwarded-Proto https`: the LB still reports the truth when
nothing in front of it claims otherwise.

private_ranges covers 127.0.0.1/8 and the Docker bridge ranges, which is where
the host Caddy reaches this container from. The LB binds loopback only,
precisely so it cannot be reached un-fronted, so the only party that can set
these headers is the host Caddy -- trusting the local peer widens nothing.

Applied to beta's LB too. Beta is where a provider change gets tested before
it reaches thermograph.org, so it needs the same behaviour or the test is not
a test.
registry: move image and repo references to the Jinemi namespace
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 4s
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / validate-observability (pull_request) Successful in 20s
PR build (required check) / build-frontend (pull_request) Successful in 37s
PR build (required check) / build-backend (pull_request) Successful in 51s
PR build (required check) / gate (pull_request) Successful in 1s
df409f88b3
Repos moved to the Jinemi org; the container packages did not follow, since
Forgejo does not transfer packages with a repo. The deploy path still resolved
`emi/thermograph/*` — a user_redirect to admin_emi — while build-push.yml
derives its push path from ${github.repository}, now jinemi/thermograph. The
next backend or frontend build would have published somewhere no deploy looks.

Point the image paths at jinemi/thermograph/* (lowercase: OCI references admit
no uppercase, which is why build-push.yml already pipes through tr), the clone
URLs at Jinemi/thermograph, and the registry logins at admin_emi — the account
that actually owns the tokens, rather than the redirect.

The live tags and both ci-runner tags were copied into the Jinemi namespace
first, so the switch has something to pull. thermograph-infra,
thermograph-observability and the retired */app packages stay under admin_emi;
they did not move.
Merge dev into feat/forgejo-jinemi-alias (dev moved while #148 was open)
All checks were successful
PR build (required check) / changes (pull_request) Successful in 7s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 1s
000a8a1048
Merge pull request 'forgejo: serve dev.jinemi.com alongside git.thermograph.org' (#148) from feat/forgejo-jinemi-alias into dev
All checks were successful
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 5s
Sync infra to hosts / sync-centralis (push) Has been skipped
secrets-guard / encrypted (push) Successful in 4s
shell-lint / shellcheck (push) Successful in 7s
7fca3556e8
Merge remote-tracking branch 'origin/dev' into fix/registry-namespace-admin-emi
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 4s
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / build-frontend (pull_request) Successful in 21s
PR build (required check) / validate-observability (pull_request) Successful in 22s
PR build (required check) / build-backend (pull_request) Successful in 44s
PR build (required check) / gate (pull_request) Successful in 1s
6745f8db27
forgejo: make dev.jinemi.com canonical (ROOT_URL and OAuth callback)
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
secrets-guard / encrypted (pull_request) Successful in 9s
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 1s
a83a0b7d32
Forgejo derives every absolute URL from its single ROOT_URL, so this is what
moves clone URLs, the Google OAuth callback, webhook payload URLs and mail
links onto the new domain. Both redirect URIs are registered on the Google
OAuth client and were verified against Google before the flip — login is
SSO-only, so a mismatch locks everyone out of the UI including out of undoing
it. README documents the probe so the check is repeatable rather than assumed.

git.thermograph.org stays on the same Caddy site block and stays load-bearing:
the registry host baked into image names, the mesh /etc/hosts pins that make
the /v2/* matcher see a mesh source IP, and the runners' registered instance
URL all still use it. Renaming the registry is a separate migration; the docs
now say so where someone would otherwise retire the name as dead weight.

The stack has no auto-deploy, so this takes effect on a manual
`docker stack deploy` from the manager.
Merge pull request 'forgejo: make dev.jinemi.com canonical (ROOT_URL and OAuth callback)' (#151) from feat/forgejo-oauth-jinemi into dev
All checks were successful
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 6s
Sync infra to hosts / sync-centralis (push) Has been skipped
secrets-guard / encrypted (push) Successful in 4s
shell-lint / shellcheck (push) Successful in 6s
209cce1cb5
Merge remote-tracking branch 'origin/dev' into fix/registry-namespace-admin-emi
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / validate-observability (pull_request) Successful in 16s
PR build (required check) / build-frontend (pull_request) Successful in 21s
PR build (required check) / build-backend (pull_request) Successful in 42s
PR build (required check) / gate (pull_request) Successful in 1s
73fcb174ef
Merge pull request 'registry: move image and repo references to the Jinemi namespace' (#150) from fix/registry-namespace-admin-emi into dev
Some checks failed
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 6s
Sync infra to hosts / sync-centralis (push) Has been skipped
Build + push images (Forgejo registry) / build-push (backend) (push) Failing after 19s
Build + push images (Forgejo registry) / build-push (frontend) (push) Failing after 18s
secrets-guard / encrypted (push) Successful in 4s
shell-lint / shellcheck (push) Successful in 8s
Validate observability stack / validate (push) Successful in 12s
Deploy / deploy (backend) (push) Failing after 5m19s
PR build (required check) / changes (pull_request) Successful in 7s
secrets-guard / encrypted (pull_request) Successful in 6s
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / validate-observability (pull_request) Successful in 20s
PR build (required check) / build-frontend (pull_request) Successful in 23s
Deploy / deploy (frontend) (push) Failing after 10m9s
PR build (required check) / build-backend (pull_request) Successful in 4m29s
PR build (required check) / gate (pull_request) Successful in 3s
9e0128bf01
forgejo: record why the runner is on vps2, rather than leaving a contradiction
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 1s
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 5s
Sync infra to hosts / sync-centralis (push) Has been skipped
secrets-guard / encrypted (push) Successful in 4s
shell-lint / shellcheck (push) Successful in 7s
486a194086
This README said, in bold, that a runner must never go on prod or beta. A
runner has been registered and running on vps2 since 2026-08-01. Both
statements cannot stand: an instruction the estate visibly ignores teaches
readers to ignore the next one too.

The original objection is kept rather than deleted, because it is correct on
its merits -- docker_host: automount hands job containers the host's Docker
socket, which on vps2 is root over both the prod and beta stacks. What changed
is that the alternative proved worse. The desktop was the ONLY registered
runner in the estate; when it dropped on 2026-07-31 nothing merged, nothing
deployed and the nightly backup did not fire for 21 hours. On a repo where
every branch is protected and every change is a PR, one absent runner freezes
everything, and the backup hangs off the same path.

So the section now records the trade and the bounds actually applied --
capacity 1, --cpus=2/--memory=4g on job containers, valid_volumes empty, no
thermograph network -- and states plainly that this is defence against
accident, not against a hostile workflow author. Forgejo's database already
stored VPS2_SSH_KEY, which is root on that box; what changed is that the
material is now reachable by a job rather than only at rest.

It also keeps the original advice for the case it was actually written for:
if you are adding CAPACITY rather than REDUNDANCY, raise capacity or use a box
that hosts nothing.
build-push: stop a whitespace-tainted REGISTRY_TOKEN failing as "denied"
All checks were successful
shell-lint / shellcheck (pull_request) Successful in 15s
PR build (required check) / changes (pull_request) Successful in 20s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
secrets-guard / encrypted (pull_request) Successful in 27s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 4s
shell-lint / shellcheck (push) Successful in 7s
secrets-guard / encrypted (push) Successful in 15s
d6553a7a05
`docker login --password-stdin` strips exactly ONE trailing newline. `echo
"$T"` adds one. So a secret pasted WITH a trailing newline -- which is what
copying from a terminal, or an editor that terminates files with one, produces
-- arrives at the registry as "<token>\n" and gets back:

    Error response from daemon: Get "https://.../v2/": denied:

with no further detail. That is indistinguishable from a revoked or wrong
token. On 2026-08-01 it stopped every build and deploy in the estate and cost
an afternoon of diagnosis on a credential that was in fact valid: the same
token, tested by hand, returned 200 on /v2/ and was issued a pull,push-scoped
registry token for jinemi/thermograph/backend.

So: strip leading/trailing whitespace and CR/LF before the pipe, and never
`echo` a credential into stdin.

The token now travels through the ENVIRONMENT rather than being interpolated
into the script text. `${{ }}` is substituted before bash parses the line, so a
value containing a quote or a newline changes the shape of the command itself,
not merely its arguments.

Two diagnostics, deliberately asymmetric:

  * empty/unset  -> HARD FAILURE naming where to set it. There is no case where
    proceeding helps.
  * wrong shape  -> WARNING only (length, and whether it is outside [0-9a-f]),
    then attempt the login anyway. A hard assertion on token format would block
    every build the day Forgejo changes that format, which is a worse failure
    than the one being prevented. The warning is enough to turn the registry's
    opaque "denied:" into a diagnosis.

Neither diagnostic prints the value; only its length and character class.

Note the username is not a factor: tested against the live registry, all of
admin_emi, emi and jinemi authenticate identically with a valid token and are
each issued a push-scoped token. Forgejo's container registry authenticates on
the token, not the username.
build-push: drop the empty ${{ }} that silenced the registry login
All checks were successful
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / changes (pull_request) Successful in 16s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 1s
5f7075e2cf
Forgejo evaluates ${{ }} expressions inside a step's `run:` script, comments
included. An EMPTY expression is a parse error, and the runner's response is to
drop the step -- no log line, no error, no failed status. The login step simply
never ran, so every subsequent `docker push` went out anonymous.

The registry then answers `unauthorized: reqPackageAccess` (or, depending on
the client path, `no basic auth credentials`), which reads exactly like a
revoked token or a missing scope. It is neither. Both are downstream of a
comment.

Isolated on one branch, one variable, back to back:

  * empty expression present -> both legs fail, no `Login Succeeded` in the log
  * empty expression removed -> both legs pass, `Login Succeeded` present,
    sha-df409f88b3fd published for backend and frontend

Nothing was wrong with the credential. The token, its scope and whether it sat
at repo or organization level were all ruled out first: pushes to
jinemi/thermograph/* succeed by hand from vps1 and from the runner host, with
matching and mismatched usernames, and the run log shows REGISTRY_TOKEN
arriving in the job environment.

Do not write a bare ${{ }} in a run block, in a comment or otherwise.
Merge pull request 'build-push: drop the empty ${{ }} that silenced the registry login' (#155) from fix/build-push-empty-expression into dev
All checks were successful
secrets-guard / encrypted (push) Successful in 5s
PR build (required check) / changes (pull_request) Successful in 7s
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 7s
shell-lint / shellcheck (push) Successful in 19s
PR build (required check) / validate-observability (pull_request) Successful in 19s
PR build (required check) / build-frontend (pull_request) Successful in 23s
PR build (required check) / build-backend (pull_request) Successful in 44s
PR build (required check) / gate (pull_request) Successful in 1s
62ba381d06
Merge pull request 'promote: dev → main (Forgejo on dev.jinemi.com; jinemi registry namespace)' (#152) from dev into main
All checks were successful
Sync infra to hosts / sync-dev (push) Has been skipped
Sync infra to hosts / sync-beta (push) Successful in 8s
Sync infra to hosts / sync-prod (push) Successful in 9s
Sync infra to hosts / sync-centralis (push) Successful in 7s
secrets-guard / encrypted (push) Successful in 7s
Build + push images (Forgejo registry) / build-push (frontend) (push) Successful in 25s
Validate observability stack / validate (push) Successful in 14s
shell-lint / shellcheck (push) Successful in 10s
Build + push images (Forgejo registry) / build-push (backend) (push) Successful in 1m34s
Deploy / deploy (backend) (push) Successful in 1m40s
Deploy / deploy (frontend) (push) Successful in 1m54s
955a64a18e
docs: correct file references and the dev reachability claim
All checks were successful
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 10s
PR build (required check) / changes (pull_request) Successful in 16s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Successful in 20s
PR build (required check) / build-backend (pull_request) Successful in 45s
PR build (required check) / gate (pull_request) Successful in 5s
af21d8e477
Audited the five CLAUDE.md files and all twenty-one README.md files against the
tree, machine-checking every in-repo path they name and verifying the testable
claims against the live hosts.

The one that matters is in the root file: dev was documented as reachable on
the mesh at 10.10.0.2:8137. It is not, and never was from anywhere but vps1 —
infra/docker-compose.yml binds the port to 127.0.0.1, and the address answers
from neither vps2 nor vps1 itself. Anyone following it gets a connection
refused with nothing to explain it.

The rest are stale paths, several from the reunification:

  * assetlinks.json moved under frontend/static/ in the subtree merge; the TWA
    README kept the pre-merge path in both places it names it. Following it
    would put the file where nothing serves it and Android app-link
    verification would fail silently.
  * push.py and notify.py now live in backend/notifications/.
  * INFRA.md and deploy/stack/README have never existed in this repo, in any
    branch.
  * the Caddyfile is at deploy/stack/lb/Caddyfile.
  * three bare relative paths that resolve for a reader but not from the
    directory the file sits in: units.js is the frontend's, deploy.sh is
    infra's, entrypoint.sh is the backend's.

Also records why mesh clients must pin the ROOT_URL host and not only the image
host: the registry's bearer-token realm follows ROOT_URL, so pinning
git.thermograph.org alone still sends the token request out the public route,
where the /v2/* matcher returns 403 and docker falls back to anonymous. That
surfaces as `unauthorized: reqPackageAccess`, indistinguishable from a bad
credential.

Verified true and left alone: the four-domain layout, both .claude runbooks,
the absence of any domain-level .forgejo directory, the pinned compose project
name, the deploy contract, prod's eight stack services, beta's five prefixed
ones with no db of its own, dev's five, and every documented make target.
forgejo: make the desktop runner unit restart-always, fix the vps2 config path
All checks were successful
PR build (required check) / changes (pull_request) Successful in 9s
shell-lint / shellcheck (pull_request) Successful in 9s
secrets-guard / encrypted (pull_request) Successful in 11s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
f25466ca76
Both runners exist and are up; these are the two defects found while verifying
that.

The desktop's systemd --user unit had Restart=on-failure. That is the same
distinction that took Forgejo down for 27 hours on 2026-07-29 (docker-stack.yml
records it for the db service): a daemon that exits 0 is not "finished
successfully", and on-failure cannot tell that from a clean shutdown. The
failure mode is silent — unit inactive (dead), runner offline in the UI, and
every protected-branch merge blocked on a check nothing will produce.

The StartLimit directives that bound that retry loop were in [Service], where
systemd accepts them without complaint and ignores them; the live unit was
running the 10s default rather than the intended 300s. Moved to [Unit], which
is where they are read.

runner-vps2/README told you to copy config.yaml next to docker-compose.yml, but
the compose file mounts ./data:/data and loads --config /data/config.yaml, so a
config there is invisible to the container — the daemon starts on defaults with
no --add-host, and every registry push then fails as if the credential were
wrong. vps2 had a stray copy at the documented path proving the instruction had
been followed.
Merge pull request 'docs: correct file references and the dev reachability claim' (#156) from fix/doc-accuracy into dev
All checks were successful
Validate observability stack / validate (push) Successful in 12s
Sync infra to hosts / sync-beta (push) Has been skipped
Build + push images (Forgejo registry) / build-push (frontend) (push) Successful in 8s
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-centralis (push) Has been skipped
Deploy / deploy (frontend) (push) Successful in 8s
Sync infra to hosts / sync-dev (push) Successful in 8s
secrets-guard / encrypted (push) Successful in 5s
Build + push images (Forgejo registry) / build-push (backend) (push) Successful in 29s
shell-lint / shellcheck (push) Successful in 19s
Deploy / deploy (backend) (push) Successful in 50s
83b074fa9e
Merge dev into feat/forgejo-runners (dev moved while #157 was open)
All checks were successful
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / changes (pull_request) Successful in 11s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 1s
fe844a3193
Merge pull request 'forgejo: make the desktop runner unit restart-always, fix the vps2 config path' (#157) from feat/forgejo-runners into dev
All checks were successful
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-centralis (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 6s
secrets-guard / encrypted (push) Successful in 5s
shell-lint / shellcheck (push) Successful in 18s
PR build (required check) / changes (pull_request) Successful in 7s
shell-lint / shellcheck (pull_request) Successful in 7s
secrets-guard / encrypted (pull_request) Successful in 11s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / build-backend (pull_request) Successful in 38s
PR build (required check) / validate-observability (pull_request) Successful in 45s
PR build (required check) / gate (pull_request) Successful in 1s
9f9c436424
Merge pull request 'promote: dev → main (doc accuracy pass, runner unit fix)' (#158) from dev into main
All checks were successful
Sync infra to hosts / sync-dev (push) Has been skipped
Deploy / deploy (frontend) (push) Successful in 8s
Build + push images (Forgejo registry) / build-push (frontend) (push) Successful in 14s
Sync infra to hosts / sync-beta (push) Successful in 6s
Sync infra to hosts / sync-prod (push) Successful in 5s
Sync infra to hosts / sync-centralis (push) Successful in 5s
Build + push images (Forgejo registry) / build-push (backend) (push) Successful in 24s
shell-lint / shellcheck (push) Successful in 9s
secrets-guard / encrypted (push) Successful in 11s
Validate observability stack / validate (push) Successful in 12s
Deploy / deploy (backend) (push) Successful in 1m14s
shell-lint / shellcheck (pull_request) Successful in 7s
secrets-guard / encrypted (pull_request) Successful in 11s
880ef395de
admin_emi merged commit 1afe6fbaac into release 2026-08-01 19:38:50 +00:00
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference: Jinemi/thermograph#159
No description provided.