thermograph/infra/DEPLOY-DEV.md
Emi Griffith 43c9e2052a
Some checks failed
PR build (required check) / changes (pull_request) Successful in 7s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Failing after 9s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
secrets: vault dev and Centralis; render dev without common.yaml
Brings the last two credential stores under SOPS, so all three environments and
the control plane are managed the same way.

dev.yaml — dev was not "intentionally not wired in", it was broken

The vault README claimed dev had no real secrets. Both halves were false.

Dev needed secrets it did not have: thermograph-dev-daemon-1 has been in a crash
loop, restarting every ~60s with "neither THERMOGRAPH_INTERNAL_TOKEN nor
THERMOGRAPH_AUTH_SECRET is set; refusing to start" — so dev has had no Discord
gateway and no job timers. With THERMOGRAPH_BASE_URL unset the app also falls
back to https://thermograph.org, so dev's IndexNow pings and verification links
claimed to be prod. dev.yaml gives it its own generated AUTH_SECRET and a real
base URL, and 12 values total.

And dev already held production credentials: backend-deploy-dev.yml injects
S3_ACCESS_KEY/S3_SECRET_KEY into every dev deploy, and the only provisioned
keypair for that bucket is read-write — on the bucket holding prod's backups.
Dev does not need them; without bucket creds the lake service 503s and history
falls through to the archive. Those workflow lines still need removing.

dev renders dev.yaml ALONE, via THERMOGRAPH_SECRETS_SKIP_COMMON=1. Eleven of
common.yaml's sixteen values are live production credentials, the render is a
plaintext concatenation consumed through env_file:, and dev is the operator's
desktop AND the Forgejo runner executing unreviewed dev-branch code with the
docker socket mounted. An override in dev.yaml would not help — last-wins
governs consumers, but prod's value is still physically a line in the file.
Verified: current renderer leaks 28 lines onto dev including both S3 keypairs
and the VAPID private key; patched gives exactly 12 keys and no prod credential.
Beta's render is byte-identical either way.

centralis.prod.yaml — its own file, its own renderer

/etc/centralis.env was hand-edited, which is how a JSON registry got written
unquoted into a shell-sourced file tonight and silently collapsed three
identities to one. The renderer now owns the file.

It is service- AND environment-scoped, not folded into prod.yaml, because
/etc/thermograph.env is loaded into every container in the app stack. Folding
these in would put CENTRALIS_AUTH_TOKEN and CENTRALIS_TOKENS — which
authenticate an endpoint carrying run_on_host and sql_query(write) — into the
web backend's environment, turning a read-anything bug in the app into a
foothold on the control plane.

Quoting is guaranteed three ways rather than assumed. Critically, `sops -d
--output-type dotenv` emits values verbatim, so reusing the existing render path
would have reproduced tonight's bug exactly. The new function decrypts to JSON
and emits POSIX single-quoted assignments; it then sources its own output in a
clean shell and compares every value before touching /etc; and it refuses to
write unless CENTRALIS_TOKENS parses as a non-empty subject->token object after
sourcing. Tested against 16 hostile values including embedded quotes, newlines
and ;rm -rf /.

Value-identity proven, not asserted: rendered vs live compared as *effective*
values on prod (a textual diff would report a false difference, and would report
a false match if the vault had captured quote characters as part of the value),
plus both env files producing a byte-identical `docker compose config` digest.
9 keys, both token subjects preserved. Nothing deployed.

Note for later: /etc/thermograph.env is also shell-sourced, and survives on raw
dotenv only because none of its 32 current values contains a shell-special
character. It is one quoted secret away from the same bug.

Claude-Session: https://claude.ai/code/session_0182KTMrsTHJc3TcewCatJFY
2026-07-24 18:10:51 -07:00

12 KiB

Dev CI/CD → LAN server on this machine

Parallel to the prod pipeline (main → VPS, see DEPLOY.md), the dev branch continuously deploys to a LAN server running on this computer. Git hosting and CI are self-hosted Forgejo (git.thermograph.org, reachable at http://10.10.0.2:3080 over the WireGuard mesh) — GitHub is retired.

open PR ──▶ CI (build + boot/health, on the LAN Forgejo runner)
              │ green
              ▼
        merge into dev (explicit — Forgejo does not auto-merge on its own;
                         see "Landing a PR" below)
              │
              ▼
        LAN runner (thermograph-lan label) runs deploy/deploy-dev.sh
              │
              ▼
        ~/thermograph-dev updated ─▶ docker compose stack (backend + frontend +
                                     Postgres 18) brought up, backend on 0.0.0.0:8137

Reachable at http://<lan-ip>:8137/ from any device on your Wi-Fi.

The dev server runs the same containerized stack as prod (docker-compose.yml), overlaid with docker-compose.dev.yml: uncapped (no CPU limits, unlike prod's backend=4 / frontend=2 / db=2) and backend published on 0.0.0.0:8137 for the LAN (prod binds loopback behind Caddy; frontend has no published port here either way, no Caddy to reach it directly, so it's only reached through backend's own reverse-proxy fallback). Bring it up by hand with make dev-up.

Why it's built this way

Forgejo has native branch protection + required-status-checks + auto-merge (unlike GitHub on a private free-tier repo, where those are paywalled). But "auto-merge" here still means opting a PR in — either clicking "Auto merge when checks succeed" in the web UI, or merging explicitly once CI is green via the API (POST .../pulls/{n}/merge, {"Do": "squash"}). Nothing merges a ready PR for you automatically just because checks pass, unlike GitHub's old in-workflow ci-cd.yml (retired along with the rest of .github/).

Moving parts

File Purpose
.forgejo/workflows/pr-build.yml required status check for PRs into dev (calls build.yml)
.forgejo/workflows/deploy-dev.yml build + deploy on pushes to dev (direct, or a PR merge)
.forgejo/workflows/build.yml shared build gate: deps, backend tests, JS syntax check, boot/health check
deploy/deploy-dev.sh pull dev into ~/thermograph-dev, docker compose up the uncapped LAN stack, health check
deploy/secrets/dev.yaml dev's SOPS vault — its own secrets and config, rendered to /etc/thermograph.env at deploy time (see "Config and secrets")
docker-compose.dev.yml dev overlay: no CPU limits, backend published on 0.0.0.0:8137 (frontend unpublished)
deploy/provision-dev-lan.sh one-time sudo-free bootstrap of the checkout + runner
deploy/forgejo/register-lan-runner.sh registers the Forgejo Actions runner on this machine

deploy-dev.yml's deploy job runs on the thermograph-lan label — not [self-hosted, thermograph-lan] (the array form is what the original GitHub version used; GitHub implicitly tags every self-hosted runner self-hosted, but Forgejo's runner has only the labels it was explicitly registered with, so the array form is permanently unschedulable there — a real bug found and fixed once, worth knowing about if a "deploy" job ever silently stops appearing in the Actions history again).

deploy-dev.sh's git auth (GH_TOKEN/GITHUB_TOKEN, Forgejo's per-job token exposed under that name for compatibility) is scoped to REPO_URL's own host via a per-command http.<scheme>://<host>/.extraheader, not a hardcoded one — important if the Forgejo host ever changes, since a stale hardcoded host would make git silently skip the header (no error) and fall through to whatever ambient credentials happen to be available, not fail loudly.

The self-hosted runner

A Forgejo Actions runner (forgejo-runner) runs on this machine, labels docker (containerized jobs, node:20-bookworm, with the host's docker.sock automounted in — container.docker_host: automount in ~/forgejo-runner/config.yaml — so docker build/push work without privileged Docker-in-Docker) and thermograph-lan (host-native jobs, for deploy-dev.sh's systemctl/docker-compose calls). Installed under ~/forgejo-runner, runs as the forgejo-runner systemd --user service.

# status / logs
systemctl --user status forgejo-runner
journalctl --user -u forgejo-runner -f

# re-register if the token/host ever changes
bash deploy/forgejo/register-lan-runner.sh <forgejo_url> <registration_token>

git.thermograph.org's public DNS resolves to beta's public IP, which the registry's mesh-only ACL (see deploy/forgejo/README.md) rejects for /v2/ paths — docker build-push.yml pushes run against the host's automounted daemon, so it's this host's own resolver that needs to route git.thermograph.org over the mesh, not anything settable from a job container. If registry pushes ever start failing with an HTTP/TLS mismatch or a 403 on /v2/, check getent hosts git.thermograph.org here — it needs to resolve to beta's WireGuard IP (10.10.0.2), typically via a /etc/hosts line, not the public one.

The LAN dev service (Docker Compose)

Runs as a Docker Compose stack, not a bare systemd unit — docker ps --filter name=thermograph-dev shows thermograph-dev-backend-1, thermograph-dev-frontend-1 (repo-split Stage 4 split the single "app" container in two — frontend has no published port, reached only through backend's own reverse-proxy fallback) and thermograph-dev-db-1 (TimescaleDB). A legacy pre-container thermograph-dev.service systemd --user unit may still exist from before this cutover; it's stale and deploy-dev.sh stops/disables it on every run — don't trust systemctl --user status thermograph-dev as a liveness check, use docker ps instead.

docker ps --filter name=thermograph-dev                                          # is it up?
docker compose -f docker-compose.yml -f docker-compose.dev.yml logs -f backend frontend   # app logs (from ~/thermograph-dev)

Bootstrap (or rebuild) it by hand:

bash deploy/provision-dev-lan.sh

The dedicated checkout at ~/thermograph-dev is separate from your working tree, so a deploy never touches uncommitted edits in ~/Code/Thermograph.

Config and secrets

Dev's configuration lives in the same SOPS vault as prod's and beta's — deploy/secrets/dev.yaml, encrypted to the same age recipient, rendered to /etc/thermograph.env by the same render-secrets.sh at deploy time. Changing a dev value is sops deploy/secrets/dev.yaml → commit → deploy, exactly as for the VPS boxes. deploy/secrets/README.md is the reference; what follows is only what is specific to this machine.

Dev renders dev.yaml alone. It does not layer common.yaml, which is the fleet's shared production credential set — registry token, both S3 keypairs (one read-write pair, on the bucket that also holds prod's backups), the VAPID private key that signs push to real subscribers, the IndexNow key, the metrics token. This box is the CI runner and runs unreviewed dev-branch code; a plaintext render of that set into /etc/thermograph.env would put all of it in every dev container's environment. deploy-dev.sh exports THERMOGRAPH_SECRETS_SKIP_COMMON=1 to prevent it, and aborts the deploy if the renderer in the checkout cannot honour that.

So dev holds its own THERMOGRAPH_AUTH_SECRET (its own, not prod's — the daemon service refuses to start without one), its own POSTGRES_PASSWORD, a LAN THERMOGRAPH_BASE_URL, THERMOGRAPH_COOKIE_SECURE=0 (plain HTTP — the shared value is 1 and would silently break login here), and the non-secret config it would otherwise have inherited. It holds no S3 keys, no VAPID keypair, no IndexNow key, no Discord or mail credentials: the lake falls through to the Open-Meteo archive without bucket creds, and the app generates and persists its own VAPID and IndexNow keys into the appdata volume.

Two dev-only mechanics, both because this box has no passwordless sudo and the runner is a systemd --user service with no tty:

  • the age key is read from ~/.config/sops/age/keys.txt, not /etc/thermograph/age.key — a root-owned 0400 key would send the renderer down its sudo cat path, where it would hang on a prompt no CI job can answer;
  • /etc/thermograph.env is pre-created owned by the deploying user, so the renderer takes its in-place write path and no deploy needs sudo at all.

Setting this up on a fresh dev machine is four typed commands — see "Setting up a new dev machine" in deploy/secrets/README.md. Until the /etc/thermograph/secrets-env marker exists the box renders nothing and falls back to deploy-dev.sh's built-in POSTGRES_PASSWORD default, which is the safe state and the one to return to (delete the marker) if the vault ever gets in the way.

REGISTRY_TOKEN is deliberately not in the vault: dev pulls with the persistent docker login credential already in this host's docker config, like the prod/beta CI paths do.

Before monitoring a PR to land

A PR needs an explicit merge once CI is green — but first confirm it's actually set up to land, or you'll monitor a PR that can never merge:

  1. Not conflictingGET /api/v1/repos/{owner}/{repo}/pulls/{n} should show "mergeable": true. If not, merge forgejo/dev in and resolve now — don't monitor a PR that can't merge.
  2. Required check actually configuredGET /api/v1/repos/{owner}/{repo}/branch_protections/dev should list status_check_contexts containing the exact string Forgejo reports for that PR's check (workflow display name + job name + event, e.g. "PR build (required check) / build (pull_request)" — not just the job name "build"). A mismatch here means no PR can ever satisfy the check even with fully green CI; this bit us once already.
  3. CI actually runningGET /api/v1/repos/{owner}/{repo}/commits/{sha}/status shows a status, not an empty list. If empty, check the runner is up (systemctl --user status forgejo-runner).

Only start monitoring once all three pass. Once CI is green and mergeable is true, merge explicitly: POST /api/v1/repos/{owner}/{repo}/pulls/{n}/merge with {"Do": "squash", "delete_branch_after_merge": true}.

Day-to-day

  • Ship: open a PR into dev, confirm CI is green, merge it explicitly — the merge triggers the deploy here.
  • Manual redeploy: trigger deploy-dev.yml via workflow_dispatch (POST /api/v1/repos/{owner}/{repo}/actions/workflows/deploy-dev.yml/dispatches, {"ref": "dev"}), or locally APP_DIR=~/thermograph-dev bash deploy/deploy-dev.sh.
  • Promote to prod: merge/fast-forward dev into main and push — that triggers the VPS pipeline in DEPLOY.md.

Monitoring

Logs and dashboards moved to a separate project — a central Grafana + Loki stack fed by a Grafana Alloy agent on every node (including this LAN dev box), over the WireGuard mesh (thermograph-observability on Forgejo; UI at https://dashboard.thermograph.org). The dev node's Alloy agent ships its container/app logs under host="dev", so LAN dev shows up alongside prod and beta in the same dashboards. This replaced the old scripts/dashboard.py.

For a quick terminal check without Grafana, the app still serves raw counters at the gated GET /api/v2/metrics route. Dev sets no THERMOGRAPH_METRICS_TOKEN, so the route is direct-loopback-only — curl localhost:8137/api/v2/metrics from this machine, which is where you already are. (A token exists on the VPS boxes to allow reads through an SSH tunnel; dev has no reason to carry a copy, and the estate's copy must not land here — see deploy/secrets/README.md.)