Some checks failed
PR build (required check) / changes (pull_request) Successful in 7s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Failing after 9s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
Brings the last two credential stores under SOPS, so all three environments and the control plane are managed the same way. dev.yaml — dev was not "intentionally not wired in", it was broken The vault README claimed dev had no real secrets. Both halves were false. Dev needed secrets it did not have: thermograph-dev-daemon-1 has been in a crash loop, restarting every ~60s with "neither THERMOGRAPH_INTERNAL_TOKEN nor THERMOGRAPH_AUTH_SECRET is set; refusing to start" — so dev has had no Discord gateway and no job timers. With THERMOGRAPH_BASE_URL unset the app also falls back to https://thermograph.org, so dev's IndexNow pings and verification links claimed to be prod. dev.yaml gives it its own generated AUTH_SECRET and a real base URL, and 12 values total. And dev already held production credentials: backend-deploy-dev.yml injects S3_ACCESS_KEY/S3_SECRET_KEY into every dev deploy, and the only provisioned keypair for that bucket is read-write — on the bucket holding prod's backups. Dev does not need them; without bucket creds the lake service 503s and history falls through to the archive. Those workflow lines still need removing. dev renders dev.yaml ALONE, via THERMOGRAPH_SECRETS_SKIP_COMMON=1. Eleven of common.yaml's sixteen values are live production credentials, the render is a plaintext concatenation consumed through env_file:, and dev is the operator's desktop AND the Forgejo runner executing unreviewed dev-branch code with the docker socket mounted. An override in dev.yaml would not help — last-wins governs consumers, but prod's value is still physically a line in the file. Verified: current renderer leaks 28 lines onto dev including both S3 keypairs and the VAPID private key; patched gives exactly 12 keys and no prod credential. Beta's render is byte-identical either way. centralis.prod.yaml — its own file, its own renderer /etc/centralis.env was hand-edited, which is how a JSON registry got written unquoted into a shell-sourced file tonight and silently collapsed three identities to one. The renderer now owns the file. It is service- AND environment-scoped, not folded into prod.yaml, because /etc/thermograph.env is loaded into every container in the app stack. Folding these in would put CENTRALIS_AUTH_TOKEN and CENTRALIS_TOKENS — which authenticate an endpoint carrying run_on_host and sql_query(write) — into the web backend's environment, turning a read-anything bug in the app into a foothold on the control plane. Quoting is guaranteed three ways rather than assumed. Critically, `sops -d --output-type dotenv` emits values verbatim, so reusing the existing render path would have reproduced tonight's bug exactly. The new function decrypts to JSON and emits POSIX single-quoted assignments; it then sources its own output in a clean shell and compares every value before touching /etc; and it refuses to write unless CENTRALIS_TOKENS parses as a non-empty subject->token object after sourcing. Tested against 16 hostile values including embedded quotes, newlines and ;rm -rf /. Value-identity proven, not asserted: rendered vs live compared as *effective* values on prod (a textual diff would report a false difference, and would report a false match if the vault had captured quote characters as part of the value), plus both env files producing a byte-identical `docker compose config` digest. 9 keys, both token subjects preserved. Nothing deployed. Note for later: /etc/thermograph.env is also shell-sourced, and survives on raw dotenv only because none of its 32 current values contains a shell-special character. It is one quoted secret away from the same bug. Claude-Session: https://claude.ai/code/session_0182KTMrsTHJc3TcewCatJFY
220 lines
12 KiB
Markdown
220 lines
12 KiB
Markdown
# Dev CI/CD → LAN server on this machine
|
|
|
|
Parallel to the prod pipeline (`main` → VPS, see `DEPLOY.md`), the **`dev`**
|
|
branch continuously deploys to a **LAN server running on this computer**. Git
|
|
hosting and CI are self-hosted **Forgejo** (`git.thermograph.org`, reachable
|
|
at `http://10.10.0.2:3080` over the WireGuard mesh) — GitHub is retired.
|
|
|
|
```
|
|
open PR ──▶ CI (build + boot/health, on the LAN Forgejo runner)
|
|
│ green
|
|
▼
|
|
merge into dev (explicit — Forgejo does not auto-merge on its own;
|
|
see "Landing a PR" below)
|
|
│
|
|
▼
|
|
LAN runner (thermograph-lan label) runs deploy/deploy-dev.sh
|
|
│
|
|
▼
|
|
~/thermograph-dev updated ─▶ docker compose stack (backend + frontend +
|
|
Postgres 18) brought up, backend on 0.0.0.0:8137
|
|
```
|
|
|
|
Reachable at `http://<lan-ip>:8137/` from any device on your Wi-Fi.
|
|
|
|
The dev server runs the **same containerized stack as prod** (`docker-compose.yml`),
|
|
overlaid with `docker-compose.dev.yml`: **uncapped** (no CPU limits, unlike prod's
|
|
backend=4 / frontend=2 / db=2) and backend published on `0.0.0.0:8137` for the
|
|
LAN (prod binds loopback behind Caddy; frontend has no published port here
|
|
either way, no Caddy to reach it directly, so it's only reached through
|
|
backend's own reverse-proxy fallback). Bring it up by hand with `make dev-up`.
|
|
|
|
## Why it's built this way
|
|
|
|
Forgejo has native branch protection + required-status-checks + auto-merge
|
|
(unlike GitHub on a private free-tier repo, where those are paywalled). But
|
|
"auto-merge" here still means opting a PR in — either clicking **"Auto merge
|
|
when checks succeed"** in the web UI, or merging explicitly once CI is green
|
|
via the API (`POST .../pulls/{n}/merge`, `{"Do": "squash"}`). Nothing merges
|
|
a ready PR for you automatically just because checks pass, unlike GitHub's
|
|
old in-workflow `ci-cd.yml` (retired along with the rest of `.github/`).
|
|
|
|
## Moving parts
|
|
|
|
| File | Purpose |
|
|
|------|---------|
|
|
| `.forgejo/workflows/pr-build.yml` | required status check for PRs into `dev` (calls `build.yml`) |
|
|
| `.forgejo/workflows/deploy-dev.yml` | build + deploy on pushes to `dev` (direct, or a PR merge) |
|
|
| `.forgejo/workflows/build.yml` | shared build gate: deps, backend tests, JS syntax check, boot/health check |
|
|
| `deploy/deploy-dev.sh` | pull `dev` into `~/thermograph-dev`, `docker compose up` the uncapped LAN stack, health check |
|
|
| `deploy/secrets/dev.yaml` | dev's SOPS vault — its own secrets and config, rendered to `/etc/thermograph.env` at deploy time (see "Config and secrets") |
|
|
| `docker-compose.dev.yml` | dev overlay: no CPU limits, backend published on `0.0.0.0:8137` (frontend unpublished) |
|
|
| `deploy/provision-dev-lan.sh` | one-time sudo-free bootstrap of the checkout + runner |
|
|
| `deploy/forgejo/register-lan-runner.sh` | registers the Forgejo Actions runner on this machine |
|
|
|
|
`deploy-dev.yml`'s `deploy` job runs on the `thermograph-lan` label — **not**
|
|
`[self-hosted, thermograph-lan]` (the array form is what the original GitHub
|
|
version used; GitHub implicitly tags every self-hosted runner `self-hosted`,
|
|
but Forgejo's runner has only the labels it was explicitly registered with,
|
|
so the array form is permanently unschedulable there — a real bug found and
|
|
fixed once, worth knowing about if a "deploy" job ever silently stops
|
|
appearing in the Actions history again).
|
|
|
|
`deploy-dev.sh`'s git auth (`GH_TOKEN`/`GITHUB_TOKEN`, Forgejo's per-job
|
|
token exposed under that name for compatibility) is scoped to `REPO_URL`'s
|
|
own host via a per-command `http.<scheme>://<host>/.extraheader`, not a
|
|
hardcoded one — important if the Forgejo host ever changes, since a stale
|
|
hardcoded host would make git silently skip the header (no error) and fall
|
|
through to whatever ambient credentials happen to be available, not fail
|
|
loudly.
|
|
|
|
## The self-hosted runner
|
|
|
|
A Forgejo Actions **runner** (`forgejo-runner`) runs on this machine, labels
|
|
`docker` (containerized jobs, node:20-bookworm, with the host's docker.sock
|
|
automounted in — `container.docker_host: automount` in
|
|
`~/forgejo-runner/config.yaml` — so `docker build`/`push` work without
|
|
privileged Docker-in-Docker) and `thermograph-lan` (host-native jobs, for
|
|
`deploy-dev.sh`'s systemctl/docker-compose calls). Installed under
|
|
`~/forgejo-runner`, runs as the `forgejo-runner` systemd `--user` service.
|
|
|
|
```bash
|
|
# status / logs
|
|
systemctl --user status forgejo-runner
|
|
journalctl --user -u forgejo-runner -f
|
|
|
|
# re-register if the token/host ever changes
|
|
bash deploy/forgejo/register-lan-runner.sh <forgejo_url> <registration_token>
|
|
```
|
|
|
|
`git.thermograph.org`'s public DNS resolves to beta's **public** IP, which
|
|
the registry's mesh-only ACL (see `deploy/forgejo/README.md`) rejects for
|
|
`/v2/` paths — `docker build-push.yml` pushes run against the *host's*
|
|
automounted daemon, so it's this **host's** own resolver that needs to route
|
|
git.thermograph.org over the mesh, not anything settable from a job
|
|
container. If registry pushes ever start failing with an HTTP/TLS mismatch
|
|
or a 403 on `/v2/`, check `getent hosts git.thermograph.org` here — it needs
|
|
to resolve to beta's WireGuard IP (`10.10.0.2`), typically via a `/etc/hosts`
|
|
line, not the public one.
|
|
|
|
## The LAN dev service (Docker Compose)
|
|
|
|
Runs as a Docker Compose stack, not a bare systemd unit — `docker ps --filter
|
|
name=thermograph-dev` shows `thermograph-dev-backend-1`, `thermograph-dev-frontend-1`
|
|
(repo-split Stage 4 split the single "app" container in two — frontend has no
|
|
published port, reached only through backend's own reverse-proxy fallback) and
|
|
`thermograph-dev-db-1` (TimescaleDB). A legacy pre-container
|
|
`thermograph-dev.service` systemd --user unit may still exist from before this
|
|
cutover; it's stale and `deploy-dev.sh` stops/disables it on every run — don't
|
|
trust `systemctl --user status thermograph-dev` as a liveness check, use `docker
|
|
ps` instead.
|
|
|
|
```bash
|
|
docker ps --filter name=thermograph-dev # is it up?
|
|
docker compose -f docker-compose.yml -f docker-compose.dev.yml logs -f backend frontend # app logs (from ~/thermograph-dev)
|
|
```
|
|
|
|
Bootstrap (or rebuild) it by hand:
|
|
|
|
```bash
|
|
bash deploy/provision-dev-lan.sh
|
|
```
|
|
|
|
The dedicated checkout at `~/thermograph-dev` is separate from your working tree,
|
|
so a deploy never touches uncommitted edits in `~/Code/Thermograph`.
|
|
|
|
## Config and secrets
|
|
|
|
Dev's configuration lives in the same SOPS vault as prod's and beta's —
|
|
`deploy/secrets/dev.yaml`, encrypted to the same age recipient, rendered to
|
|
`/etc/thermograph.env` by the same `render-secrets.sh` at deploy time. Changing a
|
|
dev value is `sops deploy/secrets/dev.yaml` → commit → deploy, exactly as for the
|
|
VPS boxes. **`deploy/secrets/README.md` is the reference**; what follows is only
|
|
what is specific to this machine.
|
|
|
|
Dev renders `dev.yaml` **alone**. It does *not* layer `common.yaml`, which is the
|
|
fleet's shared production credential set — registry token, both S3 keypairs (one
|
|
read-write pair, on the bucket that also holds prod's backups), the VAPID private
|
|
key that signs push to real subscribers, the IndexNow key, the metrics token. This
|
|
box is the CI runner and runs unreviewed `dev`-branch code; a plaintext render of
|
|
that set into `/etc/thermograph.env` would put all of it in every dev container's
|
|
environment. `deploy-dev.sh` exports `THERMOGRAPH_SECRETS_SKIP_COMMON=1` to prevent
|
|
it, and **aborts the deploy** if the renderer in the checkout cannot honour that.
|
|
|
|
So dev holds its own `THERMOGRAPH_AUTH_SECRET` (its own, not prod's — the `daemon`
|
|
service refuses to start without one), its own `POSTGRES_PASSWORD`, a LAN
|
|
`THERMOGRAPH_BASE_URL`, `THERMOGRAPH_COOKIE_SECURE=0` (plain HTTP — the shared
|
|
value is `1` and would silently break login here), and the non-secret config it
|
|
would otherwise have inherited. It holds **no** S3 keys, no VAPID keypair, no
|
|
IndexNow key, no Discord or mail credentials: the lake falls through to the
|
|
Open-Meteo archive without bucket creds, and the app generates and persists its own
|
|
VAPID and IndexNow keys into the `appdata` volume.
|
|
|
|
Two dev-only mechanics, both because this box has **no passwordless sudo** and the
|
|
runner is a `systemd --user` service with no tty:
|
|
|
|
- the age key is read from `~/.config/sops/age/keys.txt`, not
|
|
`/etc/thermograph/age.key` — a root-owned `0400` key would send the renderer down
|
|
its `sudo cat` path, where it would hang on a prompt no CI job can answer;
|
|
- `/etc/thermograph.env` is **pre-created** owned by the deploying user, so the
|
|
renderer takes its in-place write path and no deploy needs `sudo` at all.
|
|
|
|
Setting this up on a fresh dev machine is four typed commands — see **"Setting up a
|
|
new dev machine"** in `deploy/secrets/README.md`. Until the
|
|
`/etc/thermograph/secrets-env` marker exists the box renders nothing and falls back
|
|
to `deploy-dev.sh`'s built-in `POSTGRES_PASSWORD` default, which is the safe state
|
|
and the one to return to (delete the marker) if the vault ever gets in the way.
|
|
|
|
`REGISTRY_TOKEN` is deliberately not in the vault: dev pulls with the persistent
|
|
`docker login` credential already in this host's docker config, like the prod/beta
|
|
CI paths do.
|
|
|
|
## Before monitoring a PR to land
|
|
|
|
A PR needs an explicit merge once CI is green — but first confirm it's
|
|
actually set up to land, or you'll monitor a PR that can never merge:
|
|
|
|
1. **Not conflicting** — `GET /api/v1/repos/{owner}/{repo}/pulls/{n}` should
|
|
show `"mergeable": true`. If not, merge `forgejo/dev` in and resolve now —
|
|
don't monitor a PR that can't merge.
|
|
2. **Required check actually configured** — `GET
|
|
/api/v1/repos/{owner}/{repo}/branch_protections/dev` should list
|
|
`status_check_contexts` containing the *exact* string Forgejo reports for
|
|
that PR's check (workflow display name + job name + event, e.g. `"PR build
|
|
(required check) / build (pull_request)"` — not just the job name
|
|
`"build"`). A mismatch here means no PR can ever satisfy the check even
|
|
with fully green CI; this bit us once already.
|
|
3. **CI actually running** — `GET
|
|
/api/v1/repos/{owner}/{repo}/commits/{sha}/status` shows a status, not an
|
|
empty list. If empty, check the runner is up
|
|
(`systemctl --user status forgejo-runner`).
|
|
|
|
Only start monitoring once all three pass. Once CI is green and mergeable is
|
|
true, merge explicitly: `POST /api/v1/repos/{owner}/{repo}/pulls/{n}/merge`
|
|
with `{"Do": "squash", "delete_branch_after_merge": true}`.
|
|
|
|
## Day-to-day
|
|
|
|
- **Ship:** open a PR into `dev`, confirm CI is green, merge it explicitly —
|
|
the merge triggers the deploy here.
|
|
- **Manual redeploy:** trigger `deploy-dev.yml` via `workflow_dispatch`
|
|
(`POST /api/v1/repos/{owner}/{repo}/actions/workflows/deploy-dev.yml/dispatches`,
|
|
`{"ref": "dev"}`), or locally `APP_DIR=~/thermograph-dev bash deploy/deploy-dev.sh`.
|
|
- **Promote to prod:** merge/fast-forward `dev` into `main` and push — that
|
|
triggers the VPS pipeline in `DEPLOY.md`.
|
|
|
|
## Monitoring
|
|
|
|
Logs and dashboards moved to a **separate project** — a central Grafana + Loki
|
|
stack fed by a Grafana Alloy agent on every node (including this LAN dev box),
|
|
over the WireGuard mesh (`thermograph-observability` on Forgejo; UI at
|
|
`https://dashboard.thermograph.org`). The dev node's Alloy agent ships its
|
|
container/app logs under `host="dev"`, so LAN dev shows up alongside prod and
|
|
beta in the same dashboards. This replaced the old `scripts/dashboard.py`.
|
|
|
|
For a quick terminal check without Grafana, the app still serves raw counters at
|
|
the gated `GET /api/v2/metrics` route. Dev sets no `THERMOGRAPH_METRICS_TOKEN`, so
|
|
the route is direct-loopback-only — `curl localhost:8137/api/v2/metrics` from this
|
|
machine, which is where you already are. (A token exists on the VPS boxes to allow
|
|
reads through an SSH tunnel; dev has no reason to carry a copy, and the estate's
|
|
copy must not land here — see `deploy/secrets/README.md`.)
|