All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
shell-lint / shellcheck (pull_request) Successful in 7s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
PR build (required check) / validate-observability (pull_request) Successful in 18s
Environments stop being machines. Until now each box WAS an environment --
"beta" named both a deploy target and a host, "the desktop" named both the
operator's computer and the dev server -- so every path could assume one
environment per host and hardcode /opt/thermograph, /etc/thermograph.env and
ports 8137/8080. That assumption ends here:
vps1 75.119.132.91 Forgejo, Grafana/Loki, the portfolio site, and DEV
(own Postgres, mesh-only on 10.10.0.2:8137)
vps2 169.58.46.181 PROD and BETA as two Swarm stacks sharing one
TimescaleDB instance, plus Centralis, Postfix, backups
desktop AI model hosting + flex Swarm capacity, no environment
Nothing live has moved; the ordered cutover is in
infra/deploy/RUNBOOK-vps1-vps2-cutover.md.
deploy/env-topology.sh is the single source of truth: env -> host, checkout,
branch, deploy mode, stack name, env file, LB ports, DB role/database, service
prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing
distinguishing a beta deploy from a prod one -- deploy.sh refuses to run when it
disagrees with the checkout it was invoked from. The host-wide secrets-env and
deploy-mode markers survive only as a fallback for a single-environment box.
Beta gets its own stack file rather than an overlay (a merge cannot REMOVE the
db service, and beta having no database of its own is the point). Its services
are prefixed beta-*: Swarm registers a service's short name as a DNS alias on
every network it joins, so two stacks both calling a service `web` on the shared
data network would let prod's frontend resolve a beta task. Prod's stack, env
and LB config are untouched.
One Postgres, two databases with two roles: deploy/db/provision-env-db.sh
creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT
revoked from PUBLIC) plus a <role>_ro for ad-hoc queries, and refuses to run for
the environment that owns the instance -- doing so would demote prod's bootstrap
superuser.
Fixes that co-residency would otherwise have broken silently:
- CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and
also "the Forgejo box" because those shared a machine; that conflation is what
once had the prod backup dumping beta.
- The nightly backup dumps BOTH application databases to separate off-box
prefixes, and fails loudly on a missing one instead of skipping it.
- Alloy stopped deriving the `host` label from the node -- on vps2 that would
have filed every beta line as prod, feeding prod's alert rules with beta's
traffic. It is now derived per source, with a new `node` label for the machine,
and beta's log volume is mounted via a vps2-only overlay.
- ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derived their SSH
target and env-file path from the topology instead of hardcoding beta to
75.119.132.91 -- which is vps1's address now.
- Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on
vps1) instead of 0.0.0.0: correct for a home LAN box, a public exposure of
unreviewed branches on a VPS.
Terraform's hosts variable is keyed by machine with a nested environments map,
and main.tf flattens (host, environment) pairs so two environments on one box
cannot share a checkout. Caddy's single reference config is split per host.
Docs, onboarding and the runbooks are updated throughout, including the security
rationales that co-residency makes false: separate SSH credentials no longer put
a host boundary between beta and prod, and the boundary that remains is the
database and the filesystem.
294 lines
14 KiB
Markdown
294 lines
14 KiB
Markdown
# 11. Traps and stale docs
|
||
|
||
Two kinds of thing here: **live traps** (correct code that will surprise you)
|
||
and **stale docs** (files in this repo that currently describe a system that no
|
||
longer exists). The second list matters as much as the first — this repo has a
|
||
lot of documentation, most of it excellent, and some of it describing the world
|
||
before a rewrite.
|
||
|
||
Everything below was checked against the tree at the time of writing.
|
||
|
||
---
|
||
|
||
## Stale docs in this repo
|
||
|
||
### ~~`frontend/CLAUDE.md` and `frontend/README.md` describe a service that is not deployed~~ — fixed
|
||
|
||
**Fixed.** Both now describe the Go service (`frontend/server/`), state plainly
|
||
that the Python files at that level are the superseded original, and correct
|
||
`CLAUDE.md`'s claim that `make test-unit` is *"the tier CI runs"* — it isn't.
|
||
CI builds the frontend image, and `frontend/Dockerfile`'s builder stage runs
|
||
`gofmt -l` + `go vet` + `go test ./...` before compiling, so a failing **Go**
|
||
test is what fails CI. The Python tiers are local-only.
|
||
|
||
Still worth knowing: the superseded Python files carry **duplicate copies of
|
||
some contracts** — `api_client.py` has a third `API_VERSION` pin and
|
||
`format.py` a fourth `F_COUNTRIES`. Nothing deploys them, but if you run them
|
||
locally, keep them in step or you'll debug a phantom.
|
||
|
||
Whether to delete them outright is an open question, not a settled one: the
|
||
Python test tiers still run, against fixtures the Go tests share. Raise it
|
||
rather than doing it in passing.
|
||
|
||
### `infra/DEPLOY.md` describes the pre-Swarm, pre-monorepo deploy
|
||
|
||
Three separate layers of drift:
|
||
|
||
- It says **prod is deployed with `terraform apply`**. Prod is deployed by a
|
||
push to `release` firing `deploy.yml`, exactly like beta's `main` path — see
|
||
[07](07-ci-and-release.md).
|
||
- It says prod runs a **compose** stack. Prod runs **Swarm**
|
||
(`infra/deploy/stack/`).
|
||
- Its file table points at `/opt/thermograph/deploy/deploy.sh` and
|
||
`/opt/thermograph/docker-compose.yml`. Both moved under `infra/` at the
|
||
monorepo cutover.
|
||
|
||
The document's own header flags the legacy `provision.sh` / systemd / venv
|
||
walkthrough as superseded — but the drift goes further than that note admits.
|
||
The **key-management** sections (Part 1) and the `/etc/hosts` registry note are
|
||
still accurate and useful.
|
||
|
||
### `infra/DEPLOY-DEV.md` and `infra/ACCESS.md` reference deleted workflows — and predate dev becoming a real environment
|
||
|
||
Both talk about `.forgejo/workflows/deploy-dev.yml`. It does not exist under
|
||
that name. The two original LAN-dev deploy workflows were deleted rather than
|
||
ported at the CI consolidation — at the time they were genuinely inert (they
|
||
targeted a monorepo layout at `~/thermograph-dev` on a box still holding a
|
||
split-era checkout), which is where the old claim **"LAN dev is a local `make
|
||
dev-up` concern, not a CI environment"** came from.
|
||
|
||
**That claim is now false, and worth flagging precisely because it used to be
|
||
true.** `dev` is a first-class hosted environment on vps1 — its own Postgres
|
||
container, mesh-only exposure, deployed by CI like beta and prod, as the third
|
||
leg of `deploy.yml` (see [07](07-ci-and-release.md) and
|
||
[08](08-infra-secrets.md)). The desktop's `make dev-up` is still a real,
|
||
useful thing — a laptop-local rehearsal — but it is no longer the *only* way
|
||
`dev` runs, and neither doc has been updated to say so.
|
||
|
||
`ACCESS.md` §3c also lists `{build,pr-build,deploy-dev,deploy}.yml` as the
|
||
workflow set; there are now nine files and the deploy/build ones were collapsed.
|
||
It also still describes the host table from before the vps1/vps2 rename — see
|
||
`CUTOVER-NOTES.md`'s vps1/vps2 section for the current one.
|
||
|
||
Everything else in `ACCESS.md` — the agent-access model, the key rotation
|
||
procedure, the Swarm/Forgejo tracks — is current and worth reading; just
|
||
mentally translate host names.
|
||
|
||
### `backend/README.md` is written for the split-repo era
|
||
|
||
It describes `thermograph-backend` as its own repo with sibling repos, uses the
|
||
retired image path `git.thermograph.org/emi/thermograph-backend/app`, and says
|
||
*"there is no Makefile in this repo yet"* — there is (`make test`, `make smoke`,
|
||
`make clean-venv`).
|
||
|
||
The **"How it works"** section (grid → data → percentiles → caching) is one of
|
||
the best explanations of the domain model anywhere in the repo. Read it; just
|
||
mentally translate the topology.
|
||
|
||
### `CUTOVER-NOTES.md` carries a superseded block, correctly labelled
|
||
|
||
Its per-domain workflow names (`backend-build-push.yml`, `backend-deploy.yml`, …)
|
||
no longer exist — the six deploy workflows collapsed into `deploy.yml` and the
|
||
two build-push workflows into `build-push.yml`. The file says so in an explicit
|
||
"Superseded 2026-07-25" note, and everything it describes about *behaviour* is
|
||
unchanged.
|
||
|
||
It also references `backend/MONOREPO-CUTOVER-PLAN.md`, which is not in the tree.
|
||
|
||
`CUTOVER-NOTES.md` remains the source of truth for what is and isn't live.
|
||
|
||
### The domain `CLAUDE.md` files still use repo-era wording
|
||
|
||
They say "this repo" and link to sibling repo names. Mentally map
|
||
`thermograph-backend` → `backend/`, and fix the wording as you touch files near
|
||
it — that's the standing instruction in the root `CLAUDE.md`.
|
||
|
||
### `observability/README.md`: merging is not deploying
|
||
|
||
`/opt/observability` on vps1 and vps2 is a checkout of the **archived**
|
||
`emi/thermograph-observability` repo. It can never `git pull` again. A previous
|
||
observability PR merged and had **zero live effect** for exactly this reason.
|
||
The README documents the hand-ship procedure; until the checkouts are re-pointed
|
||
at the monorepo, that's the only path. Flagged again in
|
||
[09](09-observability.md).
|
||
|
||
---
|
||
|
||
## Live traps
|
||
|
||
### ~~The `docker-compose.test.yml` files pull retired image paths~~ — fixed
|
||
|
||
**Fixed.** Both now use `emi/thermograph/backend`, the path `build-push.yml`
|
||
actually publishes.
|
||
|
||
The frontend harness had the worse half of this: it also pinned
|
||
`THERMOGRAPH_BACKEND_TEST_TAG` to the split-era `v0.0.2-split-ci`. Rather than
|
||
swap one hardcoded pin for another, `scripts/backend-for-tests.sh` now
|
||
**derives** the tag from the checkout — `sha-<12hex of git log -1 -- backend/>`,
|
||
the same domain-keyed rule `build-push.yml` and `deploy.yml` both use — so the
|
||
harness follows the tree instead of drifting behind a pin. `docker-compose.test.yml`
|
||
requires the variable (`:?`) rather than defaulting it, so a stale pin can't
|
||
creep back in silently.
|
||
|
||
If the derived tag has no published image (backend commits that are still
|
||
local-only), the script says so and tells you to pin one:
|
||
|
||
```bash
|
||
THERMOGRAPH_BACKEND_TEST_TAG=sha-<12hex> make backend-up
|
||
```
|
||
|
||
> Unrelated but adjacent: `make test-integration` against a freshly-booted
|
||
> throwaway backend currently fails 7 of 16 tests with `503`, because nothing
|
||
> in that empty database is warm. Verified identical on the old image and the
|
||
> new one, so it predates these changes and is **not** caused by them. Not
|
||
> investigated further here.
|
||
|
||
### ~~`static/units.js`'s `F_REGIONS` is guarded by nothing~~ — fixed
|
||
|
||
**Fixed.** `format_test.go::TestFCountriesMatchesUnitsJS` now parses
|
||
`static/units.js` and diffs it against the Go set in both directions.
|
||
|
||
The subtlety worth knowing: the *existing* backend cross-check
|
||
(`TestFCountriesMatchesBackend`) **skips in CI**. The frontend image's build
|
||
context is `frontend/`, so `backend/` is structurally unreachable from the
|
||
builder stage — and the builder stage is the only place CI runs these tests. It
|
||
is a checkout-only guard and always was.
|
||
|
||
The new `units.js` check does not have that problem, because `static/` *is*
|
||
inside the context: `frontend/Dockerfile` copies that one file into the builder
|
||
stage specifically so the assertion runs during the image build. Verified by
|
||
mutating `units.js` and confirming the **image build fails**, not just the local
|
||
test run.
|
||
|
||
### The compose project name is load-bearing
|
||
|
||
`infra/docker-compose.yml` pins `name: thermograph`. Without it, running compose
|
||
from `infra/` derives project `infra` — a silently **new** stack with empty
|
||
volumes beside the running one. `deploy-dev.sh` exports
|
||
`COMPOSE_PROJECT_NAME=thermograph-dev` to keep dev separate on purpose.
|
||
Keep both halves.
|
||
|
||
### `deploy.sh` hard-resets the host checkout
|
||
|
||
`git reset --hard origin/$BRANCH` runs before anything else. Uncommitted edits
|
||
in `/opt/thermograph` evaporate. `infra-sync.yml` does the same reset without
|
||
deploying. The one deliberate exception is the untracked
|
||
`deploy/.image-tags.env` — losing *that* makes the next single-service deploy
|
||
roll the sibling onto `local`.
|
||
|
||
### The deploy scripts eat your rollback image
|
||
|
||
Both `deploy.sh` and `deploy-stack.sh` **end** by deleting images outside the
|
||
running pair. Beta now runs Swarm co-resident with prod on vps2, so both
|
||
usually fail to fully prune only because Swarm's stopped task containers still
|
||
hold references — dev, the one remaining compose environment (on vps1), is the
|
||
one that typically holds **no local rollback target at all**. `docker system
|
||
prune -a` is forbidden on a live box for the same reason — but the deploy
|
||
script is the thing that actually eats the image, not prune. Verify a tag
|
||
exists before promising a rollback.
|
||
|
||
### `daemon` and `lake` are never deploy targets on their own
|
||
|
||
They ride with `backend`. All three run the same image at the same tag — a
|
||
second binary and a second role inside one image — and the `/internal/*`
|
||
contract between daemon and web has no version negotiation. There's also a
|
||
mechanical reason: a single-service deploy runs `up -d --no-deps <targets>`, so
|
||
a service in compose but absent from `TARGETS` would simply never be created.
|
||
|
||
### `docker compose config --images <service>` does not filter
|
||
|
||
Confirmed live on Compose v5.3.1: it prints **every** service's image, one per
|
||
line, in file order. A `| head -1` therefore grabbed `db`'s image, and
|
||
`deploy.sh`'s daemon capability probe found no daemon binary in a Postgres image
|
||
and dropped the daemon from *every* backend deploy. The script now builds the
|
||
image reference from the same variables compose interpolates.
|
||
|
||
### Prod's (and beta's) Loki streams carry no `service` label
|
||
|
||
Prod and beta both run Swarm, co-resident on vps2; their `job="docker"` streams
|
||
have no `service` label. A raw `{host="prod", service="backend"}` query matches
|
||
nothing and reads as "no logs" when it means "wrong selector". Use
|
||
`logs_query`'s `service` argument. Prod and beta container names also change on
|
||
every redeploy (the task suffix) — resolve via `docker ps --filter name=` at
|
||
call time, never hardcode. Dev, on compose, doesn't have this problem.
|
||
|
||
### A Grafana alert rule can provision cleanly and never fire
|
||
|
||
A rule whose `condition` names a non-existent refId is valid YAML, provisions
|
||
without complaint, and pages nobody. Likewise a contact point holding the
|
||
literal string `$DISCORD_ALERT_WEBHOOK_URL` looks completely healthy from the
|
||
API — and the API **redacts** the URL, so you cannot confirm interpolation by
|
||
reading it back. The only proof is a message arriving in `#ops-alerts`.
|
||
|
||
### Terraform has no state anywhere
|
||
|
||
A casual `terraform apply` would attempt full re-provisioning of live hosts.
|
||
It's executable documentation until state is bootstrapped.
|
||
|
||
### The registry is mesh-only
|
||
|
||
`git.thermograph.org` resolves publicly to vps1's IP, but vps1's Caddy rejects
|
||
`/v2/*` from outside the WireGuard mesh. Any host that pulls needs
|
||
`10.10.0.2 git.thermograph.org` in `/etc/hosts`. vps1 itself doesn't — it *is*
|
||
the box.
|
||
|
||
### Forgejo runners only have the labels they registered with
|
||
|
||
GitHub implicitly tags every self-hosted runner `self-hosted`; Forgejo does not.
|
||
`runs-on: [self-hosted, thermograph-lan]` is permanently unschedulable there —
|
||
a real bug, found and fixed once. If a job silently stops appearing in the
|
||
Actions history, check this first.
|
||
|
||
### Forgejo's `needs:` doesn't cross workflow files
|
||
|
||
`build-push.yml` and `deploy.yml` fire from the same push with no ordering
|
||
guarantee, and a deploy racing ahead of its build and failing "not found" is
|
||
confirmed-live. That's why `deploy.sh`'s pull has a bounded ~5-minute retry.
|
||
|
||
### `/api/v2/cell?prefetch=1` must never spend quota
|
||
|
||
The hard guarantee: a cold cell answers `204` with no body. Anything named
|
||
"prefetch" or "warm" calls the cache-only loaders. Breaking this gets the whole
|
||
service rate-limited.
|
||
|
||
### Background timers must go through leader election
|
||
|
||
`singleton.claim_leader()` — flock per host, Postgres advisory lock per cluster.
|
||
A timer-driven sweep running in N uvicorn workers × M Swarm hosts multiplies
|
||
Open-Meteo quota use N×M. That is the documented cause of overnight 429/503s
|
||
after prod went to 3 workers. The heartbeat is the deliberate exception — a beat
|
||
is one local file append, so duplicates are free.
|
||
|
||
### `THERMOGRAPH_DATA_DIR` exists because of a real collision
|
||
|
||
The backend's Python package `data/` and its runtime data dir `<root>/data`
|
||
collide at the same path once the code lives at the image root. Mounting a
|
||
volume there shadows and erases `data/*.py`, so `import data.climate` fails at
|
||
boot. Point the env var **outside the code tree** in a container. And never
|
||
reintroduce `__file__`-relative paths in a module — `paths.py` resolves
|
||
everything from the repo root for this reason.
|
||
|
||
### Alloy can silently stop tailing a container
|
||
|
||
Observed on prod, 13½ hours of `thermograph_worker` logs never reaching Loki
|
||
while the container was healthy and writing. Not root-caused. Workaround:
|
||
`sudo docker restart alloy-alloy-1`. `ProdWorkerContainerSilent` now catches the
|
||
class.
|
||
|
||
### Nominatim is ~1 req/s, on one thread
|
||
|
||
All reverse geocoding funnels through a single dedicated worker draining a
|
||
queue. The earlier design slept inside the caller's thread — which in the server
|
||
is one of Starlette's shared sync-threadpool threads.
|
||
|
||
---
|
||
|
||
## If you fix one of these
|
||
|
||
Docs are cheap to correct, and the root `CLAUDE.md` already asks you to fix
|
||
repo-era wording as you touch files near it. The three entries struck through
|
||
above were fixed in the same change that added this guide; the rest are still
|
||
open, and `infra/DEPLOY.md` is probably the highest-value one left — it
|
||
describes a deploy model that has since changed twice.
|
||
|
||
Back to the [index](README.md).
|