thermograph/docs/onboarding/11-traps.md
Emi Griffith df409f88b3
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 4s
shell-lint / shellcheck (pull_request) Successful in 6s
PR build (required check) / validate-observability (pull_request) Successful in 20s
PR build (required check) / build-frontend (pull_request) Successful in 37s
PR build (required check) / build-backend (pull_request) Successful in 51s
PR build (required check) / gate (pull_request) Successful in 1s
registry: move image and repo references to the Jinemi namespace
Repos moved to the Jinemi org; the container packages did not follow, since
Forgejo does not transfer packages with a repo. The deploy path still resolved
`emi/thermograph/*` — a user_redirect to admin_emi — while build-push.yml
derives its push path from ${github.repository}, now jinemi/thermograph. The
next backend or frontend build would have published somewhere no deploy looks.

Point the image paths at jinemi/thermograph/* (lowercase: OCI references admit
no uppercase, which is why build-push.yml already pipes through tr), the clone
URLs at Jinemi/thermograph, and the registry logins at admin_emi — the account
that actually owns the tokens, rather than the redirect.

The live tags and both ci-runner tags were copied into the Jinemi namespace
first, so the switch has something to pull. thermograph-infra,
thermograph-observability and the retired */app packages stay under admin_emi;
they did not move.
2026-08-01 09:25:02 -07:00

14 KiB
Raw Blame History

11. Traps and stale docs

Two kinds of thing here: live traps (correct code that will surprise you) and stale docs (files in this repo that currently describe a system that no longer exists). The second list matters as much as the first — this repo has a lot of documentation, most of it excellent, and some of it describing the world before a rewrite.

Everything below was checked against the tree at the time of writing.


Stale docs in this repo

frontend/CLAUDE.md and frontend/README.md describe a service that is not deployed — fixed

Fixed. Both now describe the Go service (frontend/server/), state plainly that the Python files at that level are the superseded original, and correct CLAUDE.md's claim that make test-unit is "the tier CI runs" — it isn't. CI builds the frontend image, and frontend/Dockerfile's builder stage runs gofmt -l + go vet + go test ./... before compiling, so a failing Go test is what fails CI. The Python tiers are local-only.

Still worth knowing: the superseded Python files carry duplicate copies of some contractsapi_client.py has a third API_VERSION pin and format.py a fourth F_COUNTRIES. Nothing deploys them, but if you run them locally, keep them in step or you'll debug a phantom.

Whether to delete them outright is an open question, not a settled one: the Python test tiers still run, against fixtures the Go tests share. Raise it rather than doing it in passing.

infra/DEPLOY.md describes the pre-Swarm, pre-monorepo deploy

Three separate layers of drift:

  • It says prod is deployed with terraform apply. Prod is deployed by a push to release firing deploy.yml, exactly like beta's main path — see 07.
  • It says prod runs a compose stack. Prod runs Swarm (infra/deploy/stack/).
  • Its file table points at /opt/thermograph/deploy/deploy.sh and /opt/thermograph/docker-compose.yml. Both moved under infra/ at the monorepo cutover.

The document's own header flags the legacy provision.sh / systemd / venv walkthrough as superseded — but the drift goes further than that note admits. The key-management sections (Part 1) and the /etc/hosts registry note are still accurate and useful.

infra/DEPLOY-DEV.md and infra/ACCESS.md reference deleted workflows — and predate dev becoming a real environment

Both talk about .forgejo/workflows/deploy-dev.yml. It does not exist under that name. The two original LAN-dev deploy workflows were deleted rather than ported at the CI consolidation — at the time they were genuinely inert (they targeted a monorepo layout at ~/thermograph-dev on a box still holding a split-era checkout), which is where the old claim "LAN dev is a local make dev-up concern, not a CI environment" came from.

That claim is now false, and worth flagging precisely because it used to be true. dev is a first-class hosted environment on vps1 — its own Postgres container, mesh-only exposure, deployed by CI like beta and prod, as the third leg of deploy.yml (see 07 and 08). The desktop's make dev-up is still a real, useful thing — a laptop-local rehearsal — but it is no longer the only way dev runs, and neither doc has been updated to say so.

ACCESS.md §3c also lists {build,pr-build,deploy-dev,deploy}.yml as the workflow set; there are now nine files and the deploy/build ones were collapsed. It also still describes the host table from before the vps1/vps2 rename — see CUTOVER-NOTES.md's vps1/vps2 section for the current one.

Everything else in ACCESS.md — the agent-access model, the key rotation procedure, the Swarm/Forgejo tracks — is current and worth reading; just mentally translate host names.

backend/README.md is written for the split-repo era

It describes thermograph-backend as its own repo with sibling repos, uses the retired image path git.thermograph.org/admin_emi/thermograph-backend/app, and says "there is no Makefile in this repo yet" — there is (make test, make smoke, make clean-venv).

The "How it works" section (grid → data → percentiles → caching) is one of the best explanations of the domain model anywhere in the repo. Read it; just mentally translate the topology.

CUTOVER-NOTES.md carries a superseded block, correctly labelled

Its per-domain workflow names (backend-build-push.yml, backend-deploy.yml, …) no longer exist — the six deploy workflows collapsed into deploy.yml and the two build-push workflows into build-push.yml. The file says so in an explicit "Superseded 2026-07-25" note, and everything it describes about behaviour is unchanged.

It also references backend/MONOREPO-CUTOVER-PLAN.md, which is not in the tree.

CUTOVER-NOTES.md remains the source of truth for what is and isn't live.

The domain CLAUDE.md files still use repo-era wording

They say "this repo" and link to sibling repo names. Mentally map thermograph-backendbackend/, and fix the wording as you touch files near it — that's the standing instruction in the root CLAUDE.md.

observability/README.md: merging is not deploying

/opt/observability on vps1 and vps2 is a checkout of the archived admin_emi/thermograph-observability repo. It can never git pull again. A previous observability PR merged and had zero live effect for exactly this reason. The README documents the hand-ship procedure; until the checkouts are re-pointed at the monorepo, that's the only path. Flagged again in 09.


Live traps

The docker-compose.test.yml files pull retired image paths — fixed

Fixed. Both now use jinemi/thermograph/backend, the path build-push.yml actually publishes.

The frontend harness had the worse half of this: it also pinned THERMOGRAPH_BACKEND_TEST_TAG to the split-era v0.0.2-split-ci. Rather than swap one hardcoded pin for another, scripts/backend-for-tests.sh now derives the tag from the checkout — sha-<12hex of git log -1 -- backend/>, the same domain-keyed rule build-push.yml and deploy.yml both use — so the harness follows the tree instead of drifting behind a pin. docker-compose.test.yml requires the variable (:?) rather than defaulting it, so a stale pin can't creep back in silently.

If the derived tag has no published image (backend commits that are still local-only), the script says so and tells you to pin one:

THERMOGRAPH_BACKEND_TEST_TAG=sha-<12hex> make backend-up

Unrelated but adjacent: make test-integration against a freshly-booted throwaway backend currently fails 7 of 16 tests with 503, because nothing in that empty database is warm. Verified identical on the old image and the new one, so it predates these changes and is not caused by them. Not investigated further here.

static/units.js's F_REGIONS is guarded by nothing — fixed

Fixed. format_test.go::TestFCountriesMatchesUnitsJS now parses static/units.js and diffs it against the Go set in both directions.

The subtlety worth knowing: the existing backend cross-check (TestFCountriesMatchesBackend) skips in CI. The frontend image's build context is frontend/, so backend/ is structurally unreachable from the builder stage — and the builder stage is the only place CI runs these tests. It is a checkout-only guard and always was.

The new units.js check does not have that problem, because static/ is inside the context: frontend/Dockerfile copies that one file into the builder stage specifically so the assertion runs during the image build. Verified by mutating units.js and confirming the image build fails, not just the local test run.

The compose project name is load-bearing

infra/docker-compose.yml pins name: thermograph. Without it, running compose from infra/ derives project infra — a silently new stack with empty volumes beside the running one. deploy-dev.sh exports COMPOSE_PROJECT_NAME=thermograph-dev to keep dev separate on purpose. Keep both halves.

deploy.sh hard-resets the host checkout

git reset --hard origin/$BRANCH runs before anything else. Uncommitted edits in /opt/thermograph evaporate. infra-sync.yml does the same reset without deploying. The one deliberate exception is the untracked deploy/.image-tags.env — losing that makes the next single-service deploy roll the sibling onto local.

The deploy scripts eat your rollback image

Both deploy.sh and deploy-stack.sh end by deleting images outside the running pair. Beta now runs Swarm co-resident with prod on vps2, so both usually fail to fully prune only because Swarm's stopped task containers still hold references — dev, the one remaining compose environment (on vps1), is the one that typically holds no local rollback target at all. docker system prune -a is forbidden on a live box for the same reason — but the deploy script is the thing that actually eats the image, not prune. Verify a tag exists before promising a rollback.

daemon and lake are never deploy targets on their own

They ride with backend. All three run the same image at the same tag — a second binary and a second role inside one image — and the /internal/* contract between daemon and web has no version negotiation. There's also a mechanical reason: a single-service deploy runs up -d --no-deps <targets>, so a service in compose but absent from TARGETS would simply never be created.

docker compose config --images <service> does not filter

Confirmed live on Compose v5.3.1: it prints every service's image, one per line, in file order. A | head -1 therefore grabbed db's image, and deploy.sh's daemon capability probe found no daemon binary in a Postgres image and dropped the daemon from every backend deploy. The script now builds the image reference from the same variables compose interpolates.

Prod's (and beta's) Loki streams carry no service label

Prod and beta both run Swarm, co-resident on vps2; their job="docker" streams have no service label. A raw {host="prod", service="backend"} query matches nothing and reads as "no logs" when it means "wrong selector". Use logs_query's service argument. Prod and beta container names also change on every redeploy (the task suffix) — resolve via docker ps --filter name= at call time, never hardcode. Dev, on compose, doesn't have this problem.

A Grafana alert rule can provision cleanly and never fire

A rule whose condition names a non-existent refId is valid YAML, provisions without complaint, and pages nobody. Likewise a contact point holding the literal string $DISCORD_ALERT_WEBHOOK_URL looks completely healthy from the API — and the API redacts the URL, so you cannot confirm interpolation by reading it back. The only proof is a message arriving in #ops-alerts.

Terraform has no state anywhere

A casual terraform apply would attempt full re-provisioning of live hosts. It's executable documentation until state is bootstrapped.

The registry is mesh-only

git.thermograph.org resolves publicly to vps1's IP, but vps1's Caddy rejects /v2/* from outside the WireGuard mesh. Any host that pulls needs 10.10.0.2 git.thermograph.org in /etc/hosts. vps1 itself doesn't — it is the box.

Forgejo runners only have the labels they registered with

GitHub implicitly tags every self-hosted runner self-hosted; Forgejo does not. runs-on: [self-hosted, thermograph-lan] is permanently unschedulable there — a real bug, found and fixed once. If a job silently stops appearing in the Actions history, check this first.

Forgejo's needs: doesn't cross workflow files

build-push.yml and deploy.yml fire from the same push with no ordering guarantee, and a deploy racing ahead of its build and failing "not found" is confirmed-live. That's why deploy.sh's pull has a bounded ~5-minute retry.

/api/v2/cell?prefetch=1 must never spend quota

The hard guarantee: a cold cell answers 204 with no body. Anything named "prefetch" or "warm" calls the cache-only loaders. Breaking this gets the whole service rate-limited.

Background timers must go through leader election

singleton.claim_leader() — flock per host, Postgres advisory lock per cluster. A timer-driven sweep running in N uvicorn workers × M Swarm hosts multiplies Open-Meteo quota use N×M. That is the documented cause of overnight 429/503s after prod went to 3 workers. The heartbeat is the deliberate exception — a beat is one local file append, so duplicates are free.

THERMOGRAPH_DATA_DIR exists because of a real collision

The backend's Python package data/ and its runtime data dir <root>/data collide at the same path once the code lives at the image root. Mounting a volume there shadows and erases data/*.py, so import data.climate fails at boot. Point the env var outside the code tree in a container. And never reintroduce __file__-relative paths in a module — paths.py resolves everything from the repo root for this reason.

Alloy can silently stop tailing a container

Observed on prod, 13½ hours of thermograph_worker logs never reaching Loki while the container was healthy and writing. Not root-caused. Workaround: sudo docker restart alloy-alloy-1. ProdWorkerContainerSilent now catches the class.

Nominatim is ~1 req/s, on one thread

All reverse geocoding funnels through a single dedicated worker draining a queue. The earlier design slept inside the caller's thread — which in the server is one of Starlette's shared sync-threadpool threads.


If you fix one of these

Docs are cheap to correct, and the root CLAUDE.md already asks you to fix repo-era wording as you touch files near it. The three entries struck through above were fixed in the same change that added this guide; the rest are still open, and infra/DEPLOY.md is probably the highest-value one left — it describes a deploy model that has since changed twice.

Back to the index.