13 KiB
11. Traps and stale docs
Two kinds of thing here: live traps (correct code that will surprise you) and stale docs (files in this repo that currently describe a system that no longer exists). The second list matters as much as the first — this repo has a lot of documentation, most of it excellent, and some of it describing the world before a rewrite.
Everything below was checked against the tree at the time of writing.
Stale docs in this repo
frontend/CLAUDE.md and frontend/README.md describe a service that is not deployed — fixed
frontend/CLAUDE.md and frontend/README.md describe a service that is not deployedFixed. Both now describe the Go service (frontend/server/), state plainly
that the Python files at that level are the superseded original, and correct
CLAUDE.md's claim that make test-unit is "the tier CI runs" — it isn't.
CI builds the frontend image, and frontend/Dockerfile's builder stage runs
gofmt -l + go vet + go test ./... before compiling, so a failing Go
test is what fails CI. The Python tiers are local-only.
Still worth knowing: the superseded Python files carry duplicate copies of
some contracts — api_client.py has a third API_VERSION pin and
format.py a fourth F_COUNTRIES. Nothing deploys them, but if you run them
locally, keep them in step or you'll debug a phantom.
Whether to delete them outright is an open question, not a settled one: the Python test tiers still run, against fixtures the Go tests share. Raise it rather than doing it in passing.
infra/DEPLOY.md describes the pre-Swarm, pre-monorepo deploy
Three separate layers of drift:
- It says prod is deployed with
terraform apply. Prod is deployed by a push toreleasefiringdeploy.yml, exactly like beta'smainpath — see 07. - It says prod runs a compose stack. Prod runs Swarm
(
infra/deploy/stack/). - Its file table points at
/opt/thermograph/deploy/deploy.shand/opt/thermograph/docker-compose.yml. Both moved underinfra/at the monorepo cutover.
The document's own header flags the legacy provision.sh / systemd / venv
walkthrough as superseded — but the drift goes further than that note admits.
The key-management sections (Part 1) and the /etc/hosts registry note are
still accurate and useful.
infra/DEPLOY-DEV.md and infra/ACCESS.md reference deleted workflows
Both talk about .forgejo/workflows/deploy-dev.yml. It does not exist. The two
LAN-dev deploy workflows were deleted rather than ported at the CI
consolidation — they were documented as inert (they targeted a monorepo layout
at ~/thermograph-dev on a box still holding a split-era checkout). LAN dev
is a local make dev-up concern, not a CI environment.
ACCESS.md §3c also lists {build,pr-build,deploy-dev,deploy}.yml as the
workflow set; there are now nine files and the deploy/build ones were collapsed.
Everything else in ACCESS.md — the host table, the agent-access model, the key
rotation procedure, the Swarm/Forgejo tracks — is current and worth reading.
backend/README.md is written for the split-repo era
It describes thermograph-backend as its own repo with sibling repos, uses the
retired image path git.thermograph.org/emi/thermograph-backend/app, and says
"there is no Makefile in this repo yet" — there is (make test, make smoke,
make clean-venv).
The "How it works" section (grid → data → percentiles → caching) is one of the best explanations of the domain model anywhere in the repo. Read it; just mentally translate the topology.
CUTOVER-NOTES.md carries a superseded block, correctly labelled
Its per-domain workflow names (backend-build-push.yml, backend-deploy.yml, …)
no longer exist — the six deploy workflows collapsed into deploy.yml and the
two build-push workflows into build-push.yml. The file says so in an explicit
"Superseded 2026-07-25" note, and everything it describes about behaviour is
unchanged.
It also references backend/MONOREPO-CUTOVER-PLAN.md, which is not in the tree.
CUTOVER-NOTES.md remains the source of truth for what is and isn't live.
The domain CLAUDE.md files still use repo-era wording
They say "this repo" and link to sibling repo names. Mentally map
thermograph-backend → backend/, and fix the wording as you touch files near
it — that's the standing instruction in the root CLAUDE.md.
observability/README.md: merging is not deploying
/opt/observability on beta and prod is a checkout of the archived
emi/thermograph-observability repo. It can never git pull again. A previous
observability PR merged and had zero live effect for exactly this reason.
The README documents the hand-ship procedure; until the checkouts are re-pointed
at the monorepo, that's the only path. Flagged again in
09.
Live traps
The docker-compose.test.yml files pull retired image paths — fixed
docker-compose.test.yml files pull retired image pathsFixed. Both now use emi/thermograph/backend, the path build-push.yml
actually publishes.
The frontend harness had the worse half of this: it also pinned
THERMOGRAPH_BACKEND_TEST_TAG to the split-era v0.0.2-split-ci. Rather than
swap one hardcoded pin for another, scripts/backend-for-tests.sh now
derives the tag from the checkout — sha-<12hex of git log -1 -- backend/>,
the same domain-keyed rule build-push.yml and deploy.yml both use — so the
harness follows the tree instead of drifting behind a pin. docker-compose.test.yml
requires the variable (:?) rather than defaulting it, so a stale pin can't
creep back in silently.
If the derived tag has no published image (backend commits that are still local-only), the script says so and tells you to pin one:
THERMOGRAPH_BACKEND_TEST_TAG=sha-<12hex> make backend-up
Unrelated but adjacent:
make test-integrationagainst a freshly-booted throwaway backend currently fails 7 of 16 tests with503, because nothing in that empty database is warm. Verified identical on the old image and the new one, so it predates these changes and is not caused by them. Not investigated further here.
static/units.js's F_REGIONS is guarded by nothing — fixed
static/units.js's F_REGIONS is guarded by nothingFixed. format_test.go::TestFCountriesMatchesUnitsJS now parses
static/units.js and diffs it against the Go set in both directions.
The subtlety worth knowing: the existing backend cross-check
(TestFCountriesMatchesBackend) skips in CI. The frontend image's build
context is frontend/, so backend/ is structurally unreachable from the
builder stage — and the builder stage is the only place CI runs these tests. It
is a checkout-only guard and always was.
The new units.js check does not have that problem, because static/ is
inside the context: frontend/Dockerfile copies that one file into the builder
stage specifically so the assertion runs during the image build. Verified by
mutating units.js and confirming the image build fails, not just the local
test run.
The compose project name is load-bearing
infra/docker-compose.yml pins name: thermograph. Without it, running compose
from infra/ derives project infra — a silently new stack with empty
volumes beside the running one. deploy-dev.sh exports
COMPOSE_PROJECT_NAME=thermograph-dev to keep LAN dev separate on purpose.
Keep both halves.
deploy.sh hard-resets the host checkout
git reset --hard origin/$BRANCH runs before anything else. Uncommitted edits
in /opt/thermograph evaporate. infra-sync.yml does the same reset without
deploying. The one deliberate exception is the untracked
deploy/.image-tags.env — losing that makes the next single-service deploy
roll the sibling onto local.
The deploy scripts eat your rollback image
Both deploy.sh and deploy-stack.sh end by deleting images outside the
running pair. On beta that succeeds, so beta typically holds no local rollback
target at all; on prod it usually fails only because Swarm's stopped task
containers still hold references. docker system prune -a is forbidden on a
live box for the same reason — but the deploy script is the thing that actually
eats the image, not prune. Verify a tag exists before promising a rollback.
daemon and lake are never deploy targets on their own
They ride with backend. All three run the same image at the same tag — a
second binary and a second role inside one image — and the /internal/*
contract between daemon and web has no version negotiation. There's also a
mechanical reason: a single-service deploy runs up -d --no-deps <targets>, so
a service in compose but absent from TARGETS would simply never be created.
docker compose config --images <service> does not filter
Confirmed live on Compose v5.3.1: it prints every service's image, one per
line, in file order. A | head -1 therefore grabbed db's image, and
deploy.sh's daemon capability probe found no daemon binary in a Postgres image
and dropped the daemon from every backend deploy. The script now builds the
image reference from the same variables compose interpolates.
Prod's Loki streams carry no service label
Prod runs Swarm; its job="docker" streams have no service label. A raw
{host="prod", service="backend"} query matches nothing and reads as "no logs"
when it means "wrong selector". Use logs_query's service argument. Prod
container names also change on every redeploy (the task suffix) — resolve via
docker ps --filter name= at call time, never hardcode.
A Grafana alert rule can provision cleanly and never fire
A rule whose condition names a non-existent refId is valid YAML, provisions
without complaint, and pages nobody. Likewise a contact point holding the
literal string $DISCORD_ALERT_WEBHOOK_URL looks completely healthy from the
API — and the API redacts the URL, so you cannot confirm interpolation by
reading it back. The only proof is a message arriving in #ops-alerts.
Terraform has no state anywhere
A casual terraform apply would attempt full re-provisioning of live hosts.
It's executable documentation until state is bootstrapped.
The registry is mesh-only
git.thermograph.org resolves publicly to beta's IP, but beta's Caddy rejects
/v2/* from outside the WireGuard mesh. Any host that pulls needs
10.10.0.2 git.thermograph.org in /etc/hosts. Beta itself doesn't — it is
the box.
Forgejo runners only have the labels they registered with
GitHub implicitly tags every self-hosted runner self-hosted; Forgejo does not.
runs-on: [self-hosted, thermograph-lan] is permanently unschedulable there —
a real bug, found and fixed once. If a job silently stops appearing in the
Actions history, check this first.
Forgejo's needs: doesn't cross workflow files
build-push.yml and deploy.yml fire from the same push with no ordering
guarantee, and a deploy racing ahead of its build and failing "not found" is
confirmed-live. That's why deploy.sh's pull has a bounded ~5-minute retry.
/api/v2/cell?prefetch=1 must never spend quota
The hard guarantee: a cold cell answers 204 with no body. Anything named
"prefetch" or "warm" calls the cache-only loaders. Breaking this gets the whole
service rate-limited.
Background timers must go through leader election
singleton.claim_leader() — flock per host, Postgres advisory lock per cluster.
A timer-driven sweep running in N uvicorn workers × M Swarm hosts multiplies
Open-Meteo quota use N×M. That is the documented cause of overnight 429/503s
after prod went to 3 workers. The heartbeat is the deliberate exception — a beat
is one local file append, so duplicates are free.
THERMOGRAPH_DATA_DIR exists because of a real collision
The backend's Python package data/ and its runtime data dir <root>/data
collide at the same path once the code lives at the image root. Mounting a
volume there shadows and erases data/*.py, so import data.climate fails at
boot. Point the env var outside the code tree in a container. And never
reintroduce __file__-relative paths in a module — paths.py resolves
everything from the repo root for this reason.
Alloy can silently stop tailing a container
Observed on prod, 13½ hours of thermograph_worker logs never reaching Loki
while the container was healthy and writing. Not root-caused. Workaround:
sudo docker restart alloy-alloy-1. ProdWorkerContainerSilent now catches the
class.
Nominatim is ~1 req/s, on one thread
All reverse geocoding funnels through a single dedicated worker draining a queue. The earlier design slept inside the caller's thread — which in the server is one of Starlette's shared sync-threadpool threads.
If you fix one of these
Docs are cheap to correct, and the root CLAUDE.md already asks you to fix
repo-era wording as you touch files near it. The three entries struck through
above were fixed in the same change that added this guide; the rest are still
open, and infra/DEPLOY.md is probably the highest-value one left — it
describes a deploy model that has since changed twice.
Back to the index.