thermograph/docs/onboarding/02-setup.md
Emi Griffith 53e2eb7c84
All checks were successful
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-centralis (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 9s
secrets-guard / encrypted (push) Successful in 5s
Validate observability stack / validate (push) Successful in 12s
shell-lint / shellcheck (push) Successful in 8s
Build + push images (Forgejo registry) / build-push (frontend) (push) Successful in 26s
Build + push images (Forgejo registry) / build-push (backend) (push) Successful in 1m10s
Deploy / deploy (backend) (push) Successful in 1m41s
Deploy / deploy (frontend) (push) Successful in 1m48s
secrets-guard / encrypted (pull_request) Successful in 8s
shell-lint / shellcheck (pull_request) Successful in 9s
PR build (required check) / changes (pull_request) Successful in 20s
PR build (required check) / validate-observability (pull_request) Successful in 15s
PR build (required check) / build-frontend (pull_request) Successful in 17s
PR build (required check) / build-backend (pull_request) Successful in 2m26s
PR build (required check) / gate (pull_request) Successful in 1s
registry: move the registry host to dev.jinemi.com, MCP to mcp.jinemi.com
Completes the Forgejo domain migration. ROOT_URL moved to dev.jinemi.com
earlier; the registry half was deliberately deferred. Every image name,
REGISTRY_HOST default, runner label and --add-host pin now names
dev.jinemi.com, so the registry host and the bearer-token realm agree again.

Both names address the same Forgejo, so no image needs re-pushing and a
rollback to a tag pushed under the old prefix still resolves.
git.thermograph.org therefore stays served off the same Caddy site block --
one block, so the /v2/* mesh-only matcher keeps covering both names -- for
pre-migration tags and for runners holding it as their registered instance
URL.

Mesh clients now pin both names in /etc/hosts: the new one as registry host
and token realm, the old one for pre-migration tags. runner-vps2/config.yaml
carries both --add-host entries for the same reason.

Also renames Centralis' endpoint to mcp.jinemi.com in the two places this
repo names it; Centralis itself is provisioned outside this repo.

Host-side steps this cannot do (documented in deploy/forgejo/README.md,
"Host-side steps"): the Forgejo Actions variable REGISTRY_HOST, docker login
against the new host, and the /etc/hosts pins.
2026-08-01 16:06:22 -07:00

250 lines
10 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# 2. Local setup
Every command in this document was run against this checkout and produced the
output shown. If one fails for you, the difference is your machine, not the doc.
## Toolchain
| Tool | Why | Notes |
|---|---|---|
| **Python 3.12** | backend runtime and test suite | Must have the `_sqlite3` extension. A pyenv 3.10 without it will fail at `conftest` import. `scripts/test.sh` looks for `python3.12` (or uses `uv`) rather than whatever `python3` is on `PATH` — deliberately. |
| **Go 1.26** | frontend service and the backend daemon | Both are static `CGO_ENABLED=0` builds. |
| **Docker** | images, the compose/Swarm stacks, smoke tests | Docker CLI ≥ 27 is fine. |
| **uv** *(recommended)* | fast venv creation | `scripts/test.sh` uses it when present. |
| **shellcheck v0.11.0** | matches CI's pin exactly | Install to `~/.local/bin/shellcheck`. A *different* version can invent new findings and fail CI on an unrelated push. |
| **sops + age** | reading/editing the secrets vault | Only needed if you touch `infra/deploy/secrets/`. |
| **jq** | the repo's `.claude` hooks use it | Without it the prod-guard hook fails toward *asking*, which is safe but noisy. |
Check what you have:
```bash
which go python3.12 uv docker sops age jq shellcheck
go version && python3.12 --version && docker --version
```
## Backend
### Test suite (hermetic, no network, no Docker)
```bash
cd backend
make test # or: ./scripts/test.sh
make test ARGS='tests/data -q' # pass pytest args through
```
Verified: **429 passed, 8 skipped in 11.57s** on a cold venv build.
The first run builds `.venv-test/` on Python 3.12 and installs
`requirements-dev.txt`. `tests/conftest.py` is what makes the suite hermetic:
- no real Open-Meteo, Nominatim or GeoNames calls (the places index is marked
as already-loaded so no background download starts);
- a throwaway SQLite accounts DB and derived store in `/tmp`, never the repo's
`data/`;
- audit/error/access/activity/heartbeat log dirs redirected to `/tmp`;
- notifier and heartbeat threads disabled;
- `THERMOGRAPH_FRONTEND_BASE_INTERNAL` pointed at an unreachable placeholder
(the app fails loud at import without it).
### Boot it locally, no Docker, no Postgres
The backend falls back to SQLite when `THERMOGRAPH_DATABASE_URL` isn't a
Postgres URL, so this is enough:
```bash
cd backend
THERMOGRAPH_FRONTEND_BASE_INTERNAL=http://127.0.0.1:8080 \
THERMOGRAPH_BASE=/thermograph \
THERMOGRAPH_ENABLE_NOTIFIER=0 \
THERMOGRAPH_ENABLE_HEARTBEAT=0 \
.venv-test/bin/python -m uvicorn app:app --host 127.0.0.1 --port 8137
```
Verified responses:
```
GET /healthz → {"status":"ok","role":"all"}
GET /thermograph/api/version → {"backend_version":"2","min_frontend":"1","payload_ver":"p2"}
```
`app:app` is a one-line re-export shim for `web/app.py`. Keep it — systemd, CI
and the container entrypoint all target that name.
Note the port: **8137** is the backend everywhere in this project (compose,
Caddy, the smoke harness at 18137, the frontend's default internal base).
### Image smoke test
```bash
cd backend
make smoke # builds the image, boots it + a throwaway TimescaleDB (tmpfs),
# asserts /healthz and /api/version
```
Uses `docker-compose.test.yml` on host port 18137 so it can't collide with a
dev server on 8137.
## Frontend
**The live frontend is Go.** `frontend/server/` is what builds, tests, ships
and runs. The Python files one level up (`app.py`, `content.py`,
`api_client.py`, `format.py`) are the superseded original — see
[traps](11-traps.md).
### Test
```bash
cd frontend/server
go build ./... && go vet ./... && go test ./...
```
Verified: all seven packages `ok` (`config`, `content`, `contentapi`,
`contentdata`, `format`, `handlers`, `render`).
These same commands run inside `frontend/Dockerfile`'s builder stage — plus a
`gofmt -l` check that must come back empty — so **a failing Go test fails the
image build**, which is how CI catches it. There is no separate frontend test
step in the workflows.
### Run it
The process resolves `static/` and `content/` **relative to its working
directory**, so build in `server/` and run from `frontend/`:
```bash
cd frontend/server && go build -o thermograph-frontend .
cd ..
THERMOGRAPH_API_BASE_INTERNAL=http://127.0.0.1:8137 \
THERMOGRAPH_BASE=/thermograph \
PORT=8080 \
./server/thermograph-frontend
```
Verified: `GET /healthz``{"status":"ok"}`, with a structured JSON log line
per request.
`THERMOGRAPH_API_BASE_INTERNAL` is **required** — boot fails loudly without it,
by design. Optional: `THERMOGRAPH_BASE` (default `/thermograph`; the image sets
`/`), `THERMOGRAPH_API_VERSION` (default `v2` — only ever change it per the
[API-version contract](06-contracts.md)), `THERMOGRAPH_API_BASE_PUBLIC`,
`THERMOGRAPH_SSR_CACHE_TTL` (seconds, default 600), `THERMOGRAPH_GOOGLE_VERIFY`,
`THERMOGRAPH_BING_VERIFY`, `PORT` (default 8080).
### Frontend against a real backend container
```bash
cd frontend
make backend-up # pulls + runs the published backend image + throwaway db
# on 127.0.0.1:18137, waits for /healthz, prints the URL
make backend-down
```
The image tag is **derived from your checkout**`sha-<12hex of git log -1 --
backend/>`, the same domain-keyed rule `build-push.yml` and `deploy.yml` use —
so the harness follows the tree. If those backend commits are still local-only,
no image exists yet and the script says so; pin a published build with
`THERMOGRAPH_BACKEND_TEST_TAG=sha-<12hex>`.
`make test-integration` runs the Python integration tier against that. CI does
**not** run it — it needs a live backend container.
> Heads-up: against a freshly-booted throwaway backend, that tier currently
> fails 7 of 16 with `503` — the database is empty, so nothing is warm.
> Pre-existing, and unrelated to which image tag you use.
## The daemon
```bash
cd backend/daemon
go build ./... && go vet ./... && go test ./...
THERMOGRAPH_INTERNAL_TOKEN=dev-token \
THERMOGRAPH_API_BASE_INTERNAL=http://localhost:8137 \
go run .
```
It refuses to start without `THERMOGRAPH_INTERNAL_TOKEN` — and the backend
answers `404` on the whole `/internal/*` surface when that token is unset. Both
ends fail closed. With Discord unconfigured it logs once and runs cron-only.
## The full stack, locally
From the repo root:
```bash
make install # build both images as :local -- required once, see below
make up # docker-compose.yml + docker-compose.dev.yml overlay:
# uncapped CPU, backend on 127.0.0.1:8137
make down
```
`make install` is not optional on a fresh checkout. Neither compose file
carries a `build:` for backend or frontend — each ships as its own registry
image — so with no image present compose tries to **pull**
`dev.jinemi.com/emi/thermograph/backend:local`, a tag published nowhere,
and fails with `403 Forbidden` against a registry you have no login for.
`install` builds that tag locally from `backend/Dockerfile` and
`frontend/Dockerfile`. `make up` checks both images exist and points you here
if they don't.
`make -C infra dev-up` / `dev-down` remain equivalent to `up` / `down` — same
overlay, same project name, same stack.
Backend publishes on `${DEV_BIND_ADDR}`, **loopback by default**. Pass
`DEV_BIND_ADDR=0.0.0.0 make up` to reach it from phones on your own Wi-Fi —
fine on a laptop LAN, never on a fleet host.
This is a **laptop convenience**, not the hosted dev environment: the actual
`dev` environment now runs on vps1 with its own Postgres, deployed by CI, and
bound only to the WireGuard mesh (`10.10.0.2:8137`) — never `0.0.0.0`, since
vps1 is a public box. See [Infra and secrets](08-infra-secrets.md).
Both entry points set `COMPOSE_PROJECT_NAME=thermograph-dev` so this stack
keeps volumes separate from anything else. **Do not remove either half of the
project-name pinning** — `infra/docker-compose.yml` pins `name: thermograph`,
and without it running compose from `infra/` derives the project name `infra`,
silently creating a *new* stack beside the running one with fresh volumes.
Other `infra/Makefile` targets: `up`/`down` (pull + run the published images),
`db-up`/`db-down` (just Postgres, e.g. to run the app from a venv against it),
`om-up`/`om-down`/`om-backfill` (the self-hosted Open-Meteo overlay — the
backfill writes ~11.5 TB and takes hours).
## Connectors (Centralis and friends)
The fleet is not reachable from a laptop off the WireGuard mesh, and Forgejo is
mesh-only. **Centralis** is the control plane that fronts all of it — the app
database, the ERA5 lake, fleet logs, Grafana, Forgejo, docs and notes, and
Discord.
```bash
claude mcp add --transport http centralis https://mcp.jinemi.com/mcp \
--header "Authorization: Bearer $CENTRALIS_TOKEN"
```
Ask the operator for a token; don't share it. Verify with *"what's running on
prod right now?"* — it should call `fleet_status` and list the Swarm services.
Also worth installing locally: **Chrome DevTools MCP** (design verification
needs a real browser on your machine) and **Figma**. Grafana's official MCP
server is optional and read-only — but remember dashboards are provisioned from
repo JSON, so a durable change is still a PR via `dashboard_write`.
Run `mcp__centralis__onboarding` for the current, authoritative connector list;
it will be fresher than this page.
## Repo-local guardrails
`.claude/settings.json` wires three hooks that travel with the checkout:
| Hook | When | What |
|---|---|---|
| `prod-guard.sh` | before Bash / live-host MCP calls | Classifies by **allowlist**: only positively-recognised read-only commands pass; everything else asks. vps1 is guarded as strictly as vps2 — it hosts Forgejo, Grafana *and* the mesh-only `dev` environment, so a destructive command there takes out git, CI and the registry at once, not just a dev sandbox. |
| `secrets-guard.sh` | before Write/Edit | Denies any direct write to `infra/deploy/secrets/*.yaml`. Use `sops edit`. |
| `lint-after-edit.sh` | after Write/Edit | shellchecks an edited `*.sh` and feeds findings straight back. Exits quietly if shellcheck is missing. |
They enforce what `CLAUDE.md` can only ask for. If you change `prod-guard.sh`'s
classifier, re-read `.claude/hooks/README.md` first — it documents a real
silent-total-bypass failure mode in the parsing loop.
Next: [Repo map](03-repo-map.md).