thermograph/infra
Emi Griffith 9eecfc8eef
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 7s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Successful in 1m0s
PR build (required check) / gate (pull_request) Successful in 2s
frontend: rewrite the SSR content service in Go
Ports frontend/ (Jinja2/FastAPI, ~1180 LOC) to Go with html/template.
No climate math, no DB, no auth here -- every route fetches from the
backend's /content/* API, so this is I/O-bound glue with no hard-porting
wall; the risk was always in reproducing the rendering exactly, not the
language.

Verified with a golden-HTML diff, not just unit tests: both the Python
original and the Go rewrite were run against the same committed fixtures
(frontend/tests/fixtures/) and every one of the 11 routes compared
byte-for-byte. The only surviving differences after that process are
insignificant inter-tag whitespace and one attribute where Go's stricter
escaper HTML-encodes an apostrophe Jinja left literal (functionally
identical in every browser) -- confirmed programmatically by normalizing
whitespace and unescaping before diffing, not by eyeballing.

That process caught defects unit tests alone would have missed, because
map[string]any has no compile-time field check:

- Render-context keys were snake_case throughout (content.py's Jinja
  convention, ported verbatim) while the templates -- written
  independently -- read PascalCase fields. A missing map key doesn't
  error in html/template, it silently renders empty, so this was invisible
  in every status code and every "it built" signal: title, meta
  description, canonical URL, OpenGraph tags, the homepage's entire ranked
  list, and the brand-tag/nav-active state were all blank across every
  page. Fixed by renaming every key to match each template's own header
  comment (the authoritative per-page field contract) and, where an
  API struct's exported fields already matched what a template needed
  (contentapi.CityInfo, Crumb, HomeRanked, HubCountry, ...), passing the
  struct straight through instead of hand-rewrapping it in a map --
  removes a whole layer of future drift risk, not just this instance of it.
- Three pages 500'd outright: `.ToolHref` needed a fully-composed href
  string, not the bare "lat,lon" fragment the handlers were building; the
  all-time-records table needed the raw contentapi.AllTimeRecords struct,
  not a re-wrapped map.
- JSON-LD was being double-encoded: `<script type="application/ld+json">`
  is JAVASCRIPT context to html/template's contextual escaper regardless
  of the script's `type` attribute, so a template.HTML-typed value placed
  there gets re-escaped as a quoted JS string instead of emitted raw --
  the entire structured-data payload shipped as a JSON string containing
  JSON, which no crawler would parse as the intended object. Needed
  template.JS instead, the type that actually means "trusted JS source."
  The glossary term page's JSON-LD was simply never built at all (the
  Jinja original assembled it inline in the template rather than through
  content.py's context dict, and that got lost in translation) -- added.
- html/template silently strips literal HTML comments AND JavaScript
  comments from the parsed output (verified in isolation, zero template
  actions involved) -- confirmed as real engine behavior, not a bug in
  either port, so both need a FuncMap function returning template.HTML /
  template.JS respectively to survive parsing rather than a literal
  `<!-- -->` or `//` in the template source.

Packaging: multi-stage Go build, final image alpine (not distroless -- the
Swarm stack's env-entrypoint.sh shim needs bash), 187MB -> 22.6MB. Two
defects caught before they reached a host:
- The Swarm stack overrides `entrypoint:` with no `command:`, which drops
  the image's own CMD entirely (Docker/Swarm semantics, not merged) --
  env-entrypoint.sh then fell through to its hardcoded `exec uvicorn
  app:app` fallback, which doesn't exist in this image. Every deploy
  would have exited 127. Fixed with an explicit `command:` on the stack's
  frontend service, and corrected the shim's stale comment claiming CMD
  passes through automatically.
- `COPY --chown=thermograph` resolves the group by NAME at copy time;
  Alpine's `adduser -S` with no `-G` doesn't create a same-named group, so
  the classic (non-BuildKit) Docker builder -- which this CI runner falls
  back to, since it installs plain `docker.io` with no buildx plugin --
  failed outright. Fixed with an explicit group and numeric --chown.

Verification: go build/vet/test -race clean across all packages; the
Docker image builds and passes its embedded go test step under both
BuildKit and the classic builder; shellcheck 0 findings on the one script
touched; rebased onto current main (the ERA5 lake stack landed on both
main and dev during this work -- confirmed additive, no overlap with
frontend/daemon).
2026-07-23 17:51:31 -07:00
..
.claude/skills/key-gaps Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
deploy frontend: rewrite the SSR content service in Go 2026-07-23 17:51:31 -07:00
lake-iceberg Iceberg conversion container for the ERA5 lake (infra/lake-iceberg) (#24) 2026-07-23 22:56:27 +00:00
terraform Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
.env.example daemon: move the Discord gateway and scheduler out of the web process into Go (#21) 2026-07-23 22:49:54 +00:00
.gitignore Iceberg conversion container for the ERA5 lake (infra/lake-iceberg) (#24) 2026-07-23 22:56:27 +00:00
.sops.yaml Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
ACCESS.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
CLAUDE.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
DEPLOY-DEV.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
DEPLOY.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
docker-compose.dev.yml Reconcile: merge main (shellcheck guard, Go daemon) into dev 2026-07-23 17:06:52 -07:00
docker-compose.openmeteo.yml Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
docker-compose.yml frontend: rewrite the SSR content service in Go 2026-07-23 17:51:31 -07:00
docker-stack.yml Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
Makefile Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00
README.md Subtree-merge thermograph-infra (origin/main) into infra/ 2026-07-22 22:01:11 -07:00

thermograph-infra

Infrastructure for Thermograph: Terraform host provisioning, the SOPS+age secrets vault, Docker Swarm/WireGuard networking, Forgejo, Caddy, and the deploy scripts that run the already-built app image on each host. Extracted from the app monorepo (emi/thermograph) — the app repo owns building and testing the app; this repo owns running it.

  • terraform/ — provisions/configures hosts (SSH-driven by default; an optional GCP-creating module is scaffolded, no live resources yet) and triggers each deploy. See terraform/README.md.
  • deploy/secrets/ — the git-native SOPS+age secrets vault (every app secret, encrypted at rest, rendered at deploy time). See deploy/secrets/README.md.
  • deploy/swarm/, deploy/forgejo/ — the 3-node WireGuard/Swarm cluster that hosts Forgejo (git + CI + registry); does not run the app itself. See ACCESS.md and the READMEs under each directory.
  • deploy/deploy.sh — pulls the pinned app image (IMAGE_TAG) and rolls the compose stack; invoked by Terraform and by the app repo's .forgejo/workflows/deploy.yml over SSH.
  • docker-compose*.yml, docker-stack.yml — how the app image runs (compose in production today; docker-stack.yml is a design record for a possible future Swarm-based app deploy, not currently live).

The app's own source, Dockerfile, and build/test CI stay in the app repo — this repo never checks out app source; hosts only pull tagged images from the registry. See ACCESS.md for host access and the Swarm/Forgejo topology, and terraform/README.md for the day-to-day plan/apply workflow.

Branches & how changes reach each environment

  • main — what prod and beta run: their /opt/thermograph checkouts git reset --hard origin/main at the start of every deploy (deploy/deploy.sh). A merge to main reaches those hosts on the next app deploy (or a by-hand deploy.sh run); there is no separate infra deploy trigger.
  • dev — what LAN dev runs: ~/thermograph-dev resets to it via deploy/deploy-dev.sh. Keep it fast-forwarded to main (infra changes are not environment-staged today; the branches exist so LAN dev can trail or lead when needed).
  • release — currently consumed by nothing (prod tracks main, not release). It exists to mirror the app repos' dev→main→release promotion shape if per-environment infra staging is ever wanted; until then, treat main as live-everywhere.

Note the asymmetry with the app repos: app code IS environment-staged (dev→main→release maps to LAN→beta→prod via image tags), infra is not.