thermograph/frontend/server/internal/handlers/shells.go
Emi Griffith 9eecfc8eef
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 7s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Successful in 1m0s
PR build (required check) / gate (pull_request) Successful in 2s
frontend: rewrite the SSR content service in Go
Ports frontend/ (Jinja2/FastAPI, ~1180 LOC) to Go with html/template.
No climate math, no DB, no auth here -- every route fetches from the
backend's /content/* API, so this is I/O-bound glue with no hard-porting
wall; the risk was always in reproducing the rendering exactly, not the
language.

Verified with a golden-HTML diff, not just unit tests: both the Python
original and the Go rewrite were run against the same committed fixtures
(frontend/tests/fixtures/) and every one of the 11 routes compared
byte-for-byte. The only surviving differences after that process are
insignificant inter-tag whitespace and one attribute where Go's stricter
escaper HTML-encodes an apostrophe Jinja left literal (functionally
identical in every browser) -- confirmed programmatically by normalizing
whitespace and unescaping before diffing, not by eyeballing.

That process caught defects unit tests alone would have missed, because
map[string]any has no compile-time field check:

- Render-context keys were snake_case throughout (content.py's Jinja
  convention, ported verbatim) while the templates -- written
  independently -- read PascalCase fields. A missing map key doesn't
  error in html/template, it silently renders empty, so this was invisible
  in every status code and every "it built" signal: title, meta
  description, canonical URL, OpenGraph tags, the homepage's entire ranked
  list, and the brand-tag/nav-active state were all blank across every
  page. Fixed by renaming every key to match each template's own header
  comment (the authoritative per-page field contract) and, where an
  API struct's exported fields already matched what a template needed
  (contentapi.CityInfo, Crumb, HomeRanked, HubCountry, ...), passing the
  struct straight through instead of hand-rewrapping it in a map --
  removes a whole layer of future drift risk, not just this instance of it.
- Three pages 500'd outright: `.ToolHref` needed a fully-composed href
  string, not the bare "lat,lon" fragment the handlers were building; the
  all-time-records table needed the raw contentapi.AllTimeRecords struct,
  not a re-wrapped map.
- JSON-LD was being double-encoded: `<script type="application/ld+json">`
  is JAVASCRIPT context to html/template's contextual escaper regardless
  of the script's `type` attribute, so a template.HTML-typed value placed
  there gets re-escaped as a quoted JS string instead of emitted raw --
  the entire structured-data payload shipped as a JSON string containing
  JSON, which no crawler would parse as the intended object. Needed
  template.JS instead, the type that actually means "trusted JS source."
  The glossary term page's JSON-LD was simply never built at all (the
  Jinja original assembled it inline in the template rather than through
  content.py's context dict, and that got lost in translation) -- added.
- html/template silently strips literal HTML comments AND JavaScript
  comments from the parsed output (verified in isolation, zero template
  actions involved) -- confirmed as real engine behavior, not a bug in
  either port, so both need a FuncMap function returning template.HTML /
  template.JS respectively to survive parsing rather than a literal
  `<!-- -->` or `//` in the template source.

Packaging: multi-stage Go build, final image alpine (not distroless -- the
Swarm stack's env-entrypoint.sh shim needs bash), 187MB -> 22.6MB. Two
defects caught before they reached a host:
- The Swarm stack overrides `entrypoint:` with no `command:`, which drops
  the image's own CMD entirely (Docker/Swarm semantics, not merged) --
  env-entrypoint.sh then fell through to its hardcoded `exec uvicorn
  app:app` fallback, which doesn't exist in this image. Every deploy
  would have exited 127. Fixed with an explicit `command:` on the stack's
  frontend service, and corrected the shim's stale comment claiming CMD
  passes through automatically.
- `COPY --chown=thermograph` resolves the group by NAME at copy time;
  Alpine's `adduser -S` with no `-G` doesn't create a same-named group, so
  the classic (non-BuildKit) Docker builder -- which this CI runner falls
  back to, since it installs plain `docker.io` with no buildx plugin --
  failed outright. Fixed with an explicit group and numeric --chown.

Verification: go build/vet/test -race clean across all packages; the
Docker image builds and passes its embedded go test step under both
BuildKit and the classic builder; shellcheck 0 findings on the one script
touched; rebased onto current main (the ERA5 lake stack landed on both
main and dev during this work -- confirmed additive, no overlap with
frontend/daemon).
2026-07-23 17:51:31 -07:00

131 lines
4.5 KiB
Go

package handlers
import (
"html/template"
"io"
"net/http"
"os"
"path/filepath"
"strings"
"sync"
"thermograph/frontend/internal/render"
)
// headVerifyHTML is content.head_verify_html for the SPA shells: the
// search-engine ownership-verification <meta> tags, from env (empty when
// unset). The SSR pages carry the identical markup via the template FuncMap's
// head_verify entry (internal/content owns that copy); the shells need it
// here because their HTML never passes through the template engine.
// template.HTMLEscapeString emits the same five entities markupsafe escaped
// (&amp; &lt; &gt; &#39; &#34;).
func headVerifyHTML(google, bing string) template.HTML {
var metas []string
if google != "" {
metas = append(metas, `<meta name="google-site-verification" content="`+
template.HTMLEscapeString(google)+`">`)
}
if bing != "" {
metas = append(metas, `<meta name="msvalidate.01" content="`+
template.HTMLEscapeString(bing)+`">`)
}
return template.HTML(strings.Join(metas, "\n "))
}
// shellMemoMax bounds the per-origin memo. The origin is client-controlled
// (Host/X-Forwarded-Host), so an unbounded map is the same cheap
// memory-exhaustion vector api_client.py's LRU cap closed — the Python's
// _by_origin dict had no bound (one canonical origin in every real topology
// made it moot); a flat cap keeps that property for legit traffic and just
// resets the memo under abuse instead of growing forever.
const shellMemoMax = 1024
// shell is one SPA-shell route's state — the Go port of app.py's _page():
// the file (and the verification <meta> tags, both constant for the process's
// lifetime) is read and prepped once, not on every request; only the
// __ORIGIN__ substitution actually varies per request, and even that repeats
// across requests (one canonical origin in the common topology), so the
// substituted HTML + its ETag are memoized per origin instead of
// re-read-and-re-sha1'd every time.
type shell struct {
srv *Server
file string
mu sync.Mutex
template string // "" = not loaded yet; a failed read retries next request
byOrigin map[string]shellEntry
}
type shellEntry struct {
html string
etag string
}
// shellHandler builds the handler for one SPA-shell HTML page (the
// interactive tool's calendar/day/score/compare/legend/alerts views). It
// serves the file with its __ORIGIN__ placeholder (the link-preview/Open
// Graph tags) filled in — preview crawlers need absolute URLs, and the host
// differs between LAN and prod. originOf (not a simpler duplicate) matters
// here: this route is reached both directly (Caddy) and through backend's
// proxy fallback, and only that version prefers X-Forwarded-Host over Host —
// required for the proxied case to resolve the real browser-facing host
// instead of this internal hop's own address.
func (s *Server) shellHandler(file string) http.HandlerFunc {
sh := &shell{srv: s, file: file, byOrigin: make(map[string]shellEntry)}
return sh.serve
}
// load reads and preps the shell file; the caller holds sh.mu. Search-engine
// verification <meta> tags (same source as the SSR content pages) are folded
// in here, so the interactive tool's own pages carry them too.
func (sh *shell) load() (string, error) {
if sh.template != "" {
return sh.template, nil
}
raw, err := os.ReadFile(filepath.Join(sh.srv.staticDir, sh.file))
if err != nil {
return "", err
}
html := string(raw)
if verify := string(sh.srv.headVerify); verify != "" {
html = strings.Replace(html, "<head>", "<head>\n "+verify, 1)
}
sh.template = html
return html, nil
}
func (sh *shell) serve(w http.ResponseWriter, r *http.Request) {
originPrefix := originOf(r) + sh.srv.base
sh.mu.Lock()
ent, ok := sh.byOrigin[originPrefix]
if !ok {
tpl, err := sh.load()
if err != nil {
sh.mu.Unlock()
sh.srv.serverError(w, r, err)
return
}
html := strings.ReplaceAll(tpl, "__ORIGIN__", originPrefix)
ent = shellEntry{html: html, etag: render.ETag([]byte(html))}
if len(sh.byOrigin) >= shellMemoMax {
clear(sh.byOrigin)
}
sh.byOrigin[originPrefix] = ent
}
sh.mu.Unlock()
// app.py's _page compared If-None-Match to the ETag verbatim (no
// comma-splitting) — NotModifiedExact keeps that precise behavior.
if render.NotModifiedExact(r.Header.Get("If-None-Match"), ent.etag) {
w.Header().Set("ETag", ent.etag)
w.WriteHeader(http.StatusNotModified)
return
}
w.Header().Set("ETag", ent.etag)
w.Header().Set("Content-Type", "text/html; charset=utf-8")
w.WriteHeader(http.StatusOK)
if r.Method != http.MethodHead {
_, _ = io.WriteString(w, ent.html)
}
}