thermograph/frontend/server/internal/content/seo.go
Emi Griffith 9eecfc8eef
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 7s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Successful in 1m0s
PR build (required check) / gate (pull_request) Successful in 2s
frontend: rewrite the SSR content service in Go
Ports frontend/ (Jinja2/FastAPI, ~1180 LOC) to Go with html/template.
No climate math, no DB, no auth here -- every route fetches from the
backend's /content/* API, so this is I/O-bound glue with no hard-porting
wall; the risk was always in reproducing the rendering exactly, not the
language.

Verified with a golden-HTML diff, not just unit tests: both the Python
original and the Go rewrite were run against the same committed fixtures
(frontend/tests/fixtures/) and every one of the 11 routes compared
byte-for-byte. The only surviving differences after that process are
insignificant inter-tag whitespace and one attribute where Go's stricter
escaper HTML-encodes an apostrophe Jinja left literal (functionally
identical in every browser) -- confirmed programmatically by normalizing
whitespace and unescaping before diffing, not by eyeballing.

That process caught defects unit tests alone would have missed, because
map[string]any has no compile-time field check:

- Render-context keys were snake_case throughout (content.py's Jinja
  convention, ported verbatim) while the templates -- written
  independently -- read PascalCase fields. A missing map key doesn't
  error in html/template, it silently renders empty, so this was invisible
  in every status code and every "it built" signal: title, meta
  description, canonical URL, OpenGraph tags, the homepage's entire ranked
  list, and the brand-tag/nav-active state were all blank across every
  page. Fixed by renaming every key to match each template's own header
  comment (the authoritative per-page field contract) and, where an
  API struct's exported fields already matched what a template needed
  (contentapi.CityInfo, Crumb, HomeRanked, HubCountry, ...), passing the
  struct straight through instead of hand-rewrapping it in a map --
  removes a whole layer of future drift risk, not just this instance of it.
- Three pages 500'd outright: `.ToolHref` needed a fully-composed href
  string, not the bare "lat,lon" fragment the handlers were building; the
  all-time-records table needed the raw contentapi.AllTimeRecords struct,
  not a re-wrapped map.
- JSON-LD was being double-encoded: `<script type="application/ld+json">`
  is JAVASCRIPT context to html/template's contextual escaper regardless
  of the script's `type` attribute, so a template.HTML-typed value placed
  there gets re-escaped as a quoted JS string instead of emitted raw --
  the entire structured-data payload shipped as a JSON string containing
  JSON, which no crawler would parse as the intended object. Needed
  template.JS instead, the type that actually means "trusted JS source."
  The glossary term page's JSON-LD was simply never built at all (the
  Jinja original assembled it inline in the template rather than through
  content.py's context dict, and that got lost in translation) -- added.
- html/template silently strips literal HTML comments AND JavaScript
  comments from the parsed output (verified in isolation, zero template
  actions involved) -- confirmed as real engine behavior, not a bug in
  either port, so both need a FuncMap function returning template.HTML /
  template.JS respectively to survive parsing rather than a literal
  `<!-- -->` or `//` in the template source.

Packaging: multi-stage Go build, final image alpine (not distroless -- the
Swarm stack's env-entrypoint.sh shim needs bash), 187MB -> 22.6MB. Two
defects caught before they reached a host:
- The Swarm stack overrides `entrypoint:` with no `command:`, which drops
  the image's own CMD entirely (Docker/Swarm semantics, not merged) --
  env-entrypoint.sh then fell through to its hardcoded `exec uvicorn
  app:app` fallback, which doesn't exist in this image. Every deploy
  would have exited 127. Fixed with an explicit `command:` on the stack's
  frontend service, and corrected the shim's stale comment claiming CMD
  passes through automatically.
- `COPY --chown=thermograph` resolves the group by NAME at copy time;
  Alpine's `adduser -S` with no `-G` doesn't create a same-named group, so
  the classic (non-BuildKit) Docker builder -- which this CI runner falls
  back to, since it installs plain `docker.io` with no buildx plugin --
  failed outright. Fixed with an explicit group and numeric --chown.

Verification: go build/vet/test -race clean across all packages; the
Docker image builds and passes its embedded go test step under both
BuildKit and the classic builder; shellcheck 0 findings on the one script
touched; rebased onto current main (the ERA5 lake stack landed on both
main and dev during this work -- confirmed additive, no overlap with
frontend/daemon).
2026-07-23 17:51:31 -07:00

119 lines
4.5 KiB
Go

package content
import (
"net/http"
"strings"
)
// RobotsTxt is the port of content.py's robots_txt — the body is built the
// same way, byte for byte.
func (h *Handlers) RobotsTxt(w http.ResponseWriter, r *http.Request) {
base := h.cfg.Base
baseURL := origin(r) + base
body := "User-agent: *\n" +
"Allow: /\n" +
"Disallow: " + base + "/api/\n" +
"Disallow: " + base + "/alerts\n" +
"Sitemap: " + baseURL + "/sitemap.xml\n"
w.Header().Set("Content-Type", "text/plain; charset=utf-8")
w.Write([]byte(body))
}
// SitemapXML is the port of content.py's sitemap_xml: one <url> line per
// backend sitemap entry, <lastmod> pinned to the per-process boot date.
// ~2.35 MB and crawled repeatedly — the short Cache-Control saves every
// crawler hit within the window a full Sitemap() call plus the string build,
// on top of the response body itself (the backend client's TTL cache absorbs
// the rest).
func (h *Handlers) SitemapXML(w http.ResponseWriter, r *http.Request) {
baseURL := origin(r) + h.cfg.Base
entries, err := h.api.Sitemap()
if err != nil {
h.apiError(w, err)
return
}
var b strings.Builder
b.WriteString(`<?xml version="1.0" encoding="UTF-8"?>` + "\n")
b.WriteString(`<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">`)
for _, e := range entries {
b.WriteString("\n<url><loc>")
b.WriteString(baseURL)
b.WriteString(e.Path)
b.WriteString("</loc><lastmod>")
b.WriteString(h.bootDate)
b.WriteString("</lastmod><changefreq>")
b.WriteString(e.Changefreq)
b.WriteString("</changefreq><priority>")
b.WriteString(e.Priority)
b.WriteString("</priority></url>")
}
b.WriteString("\n</urlset>")
w.Header().Set("Content-Type", "application/xml")
w.Header().Set("Cache-Control", "public, max-age=300")
w.Write([]byte(b.String()))
}
// registerIndexNow wires the IndexNow key file. IndexNow verification works
// by serving a file at /<key>.txt — the key isn't just response *content*,
// it's baked into the route *path*, which the mux needs at registration
// time. That used to mean an eager, unretried key fetch at boot — so if the
// backend was unreachable (down, mid-restart, a network blip) the frontend
// would never finish booting at all. That's exactly the coupling the
// repo-split removed: frontend and backend deploy asynchronously, so
// "backend happens to be briefly unreachable" must be a normal, survivable
// condition at frontend boot, not a crash.
//
// So: try the eager fetch once (preserving the plain static route for the
// overwhelmingly common case where the backend IS reachable at boot). If it
// fails, log a warning and fall back to a route that lazily (re)fetches the
// key on each request via the client's TTL cache — once the backend comes
// back up the correct key starts being served automatically, no frontend
// restart required, and no request ever gets a permanently wrong/empty key.
//
// Mux mechanics differ from Starlette here: Go's patterns cannot express the
// Python's "/{token}.txt" suffix wildcard, so the lazy fallback claims the
// whole single-segment slot ({BASE}/{token}) and forwards anything that
// isn't a .txt request to the static file server it displaced (explicit page
// routes still win on specificity).
func (h *Handlers) registerIndexNow(mux *http.ServeMux, static http.Handler) {
base := h.cfg.Base
key, err := h.api.IndexNowKey()
if err == nil {
mux.HandleFunc("GET "+base+"/"+key+".txt", func(w http.ResponseWriter, r *http.Request) {
writeKey(w, key)
})
return
}
h.log.Printf("indexnow_key fetch failed at boot (backend unreachable?) -- "+
"continuing boot without it; falling back to lazy per-request lookup at "+
"/<key>.txt so the frontend doesn't depend on backend liveness to start up: %v", err)
mux.HandleFunc("GET "+base+"/{token}", func(w http.ResponseWriter, r *http.Request) {
token := r.PathValue("token")
if !strings.HasSuffix(token, ".txt") {
// Not a key-file request: this pattern displaced the static
// subtree for single-segment paths, so hand it back.
if static != nil {
static.ServeHTTP(w, r)
return
}
http.NotFound(w, r)
return
}
key, err := h.api.IndexNowKey()
if err != nil {
writeDetail(w, http.StatusServiceUnavailable, "backend unavailable")
return
}
if strings.TrimSuffix(token, ".txt") != key {
writeDetail(w, http.StatusNotFound, "Not Found")
return
}
writeKey(w, key)
})
}
func writeKey(w http.ResponseWriter, key string) {
w.Header().Set("Content-Type", "text/plain; charset=utf-8")
w.Write([]byte(key + "\n"))
}