Pre-warm the SEO content derived-store off-request #49
No reviewers
Labels
No labels
Compat/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference: Jinemi/thermograph#49
Loading…
Reference in a new issue
No description provided.
Delete branch "fix/content-prewarm"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Fix #3 of 3 for the content-page latency problem. Stacks on the content-token work (#47).
Problem
The
/climate/<city>[/month|/records]pages render from the content derived-store, keyed bycontent_token=PAYLOAD_VER:CONTENT_VER:archive-last-date. That token turns over ~daily as a cell's archive gains a day, so the first request after each advance recomputes a 45-year payload cold — the latency users and crawlers hit.Change
warm_cities.warm_content(limit=None, origin=None, pace=0.05)— for each curated city, compute the cheap content token (no history load) and check whether every content kind (city, 12 months, records) already has a fresh row. Cities wholly fresh are skipped without loading the 45-year archive, so a re-run is near-instant and only cells whose archive genuinely advanced do work. Cache-only exactly likemain(): a cell with no cached archive is skipped, never fetched, so warming spends no upstream quota. Rows are written under the exact(kind, key, token)content_routes.pyreads —content-city/content-recordson{slug}:{origin},content-monthon{slug}:{month}— withorigindefaulting to the canonical prod origin. A per-calllimitbounds one pass's wall-clock; the idempotent skip resumes on the next call.notify.run_loop, started only on the elected leader) next to the homepage sweep._maybe_warm_contentruns at most every 30 min, capped at 50 cities/tick, so a tick can't stall the notifier and the ~1000-city set refreshes across a handful of ticks — well inside the <=1-day staleness the token already tolerates. Chose the notifier loop over ops-cron because it is the in-app leader-gated background worker here (there is no APScheduler module in this tree).Tests
backend/tests/test_warm_content.py: populates all three kinds under the right(kind, key, token); skips a city with no cached history and never triggers an upstream fetch; idempotent — a second run writes zero payloads;limitcaps cities built per call. Full suite: 390 passed, 8 skipped.Content SEO pages keyed derived-payload cache validity on history_token = PAYLOAD_VER:hist_end, computed from the fully-loaded ~45-year archive. Every hourly tail top-up advanced hist_end and invalidated the whole content cache for a cell, and computing the token at all required loading the full history first — so even cache hits paid the full load. Add content_token(cell_id) = PAYLOAD_VER:CONTENT_VER:max_date, keyed on the archive's newest DATE read cheaply without loading history: climate_store.history_max_date does an indexed MAX(date) over the (cell_id, date) PK on Postgres; climate.history_max_date dispatches to it or to a single-column scan of the cached parquet on the dev backend. The token survives intra-day top-ups and turns over only when the last archived day advances (~1x/day), keeping content pages <=1 day stale. Fail-soft: a store/DB error buckets to 'none' rather than raising. CONTENT_VER ("c1") is a separate content-shape version so a content-only change need not orphan every other kind's cache.