Pre-warm the SEO content derived-store off-request #49

Merged
admin_emi merged 2 commits from fix/content-prewarm into dev 2026-07-24 19:31:10 +00:00
Owner

Fix #3 of 3 for the content-page latency problem. Stacks on the content-token work (#47).

Problem

The /climate/<city>[/month|/records] pages render from the content derived-store, keyed by content_token = PAYLOAD_VER:CONTENT_VER:archive-last-date. That token turns over ~daily as a cell's archive gains a day, so the first request after each advance recomputes a 45-year payload cold — the latency users and crawlers hit.

Change

  • warm_cities.warm_content(limit=None, origin=None, pace=0.05) — for each curated city, compute the cheap content token (no history load) and check whether every content kind (city, 12 months, records) already has a fresh row. Cities wholly fresh are skipped without loading the 45-year archive, so a re-run is near-instant and only cells whose archive genuinely advanced do work. Cache-only exactly like main(): a cell with no cached archive is skipped, never fetched, so warming spends no upstream quota. Rows are written under the exact (kind, key, token) content_routes.py reads — content-city/content-records on {slug}:{origin}, content-month on {slug}:{month} — with origin defaulting to the canonical prod origin. A per-call limit bounds one pass's wall-clock; the idempotent skip resumes on the next call.
  • Scheduler wiring — ride the existing leader-gated notifier loop (notify.run_loop, started only on the elected leader) next to the homepage sweep. _maybe_warm_content runs at most every 30 min, capped at 50 cities/tick, so a tick can't stall the notifier and the ~1000-city set refreshes across a handful of ticks — well inside the <=1-day staleness the token already tolerates. Chose the notifier loop over ops-cron because it is the in-app leader-gated background worker here (there is no APScheduler module in this tree).

Tests

backend/tests/test_warm_content.py: populates all three kinds under the right (kind, key, token); skips a city with no cached history and never triggers an upstream fetch; idempotent — a second run writes zero payloads; limit caps cities built per call. Full suite: 390 passed, 8 skipped.

Fix #3 of 3 for the content-page latency problem. Stacks on the content-token work (#47). ## Problem The `/climate/<city>[/month|/records]` pages render from the content derived-store, keyed by `content_token` = `PAYLOAD_VER:CONTENT_VER:archive-last-date`. That token turns over ~daily as a cell's archive gains a day, so the first request after each advance recomputes a 45-year payload cold — the latency users and crawlers hit. ## Change - **`warm_cities.warm_content(limit=None, origin=None, pace=0.05)`** — for each curated city, compute the cheap content token (no history load) and check whether every content kind (city, 12 months, records) already has a fresh row. Cities wholly fresh are skipped *without* loading the 45-year archive, so a re-run is near-instant and only cells whose archive genuinely advanced do work. Cache-only exactly like `main()`: a cell with no cached archive is skipped, never fetched, so warming spends no upstream quota. Rows are written under the exact `(kind, key, token)` `content_routes.py` reads — `content-city`/`content-records` on `{slug}:{origin}`, `content-month` on `{slug}:{month}` — with `origin` defaulting to the canonical prod origin. A per-call `limit` bounds one pass's wall-clock; the idempotent skip resumes on the next call. - **Scheduler wiring** — ride the existing leader-gated notifier loop (`notify.run_loop`, started only on the elected leader) next to the homepage sweep. `_maybe_warm_content` runs at most every 30 min, capped at 50 cities/tick, so a tick can't stall the notifier and the ~1000-city set refreshes across a handful of ticks — well inside the <=1-day staleness the token already tolerates. Chose the notifier loop over ops-cron because it is the in-app leader-gated background worker here (there is no APScheduler module in this tree). ## Tests `backend/tests/test_warm_content.py`: populates all three kinds under the right `(kind, key, token)`; skips a city with no cached history and never triggers an upstream fetch; idempotent — a second run writes zero payloads; `limit` caps cities built per call. Full suite: 390 passed, 8 skipped.
admin_emi added 2 commits 2026-07-24 19:27:16 +00:00
Add cheap, stable content-page cache token
All checks were successful
PR build (required check) / changes (pull_request) Successful in 8s
secrets-guard / encrypted (pull_request) Successful in 6s
PR build (required check) / build-frontend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 10s
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-backend (pull_request) Successful in 1m2s
PR build (required check) / gate (pull_request) Successful in 3s
85810c21c5
Content SEO pages keyed derived-payload cache validity on
history_token = PAYLOAD_VER:hist_end, computed from the fully-loaded
~45-year archive. Every hourly tail top-up advanced hist_end and
invalidated the whole content cache for a cell, and computing the token
at all required loading the full history first — so even cache hits paid
the full load.

Add content_token(cell_id) = PAYLOAD_VER:CONTENT_VER:max_date, keyed on
the archive's newest DATE read cheaply without loading history:
climate_store.history_max_date does an indexed MAX(date) over the
(cell_id, date) PK on Postgres; climate.history_max_date dispatches to it
or to a single-column scan of the cached parquet on the dev backend. The
token survives intra-day top-ups and turns over only when the last
archived day advances (~1x/day), keeping content pages <=1 day stale.
Fail-soft: a store/DB error buckets to 'none' rather than raising.

CONTENT_VER ("c1") is a separate content-shape version so a content-only
change need not orphan every other kind's cache.
Pre-warm the SEO content derived-store off-request
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-backend (pull_request) Successful in 49s
PR build (required check) / gate (pull_request) Successful in 2s
0841390aaf
The /climate/<city>[/month|/records] pages render from the content
derived-store keyed by content_token (PAYLOAD_VER:CONTENT_VER:archive
last-date), which turns over ~daily when a cell's archive gains a day.
On the first request after each advance the page recomputed a 45-year
payload cold.

Add warm_cities.warm_content: for each curated city, compute the cheap
content token and check whether every content kind (city, 12 months,
records) already has a fresh row; cities wholly fresh are skipped without
loading the archive, so a re-run is near-instant and only cells whose
archive advanced do work. Cache-only like main() — a cell with no cached
archive is skipped, never fetched, so warming spends no upstream quota.
A per-call limit caps cities (re)built so one pass's wall-clock is
bounded; the idempotent skip resumes on the next call.

Ride it on the leader-gated notifier loop next to the homepage sweep:
_maybe_warm_content runs at most every 30 min, capped at 50 cities per
tick, so a tick can't stall the notifier and the ~1000-city set refreshes
across a handful of ticks — well inside the <=1-day staleness the token
already tolerates.
admin_emi merged commit bf0aaf1ec0 into dev 2026-07-24 19:31:10 +00:00
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference: Jinemi/thermograph#49
No description provided.