thermograph/warm_cities.py

49 lines
1.8 KiB
Python
Raw Normal View History

SEO: crawlable programmatic climate pages + technical hygiene (#96) * SEO: generate curated city set for crawlable climate pages gen_cities.py reuses the GeoNames index places.py already parses to select the top ~500 metros by population, assigns each a stable URL-safe slug (dropping admin1 when it repeats the city name), and writes committed backend/cities.json. cities.py loads it lazily with slug lookup, all_slugs(), display_name(), and by_country() grouping for the upcoming hub + sitemap. * SEO: rendering core, robots.txt, sitemap.xml, and metadata hygiene - content.py: Jinja2 environment + HTML responder (ETag/304), dynamic /robots.txt (disallows /api and /alerts, points at the sitemap) and /sitemap.xml (enumerates the home/static pages plus every city, month, and records URL from cities.py). Registered on the app before the StaticFiles mount so the routes win. - templates/base.html.j2: shared layout with unique title/description, self- referential canonical, Open Graph, favicon/manifest, header nav (adds a Climate link) and a footer link graph. - Give each existing page a unique <meta description> (were 5x identical) and a self-referential <link rel=canonical>; add WebApplication JSON-LD to the home page. - Pin jinja2. * SEO: server-rendered per-city climate page (/climate/{slug}) The keystone crawlable page: for a city it snaps to the grid cell, loads the archive (fetching once if missing, self-healing), and renders as real HTML — a 'how today compares' block (grade + percentile per metric from grade_day, tinted by tier), a monthly normals table (climatology at each month's 15th, shown in °F and °C), all-time records (new grading.all_time_records helper), a breadcrumb, Dataset+Place+BreadcrumbList JSON-LD, self-referential canonical, and links into the interactive tool + month/records pages. Content-page CSS added to style.css (renamed the table class to avoid colliding with the app's .normals flex row). * SEO: month (/climate/{slug}/{month}) and records (/climate/{slug}/records) pages Month pages render the exact-month long-tail ('average weather in {city} in {month}') with that month's average high/low, typical p10-p90 range, month-specific records, and prev/next month links. Records pages show all-time record highs/lows per metric with dates (grading.all_time_records). Shared _resolve_city helper; the literal /records route is registered before the {month} param and month names are validated (unknown month -> 404). * SEO: climate hub, weather glossary, and about/methodology pages - /climate: crawlable directory of all ~500 cities grouped by country — the internal-link graph that lets search engines discover every city page. - /glossary + /glossary/{term}: plain-language definitions (climate normal, percentile, temperature anomaly, feels-like, heat index, wind chill, humidity, reanalysis) with DefinedTerm JSON-LD and cross-links into the tool. - /about: methodology page (ERA5 data source, 45-year baseline, +/-7-day window, percentile grading) for E-E-A-T. All linked from the shared footer. * SEO: archive warmer, content-page tests, and deploy docs - warm_cities.py: paced, idempotent offline warmer that pre-fetches each city cell's archive so /climate pages serve from cache and a crawl can't burst the archive quota (pages self-heal if hit before warming). - tests/test_content.py: city-set slug uniqueness/lookup, robots.txt, sitemap enumerating city/month/records URLs, and that a rendered city page carries the stats + canonical + Dataset JSON-LD in the HTML; plus month/records/hub/glossary/ about routing and 404s. - DEPLOY.md: document the content pages, the warm step, and submitting the sitemap.
2026-07-15 23:53:11 +00:00
"""Pre-warm the archives for the curated city set (backend/cities.json) so the
crawlable /climate pages render from cache and a search-engine crawl never bursts
the archive API quota. Run at/after deploy:
python warm_cities.py [--limit N] [--pace SECONDS]
Idempotent: a cell whose archive is already cached is skipped. Fetches are paced
(default 2s) to stay well under the archive API's rate limit. A cell that still
has no cached archive when its page is first requested self-heals via get_history,
so this is an optimization, not a hard dependency.
"""
import sys
import time
import cities
import climate
import grid
def main(limit: int | None = None, pace: float = 2.0) -> None:
todo = cities.all_cities()
if limit:
todo = todo[:limit]
fetched = skipped = failed = 0
for i, c in enumerate(todo, 1):
cell = grid.snap(c["lat"], c["lon"])
cached = climate.load_cached_history(cell)
if cached is not None and not cached.is_empty():
skipped += 1
continue
try:
climate.get_history(cell) # fetch + cache the ~45-yr archive
climate.get_recent_forecast(cell) # + the recent/forecast bundle (today block)
fetched += 1
print(f"[{i}/{len(todo)}] warmed {c['slug']} ({cell['id']})")
time.sleep(pace)
except Exception as e: # noqa: BLE001 - keep going; the page self-heals later
failed += 1
print(f"[{i}/{len(todo)}] FAILED {c['slug']}: {e}")
time.sleep(pace)
print(f"done: fetched={fetched} skipped(cached)={skipped} failed={failed}")
if __name__ == "__main__":
args = sys.argv[1:]
lim = int(args[args.index("--limit") + 1]) if "--limit" in args else None
pc = float(args[args.index("--pace") + 1]) if "--pace" in args else 2.0
main(limit=lim, pace=pc)