thermograph/gen_cities.py

83 lines
2.9 KiB
Python
Raw Normal View History

SEO: crawlable programmatic climate pages + technical hygiene (#96) * SEO: generate curated city set for crawlable climate pages gen_cities.py reuses the GeoNames index places.py already parses to select the top ~500 metros by population, assigns each a stable URL-safe slug (dropping admin1 when it repeats the city name), and writes committed backend/cities.json. cities.py loads it lazily with slug lookup, all_slugs(), display_name(), and by_country() grouping for the upcoming hub + sitemap. * SEO: rendering core, robots.txt, sitemap.xml, and metadata hygiene - content.py: Jinja2 environment + HTML responder (ETag/304), dynamic /robots.txt (disallows /api and /alerts, points at the sitemap) and /sitemap.xml (enumerates the home/static pages plus every city, month, and records URL from cities.py). Registered on the app before the StaticFiles mount so the routes win. - templates/base.html.j2: shared layout with unique title/description, self- referential canonical, Open Graph, favicon/manifest, header nav (adds a Climate link) and a footer link graph. - Give each existing page a unique <meta description> (were 5x identical) and a self-referential <link rel=canonical>; add WebApplication JSON-LD to the home page. - Pin jinja2. * SEO: server-rendered per-city climate page (/climate/{slug}) The keystone crawlable page: for a city it snaps to the grid cell, loads the archive (fetching once if missing, self-healing), and renders as real HTML — a 'how today compares' block (grade + percentile per metric from grade_day, tinted by tier), a monthly normals table (climatology at each month's 15th, shown in °F and °C), all-time records (new grading.all_time_records helper), a breadcrumb, Dataset+Place+BreadcrumbList JSON-LD, self-referential canonical, and links into the interactive tool + month/records pages. Content-page CSS added to style.css (renamed the table class to avoid colliding with the app's .normals flex row). * SEO: month (/climate/{slug}/{month}) and records (/climate/{slug}/records) pages Month pages render the exact-month long-tail ('average weather in {city} in {month}') with that month's average high/low, typical p10-p90 range, month-specific records, and prev/next month links. Records pages show all-time record highs/lows per metric with dates (grading.all_time_records). Shared _resolve_city helper; the literal /records route is registered before the {month} param and month names are validated (unknown month -> 404). * SEO: climate hub, weather glossary, and about/methodology pages - /climate: crawlable directory of all ~500 cities grouped by country — the internal-link graph that lets search engines discover every city page. - /glossary + /glossary/{term}: plain-language definitions (climate normal, percentile, temperature anomaly, feels-like, heat index, wind chill, humidity, reanalysis) with DefinedTerm JSON-LD and cross-links into the tool. - /about: methodology page (ERA5 data source, 45-year baseline, +/-7-day window, percentile grading) for E-E-A-T. All linked from the shared footer. * SEO: archive warmer, content-page tests, and deploy docs - warm_cities.py: paced, idempotent offline warmer that pre-fetches each city cell's archive so /climate pages serve from cache and a crawl can't burst the archive quota (pages self-heal if hit before warming). - tests/test_content.py: city-set slug uniqueness/lookup, robots.txt, sitemap enumerating city/month/records URLs, and that a rendered city page carries the stats + canonical + Dataset JSON-LD in the HTML; plus month/records/hub/glossary/ about routing and 404s. - DEPLOY.md: document the content pages, the warm step, and submitting the sitemap.
2026-07-15 23:53:11 +00:00
"""Offline generator for backend/cities.json — the finite set of cities that get
crawlable climate pages (/climate/<slug>). Run occasionally to refresh the list:
python gen_cities.py [N] # default N=500 top metros by population
It reuses the GeoNames index that places.py already downloads/parses (calling
places._load() synchronously fills places._data), takes the top-N places by
population, and assigns each a stable, unique, URL-safe slug. Committing the output
keeps the routable city set explicit and reviewable, and decouples page-serving
from the async place-name loader.
"""
import json
import os
import re
import sys
import unicodedata
import places
OUT_PATH = os.path.join(os.path.dirname(__file__), "cities.json")
# GeoNames entry tuple layout (see places._load): the fields we keep.
_NAME, _ADMIN1, _COUNTRY, _CC, _LAT, _LON, _POP = 1, 2, 3, 4, 5, 6, 7
def slugify(*parts: str) -> str:
"""ASCII, lowercase, hyphenated slug from name/admin/country parts."""
text = " ".join(p for p in parts if p)
text = unicodedata.normalize("NFKD", text).encode("ascii", "ignore").decode()
text = re.sub(r"[^a-zA-Z0-9]+", "-", text).strip("-").lower()
return re.sub(r"-{2,}", "-", text)
def build(n: int = 500) -> list[dict]:
places._load() # synchronous parse; fills places._data (entries are pop-desc)
if not places._data:
raise SystemExit("GeoNames index failed to load (see logs); cannot generate cities.")
entries = places._data[0]
out: list[dict] = []
seen_slugs: set[str] = set()
for e in entries:
if len(out) >= n:
break
name, admin1, country, cc = e[_NAME], e[_ADMIN1], e[_COUNTRY], e[_CC]
# Drop admin1 from the slug when it just repeats the city name
# (e.g. Tokyo/Tokyo, Singapore/Singapore) to avoid "tokyo-tokyo-jp".
admin_part = admin1 if admin1 and slugify(admin1) != slugify(name) else ""
base = slugify(name, admin_part, cc or "")
if not base:
continue
slug = base
i = 2
while slug in seen_slugs: # disambiguate the rare collision
slug = f"{base}-{i}"
i += 1
seen_slugs.add(slug)
out.append({
"slug": slug,
"name": name,
"admin1": admin1,
"country": country,
"country_code": cc,
"lat": round(e[_LAT], 5),
"lon": round(e[_LON], 5),
"population": e[_POP],
})
return out
def main() -> None:
n = int(sys.argv[1]) if len(sys.argv) > 1 else 500
cities = build(n)
with open(OUT_PATH, "w", encoding="utf-8") as f:
json.dump(cities, f, ensure_ascii=False, indent=0, separators=(",", ":"))
f.write("\n")
print(f"wrote {len(cities)} cities -> {OUT_PATH}")
print("sample:", ", ".join(c["slug"] for c in cities[:8]))
if __name__ == "__main__":
main()