* SEO: generate curated city set for crawlable climate pages
gen_cities.py reuses the GeoNames index places.py already parses to select the top
~500 metros by population, assigns each a stable URL-safe slug (dropping admin1 when
it repeats the city name), and writes committed backend/cities.json. cities.py loads
it lazily with slug lookup, all_slugs(), display_name(), and by_country() grouping
for the upcoming hub + sitemap.
* SEO: rendering core, robots.txt, sitemap.xml, and metadata hygiene
- content.py: Jinja2 environment + HTML responder (ETag/304), dynamic /robots.txt
(disallows /api and /alerts, points at the sitemap) and /sitemap.xml (enumerates
the home/static pages plus every city, month, and records URL from cities.py).
Registered on the app before the StaticFiles mount so the routes win.
- templates/base.html.j2: shared layout with unique title/description, self-
referential canonical, Open Graph, favicon/manifest, header nav (adds a Climate
link) and a footer link graph.
- Give each existing page a unique <meta description> (were 5x identical) and a
self-referential <link rel=canonical>; add WebApplication JSON-LD to the home page.
- Pin jinja2.
* SEO: server-rendered per-city climate page (/climate/{slug})
The keystone crawlable page: for a city it snaps to the grid cell, loads the
archive (fetching once if missing, self-healing), and renders as real HTML — a
'how today compares' block (grade + percentile per metric from grade_day, tinted
by tier), a monthly normals table (climatology at each month's 15th, shown in °F
and °C), all-time records (new grading.all_time_records helper), a breadcrumb,
Dataset+Place+BreadcrumbList JSON-LD, self-referential canonical, and links into
the interactive tool + month/records pages. Content-page CSS added to style.css
(renamed the table class to avoid colliding with the app's .normals flex row).
* SEO: month (/climate/{slug}/{month}) and records (/climate/{slug}/records) pages
Month pages render the exact-month long-tail ('average weather in {city} in
{month}') with that month's average high/low, typical p10-p90 range, month-specific
records, and prev/next month links. Records pages show all-time record highs/lows
per metric with dates (grading.all_time_records). Shared _resolve_city helper; the
literal /records route is registered before the {month} param and month names are
validated (unknown month -> 404).
* SEO: climate hub, weather glossary, and about/methodology pages
- /climate: crawlable directory of all ~500 cities grouped by country — the
internal-link graph that lets search engines discover every city page.
- /glossary + /glossary/{term}: plain-language definitions (climate normal,
percentile, temperature anomaly, feels-like, heat index, wind chill, humidity,
reanalysis) with DefinedTerm JSON-LD and cross-links into the tool.
- /about: methodology page (ERA5 data source, 45-year baseline, +/-7-day window,
percentile grading) for E-E-A-T. All linked from the shared footer.
* SEO: archive warmer, content-page tests, and deploy docs
- warm_cities.py: paced, idempotent offline warmer that pre-fetches each city
cell's archive so /climate pages serve from cache and a crawl can't burst the
archive quota (pages self-heal if hit before warming).
- tests/test_content.py: city-set slug uniqueness/lookup, robots.txt, sitemap
enumerating city/month/records URLs, and that a rendered city page carries the
stats + canonical + Dataset JSON-LD in the HTML; plus month/records/hub/glossary/
about routing and 404s.
- DEPLOY.md: document the content pages, the warm step, and submitting the sitemap.
74 lines
2.6 KiB
Python
74 lines
2.6 KiB
Python
"""Tests for the crawlable SEO content pages: the city set, robots/sitemap, and
|
|
that the server-rendered pages contain the stats in the HTML (with the weather
|
|
layer faked, like test_api.py)."""
|
|
import pytest
|
|
from fastapi.testclient import TestClient
|
|
|
|
import app as appmod
|
|
import cities
|
|
import climate
|
|
|
|
# BASE defaults to /thermograph in tests (THERMOGRAPH_BASE unset).
|
|
B = "/thermograph"
|
|
SLUG = "london-england-gb" # present in cities.json
|
|
|
|
|
|
@pytest.fixture
|
|
def client(monkeypatch, history, recent):
|
|
monkeypatch.setattr(climate, "load_cached_history", lambda cell: history.clone())
|
|
monkeypatch.setattr(climate, "get_recent_forecast", lambda cell: recent.clone())
|
|
monkeypatch.setattr(climate, "get_history", lambda cell: (history.clone(), {"cached": True}))
|
|
return TestClient(appmod.app)
|
|
|
|
|
|
def test_city_set_slugs_unique_and_lookup():
|
|
slugs = cities.all_slugs()
|
|
assert len(slugs) == len(set(slugs)) >= 100
|
|
assert cities.get(SLUG) is not None
|
|
assert cities.get("does-not-exist") is None
|
|
|
|
|
|
def test_robots_txt(client):
|
|
r = client.get(f"{B}/robots.txt")
|
|
assert r.status_code == 200
|
|
assert "Sitemap:" in r.text and "/sitemap.xml" in r.text
|
|
assert "Disallow: /thermograph/api/" in r.text
|
|
|
|
|
|
def test_sitemap_lists_city_urls(client):
|
|
r = client.get(f"{B}/sitemap.xml")
|
|
assert r.status_code == 200
|
|
assert "<urlset" in r.text
|
|
assert f"/climate/{SLUG}</loc>" in r.text
|
|
assert f"/climate/{SLUG}/july</loc>" in r.text
|
|
assert f"/climate/{SLUG}/records</loc>" in r.text
|
|
|
|
|
|
def test_city_page_renders_stats_in_html(client):
|
|
r = client.get(f"{B}/climate/{SLUG}")
|
|
assert r.status_code == 200
|
|
b = r.text
|
|
assert "London" in b and "climate" in b.lower()
|
|
assert 'rel="canonical"' in b and f"/climate/{SLUG}" in b
|
|
assert '"@type":"Dataset"' in b
|
|
assert "average temperatures by month" in b.lower()
|
|
# a real number rendered (°F appears in the normals/today text)
|
|
assert "°F" in b
|
|
|
|
|
|
def test_city_404(client):
|
|
assert client.get(f"{B}/climate/nope-not-a-city").status_code == 404
|
|
|
|
|
|
def test_month_and_records(client):
|
|
assert client.get(f"{B}/climate/{SLUG}/july").status_code == 200
|
|
assert client.get(f"{B}/climate/{SLUG}/records").status_code == 200
|
|
assert client.get(f"{B}/climate/{SLUG}/notamonth").status_code == 404
|
|
|
|
|
|
def test_hub_glossary_about(client):
|
|
assert client.get(f"{B}/climate").status_code == 200
|
|
assert client.get(f"{B}/glossary").status_code == 200
|
|
assert client.get(f"{B}/glossary/percentile").status_code == 200
|
|
assert client.get(f"{B}/glossary/not-a-term").status_code == 404
|
|
assert client.get(f"{B}/about").status_code == 200
|