Persistent derived-data cache: SQLite store, ETag revalidation, view bundle, IndexedDB frontend (#21)
* Compress API responses and revalidate static assets instead of re-downloading - Add GZipMiddleware (min 1 KB): the 2-year calendar JSON shrinks ~6-8x. - Serve pages/assets with Cache-Control: no-cache instead of no-store, so browsers revalidate via the ETag/Last-Modified that FileResponse and StaticFiles already emit. Unchanged assets now cost an empty 304 rather than a full transfer on every page navigation, while deploys still show up immediately. * Persist derived responses in SQLite so grading is computed once per cell, not per request New data/thermograph.sqlite (WAL) holds what's derived from the raw parquet records — finished grade/calendar/day/forecast payloads and reverse-geocode labels — so the expensive work (notably the 2-year calendar grade_range, ~270ms) becomes a ~5ms database read that survives restarts and is shared across views. Raw parquet stays the source of truth; the store is a pure accelerator (every reader falls back to recomputing on a miss, and deleting the db is a safe reset). Freshness is token-driven, not clock-driven: each cached payload is validated by a token encoding what it was computed from (payload schema version, the archive record's end date, the recent-fetch stamp). The existing freshness drivers are untouched — get_history still tops up the tail hourly and get_recent_forecast still refetches hourly — and tokens are derived from what they return, so cached payloads expire exactly when their inputs change. The same tokens double as weak ETags: If-None-Match answers with an empty 304 without touching the payload. - backend/store.py: derived-payload + revgeo tables, thread-local WAL conns, every helper fail-soft. - app.py: endpoints split into pure payload builders + HTTP/caching shells; the in-memory 10-minute _CAL_CACHE is retired (superseded by the persistent store). - climate.py: revgeo persisted through the store; recent_stamp() and load_cached_history() (no-network read) helpers. - backend/migrate.py + make migrate: idempotent, resumable backfill of the store from existing parquet caches (default calendar span + latest-day detail + revgeo, ≤1 throttled Nominatim call per unlabeled cell). Never fetches weather. - grid.from_id(): rebuild a cell from its cache filename (migrate tooling). * Add /api/v2/cell: one bundle carrying every view's payload GET /api/v2/cell?lat&lon returns the grade, forecast, calendar (last 24 months) and day (today) payloads in one response, each the exact payload its per-view endpoint returns — built by the same builders and cached under the same derived-store keys/tokens — paired with the etag that endpoint would emit. The frontend can warm all views with a single request, seed its per-view cache from the slices, and later revalidate each view individually with If-None-Match. The bundle's own etag combines the slices', so an unchanged bundle is an empty 304. prefetch=1 is a warm-only mode for neighbor-cell prefetching with a hard guarantee: it never spends weather-API quota. A cell with no cached archive answers 204, and only the history-derived slices (calendar + latest-day detail) are built. At most one Nominatim lookup for a never-labeled cell. * Frontend: IndexedDB response cache with stale-while-revalidate + bundle prefetch The per-URL response cache moves from localStorage/sessionStorage (~5 MB quota, which multi-year calendar payloads regularly blew through) to IndexedDB, with an in-memory map in front. Entries carry the server's ETag, so anything stale revalidates conditionally — unchanged data costs an empty 304 and a re-stamp, never a re-transfer. On network failure the stale copy is served over an error. getJSON gains an optional onUpdate callback opting into stale-while-revalidate: the three views (weekly, day, calendar) now render a cached copy immediately — spinners are delayed 150ms so warm loads never flash-blank — and repaint only if background revalidation finds changed data. New data shows up the moment it exists instead of waiting out a TTL. Cross-view prefetch collapses from one request per view to a single /api/v2/cell bundle, whose slices (exact per-view payloads + their etags) are seeded under the URLs each view actually requests; the bundle call itself is conditional via a remembered etag. Afterwards the 8 surrounding grid cells are warmed server-side with prefetch=1 (never spends weather-API quota; staggered ≥1.1s for the one possible Nominatim lookup each), so tapping nearby lands on already-graded data. Legacy tg:* storage entries are cleared once; cache entries untouched for two weeks are pruned on page load.
This commit is contained in:
parent
28e6b5301f
commit
e01d2e0eb8
5 changed files with 716 additions and 159 deletions
555
app.py
555
app.py
|
|
@ -1,10 +1,14 @@
|
|||
"""Thermograph API — grade recent local weather against ~45 years of climatology."""
|
||||
import contextlib
|
||||
import datetime
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import time
|
||||
|
||||
import pandas as pd
|
||||
from fastapi import APIRouter, FastAPI, HTTPException, Query
|
||||
from fastapi import APIRouter, FastAPI, HTTPException, Query, Request, Response
|
||||
from fastapi.middleware.gzip import GZipMiddleware
|
||||
from fastapi.responses import FileResponse, RedirectResponse
|
||||
from fastapi.staticfiles import StaticFiles
|
||||
|
||||
|
|
@ -12,6 +16,7 @@ import audit
|
|||
import climate
|
||||
import grading
|
||||
import grid
|
||||
import store
|
||||
|
||||
FRONTEND_DIR = os.path.join(os.path.dirname(__file__), "..", "frontend")
|
||||
|
||||
|
|
@ -26,6 +31,12 @@ BASE = "/" + os.environ.get("THERMOGRAPH_BASE", "/thermograph").strip("/")
|
|||
# be absent on an older cache, so only carry the ones present.
|
||||
OBS_COLS = ("tmax", "tmin", "precip", "feels", "humid", "wind", "gust")
|
||||
|
||||
# Bump when any response payload shape changes (new metrics, renamed fields…).
|
||||
# The version is part of every derived-store validity token, so one bump
|
||||
# atomically orphans all pre-upgrade cached payloads instead of letting a stale
|
||||
# row whose history_end happens to match keep serving the old shape.
|
||||
PAYLOAD_VER = "p1"
|
||||
|
||||
|
||||
def _obs_from_row(row) -> dict:
|
||||
return {k: row[k] for k in OBS_COLS if k in row}
|
||||
|
|
@ -73,20 +84,222 @@ def _attach_dry_streaks(graded: list[dict], *precip_frames) -> None:
|
|||
|
||||
app = FastAPI(title="Thermograph", version="0.2.0")
|
||||
|
||||
# Compress every sizeable response (the 2-year calendar JSON shrinks ~6-8×).
|
||||
# Applies to API JSON and static assets alike; tiny responses are left alone.
|
||||
app.add_middleware(GZipMiddleware, minimum_size=1024)
|
||||
|
||||
|
||||
@app.middleware("http")
|
||||
async def no_store_static(request, call_next):
|
||||
"""Serve the frontend with no-store so phones/browsers always pick up the
|
||||
latest HTML/JS/CSS instead of a stale cached copy. This is a LAN dev tool, so
|
||||
skipping the cache entirely is fine and avoids "my change isn't showing" bugs."""
|
||||
async def revalidate_static(request, call_next):
|
||||
"""Serve the frontend with no-cache (NOT no-store): browsers may keep a copy
|
||||
but must revalidate it on every use. Starlette's FileResponse/StaticFiles
|
||||
already emit ETag + Last-Modified, so an unchanged asset costs one conditional
|
||||
request answered with an empty 304 instead of a full re-download — page-to-page
|
||||
navigation stops re-transferring the JS/CSS/HTML while still picking up every
|
||||
deploy immediately (no "my change isn't showing" bugs)."""
|
||||
response = await call_next(request)
|
||||
path = request.url.path
|
||||
pages = (BASE, f"{BASE}/", f"{BASE}/calendar", f"{BASE}/day", f"{BASE}/compare", f"{BASE}/legend")
|
||||
if path.endswith((".js", ".css", ".html")) or path in pages:
|
||||
response.headers["Cache-Control"] = "no-store, must-revalidate"
|
||||
response.headers["Cache-Control"] = "no-cache"
|
||||
return response
|
||||
|
||||
|
||||
# --- derived-store plumbing --------------------------------------------------
|
||||
# Every graded payload is cached in SQLite under (kind, cell, key) and validated
|
||||
# by a token that encodes what it was computed from (see store.py). The freshness
|
||||
# drivers stay where they were — climate.get_history tops up the archive tail
|
||||
# hourly and get_recent_forecast refetches hourly — and the tokens are derived
|
||||
# from what those return, so a cached payload expires exactly when its inputs
|
||||
# change and never before. The same tokens double as ETags: a client sending
|
||||
# If-None-Match gets an empty 304 without the payload even being loaded.
|
||||
|
||||
class _NullRun:
|
||||
"""audit.RunAudit stand-in for offline callers (the migrate script)."""
|
||||
|
||||
def set(self, **kw):
|
||||
return self
|
||||
|
||||
def phase(self, name):
|
||||
return contextlib.nullcontext()
|
||||
|
||||
|
||||
def _hist_end(history) -> str:
|
||||
"""Last day in the cell's archive record — the freshness token for everything
|
||||
derived purely from history. Advances via the hourly tail top-up inside
|
||||
climate.get_history, which cached payloads follow automatically."""
|
||||
return pd.Timestamp(history["date"].max()).date().isoformat()
|
||||
|
||||
|
||||
def _etag_for(kind: str, cell_id: str, key: str, token: str) -> str:
|
||||
"""Deterministic weak ETag from a payload's identity + validity token. Weak
|
||||
because the same content may be served under different encodings (gzip)."""
|
||||
h = hashlib.sha1(f"{kind}:{cell_id}:{key}:{token}".encode()).hexdigest()[:20]
|
||||
return f'W/"{h}"'
|
||||
|
||||
|
||||
def _not_modified(request: Request, etag: str) -> bool:
|
||||
"""Does the client already hold this exact payload? Answerable from the token
|
||||
alone — no payload load needed."""
|
||||
inm = request.headers.get("if-none-match")
|
||||
if not inm:
|
||||
return False
|
||||
tags = {t.strip() for t in inm.split(",")}
|
||||
return "*" in tags or etag in tags or etag.removeprefix("W/") in tags
|
||||
|
||||
|
||||
def _encode(payload: dict) -> bytes:
|
||||
"""Compact JSON bytes, matching store.put_payload's encoding (allow_nan=False
|
||||
mirrors Starlette — a NaN from grading should fail loudly, not reach a client)."""
|
||||
return json.dumps(payload, default=str, allow_nan=False, separators=(",", ":")).encode()
|
||||
|
||||
|
||||
def _json_response(body: bytes, etag: str | None = None) -> Response:
|
||||
headers = {"ETag": etag} if etag else {}
|
||||
return Response(content=body, media_type="application/json", headers=headers)
|
||||
|
||||
|
||||
# --- payload builders --------------------------------------------------------
|
||||
# Pure "inputs → response dict" functions shared by the per-view endpoints, the
|
||||
# /cell bundle, and the offline migrate script. HTTP semantics (audit runs, cache
|
||||
# lookups, etags) stay in the endpoints; `run` is the audit run or None.
|
||||
|
||||
def _build_grade(cell, target, days, history, recent, cache_meta, place, run=None) -> dict:
|
||||
"""/grade payload: the last `days` observed days graded against their own
|
||||
±7-day windows, plus the target day's climatology summary."""
|
||||
run = run or _NullRun()
|
||||
with run.phase("grading"):
|
||||
# Only grade observed days up to and including the target date (last `days`).
|
||||
recent = recent[recent["date"] <= target].sort_values("date").tail(days)
|
||||
graded = _grade_rows(history, recent)
|
||||
_attach_dry_streaks(graded, history, recent)
|
||||
graded.reverse() # newest first for display
|
||||
climo = grading.climatology(history, int(target.dayofyear))
|
||||
|
||||
# Summarize what this run actually covered. A fresh history fetch pulls the
|
||||
# full ~45-year archive (run_type "full"); a cache hit only fetched the
|
||||
# recent window (run_type "partial").
|
||||
years = pd.to_datetime(history["date"]).dt.year
|
||||
full = not cache_meta.get("cached", False)
|
||||
run.set(
|
||||
run_type="full" if full else "partial",
|
||||
history_source="fetch" if full else "cache",
|
||||
history_rows=int(len(history)),
|
||||
history_years=int(years.max() - years.min() + 1),
|
||||
history_span=[int(years.min()), int(years.max())],
|
||||
cache_age_days=cache_meta.get("cache_age_days"),
|
||||
graded_days=len(graded),
|
||||
place_found=place is not None,
|
||||
)
|
||||
return {
|
||||
"cell": cell,
|
||||
"place": place,
|
||||
"target_date": target.date().isoformat(),
|
||||
"cache": cache_meta,
|
||||
"climatology": climo,
|
||||
"recent": graded,
|
||||
}
|
||||
|
||||
|
||||
CAL_MAX_SPAN_DAYS = 732 # ~2 years — the most a single calendar request will grade
|
||||
# (the frontend splits longer spans into 2-year chunks)
|
||||
|
||||
|
||||
def _cal_span(history, start, end, months) -> tuple[pd.Timestamp, pd.Timestamp]:
|
||||
"""Clamp a requested calendar range to the available record (≤ ~2 years).
|
||||
|
||||
Without a start, defaults to whole months back landing on the SAME
|
||||
day-of-month as the end, so the two range endpoints line up
|
||||
(e.g. 2024-06-29 → 2026-06-29)."""
|
||||
first = pd.Timestamp(history["date"].min()).normalize()
|
||||
last = pd.Timestamp(history["date"].max()).normalize()
|
||||
end_ts = min(pd.Timestamp(end).normalize(), last) if end else last
|
||||
if start:
|
||||
start_ts = pd.Timestamp(start).normalize()
|
||||
else:
|
||||
start_ts = (end_ts - pd.DateOffset(months=months)).normalize()
|
||||
# Cap the graded span at ~2 years, and keep it inside the available record.
|
||||
min_start = end_ts - pd.Timedelta(days=CAL_MAX_SPAN_DAYS - 1)
|
||||
start_ts = max(start_ts, min_start, first)
|
||||
start_ts = min(start_ts, end_ts)
|
||||
return start_ts, end_ts
|
||||
|
||||
|
||||
def _build_calendar(cell, history, start_ts, end_ts, months, place, run=None) -> dict:
|
||||
"""/calendar payload: every day in [start_ts, end_ts] graded, compact shape."""
|
||||
run = run or _NullRun()
|
||||
with run.phase("grading"):
|
||||
days = grading.grade_range(history, start_ts, end_ts)
|
||||
run.set(graded_days=len(days))
|
||||
return {
|
||||
"api_version": "v2",
|
||||
"cell": cell,
|
||||
"place": place,
|
||||
"range": {"start": start_ts.date().isoformat(), "end": end_ts.date().isoformat()},
|
||||
"months": months,
|
||||
"days": days,
|
||||
}
|
||||
|
||||
|
||||
def _build_day(cell, history, target, place, run=None, recent=None) -> dict:
|
||||
"""/day payload: the full percentile breakdown for one day. Observed values
|
||||
prefer the archive record; a date newer than it comes from the recent window
|
||||
(passed pre-loaded by the bundle path, fetched here otherwise)."""
|
||||
run = run or _NullRun()
|
||||
last = pd.Timestamp(history["date"].max()).normalize()
|
||||
|
||||
obs = None
|
||||
hit = history[history["date"] == target]
|
||||
if not hit.empty:
|
||||
obs = _obs_from_row(hit.iloc[0])
|
||||
elif target > last:
|
||||
try:
|
||||
if recent is None:
|
||||
with run.phase("recent"):
|
||||
recent = climate.get_recent_forecast(cell)
|
||||
rr = recent[recent["date"] == target]
|
||||
if not rr.empty:
|
||||
obs = _obs_from_row(rr.iloc[0])
|
||||
except Exception: # noqa: BLE001 - detail page still works from climatology alone
|
||||
obs = None
|
||||
|
||||
with run.phase("detail"):
|
||||
detail = grading.day_detail(history, target, obs)
|
||||
run.set(has_observation=obs is not None)
|
||||
return {
|
||||
"api_version": "v2",
|
||||
"cell": cell,
|
||||
"place": place,
|
||||
"latest": last.date().isoformat(),
|
||||
"detail": detail,
|
||||
}
|
||||
|
||||
|
||||
def _build_forecast(cell, days, history, fc, today, place, run=None) -> dict:
|
||||
"""/forecast payload: the next `days` forecast days graded — same shape as
|
||||
/grade so the frontend renders it identically, furthest-out day first."""
|
||||
run = run or _NullRun()
|
||||
# Strictly future days (tomorrow onward), earliest→latest, capped at `days`.
|
||||
future = fc[fc["date"] > today].sort_values("date").head(days)
|
||||
with run.phase("grading"):
|
||||
graded = _grade_rows(history, future)
|
||||
_attach_dry_streaks(graded, history, fc)
|
||||
graded.reverse() # furthest-out first, so the chart reverses to L→R chronological
|
||||
climo = grading.climatology(history, int(today.dayofyear))
|
||||
run.set(graded_days=len(graded), place_found=place is not None)
|
||||
return {
|
||||
"api_version": "v2",
|
||||
"cell": cell,
|
||||
"place": place,
|
||||
"target_date": today.date().isoformat(),
|
||||
"forecast": True,
|
||||
"climatology": climo,
|
||||
"recent": graded,
|
||||
}
|
||||
|
||||
|
||||
# --- endpoints ----------------------------------------------------------------
|
||||
|
||||
def api_geocode(q: str = Query(..., min_length=1)):
|
||||
try:
|
||||
return {"results": climate.geocode(q)}
|
||||
|
|
@ -95,6 +308,7 @@ def api_geocode(q: str = Query(..., min_length=1)):
|
|||
|
||||
|
||||
def api_grade(
|
||||
request: Request,
|
||||
lat: float = Query(..., ge=-90, le=90),
|
||||
lon: float = Query(..., ge=-180, le=180),
|
||||
date: str | None = Query(None, description="target date YYYY-MM-DD (default today)"),
|
||||
|
|
@ -132,67 +346,25 @@ def api_grade(
|
|||
if history.empty:
|
||||
raise HTTPException(status_code=404, detail="No historical data for this cell.")
|
||||
|
||||
with run.phase("grading"):
|
||||
# Only grade observed days up to and including the target date (last `days`).
|
||||
recent = recent[recent["date"] <= target].sort_values("date").tail(days)
|
||||
graded = _grade_rows(history, recent)
|
||||
_attach_dry_streaks(graded, history, recent)
|
||||
graded.reverse() # newest first for display
|
||||
climo = grading.climatology(history, int(target.dayofyear))
|
||||
key = f"{target.date().isoformat()}:{days}"
|
||||
token = f"{PAYLOAD_VER}:{_hist_end(history)}:{climate.recent_stamp(cell['id'])}"
|
||||
etag = _etag_for("grade", cell["id"], key, token)
|
||||
if _not_modified(request, etag):
|
||||
run.set(run_type="cache", not_modified=True)
|
||||
return Response(status_code=304, headers={"ETag": etag})
|
||||
body = store.get_payload("grade", cell["id"], key, token)
|
||||
if body is not None:
|
||||
run.set(run_type="cache", history_source="store")
|
||||
return _json_response(body, etag)
|
||||
|
||||
with run.phase("reverse_geocode"):
|
||||
place = climate.reverse_geocode(cell["center_lat"], cell["center_lon"])
|
||||
|
||||
# Summarize what this run actually covered. A fresh history fetch pulls the
|
||||
# full ~45-year archive (run_type "full"); a cache hit only fetched the
|
||||
# recent window (run_type "partial").
|
||||
years = pd.to_datetime(history["date"]).dt.year
|
||||
full = not cache_meta.get("cached", False)
|
||||
run.set(
|
||||
run_type="full" if full else "partial",
|
||||
history_source="fetch" if full else "cache",
|
||||
history_rows=int(len(history)),
|
||||
history_years=int(years.max() - years.min() + 1),
|
||||
history_span=[int(years.min()), int(years.max())],
|
||||
cache_age_days=cache_meta.get("cache_age_days"),
|
||||
graded_days=len(graded),
|
||||
place_found=place is not None,
|
||||
)
|
||||
|
||||
return {
|
||||
"cell": cell,
|
||||
"place": place,
|
||||
"target_date": target.date().isoformat(),
|
||||
"cache": cache_meta,
|
||||
"climatology": climo,
|
||||
"recent": graded,
|
||||
}
|
||||
|
||||
|
||||
CAL_MAX_SPAN_DAYS = 732 # ~2 years — the most a single calendar request will grade
|
||||
# (the frontend splits longer spans into 2-year chunks)
|
||||
|
||||
# Short-lived in-memory cache of full calendar responses, keyed by (cell, span).
|
||||
# Lets the home page prefetch the calendar in the background so navigating to it is
|
||||
# instant, and makes repeat calendar loads free. Bounded + TTL'd.
|
||||
CAL_CACHE_TTL = 600 # seconds (10 min)
|
||||
_CAL_CACHE: dict[tuple, tuple[float, dict]] = {}
|
||||
|
||||
|
||||
def _cal_cache_get(key):
|
||||
hit = _CAL_CACHE.get(key)
|
||||
if hit and (time.time() - hit[0]) < CAL_CACHE_TTL:
|
||||
return hit[1]
|
||||
return None
|
||||
|
||||
|
||||
def _cal_cache_put(key, resp):
|
||||
_CAL_CACHE[key] = (time.time(), resp)
|
||||
if len(_CAL_CACHE) > 64: # evict the oldest entry to stay bounded
|
||||
_CAL_CACHE.pop(min(_CAL_CACHE, key=lambda k: _CAL_CACHE[k][0]), None)
|
||||
payload = _build_grade(cell, target, days, history, recent, cache_meta, place, run)
|
||||
return _json_response(store.put_payload("grade", cell["id"], key, token, payload), etag)
|
||||
|
||||
|
||||
def api_calendar(
|
||||
request: Request,
|
||||
lat: float = Query(..., ge=-90, le=90),
|
||||
lon: float = Query(..., ge=-180, le=180),
|
||||
start: str | None = Query(None, description="first date YYYY-MM-DD (overrides months)"),
|
||||
|
|
@ -202,10 +374,12 @@ def api_calendar(
|
|||
"""Grade every day over a date range (default: last 2 years) for the calendar view.
|
||||
|
||||
A custom [start, end] range is supported and capped at ~2 years per request;
|
||||
the frontend splits longer spans into successive 2-year chunks. Without a
|
||||
start it falls back to `months` back from `end`. Grades against the already-
|
||||
cached history record (which reaches to ~6 days ago), so no extra fetch is
|
||||
needed. Returns a compact per-day shape (see grading.grade_range). New in API v2.
|
||||
the frontend splits longer spans into successive 2-year chunks. Grades against
|
||||
the already-cached history record (which reaches to ~6 days ago), so no extra
|
||||
fetch is needed. The graded payload is cached in the derived store keyed by
|
||||
the clamped span and validated by the record's end date, so a repeat span is
|
||||
a database read — across restarts — until new archive days arrive. Returns a
|
||||
compact per-day shape (see grading.grade_range). New in API v2.
|
||||
"""
|
||||
if not grid.in_north_america(lat, lon):
|
||||
raise HTTPException(
|
||||
|
|
@ -229,52 +403,29 @@ def api_calendar(
|
|||
if history.empty:
|
||||
raise HTTPException(status_code=404, detail="No historical data for this cell.")
|
||||
|
||||
first = pd.Timestamp(history["date"].min()).normalize()
|
||||
last = pd.Timestamp(history["date"].max()).normalize()
|
||||
end_ts = min(pd.Timestamp(end).normalize(), last) if end else last
|
||||
if start:
|
||||
start_ts = pd.Timestamp(start).normalize()
|
||||
else:
|
||||
# Default: whole months back landing on the SAME day-of-month as the end,
|
||||
# so the two range endpoints line up (e.g. 2024-06-29 → 2026-06-29).
|
||||
start_ts = (end_ts - pd.DateOffset(months=months)).normalize()
|
||||
# Cap the graded span at ~2 years, and keep it inside the available record.
|
||||
min_start = end_ts - pd.Timedelta(days=CAL_MAX_SPAN_DAYS - 1)
|
||||
start_ts = max(start_ts, min_start, first)
|
||||
start_ts = min(start_ts, end_ts)
|
||||
|
||||
cache_key = (cell["id"], start_ts.date().isoformat(), end_ts.date().isoformat())
|
||||
cached = _cal_cache_get(cache_key)
|
||||
if cached is not None:
|
||||
run.set(run_type="cache", history_source="cache", graded_days=len(cached["days"]))
|
||||
return cached
|
||||
|
||||
with run.phase("grading"):
|
||||
days = grading.grade_range(history, start_ts, end_ts)
|
||||
start_ts, end_ts = _cal_span(history, start, end, months)
|
||||
key = f"{start_ts.date().isoformat()}:{end_ts.date().isoformat()}:{months}"
|
||||
token = f"{PAYLOAD_VER}:{_hist_end(history)}"
|
||||
etag = _etag_for("calendar", cell["id"], key, token)
|
||||
if _not_modified(request, etag):
|
||||
run.set(run_type="cache", not_modified=True)
|
||||
return Response(status_code=304, headers={"ETag": etag})
|
||||
body = store.get_payload("calendar", cell["id"], key, token)
|
||||
if body is not None:
|
||||
run.set(run_type="cache", history_source="store")
|
||||
return _json_response(body, etag)
|
||||
|
||||
with run.phase("reverse_geocode"):
|
||||
place = climate.reverse_geocode(cell["center_lat"], cell["center_lon"])
|
||||
|
||||
payload = _build_calendar(cell, history, start_ts, end_ts, months, place, run)
|
||||
full = not cache_meta.get("cached", False)
|
||||
run.set(
|
||||
run_type="full" if full else "partial",
|
||||
history_source="fetch" if full else "cache",
|
||||
graded_days=len(days),
|
||||
)
|
||||
|
||||
resp = {
|
||||
"api_version": "v2",
|
||||
"cell": cell,
|
||||
"place": place,
|
||||
"range": {"start": start_ts.date().isoformat(), "end": end_ts.date().isoformat()},
|
||||
"months": months,
|
||||
"days": days,
|
||||
}
|
||||
_cal_cache_put(cache_key, resp)
|
||||
return resp
|
||||
run.set(run_type="full" if full else "partial",
|
||||
history_source="fetch" if full else "cache")
|
||||
return _json_response(store.put_payload("calendar", cell["id"], key, token, payload), etag)
|
||||
|
||||
|
||||
def api_day(
|
||||
request: Request,
|
||||
lat: float = Query(..., ge=-90, le=90),
|
||||
lon: float = Query(..., ge=-180, le=180),
|
||||
date: str | None = Query(None, description="target date YYYY-MM-DD (default = latest available)"),
|
||||
|
|
@ -283,8 +434,9 @@ def api_day(
|
|||
in that day-of-year's ±7-day window, plus where the observed values land.
|
||||
|
||||
Powers the single-day detail page. Grades against the cached history record;
|
||||
for a date newer than the cache it pulls the recent window for the observation.
|
||||
New in API v2.
|
||||
for a date newer than the cache it pulls the recent window for the observation
|
||||
(those payloads expire hourly — the recent bundle's own cadence — while fully
|
||||
archived days stay valid until the record itself advances). New in API v2.
|
||||
"""
|
||||
if not grid.in_north_america(lat, lon):
|
||||
raise HTTPException(
|
||||
|
|
@ -308,45 +460,30 @@ def api_day(
|
|||
target = pd.Timestamp(date).normalize() if date else last
|
||||
run.set(target_date=target.date().isoformat())
|
||||
|
||||
# Observed values for the target date: prefer the cached history record; if
|
||||
# the date is newer than the cache (last few days), fetch the recent window.
|
||||
obs = None
|
||||
hit = history[history["date"] == target]
|
||||
if not hit.empty:
|
||||
obs = _obs_from_row(hit.iloc[0])
|
||||
elif target > last:
|
||||
try:
|
||||
with run.phase("recent"):
|
||||
recent = climate.get_recent_forecast(cell)
|
||||
rr = recent[recent["date"] == target]
|
||||
if not rr.empty:
|
||||
obs = _obs_from_row(rr.iloc[0])
|
||||
except Exception: # noqa: BLE001 - detail page still works from climatology alone
|
||||
obs = None
|
||||
|
||||
with run.phase("detail"):
|
||||
detail = grading.day_detail(history, target, obs)
|
||||
key = target.date().isoformat()
|
||||
token = f"{PAYLOAD_VER}:{_hist_end(history)}"
|
||||
if target > last:
|
||||
token += f":h{int(time.time() // 3600)}" # obs from the hourly recent bundle
|
||||
etag = _etag_for("day", cell["id"], key, token)
|
||||
if _not_modified(request, etag):
|
||||
run.set(run_type="cache", not_modified=True)
|
||||
return Response(status_code=304, headers={"ETag": etag})
|
||||
body = store.get_payload("day", cell["id"], key, token)
|
||||
if body is not None:
|
||||
run.set(run_type="cache", history_source="store")
|
||||
return _json_response(body, etag)
|
||||
|
||||
with run.phase("reverse_geocode"):
|
||||
place = climate.reverse_geocode(cell["center_lat"], cell["center_lon"])
|
||||
|
||||
payload = _build_day(cell, history, target, place, run)
|
||||
full = not cache_meta.get("cached", False)
|
||||
run.set(
|
||||
run_type="full" if full else "partial",
|
||||
history_source="fetch" if full else "cache",
|
||||
has_observation=obs is not None,
|
||||
)
|
||||
|
||||
return {
|
||||
"api_version": "v2",
|
||||
"cell": cell,
|
||||
"place": place,
|
||||
"latest": last.date().isoformat(),
|
||||
"detail": detail,
|
||||
}
|
||||
run.set(run_type="full" if full else "partial",
|
||||
history_source="fetch" if full else "cache")
|
||||
return _json_response(store.put_payload("day", cell["id"], key, token, payload), etag)
|
||||
|
||||
|
||||
def api_forecast(
|
||||
request: Request,
|
||||
lat: float = Query(..., ge=-90, le=90),
|
||||
lon: float = Query(..., ge=-180, le=180),
|
||||
days: int = Query(7, ge=1, le=14, description="how many forecast days ahead to grade"),
|
||||
|
|
@ -355,8 +492,9 @@ def api_forecast(
|
|||
|
||||
Same shape as /grade (so the frontend renders it identically), but `recent`
|
||||
holds the forward forecast — furthest-out day first. The forecast is fetched
|
||||
fresh and cached only ~1 hour (see climate.get_recent_forecast) to track updates.
|
||||
New in API v2.
|
||||
fresh and cached only ~1 hour (see climate.get_recent_forecast) to track
|
||||
updates; the graded payload's validity is tied to that fetch stamp, so it
|
||||
expires exactly when a new forecast lands. New in API v2.
|
||||
"""
|
||||
if not grid.in_north_america(lat, lon):
|
||||
raise HTTPException(
|
||||
|
|
@ -379,29 +517,139 @@ def api_forecast(
|
|||
raise HTTPException(status_code=404, detail="No historical data for this cell.")
|
||||
|
||||
today = pd.Timestamp(datetime.date.today())
|
||||
# Strictly future days (tomorrow onward), earliest→latest, capped at `days`.
|
||||
future = fc[fc["date"] > today].sort_values("date").head(days)
|
||||
key = f"{today.date().isoformat()}:{days}"
|
||||
token = f"{PAYLOAD_VER}:{_hist_end(history)}:{climate.recent_stamp(cell['id'])}"
|
||||
etag = _etag_for("forecast", cell["id"], key, token)
|
||||
if _not_modified(request, etag):
|
||||
run.set(run_type="cache", not_modified=True)
|
||||
return Response(status_code=304, headers={"ETag": etag})
|
||||
body = store.get_payload("forecast", cell["id"], key, token)
|
||||
if body is not None:
|
||||
run.set(run_type="cache", history_source="store")
|
||||
return _json_response(body, etag)
|
||||
|
||||
with run.phase("grading"):
|
||||
graded = _grade_rows(history, future)
|
||||
_attach_dry_streaks(graded, history, fc)
|
||||
graded.reverse() # furthest-out first, so the chart reverses to L→R chronological
|
||||
climo = grading.climatology(history, int(today.dayofyear))
|
||||
with run.phase("reverse_geocode"):
|
||||
place = climate.reverse_geocode(cell["center_lat"], cell["center_lon"])
|
||||
payload = _build_forecast(cell, days, history, fc, today, place, run)
|
||||
run.set(run_type="partial")
|
||||
return _json_response(store.put_payload("forecast", cell["id"], key, token, payload), etag)
|
||||
|
||||
|
||||
def api_cell(
|
||||
request: Request,
|
||||
lat: float = Query(..., ge=-90, le=90),
|
||||
lon: float = Query(..., ge=-180, le=180),
|
||||
prefetch: int = Query(0, ge=0, le=1,
|
||||
description="1 = warm-only: never fetch weather upstream; 204 for a cold cell"),
|
||||
):
|
||||
"""One bundle carrying every view's payload for a cell, so the frontend warms
|
||||
all views with a single request instead of four.
|
||||
|
||||
Each slice is the EXACT payload its per-view endpoint returns — built by the
|
||||
same builders and cached under the same derived-store keys/tokens — paired
|
||||
with the etag that endpoint would emit. The client seeds its per-view cache
|
||||
from the slices and later revalidates each view individually with
|
||||
If-None-Match, so the bundle and the per-view endpoints stay one cache.
|
||||
|
||||
prefetch=1 is the neighbor-warming mode with a hard guarantee: it never
|
||||
spends weather-API quota. A cell with no cached archive answers 204 (no
|
||||
body), and only the history-derived slices (calendar + latest-day detail)
|
||||
are built — grading recent/forecast days needs the hourly upstream bundle.
|
||||
Reverse geocoding makes at most one Nominatim call for a never-labeled cell;
|
||||
the client staggers neighbor prefetches to respect that service. New in API v2.
|
||||
"""
|
||||
if not grid.in_north_america(lat, lon):
|
||||
raise HTTPException(
|
||||
status_code=400,
|
||||
detail="Location is outside the supported US + Canada region.",
|
||||
)
|
||||
cell = grid.snap(lat, lon)
|
||||
today = pd.Timestamp(datetime.date.today())
|
||||
|
||||
with audit.RunAudit(endpoint="cell", lat=round(lat, 4), lon=round(lon, 4),
|
||||
cell_id=cell["id"], prefetch=bool(prefetch)) as run:
|
||||
recent = None
|
||||
if prefetch:
|
||||
history = climate.load_cached_history(cell)
|
||||
if history is None or history.empty:
|
||||
run.set(run_type="cold-skip")
|
||||
return Response(status_code=204)
|
||||
cache_meta = {"cached": True}
|
||||
else:
|
||||
try:
|
||||
with run.phase("history"):
|
||||
history, cache_meta = climate.get_history(cell)
|
||||
with run.phase("recent"):
|
||||
recent = climate.get_recent_forecast(cell)
|
||||
except Exception as e: # noqa: BLE001
|
||||
raise _weather_fetch_error(e)
|
||||
if history.empty:
|
||||
raise HTTPException(status_code=404, detail="No historical data for this cell.")
|
||||
|
||||
cid = cell["id"]
|
||||
hist_end = _hist_end(history)
|
||||
last = pd.Timestamp(history["date"].max()).normalize()
|
||||
|
||||
with run.phase("reverse_geocode"):
|
||||
place = climate.reverse_geocode(cell["center_lat"], cell["center_lon"])
|
||||
|
||||
run.set(run_type="partial", graded_days=len(graded), place_found=place is not None)
|
||||
def slice_for(kind: str, key: str, token: str, build) -> dict:
|
||||
"""The endpoint's derived row (or build + store it), paired with the
|
||||
etag that endpoint would emit for the same request."""
|
||||
etag = _etag_for(kind, cid, key, token)
|
||||
data = store.get_json(kind, cid, key, token)
|
||||
if data is None:
|
||||
data = build()
|
||||
store.put_payload(kind, cid, key, token, data)
|
||||
return {"etag": etag, "data": data}
|
||||
|
||||
return {
|
||||
slices = {}
|
||||
hist_token = f"{PAYLOAD_VER}:{hist_end}"
|
||||
# Calendar: the default last-24-months span — what the calendar view (and
|
||||
# the cross-view prefetch) requests first.
|
||||
start_ts, end_ts = _cal_span(history, None, None, 24)
|
||||
cal_key = f"{start_ts.date().isoformat()}:{end_ts.date().isoformat()}:24"
|
||||
slices["calendar"] = slice_for(
|
||||
"calendar", cal_key, hist_token,
|
||||
lambda: _build_calendar(cell, history, start_ts, end_ts, 24, place, run))
|
||||
|
||||
if prefetch:
|
||||
# Latest archived day: its observation comes from history alone, so
|
||||
# it's buildable without the (skipped) recent bundle.
|
||||
slices["day"] = slice_for(
|
||||
"day", last.date().isoformat(), hist_token,
|
||||
lambda: _build_day(cell, history, last, place, run))
|
||||
else:
|
||||
rf_token = f"{PAYLOAD_VER}:{hist_end}:{climate.recent_stamp(cid)}"
|
||||
slices["grade"] = slice_for(
|
||||
"grade", f"{today.date().isoformat()}:14", rf_token,
|
||||
lambda: _build_grade(cell, today, 14, history, recent, cache_meta, place, run))
|
||||
slices["forecast"] = slice_for(
|
||||
"forecast", f"{today.date().isoformat()}:7", rf_token,
|
||||
lambda: _build_forecast(cell, 7, history, recent, today, place, run))
|
||||
day_token = hist_token
|
||||
if today > last:
|
||||
day_token += f":h{int(time.time() // 3600)}" # obs from the hourly recent bundle
|
||||
slices["day"] = slice_for(
|
||||
"day", today.date().isoformat(), day_token,
|
||||
lambda: _build_day(cell, history, today, place, run, recent=recent))
|
||||
|
||||
# The bundle's identity is the combination of its slices' identities.
|
||||
etag = _etag_for("cell", cid, "bundle" + (":p" if prefetch else ""),
|
||||
"|".join(s["etag"] for s in slices.values()))
|
||||
if _not_modified(request, etag):
|
||||
run.set(run_type="cache", not_modified=True)
|
||||
return Response(status_code=304, headers={"ETag": etag})
|
||||
run.set(run_type="bundle", slices=sorted(slices))
|
||||
payload = {
|
||||
"api_version": "v2",
|
||||
"cell": cell,
|
||||
"place": place,
|
||||
"target_date": today.date().isoformat(),
|
||||
"forecast": True,
|
||||
"climatology": climo,
|
||||
"recent": graded,
|
||||
"today": today.date().isoformat(),
|
||||
"slices": slices,
|
||||
}
|
||||
# Not stored as its own derived row — the slices already are.
|
||||
return _json_response(_encode(payload), etag)
|
||||
|
||||
|
||||
# --- API versioning --------------------------------------------------------
|
||||
|
|
@ -419,6 +667,7 @@ v2.add_api_route("/grade", api_grade, methods=["GET"])
|
|||
v2.add_api_route("/calendar", api_calendar, methods=["GET"])
|
||||
v2.add_api_route("/day", api_day, methods=["GET"])
|
||||
v2.add_api_route("/forecast", api_forecast, methods=["GET"])
|
||||
v2.add_api_route("/cell", api_cell, methods=["GET"])
|
||||
|
||||
app.include_router(v1, prefix=f"{BASE}/api") # legacy unversioned == v1 (backward compatible)
|
||||
app.include_router(v1, prefix=f"{BASE}/api/v1")
|
||||
|
|
|
|||
52
climate.py
52
climate.py
|
|
@ -15,6 +15,7 @@ import numpy as np
|
|||
import pandas as pd
|
||||
|
||||
import audit
|
||||
import store
|
||||
|
||||
CACHE_DIR = os.path.join(os.path.dirname(__file__), "..", "data", "cache")
|
||||
MAX_ATTEMPTS = 3 # per upstream call, before giving up
|
||||
|
|
@ -362,6 +363,20 @@ def get_history(cell: dict) -> tuple[pd.DataFrame, dict]:
|
|||
return df, meta
|
||||
|
||||
|
||||
def load_cached_history(cell: dict) -> pd.DataFrame | None:
|
||||
"""History from the parquet cache ONLY — never fetches upstream and never tops
|
||||
up the tail. Powers the warm-only prefetch path (which must not spend upstream
|
||||
quota) and the offline migrate script. None when the cell has no
|
||||
(schema-complete) cached record."""
|
||||
hit = _read_history_cache(_cache_path(cell["id"]))
|
||||
if hit is None:
|
||||
return None
|
||||
df = hit[0].copy()
|
||||
df["doy"] = df["date"].dt.dayofyear.astype("int16")
|
||||
_derive_humidity(df)
|
||||
return df
|
||||
|
||||
|
||||
def _load_history(cell: dict) -> tuple[pd.DataFrame, dict]:
|
||||
"""Return (daily history frame, cache metadata) for a cell.
|
||||
|
||||
|
|
@ -431,6 +446,16 @@ def _rf_cache_path(cell_id: str) -> str:
|
|||
return os.path.join(CACHE_DIR, f"{cell_id}_rf.parquet")
|
||||
|
||||
|
||||
def recent_stamp(cell_id: str) -> int:
|
||||
"""Identity stamp (mtime, whole seconds) of the cell's cached recent+forecast
|
||||
parquet; 0 when absent. Changes exactly when the recent/forecast data does, so
|
||||
it's the freshness token for every payload that grades recent or future days."""
|
||||
try:
|
||||
return int(os.path.getmtime(_rf_cache_path(cell_id)))
|
||||
except OSError:
|
||||
return 0
|
||||
|
||||
|
||||
def get_recent_forecast(cell: dict) -> pd.DataFrame:
|
||||
"""Recent observations + forward forecast, with humidity as absolute humidity
|
||||
(g/m³). Thin wrapper over the raw loader (see below)."""
|
||||
|
|
@ -472,18 +497,32 @@ def _load_recent_forecast(cell: dict) -> pd.DataFrame:
|
|||
return df
|
||||
|
||||
|
||||
_REVGEO_CACHE: dict[tuple[float, float], str | None] = {}
|
||||
_REVGEO_CACHE: dict[str, str | None] = {}
|
||||
|
||||
|
||||
def reverse_geocode_cached(lat: float, lon: float) -> tuple[bool, str | None]:
|
||||
"""(found, label) from the in-memory + SQLite revgeo caches only — never calls
|
||||
Nominatim. Used by paths that must not add upstream traffic (prefetch)."""
|
||||
key = store.revgeo_key(lat, lon)
|
||||
if key in _REVGEO_CACHE:
|
||||
return (True, _REVGEO_CACHE[key])
|
||||
found, label = store.get_revgeo(key)
|
||||
if found:
|
||||
_REVGEO_CACHE[key] = label
|
||||
return (found, label)
|
||||
|
||||
|
||||
def reverse_geocode(lat: float, lon: float) -> str | None:
|
||||
"""Best-effort "City, State" label for a point (OpenStreetMap Nominatim).
|
||||
"""Best-effort neighbourhood/city label for a point (OpenStreetMap Nominatim).
|
||||
|
||||
Cached per ~cell so panning around doesn't hammer the service, and failures
|
||||
Cached per ~cell — in memory for the process and in SQLite across restarts —
|
||||
so panning around (or redeploying) doesn't re-hit the service, and failures
|
||||
return None so the caller can fall back to bare coordinates.
|
||||
"""
|
||||
key = (round(lat, 3), round(lon, 3))
|
||||
if key in _REVGEO_CACHE:
|
||||
return _REVGEO_CACHE[key]
|
||||
key = store.revgeo_key(lat, lon)
|
||||
found, label = reverse_geocode_cached(lat, lon)
|
||||
if found:
|
||||
return label
|
||||
label = None
|
||||
try:
|
||||
r = _request(
|
||||
|
|
@ -510,6 +549,7 @@ def reverse_geocode(lat: float, lon: float) -> str | None:
|
|||
except Exception: # noqa: BLE001 - reverse geocoding is a nicety, never fatal
|
||||
label = None
|
||||
_REVGEO_CACHE[key] = label
|
||||
store.put_revgeo(key, label) # survive restarts (a None label is retried daily)
|
||||
return label
|
||||
|
||||
|
||||
|
|
|
|||
10
grid.py
10
grid.py
|
|
@ -46,6 +46,16 @@ def snap(lat: float, lon: float) -> dict:
|
|||
}
|
||||
|
||||
|
||||
def from_id(cell_id: str) -> dict:
|
||||
"""Rebuild the full cell dict from a cache id ("i_j") — the inverse of snap().
|
||||
Lets offline tooling (the migrate script) recover a cell from its parquet
|
||||
filename alone. Raises ValueError on a malformed id."""
|
||||
i, j = (int(p) for p in cell_id.split("_"))
|
||||
center_lat = (i + 0.5) * LAT_STEP
|
||||
center_lon = (j + 0.5) * _lon_step(center_lat)
|
||||
return snap(center_lat, center_lon)
|
||||
|
||||
|
||||
def in_north_america(lat: float, lon: float) -> bool:
|
||||
"""Rough bounding box for the US (incl. Alaska/Hawaii) and Canada."""
|
||||
return 14.0 <= lat <= 84.0 and -172.0 <= lon <= -52.0
|
||||
|
|
|
|||
91
migrate.py
Normal file
91
migrate.py
Normal file
|
|
@ -0,0 +1,91 @@
|
|||
"""Backfill the SQLite derived store from the existing parquet cell caches.
|
||||
|
||||
For every cell with a cached archive record, materialize the payloads that are
|
||||
derivable offline — the default 2-year calendar and the latest-day detail — plus
|
||||
its reverse-geocode label, so a server restarted (or freshly deployed) onto an
|
||||
existing parquet cache serves those views from the database immediately instead
|
||||
of regrading each cell on first touch.
|
||||
|
||||
Idempotent and resumable: cells whose derived rows already match the current
|
||||
validity token are skipped, so re-running after an interruption (or after new
|
||||
archive days arrive) only does the missing work. Never fetches weather data —
|
||||
history comes strictly from the parquet cache (cells without one are skipped and
|
||||
will materialize lazily on first request, as always). Reverse geocoding makes at
|
||||
most one Nominatim call per unlabeled cell, throttled to ~1/s per the service's
|
||||
policy. Safe to run alongside a live server: writes go through the same WAL
|
||||
store the server reads, and readers fall back to recomputing on any miss.
|
||||
|
||||
Run via `make migrate`.
|
||||
"""
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
|
||||
import pandas as pd
|
||||
|
||||
import climate
|
||||
import grid
|
||||
import store
|
||||
from app import PAYLOAD_VER, _build_calendar, _build_day, _cal_span, _hist_end
|
||||
|
||||
|
||||
def migrate() -> int:
|
||||
if not os.path.isdir(climate.CACHE_DIR):
|
||||
print("No parquet cache directory yet — nothing to migrate.")
|
||||
return 0
|
||||
cells = sorted(
|
||||
f[: -len(".parquet")]
|
||||
for f in os.listdir(climate.CACHE_DIR)
|
||||
if f.endswith(".parquet") and not f.endswith("_rf.parquet")
|
||||
)
|
||||
built = current = skipped = 0
|
||||
for cell_id in cells:
|
||||
try:
|
||||
cell = grid.from_id(cell_id)
|
||||
except ValueError:
|
||||
print(f" {cell_id}: unrecognized cache filename — skipped")
|
||||
skipped += 1
|
||||
continue
|
||||
history = climate.load_cached_history(cell)
|
||||
if history is None or history.empty:
|
||||
# Pre-current-schema record; the server refetches it lazily on first
|
||||
# touch (see climate.NEW_COLS) — nothing to materialize offline.
|
||||
print(f" {cell_id}: no schema-complete record — skipped")
|
||||
skipped += 1
|
||||
continue
|
||||
|
||||
token = f"{PAYLOAD_VER}:{_hist_end(history)}"
|
||||
start_ts, end_ts = _cal_span(history, None, None, 24)
|
||||
cal_key = f"{start_ts.date().isoformat()}:{end_ts.date().isoformat()}:24"
|
||||
last = pd.Timestamp(history["date"].max()).normalize()
|
||||
day_key = last.date().isoformat()
|
||||
|
||||
have_cal = store.get_payload("calendar", cell_id, cal_key, token) is not None
|
||||
have_day = store.get_payload("day", cell_id, day_key, token) is not None
|
||||
if have_cal and have_day:
|
||||
current += 1
|
||||
continue
|
||||
|
||||
found, place = climate.reverse_geocode_cached(cell["center_lat"], cell["center_lon"])
|
||||
if not found:
|
||||
place = climate.reverse_geocode(cell["center_lat"], cell["center_lon"])
|
||||
time.sleep(1.1) # Nominatim usage policy: at most ~1 request/second
|
||||
|
||||
if not have_cal:
|
||||
store.put_payload("calendar", cell_id, cal_key, token,
|
||||
_build_calendar(cell, history, start_ts, end_ts, 24, place))
|
||||
if not have_day:
|
||||
store.put_payload("day", cell_id, day_key, token,
|
||||
_build_day(cell, history, last, place))
|
||||
built += 1
|
||||
print(f" {cell_id}: materialized ({place or 'no label'})")
|
||||
|
||||
s = store.stats()
|
||||
print(f"{built} cell(s) materialized, {current} already current, {skipped} skipped"
|
||||
f" · derived rows: {s['derived']} · revgeo: {s['revgeo']}"
|
||||
f" · db: {s['bytes'] / 1e6:.1f} MB → {s['db_path']}")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(migrate())
|
||||
167
store.py
Normal file
167
store.py
Normal file
|
|
@ -0,0 +1,167 @@
|
|||
"""Persistent derived-data store — SQLite (WAL) at data/thermograph.sqlite.
|
||||
|
||||
The raw per-cell parquet records stay the source of truth; this database holds
|
||||
only what's *derived* from them, so expensive work is done once per cell and
|
||||
survives restarts/deploys instead of being recomputed per request, per view:
|
||||
|
||||
* ``derived`` — finished API response payloads (grade / calendar / day /
|
||||
forecast), zlib-compressed JSON keyed by (kind, cell, request key) and
|
||||
guarded by a validity ``token``. The token encodes everything the payload
|
||||
depends on (payload schema version, history end date, recent-fetch stamp…);
|
||||
a row whose stored token doesn't match the caller's current token is simply
|
||||
a miss, so stale data can never be served — freshness is driven by the token
|
||||
advancing, never by clock-based expiry.
|
||||
|
||||
* ``revgeo`` — reverse-geocode labels ("City, State" per ~cell), previously an
|
||||
in-memory dict that re-hit Nominatim after every restart.
|
||||
|
||||
The store is a pure accelerator: every reader falls back to recomputing from
|
||||
parquet when the database is missing, locked, or corrupt (all helpers swallow
|
||||
their own errors), and deleting data/thermograph.sqlite is a safe full reset.
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import sqlite3
|
||||
import threading
|
||||
import time
|
||||
import zlib
|
||||
|
||||
DB_PATH = os.path.join(os.path.dirname(__file__), "..", "data", "thermograph.sqlite")
|
||||
|
||||
_SCHEMA = """
|
||||
CREATE TABLE IF NOT EXISTS derived (
|
||||
kind TEXT NOT NULL, -- payload family: grade | calendar | day | forecast
|
||||
cell_id TEXT NOT NULL,
|
||||
key TEXT NOT NULL, -- request identity within the kind (dates/spans/params)
|
||||
token TEXT NOT NULL, -- validity: mismatched token == miss (see module docstring)
|
||||
payload BLOB NOT NULL, -- zlib-compressed JSON bytes
|
||||
updated_at REAL NOT NULL,
|
||||
PRIMARY KEY (kind, cell_id, key)
|
||||
);
|
||||
CREATE TABLE IF NOT EXISTS revgeo (
|
||||
key TEXT PRIMARY KEY, -- "lat,lon" rounded to ~cell precision
|
||||
label TEXT, -- NULL == lookup ran but found no label
|
||||
updated_at REAL NOT NULL
|
||||
);
|
||||
"""
|
||||
|
||||
# SQLite connections aren't shareable across threads; uvicorn serves requests on
|
||||
# a thread pool, so each thread lazily opens its own connection. WAL lets those
|
||||
# readers run concurrently with the (serialized) writers.
|
||||
_local = threading.local()
|
||||
|
||||
|
||||
def _conn() -> sqlite3.Connection | None:
|
||||
conn = getattr(_local, "conn", None)
|
||||
if conn is not None:
|
||||
return conn
|
||||
try:
|
||||
os.makedirs(os.path.dirname(DB_PATH), exist_ok=True)
|
||||
conn = sqlite3.connect(DB_PATH, timeout=5.0)
|
||||
conn.execute("PRAGMA journal_mode=WAL")
|
||||
conn.execute("PRAGMA synchronous=NORMAL")
|
||||
conn.executescript(_SCHEMA)
|
||||
_local.conn = conn
|
||||
return conn
|
||||
except Exception: # noqa: BLE001 - the store is optional; readers fall back
|
||||
return None
|
||||
|
||||
|
||||
# --- derived payloads -------------------------------------------------------
|
||||
|
||||
def get_payload(kind: str, cell_id: str, key: str, token: str) -> bytes | None:
|
||||
"""Raw JSON bytes of a cached payload, or None when absent or token-invalid."""
|
||||
conn = _conn()
|
||||
if conn is None:
|
||||
return None
|
||||
try:
|
||||
row = conn.execute(
|
||||
"SELECT token, payload FROM derived WHERE kind=? AND cell_id=? AND key=?",
|
||||
(kind, cell_id, key),
|
||||
).fetchone()
|
||||
if row is None or row[0] != token:
|
||||
return None
|
||||
return zlib.decompress(row[1])
|
||||
except Exception: # noqa: BLE001
|
||||
return None
|
||||
|
||||
|
||||
def get_json(kind: str, cell_id: str, key: str, token: str) -> dict | None:
|
||||
"""Like get_payload but decoded — for callers assembling composite responses."""
|
||||
body = get_payload(kind, cell_id, key, token)
|
||||
return json.loads(body) if body is not None else None
|
||||
|
||||
|
||||
def put_payload(kind: str, cell_id: str, key: str, token: str, payload: dict) -> bytes:
|
||||
"""Encode ``payload`` to compact JSON, persist it, and return the encoded bytes
|
||||
(so the caller can serve the exact same bytes without re-serializing).
|
||||
|
||||
allow_nan=False mirrors Starlette's JSONResponse: a NaN slipping through
|
||||
grading should fail loudly here, not be cached as JS-unparseable `NaN`."""
|
||||
body = json.dumps(payload, default=str, allow_nan=False,
|
||||
separators=(",", ":")).encode()
|
||||
conn = _conn()
|
||||
if conn is not None:
|
||||
try:
|
||||
conn.execute(
|
||||
"INSERT OR REPLACE INTO derived VALUES (?,?,?,?,?,?)",
|
||||
(kind, cell_id, key, token, zlib.compress(body, 6), time.time()),
|
||||
)
|
||||
conn.commit()
|
||||
except Exception: # noqa: BLE001 - failing to cache must not fail the request
|
||||
pass
|
||||
return body
|
||||
|
||||
|
||||
# --- reverse-geocode labels --------------------------------------------------
|
||||
|
||||
REVGEO_MISS_TTL = 86400.0 # a "no label" result may be transient — retry after a day
|
||||
|
||||
|
||||
def revgeo_key(lat: float, lon: float) -> str:
|
||||
return f"{round(lat, 3)},{round(lon, 3)}"
|
||||
|
||||
|
||||
def get_revgeo(key: str) -> tuple[bool, str | None]:
|
||||
"""(found, label). A stored NULL label counts as found (a cached miss) until
|
||||
it ages past REVGEO_MISS_TTL, when the lookup becomes retryable."""
|
||||
conn = _conn()
|
||||
if conn is None:
|
||||
return (False, None)
|
||||
try:
|
||||
row = conn.execute("SELECT label, updated_at FROM revgeo WHERE key=?", (key,)).fetchone()
|
||||
if row is None:
|
||||
return (False, None)
|
||||
if row[0] is None and time.time() - row[1] > REVGEO_MISS_TTL:
|
||||
return (False, None)
|
||||
return (True, row[0])
|
||||
except Exception: # noqa: BLE001
|
||||
return (False, None)
|
||||
|
||||
|
||||
def put_revgeo(key: str, label: str | None) -> None:
|
||||
conn = _conn()
|
||||
if conn is None:
|
||||
return
|
||||
try:
|
||||
conn.execute("INSERT OR REPLACE INTO revgeo VALUES (?,?,?)", (key, label, time.time()))
|
||||
conn.commit()
|
||||
except Exception: # noqa: BLE001
|
||||
pass
|
||||
|
||||
|
||||
def stats() -> dict:
|
||||
"""Row counts + on-disk size, for the migrate script's summary output."""
|
||||
conn = _conn()
|
||||
out = {"db_path": os.path.abspath(DB_PATH), "derived": 0, "revgeo": 0, "bytes": 0}
|
||||
if conn is None:
|
||||
return out
|
||||
try:
|
||||
out["derived"] = conn.execute("SELECT COUNT(*) FROM derived").fetchone()[0]
|
||||
out["revgeo"] = conn.execute("SELECT COUNT(*) FROM revgeo").fetchone()[0]
|
||||
# Recent writes may still live in the WAL sidecar — count both files.
|
||||
out["bytes"] = sum(os.path.getsize(p) for p in (DB_PATH, DB_PATH + "-wal")
|
||||
if os.path.exists(p))
|
||||
except Exception: # noqa: BLE001
|
||||
pass
|
||||
return out
|
||||
Loading…
Reference in a new issue