# UI product-event instrumentation — design **Status: design + prototype. Feature-flagged OFF. Not deployed anywhere.** Enabling it in any environment requires the operator decisions in [§6](#6-decisions-needed-from-the-operator) first. The goal is to learn what people actually *do* in the Thermograph UI — did they ever get a location, does search work, which views earn their maintenance — with instrumentation that is *structurally incapable* of tracking a person, so it needs no consent banner and contradicts nothing on the public privacy page. --- ## 0. What already exists (build on this, don't reinvent) | Seam | What it does | Where | |---|---|---| | Request middleware | classifies + counts every inbound request, writes an access JSONL line **including the raw client IP** | `backend/web/app.py` `revalidate_static` | | `core/metrics.py` | since-start counters (inbound by category, outbound by source, heartbeats) behind `GET /api/v2/metrics` | `backend/core/metrics.py` | | **Product-event beacon** | `POST /api/v2/event`, `navigator.sendBeacon`, allowlisted names, per-IP rate limit, referrer reduced to a bare domain, always 204 | `api_event` + `metrics.record_event` | | Event counters | `(event, referrer_domain, UTC day)` → count, **7-day** window, UNLOGGED, TRUNCATEd every deploy | `metrics.EVENTS`, `event_cume` | | Client tracker | `track(event)` + `[data-event]` click delegation | `frontend/static/digest.js` | | Activity JSONL | `audit.log_activity(tag, record)` → `logs/activity/*.jsonl` | `backend/core/audit.py` | | Loki/Grafana | Alloy ships `/applogs/**/*.jsonl` fleet-wide, **30-day** retention | `observability/` | So four coarse events (`home.locate`, `home.digest_signup`, `home.records_click`, `home.share`, plus `home.nav_*`) are already live, on the SSR homepage only. **Three gaps this design closes.** 1. **No dimensions.** `home.locate` is a bare counter. It cannot say whether the tap *succeeded*, which is the only interesting part. 2. **No durability.** The counters are UNLOGGED and TRUNCATEd on every deploy, and capped at 7 days. Loki holds 30 days of JSONL. Neither survives a quarter, so no seasonal or before/after comparison is possible — which is most of the value for a site about *weather*. 3. **Nothing outside the homepage.** Calendar, Day, Compare, Score, Alerts — the whole interactive product — is uninstrumented. Two corrections to assumptions that were made when this work was scoped: Loki retention is **30 days** (`observability/loki/config.yml`, `720h`), not 24h; and the Alloy glob `/applogs/**/*.jsonl` already picks up the `activity/` stream, so a new JSONL sink needs no shipping config. --- ## 1. The events worth capturing Traffic is ~9,700 req/day of which ~60% is crawlers, so **~3,000 human req/day**. At that volume a firehose buys nothing a well-chosen aggregate doesn't, and every extra event is another thing to keep honest. Seven events, each tied to a decision someone would actually make: | Event | Question it answers | Decision it informs | |---|---|---| | `view.open` | How many sessions reach each view, and from where? | The denominator for everything else; whether Compare/Score justify their maintenance | | `place.pick` | Does anyone ever get past "no place picked", and by which route (geolocate / map / search / restore / link)? | **The activation question.** If geolocate is mostly declined, promote the map picker | | `place.search` | Does location search find anything? (`hit` / `typo` / `miss`) | Whether to invest in the geocoder/typo tolerance | | `view.control` | Which in-view controls get used — chart metric, °F/°C, date, range, season, compare add/remove, distribution metric | Which chart metrics to keep/promote; whether the unit default is wrong | | `share` | Is the permalink / PNG export used at all? | Whether to keep maintaining the SVG→PNG exporter | | `alert` | The alerts funnel: open → start → create, push on/off | Where subscription signup leaks | | `deadend` | Errors the visitor actually *saw*: grade error, no data, geolocation denied, search miss, 404, offline | Bug reports nobody filed | **Deliberately not captured** — because the request log already answers them, and duplicating a signal means two numbers that will eventually disagree: page-view counts, API latency/error rates, and asset traffic. Also deliberately not captured: hover/scroll/dwell (needs a session identifier to mean anything), map pan/zoom (very high volume, low decision value, and pan/zoom coordinates are location data), and anything at all about *which* place was chosen. ### Two events are recorded server-side, not in the browser `place.search` is recorded in `api_suggest`: the server already knows the answer, every surface that searches is covered for free, and **the query string never has to leave the request handler** — only the three-way outcome is stored. Same argument would apply to alert creation (`accounts/api_accounts.py`). Prefer this whenever the server can derive the fact: less client code, no forgeability, and strictly less data in flight. --- ## 2. Schema Small, typed, and *closed*. Three dimension slots (`view`, `prop`, `value`) plus a referrer domain on `view.open` only. Every value is an enum from a fixed allowlist (`backend/core/events.py` `SCHEMA`). Unknown event name → one `other` bucket with blank dimensions; unknown value in a declared slot → `other`; a slot an event doesn't declare → dropped entirely. **Nothing a client sends is ever persisted verbatim.** ``` view.open view= prop=direct|nav|link|restore +referrer domain place.pick view= prop=geolocate|map|search|restore|link place.search prop=hit|typo|miss view.control view= prop= value= share view= prop=link|png|permalink alert prop=open|start|create|push_on|push_off|test value=y|n deadend view= prop=grade_error|no_data|geo_denied|search_miss|not_found|offline ``` `VIEWS` = home, calendar, day, compare, score, alerts, legend, city, cities, records, month, glossary, about, privacy. `CONTROLS` = chart_metric, unit, date, range, season, totals, dist_metric, comfort, compare_add, compare_remove, filter_sheet, day_step. `VALUES` = the metric keys, `F`/`C`, `today`/`past`/`future`, count buckets (`1`…`5plus`), span buckets (`1y`…`max`), season keys, `on`/`off`, `other`. Numbers are bucketed **client-side** before they are sent (`5plus`, not `7`): a raw count is a small amount of entropy, and a bucket is all the decision needs. **Storage row:** `(hour_bucket, event, view, prop, value, referrer) → count`. There is no row per interaction, and no column an identifier could go in. Worst-case key space is a few hundred rows/hour; realistically ~50. Migration note: the four live `home.*` events keep their current unflagged, un-dimensioned behaviour so the ops endpoint's numbers don't reset when this lands. Fold them into `view.open`/`place.pick`/`share` and retire the old names as a follow-up, once there's overlap data to stitch the series. --- ## 3. Transport and storage **Chosen: the existing beacon → two sinks.** * **Transport** — extend `POST /api/v2/event` (`navigator.sendBeacon`, batched, always 204). Already built, already excluded from the traffic counters, already rate-limited. Batching (≤20 events, flushed on a 4 s timer and unconditionally on `pagehide`/`visibilitychange`) matters because chart-metric taps and slider drags arrive in bursts. * **Durable sink** — `ui_event_hourly`, a TimescaleDB hypertable (`backend/alembic/versions/0003_ui_events.py`): 30-day chunks, a native retention policy, and an `ON CONFLICT … count + 1` upsert. The DB is already TimescaleDB with hypertables and a compression/retention pattern to copy. * **Live sink** — one `audit.log_activity("ui.event", …)` JSONL line per event. Alloy already ships that path, so the last 30 days are queryable in Grafana immediately without waiting on the rollup. Same allowlisted enums, nothing else. * **The existing in-process counters stay untouched** — they remain the live "what's happening right now" view on `/api/v2/metrics`. **Why not the alternatives:** * *Loki only.* 30-day retention loses every seasonal comparison — the thing a weather product most wants. Loki is also a poor aggregation engine for "unit switches per week over a year". Kept as the live tier, rejected as the system of record. * *Extend the in-process counters.* Cheapest option and no new storage, but they are UNLOGGED, TRUNCATEd every deploy, capped at 7 days, and have exactly one dimension. Every question in §1 needs at least two. Kept, not extended. * *Per-event rows in a hypertable (a real firehose).* With no identifier, a row per interaction gives you nothing an hourly aggregate doesn't — except intra-hour timing, which no listed decision needs. It costs a table that *could* hold a behavioural sequence, which is exactly the property that would make the privacy analysis hard. Rejected on privacy grounds, not cost. * *A third-party analytics SaaS (Plausible/Umami/GA).* GA is a non-starter (US transfer, consent banner, cookies). Self-hosted Plausible/Umami would be defensible, but it means another service to run and its own identifier model (Plausible hashes IP+UA+salt — a pseudonymous identifier, which is precisely the decision being deferred to the operator in §6.1). Rejected as more moving parts for less control. --- ## 4. Abuse and bot resistance The endpoint is public and unauthenticated on a site where 60% of traffic is crawlers and scanners. Layered, in order of how much they actually buy: 1. **Closed allowlist** (the real defence). Nothing outside `SCHEMA` is stored; unknown names collapse into one `other` bucket. Storage is bounded by the *product of the enums*, not by what anyone sends. The endpoint cannot be grown into a spam sink or a data-exfil channel. 2. **Body cap, enforced while streaming** — `MAX_BODY_BYTES = 4096`; a declared `Content-Length` past it is refused without reading, and an undeclared one stops mid-stream. (The cap must comfortably exceed a full 20-event batch, or the byte cap silently eats legitimate traffic instead of abusive traffic — there is a regression test for exactly that.) 3. **Per-IP rate limit** — the existing token bucket, raised from 30/min to **120/min** (`THERMOGRAPH_EVENT_RATE_PER_MIN`) because 30 was sized for four coarse events and a batched interaction stream would silently truncate normal use, biased toward the most engaged visitors. The IP is a key in a map cleared every minute; it is never written anywhere. 4. **Soft first-party check** — `Sec-Fetch-Site` must be `same-origin`/ `same-site`/`none`; a `cross-site` value is dropped, and a POST with neither `Sec-Fetch-Site` nor `Origin` is flagged as the first bucket to look at. Every browser that can run our JS sets this; naive scanners don't. Explicitly **not** a security boundary — anything that bothers can spoof it. 5. **Uniform 204** — counted, rejected, over-cap, and rate-limited are indistinguishable, so probing tells an attacker nothing. 6. **No UA-based bot filtering.** Blocking on User-Agent is a coin flip and gives false confidence. The structural defence is that the beacon requires JS execution and is linked from nowhere. **Poisoning** (someone inflating a number rather than exhausting storage) cannot be fully prevented without identity. Mitigation is cross-checking: the server's own inbound page counts are an independent measure of `view.open`, so a divergence between them is the alarm. Any analysis should treat these numbers as *directional*, never as revenue-grade. --- ## 5. Privacy position The site is EU-hosted with EU visitors, so GDPR + ePrivacy apply. The public privacy page currently promises, verbatim: *"no analytics library, no cookies for tracking, and no per-visitor identifier… simple aggregate counts… with nothing tying any of it to an individual."* That is a commitment already made, and this design is built to stay inside it rather than to renegotiate it. **What the design does:** * **No identifier of any kind.** No cookie, no `localStorage`/`sessionStorage` id, no fingerprint, no hashed IP+UA. Not even the logged-in user id: the handler never reads the session cookie, and `sendBeacon` cannot attach credentials cross-origin anyway. * **Nothing is written to the visitor's device**, so ePrivacy Art. 5(3) (the "cookie rule", which covers *any* storage or access on terminal equipment, not just cookies) is not engaged at all. * **No IP is stored** by this feature — used as a per-minute rate-limit key and discarded. * **No free text, no coordinates, no URLs.** `place.pick` records *how* a location was chosen, never where. The referrer is reduced server-side to a bare registrable domain, never a full URL (which can carry someone's search query). * **Aggregate-only storage.** An hour bucket and a tuple of enums; no row per interaction, so no behavioural sequence exists to be reconstructed. * **GPC and DNT are honoured** — a visitor signalling either sends nothing. **Consequence to be honest about:** with no identifier there are **no unique visitors, no sessions, no funnels, no bounce rate, no returning-visitor rate**. `view.open` counts view opens, not people. Ratios across events (`place.pick / view.open`) are *rates*, not *conversion*, because the numerator and denominator can't be tied to the same visit. That is the price of not needing consent, and it is worth it — but it must not be quietly forgotten when someone later asks "how many users do we have?" **Legal reading (for the operator to confirm, not for me to decide):** with no identifier, no device storage, no cross-site data and aggregate-only retention, this is anonymous audience measurement — it fits the CNIL exemption criteria and the EDPB's "strictly necessary"/first-party-analytics reasoning, so **no consent banner should be required**. Transient IP processing for rate limiting rests on legitimate interest (Art. 6(1)(f)) and is security-necessary. No DPIA is triggered (no systematic monitoring, no profiling, no special categories). One line should be added to the Art. 30 record of processing. **A pre-existing issue this surfaced, unrelated to the new events.** `audit.log_access` writes the **raw client IP** for every non-static request into `logs/access/*.jsonl`, which Alloy ships to Loki with 30-day retention. That is personal data under GDPR, it is retained for 30 days fleet-wide, and it sits awkwardly beside the privacy page's *"not logged beyond the normal web-server request handling"*. **This design does not touch it and does not extend it** — but it should be decided on (§6.3) rather than left implicit, and it would be odd to ship privacy-first analytics while the access log keeps raw IPs for a month. --- ## 6. Decisions needed from the operator Nothing below is decided in the prototype. The flag stays off until these are answered. **6.1 — Identifier: none, ephemeral session, or per-user?** Recommended: **none** (what the prototype implements). *Consequence of "none":* no unique visitors, no funnels, no retention cohorts — counts of interactions, not of people. *Ephemeral session id* (random, `sessionStorage`, 30 min): enables funnels and sessions. But it is storage on the visitor's device → **ePrivacy consent banner required**, and it makes the data pseudonymous rather than anonymous (subject-access/erasure duties attach). Also requires rewriting the privacy page. *Per-user when logged in:* directly links behaviour to an identified person. Highest analytical value, highest duty, and hardest to justify for a free weather site. Recommended answer: **never**, regardless of what is chosen for anonymous visitors. **6.2 — Failed search queries: capture the text or not?** The single most useful piece of product data here is *what people searched for when we found nothing*. Three options: (a) never capture — prototype default; (b) capture **zero-result queries only**, server-side, normalised (lowercased, length-capped, rejected if it contains `@`, a digit run, or a URL-ish token), 30-day retention, never joined to anything; (c) capture all queries — not recommended, since successful queries are already answered by the `hit` counter and a search box is a free-text field a person can type anything into. Recommended: **(b), as a separately flagged follow-up**, with its own privacy-page sentence. Do not bundle it with the initial rollout. **6.3 — Raw IPs in the access log (pre-existing).** Keep as-is / truncate to /24 (IPv4) and /48 (IPv6) / hash with a daily-rotating salt / drop entirely. Recommended: **truncate**, which preserves the geo-scale and abuse-pattern uses while stopping single-device identification. This is independent of the new events but should be decided in the same pass. **6.4 — Consent banner: rely on the analytics exemption, or add one?** Recommended: **rely on the exemption** — it is only available while 6.1 stays "none" and 6.2 stays "(a) or (b)". Adding a banner unlocks identifiers but costs a real chunk of the (small) traffic to banner fatigue and adds a consent-state machine to every page. **6.5 — Retention for `ui_event_hourly`.** Set in the migration (`RETENTION_DAYS`, currently **400** — a full year-over-year comparison). Anonymous aggregates make this a storage/usefulness question rather than a GDPR one, but it should be a deliberate number. Options: 90 / 400 / indefinite. **6.6 — Privacy-page copy.** The page must keep matching reality *before* the flag goes on: the Analytics section needs a sentence naming what is counted and stating the retention. Who writes and approves that copy? **6.7 — Environment separation.** Dev and beta run the same code. Should events be recorded there at all, and if so, tagged with the environment so they never mix into prod's numbers? (A `host`/`env` label already exists in Loki; the hypertable has no such column.) Recommended: **flag on for dev only during validation, then off; prod-only after.** **6.8 — Referrer beyond `view.open`.** The prototype keeps a referrer domain only on `view.open`. Widening it to every event would allow "which referral source uses Compare" at the cost of a much larger key space. Recommended: **leave as-is** until a question needs it. **6.9 — GPC/DNT.** The prototype honours both, biasing counts slightly downward. Confirm that trade, or drop it. --- ## 7. Implementation plan ### Backend (`backend/`) | File | Change | Status in prototype | |---|---|---| | `core/events.py` | **new** — `SCHEMA`, `normalize()`, hourly upsert, JSONL sink, `ENABLED` flag | done | | `core/metrics.py` | expose `rate_ok()` so one event spends one token across both sinks; `record_event(rate_limit=False)`; raise the per-IP ceiling to 120/min and make it env-tunable; `reset_rate_limiter()` for tests | done | | `web/app.py` | `api_event`: batched payload, streaming body cap, `Sec-Fetch-Site` check, fan-out to both tiers, still always 204. `api_suggest`: record `place.search` server-side | done | | `alembic/versions/0003_ui_events.py` | **new** — `ui_event_hourly` hypertable + retention policy, Postgres-only (no-ops on the SQLite test bind, like `0002`) | done | | `accounts/api_accounts.py` | record `alert` steps server-side on subscription create/delete | **todo** | | `web/app.py` / `frontend/content.py` | record `deadend prop=not_found` on 404 responses | **todo** | | `tests/core/test_events.py` | **new** — allowlist, slot dropping, no-free-text-by-construction, flag-off inertness, never-raises | done (10 tests) | | `tests/web/test_event_beacon.py` | **new** — uniform 204, batch cap, oversize drop, cross-site drop, coordinates/free text never survive, legacy shape still counts | done (11 tests) | ### Frontend (`frontend/`) | File | Change | Status | |---|---|---| | `static/track.js` | **new** — flag check, GPC/DNT opt-out, batching + flush on `pagehide`/`visibilitychange`, `currentView()`, `[data-event]` delegation, `trackLegacy()` for the four live events | done | | `static/digest.js` | beacon moved out; re-exports `track`/`trackLegacy` so importers keep working | done | | `static/app.js` | `view.open`, `place.pick` (geolocate + picker), `deadend` (geo denied, grade error, offline), `share` (link + png), `view.control chart_metric` | done | | `static/mappicker.js` | thread a `method` ("search" vs "map") through `finish()` → `onPick(lat, lon, method)` | done | | `static/units.js` | declarative `data-event` attributes on the °F/°C toggle — no import needed | done | | `static/*.html` (6 SPA shells) | load `track.js` | done | | `templates/base.html.j2` | stamp `data-tg-events` on `` when enabled | done | | `content.py`, `app.py` | `EVENTS_ENABLED`, applied to both HTML paths (Jinja + memoized static shells) | done | | `static/calendar.js`, `compare.js`, `day.js`, `score.js`, `subscriptions.js` | `view.open` + their own `view.control` / `place.pick` / `alert` call sites | **todo** | | `templates/privacy.html.j2` | Analytics-section copy update — **blocked on 6.6** | **todo** | | `tests/unit/test_pages.py` | flag absent by default on every HTML path; stamped on both paths when enabled | done (2 tests) | ### Rollout 1. Land with `THERMOGRAPH_EVENTS` **unset everywhere**. The feature is inert: the beacon still answers 204, nothing reaches a sink, and the flag stamp is absent from every page (asserted by test). 2. Resolve §6. Update the privacy page. Nothing turns on before this. 3. Run the migration on dev; enable on **dev only**; verify in Grafana (`tag="ui.event"`) that the shape is right and that no free text or coordinate ever appears. 4. Enable on beta for a week. Compare `view.open` against the middleware's own page counts — a large divergence means bot noise or a client bug, and is the go/no-go for prod. 5. Enable on prod. Re-check the `other` bucket and the `Sec-Fetch-Site`-less bucket weekly for the first month; a growing `other` means either a schema gap or someone poking the endpoint. 6. Kill switch: unset the flag. No deploy, no migration, no data loss — existing rows just stop growing. ### Cost check ~3,000 human req/day → on the order of 1–3k events/day, batched ≈ a few hundred extra requests/day (~0.005 req/s) and a handful of upserts. Negligible against a 48 GB box; the hypertable grows by roughly a few hundred rows/day, well under 100 MB/year before compression.