* Split web/worker duties with THERMOGRAPH_ROLE Background work (the subscription notifier) is welded to the same process that serves requests, so scaling the web tier to N replicas would also scale notifier instances unless something restricts it further than leader election alone. Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process behavior). Every replica runs the same image; ROLE only gates whether a process is allowed to own the notifier at all, layered on top of the existing leader election: web replicas never start it even if they'd win leader election, worker replicas start it if they win. The decision is pulled into _should_run_notifier() so it's unit-testable without booting the full app (DB init, places index, neighbor warmer). Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE) so a worker replica - which serves no real traffic - still has something Swarm can health-check. * Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate Three changes toward the hop-1 interim cutover, all inert until Track B stands up the platform: docker-stack.yml: the Swarm stack file for the interim cutover, distinct from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a pre-built image (IMAGE_TAG) instead of building in place; app/worker publish no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted), reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard needs a smaller MTU or large payloads silently stall); db is placement- pinned to a labelled node; app/worker skip inline migrations (RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing that ever runs Alembic; secrets are real Swarm secrets mounted at /run/secrets, read by the entrypoint shim rather than plain env vars. TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default latest-pg18, today's behavior unchanged), wired through Terraform (timescaledb_tag, default "latest-pg18") so it can actually be pinned to an exact minor without hand-editing the host - required before any host of the stack could replicate with another (a floating tag risks mismatched extension minors, which blocks a physical replica and risks compressed- chunk corruption on restore). Caddy active health-gate: both the Terraform-rendered Caddyfile and the live deploy/Caddyfile now health-check the app on the same cheap /healthz route its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage, so it's cheap enough for a tight interval and works identically for a worker replica, which serves no public traffic at all) - Caddy won't forward into a container that's still booting or unhealthy. Verified live: built and booted the real image via docker compose - both containers report healthy via the new /healthz-based HEALTHCHECK, and GET / still renders the full SSR homepage unchanged. Both Caddyfiles validated with the real caddy binary. docker-stack.yml validated with docker compose config (required-var guards fire with clear messages; secrets correctly mount at /run/secrets/<name>, matching the entrypoint shim's mapping). docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside the existing openmeteo overlay. terraform validate + fmt clean.
34 lines
1.3 KiB
Python
34 lines
1.3 KiB
Python
"""THERMOGRAPH_ROLE: which processes are allowed to own the subscription
|
|
notifier (web/worker/all split), plus the /healthz liveness route."""
|
|
from fastapi.testclient import TestClient
|
|
|
|
from core import singleton
|
|
from web import app as appmod
|
|
|
|
|
|
def test_default_role_is_all_and_permits_the_notifier(monkeypatch):
|
|
monkeypatch.setattr(appmod, "ROLE", "all")
|
|
monkeypatch.setattr(singleton, "claim_leader", lambda: True)
|
|
assert appmod._should_run_notifier() is True
|
|
|
|
|
|
def test_web_role_never_runs_the_notifier_even_if_it_would_win_leader(monkeypatch):
|
|
monkeypatch.setattr(appmod, "ROLE", "web")
|
|
monkeypatch.setattr(singleton, "claim_leader", lambda: True)
|
|
assert appmod._should_run_notifier() is False
|
|
|
|
|
|
def test_worker_role_runs_the_notifier_only_if_it_wins_leader(monkeypatch):
|
|
monkeypatch.setattr(appmod, "ROLE", "worker")
|
|
monkeypatch.setattr(singleton, "claim_leader", lambda: False)
|
|
assert appmod._should_run_notifier() is False
|
|
monkeypatch.setattr(singleton, "claim_leader", lambda: True)
|
|
assert appmod._should_run_notifier() is True
|
|
|
|
|
|
def test_healthz_reports_status_and_configured_role(monkeypatch):
|
|
monkeypatch.setattr(appmod, "ROLE", "worker")
|
|
client = TestClient(appmod.app)
|
|
r = client.get("/healthz")
|
|
assert r.status_code == 200
|
|
assert r.json() == {"status": "ok", "role": "worker"}
|