All checks were successful
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / build-backend (pull_request) Successful in 1m18s
PR build (required check) / changes (pull_request) Successful in 6s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 3s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / validate-observability (pull_request) Has been skipped
web/app.py started two long-lived background jobs under a leader election: the
Discord gateway bot and an APScheduler. Both are stateful I/O loops -- reconnect,
RESUME, heartbeat, backoff, interval timers -- living inside an async web app
that also has to serve requests. This moves them into a single Go binary.
Go owns ONLY the stateful I/O. It owns no climate or grading logic: anything
needing data calls back into Python over a new internal-only HTTP surface
(/internal/discord/grade, /internal/jobs/warm-cities, /internal/jobs/indexnow).
Grading depends on polars and the parquet cache; reimplementing it in Go would
make the bot's grades drift from the API's, and the slash-command path
deliberately shares one grade builder so the two can never disagree. The grade
route returns gateway-ready JSON -- including the ephemeral-flag drop that
discord_bot.py used to do -- and Go relays those bytes verbatim without parsing
the embed.
Packaging: the binary is built by a golang:1.26 stage in the backend Dockerfile
and shipped in the SAME image, run as a second compose service off the SAME tag.
The daemon and backend share the /internal/* contract, so they must never skew
versions; one image makes that structural rather than a convention. Its
entrypoint bypasses entrypoint.sh -- the backend owns alembic, and two racing
migrators is a real hazard.
replicas: 1 in the Swarm stack is load-bearing. Discord permits exactly one
gateway connection per bot token; the pin replaces core/singleton.claim_leader
for this workload. update_config uses order: stop-first, since start-first would
briefly run two gateways. autoscale.sh targets ${STACK_NAME}_web only, so it
cannot scale this.
Security: the internal routes compare the token with hmac.compare_digest and the
whole router 404s when THERMOGRAPH_INTERNAL_TOKEN is unset -- fail closed, never
default open. Caddy only routes /api/*, /digest and /discord/interactions to the
backend, so /internal/* was never publicly reachable; the token is defence in
depth. The router mounts before the catch-all frontend proxy so /internal/*
cannot fall through to it. The daemon refuses to start without the token.
Behaviour preserved from the Python, with the reasoning carried into the Go
comments: non-privileged intents (no MESSAGE_CONTENT, so no portal review);
fatal close codes 4004/4010-4014 stop rather than loop; the bot-author and
self-author mention-loop guard; allowed_mentions locked to {"parse":[],
"replied_user":true} so a crafted query cannot turn a reply into an @everyone
ping; the first cron tick deferred one full interval rather than firing at boot,
since warm-cities already runs at deploy time; and no overlapping warm-cities
run, which would double-spend the archive-fetch quota.
Two deliberate improvements over the Python. A close intended for RESUME now
uses 4000 rather than 1000 -- Discord invalidates a session closed 1000/1001, so
the Python's default close silently defeated its own resume. And MESSAGE_CREATE
is handled on a bounded worker pool rather than an unbounded thread hand-off, so
a flood of mentions cannot spawn unbounded work against the backend.
A .dockerignore is added because a disposable backend/.venv was being swallowed
by COPY . /app/ and duplicated again by the chown layer, inflating the image to
1.8 GB; it builds at 578 MB.
Tests: 29 Go gateway tests covering every behaviour the deleted
test_discord_bot.py asserted, plus cron/config/apiclient suites; 10 new Python
tests for the internal routes (fail-closed, auth, flag drop, per-job 409 guard).
Full suite 359 passed / 7 skipped; go build, vet and test -race clean.
134 lines
5.6 KiB
Python
134 lines
5.6 KiB
Python
"""The internal control surface the thermograph-daemon Go binary calls back
|
|
into (api/internal_routes.py): shared-secret gating (fail closed when the token
|
|
was never provisioned), the gateway-ready /grade reply with the interactions-only
|
|
ephemeral flag dropped, the two job triggers, and the one-in-flight-run guard.
|
|
The underlying job/grade functions are stubbed — hermetic, no real network."""
|
|
import pytest
|
|
from fastapi.testclient import TestClient
|
|
|
|
from api import internal_routes
|
|
from notifications import discord_interactions as di
|
|
from web import app as appmod
|
|
|
|
TOKEN = "test-internal-token"
|
|
HDR = {"X-Thermograph-Internal-Token": TOKEN}
|
|
|
|
|
|
@pytest.fixture
|
|
def client(monkeypatch):
|
|
monkeypatch.setenv("THERMOGRAPH_INTERNAL_TOKEN", TOKEN)
|
|
return TestClient(appmod.app)
|
|
|
|
|
|
@pytest.fixture
|
|
def grade_stub(monkeypatch):
|
|
"""Reply payload for the shared grade builder, mirroring the real shape
|
|
(embeds + an ephemeral flag the gateway path must drop)."""
|
|
calls = []
|
|
monkeypatch.setattr(
|
|
di, "_grade_message",
|
|
lambda query: calls.append(query) or {"embeds": [{"title": f"grade:{query}"}],
|
|
"flags": 64})
|
|
return calls
|
|
|
|
|
|
# --- auth: fail closed, then reject before accept ----------------------------
|
|
|
|
def test_router_is_disabled_when_no_token_is_provisioned(monkeypatch):
|
|
"""No THERMOGRAPH_INTERNAL_TOKEN => the surface doesn't exist (404), even
|
|
for a caller presenting a header — never fall open to no-auth."""
|
|
monkeypatch.delenv("THERMOGRAPH_INTERNAL_TOKEN", raising=False)
|
|
client = TestClient(appmod.app)
|
|
assert client.post("/internal/discord/grade", json={"query": "x"},
|
|
headers=HDR).status_code == 404
|
|
assert client.post("/internal/jobs/warm-cities", json={}, headers=HDR).status_code == 404
|
|
assert client.post("/internal/jobs/indexnow", json={}, headers=HDR).status_code == 404
|
|
|
|
|
|
def test_missing_header_is_rejected(client):
|
|
assert client.post("/internal/jobs/indexnow", json={}).status_code == 401
|
|
|
|
|
|
def test_wrong_token_is_rejected(client):
|
|
r = client.post("/internal/jobs/indexnow", json={},
|
|
headers={"X-Thermograph-Internal-Token": "not-the-token"})
|
|
assert r.status_code == 401
|
|
|
|
|
|
def test_correct_token_is_accepted(client, monkeypatch):
|
|
import indexnow
|
|
monkeypatch.setattr(indexnow, "submit_if_changed", lambda: None)
|
|
assert client.post("/internal/jobs/indexnow", json={}, headers=HDR).status_code == 200
|
|
|
|
|
|
# --- /internal/discord/grade -------------------------------------------------
|
|
|
|
def test_grade_returns_the_shared_builder_message_without_ephemeral_flag(client, grade_stub):
|
|
r = client.post("/internal/discord/grade", json={"query": "Phoenix"}, headers=HDR)
|
|
assert r.status_code == 200
|
|
# Gateway messages can't be ephemeral — the flag must be gone, everything
|
|
# else exactly as _grade_message built it.
|
|
assert r.json() == {"embeds": [{"title": "grade:Phoenix"}]}
|
|
assert grade_stub == ["Phoenix"]
|
|
|
|
|
|
def test_grade_requires_a_query(client, grade_stub):
|
|
assert client.post("/internal/discord/grade", json={}, headers=HDR).status_code == 422
|
|
assert grade_stub == []
|
|
|
|
|
|
# --- job triggers ------------------------------------------------------------
|
|
|
|
def test_warm_cities_invokes_warm_cities_main(client, monkeypatch):
|
|
import warm_cities
|
|
calls = []
|
|
monkeypatch.setattr(warm_cities, "main", lambda: calls.append(1))
|
|
r = client.post("/internal/jobs/warm-cities", json={}, headers=HDR)
|
|
assert r.status_code == 200
|
|
assert r.json() == {"ok": True}
|
|
assert calls == [1]
|
|
|
|
|
|
def test_indexnow_invokes_submit_if_changed(client, monkeypatch):
|
|
import indexnow
|
|
calls = []
|
|
monkeypatch.setattr(indexnow, "submit_if_changed", lambda: calls.append(1))
|
|
r = client.post("/internal/jobs/indexnow", json={}, headers=HDR)
|
|
assert r.status_code == 200
|
|
assert r.json() == {"ok": True}
|
|
assert calls == [1]
|
|
|
|
|
|
# --- one-in-flight-run guard -------------------------------------------------
|
|
|
|
def test_concurrent_job_invocation_answers_409(client, monkeypatch):
|
|
"""A second trigger while a run is in flight must not stack — warm_cities
|
|
spends the Open-Meteo quota, so a stacked run double-spends it. Holding the
|
|
job's lock stands in for an in-flight run."""
|
|
import indexnow
|
|
monkeypatch.setattr(indexnow, "submit_if_changed", lambda: None)
|
|
lock = internal_routes._JOB_LOCKS["warm-cities"]
|
|
assert lock.acquire(blocking=False)
|
|
try:
|
|
assert client.post("/internal/jobs/warm-cities", json={}, headers=HDR).status_code == 409
|
|
# The guard is per job: an in-flight warm run doesn't block indexnow.
|
|
assert client.post("/internal/jobs/indexnow", json={}, headers=HDR).status_code == 200
|
|
finally:
|
|
lock.release()
|
|
# Released => the next trigger runs again.
|
|
import warm_cities
|
|
monkeypatch.setattr(warm_cities, "main", lambda: None)
|
|
assert client.post("/internal/jobs/warm-cities", json={}, headers=HDR).status_code == 200
|
|
|
|
|
|
def test_job_lock_is_released_after_a_failing_run(monkeypatch):
|
|
"""A crashed run must not wedge the job forever: the 500 surfaces, the lock
|
|
is released, and the next trigger runs."""
|
|
monkeypatch.setenv("THERMOGRAPH_INTERNAL_TOKEN", TOKEN)
|
|
client = TestClient(appmod.app, raise_server_exceptions=False)
|
|
import indexnow
|
|
monkeypatch.setattr(indexnow, "submit_if_changed",
|
|
lambda: (_ for _ in ()).throw(RuntimeError("boom")))
|
|
assert client.post("/internal/jobs/indexnow", json={}, headers=HDR).status_code == 500
|
|
monkeypatch.setattr(indexnow, "submit_if_changed", lambda: None)
|
|
assert client.post("/internal/jobs/indexnow", json={}, headers=HDR).status_code == 200
|