thermograph/backend/daemon/internal/gateway/gateway_test.go
Emi Griffith adf824b33f
All checks were successful
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-backend (pull_request) Successful in 1m33s
PR build (required check) / gate (pull_request) Successful in 3s
daemon: move the Discord gateway and scheduler out of the web process into Go
web/app.py started two long-lived background jobs under a leader election: the
Discord gateway bot and an APScheduler. Both are stateful I/O loops -- reconnect,
RESUME, heartbeat, backoff, interval timers -- living inside an async web app
that also has to serve requests. This moves them into a single Go binary.

Go owns ONLY the stateful I/O. It owns no climate or grading logic: anything
needing data calls back into Python over a new internal-only HTTP surface
(/internal/discord/grade, /internal/jobs/warm-cities, /internal/jobs/indexnow).
Grading depends on polars and the parquet cache; reimplementing it in Go would
make the bot's grades drift from the API's, and the slash-command path
deliberately shares one grade builder so the two can never disagree. The grade
route returns gateway-ready JSON -- including the ephemeral-flag drop that
discord_bot.py used to do -- and Go relays those bytes verbatim without parsing
the embed.

Packaging: the binary is built by a golang:1.26 stage in the backend Dockerfile
and shipped in the SAME image, run as a second compose service off the SAME tag.
The daemon and backend share the /internal/* contract, so they must never skew
versions; one image makes that structural rather than a convention. Its
entrypoint bypasses entrypoint.sh -- the backend owns alembic, and two racing
migrators is a real hazard.

replicas: 1 in the Swarm stack is load-bearing. Discord permits exactly one
gateway connection per bot token; the pin replaces core/singleton.claim_leader
for this workload. update_config uses order: stop-first, since start-first would
briefly run two gateways. autoscale.sh targets ${STACK_NAME}_web only, so it
cannot scale this.

Security: the internal routes compare the token with hmac.compare_digest and the
whole router 404s when THERMOGRAPH_INTERNAL_TOKEN is unset -- fail closed, never
default open. Caddy only routes /api/*, /digest and /discord/interactions to the
backend, so /internal/* was never publicly reachable; the token is defence in
depth. The router mounts before the catch-all frontend proxy so /internal/*
cannot fall through to it. The daemon refuses to start without the token.

Behaviour preserved from the Python, with the reasoning carried into the Go
comments: non-privileged intents (no MESSAGE_CONTENT, so no portal review);
fatal close codes 4004/4010-4014 stop rather than loop; the bot-author and
self-author mention-loop guard; allowed_mentions locked to {"parse":[],
"replied_user":true} so a crafted query cannot turn a reply into an @everyone
ping; the first cron tick deferred one full interval rather than firing at boot,
since warm-cities already runs at deploy time; and no overlapping warm-cities
run, which would double-spend the archive-fetch quota.

Two deliberate improvements over the Python. A close intended for RESUME now
uses 4000 rather than 1000 -- Discord invalidates a session closed 1000/1001, so
the Python's default close silently defeated its own resume. And MESSAGE_CREATE
is handled on a bounded worker pool rather than an unbounded thread hand-off, so
a flood of mentions cannot spawn unbounded work against the backend.

A .dockerignore is added because a disposable backend/.venv was being swallowed
by COPY . /app/ and duplicated again by the chown layer, inflating the image to
1.8 GB; it builds at 578 MB.

Tests: 29 Go gateway tests covering every behaviour the deleted
test_discord_bot.py asserted, plus cron/config/apiclient suites; 10 new Python
tests for the internal routes (fail-closed, auth, flag drop, per-job 409 guard).
Full suite 359 passed / 7 skipped; go build, vet and test -race clean.
2026-07-23 15:47:07 -07:00

182 lines
5.8 KiB
Go

// Connection-machinery tests, restricted to the pure parts (backoff, fatal
// close codes, READY/session bookkeeping, sequence tracking). Deliberately no
// live-gateway test — that path is exercised in staging, not CI.
package gateway
import (
"context"
"encoding/json"
"net/http"
"testing"
"time"
"github.com/coder/websocket"
)
func TestBackoffDoublesAndCaps(t *testing.T) {
want := []time.Duration{
2 * time.Second, 4 * time.Second, 8 * time.Second, 16 * time.Second,
32 * time.Second, 60 * time.Second, 60 * time.Second, 60 * time.Second,
}
d := backoffStart
for i, w := range want {
d = nextBackoff(d)
if d != w {
t.Fatalf("step %d: want %s, got %s", i, w, d)
}
}
}
func TestFatalCloseCodes(t *testing.T) {
for _, c := range []websocket.StatusCode{4004, 4010, 4011, 4012, 4013, 4014} {
if !isFatalClose(c) {
t.Errorf("close code %d must be fatal (reconnecting loops forever)", c)
}
}
// 4008 (rate limited) and 4009 (session timed out) are resumable; 4000 is
// our own zombie close; -1 is CloseStatus's not-a-close-error sentinel.
for _, c := range []websocket.StatusCode{-1, 1000, 1001, 4000, 4001, 4008, 4009} {
if isFatalClose(c) {
t.Errorf("close code %d must not be fatal", c)
}
}
}
func TestReadyCapturesResumeState(t *testing.T) {
b := New("tok", &fakeGrader{})
sess := &session{}
d := []byte(`{"session_id":"abc","resume_gateway_url":"wss://resume.example/","user":{"id":"999"}}`)
b.dispatch(context.Background(), payload{Op: opDispatch, T: "READY", D: d}, sess)
if sess.sessionID != "abc" {
t.Fatalf("want session id abc, got %q", sess.sessionID)
}
// Trailing slash trimmed, gateway query string appended — RESUME must go
// to the session's own URL with the same v10/json parameters.
if sess.resumeURL != "wss://resume.example/?v=10&encoding=json" {
t.Fatalf("bad resume url: %q", sess.resumeURL)
}
if sess.userID != botID {
t.Fatalf("want user id %s, got %q", botID, sess.userID)
}
}
func TestReadyWithoutResumeURLLeavesItEmpty(t *testing.T) {
b := New("tok", &fakeGrader{})
sess := &session{resumeURL: "wss://stale.example/?v=10&encoding=json"}
d := []byte(`{"session_id":"abc","user":{"id":"999"}}`)
b.dispatch(context.Background(), payload{Op: opDispatch, T: "READY", D: d}, sess)
if sess.resumeURL != "" {
t.Fatalf("a READY without resume_gateway_url must clear the stale url, got %q", sess.resumeURL)
}
}
func TestMessageCreateBeforeReadyIsIgnored(t *testing.T) {
// Until READY supplies our user id we cannot detect mentions (or guard
// against answering ourselves), so nothing may be handled.
g := &fakeGrader{resp: json.RawMessage(`{}`)}
b := New("tok", g)
sess := &session{} // no userID yet
d, err := json.Marshal(testMsg("<@999> Phoenix", mentioning(botID)))
if err != nil {
t.Fatalf("marshal: %v", err)
}
b.dispatch(context.Background(), payload{Op: opDispatch, T: "MESSAGE_CREATE", D: d}, sess)
time.Sleep(20 * time.Millisecond)
if g.callCount() != 0 {
t.Fatal("pre-READY message must be ignored")
}
}
func TestPayloadSequenceIsNullable(t *testing.T) {
var withSeq, withoutSeq payload
if err := json.Unmarshal([]byte(`{"op":0,"t":"X","s":7,"d":{}}`), &withSeq); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if withSeq.S == nil || *withSeq.S != 7 {
t.Fatalf("want seq 7, got %v", withSeq.S)
}
if err := json.Unmarshal([]byte(`{"op":11,"s":null,"d":null}`), &withoutSeq); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if withoutSeq.S != nil {
t.Fatalf("null seq must stay nil, got %v", *withoutSeq.S)
}
}
func TestSessionSeqTracking(t *testing.T) {
sess := &session{}
if _, ok := sess.lastSeq(); ok {
t.Fatal("fresh session must have no sequence number")
}
sess.setSeq(41)
sess.setSeq(42)
if seq, ok := sess.lastSeq(); !ok || seq != 42 {
t.Fatalf("want seq 42, got %d (ok=%v)", seq, ok)
}
}
// A malformed HELLO must be an ERROR, not a clean "reidentify". Run treats a
// clean return as a PLANNED reconnect: backoff resets to 1s and it redials with
// no sleep at all. So if these ever regress to returning nil, a gateway stuck
// sending a bad HELLO becomes a tight reconnect loop against Discord.
func TestParseHelloRejectsMalformedSoBackoffApplies(t *testing.T) {
for _, tc := range []struct {
name string
raw string
}{
{"not json", `{`},
{"wrong op", `{"op":0,"d":{"heartbeat_interval":41250}}`},
{"missing interval", `{"op":10,"d":{}}`},
{"zero interval", `{"op":10,"d":{"heartbeat_interval":0}}`},
{"negative interval", `{"op":10,"d":{"heartbeat_interval":-1}}`},
{"interval wrong type", `{"op":10,"d":{"heartbeat_interval":"soon"}}`},
} {
t.Run(tc.name, func(t *testing.T) {
got, err := parseHello([]byte(tc.raw))
if err == nil {
t.Fatalf("parseHello(%s) = %v, nil; want an error so Run applies backoff", tc.raw, got)
}
if got != 0 {
t.Fatalf("parseHello(%s) interval = %v; want 0 on error", tc.raw, got)
}
})
}
}
func TestParseHelloAcceptsValid(t *testing.T) {
got, err := parseHello([]byte(`{"op":10,"d":{"heartbeat_interval":41250}}`))
if err != nil {
t.Fatalf("parseHello: unexpected error %v", err)
}
if want := 41250 * time.Millisecond; got != want {
t.Fatalf("interval = %v; want %v", got, want)
}
}
// Discord's advertised retry_after is clamped to 5s, matching the Python REST
// helper's _MAX_BACKOFF_S. Replies run on a small fixed worker pool, so a bogus
// server value must not park a slot for minutes.
func TestRetryAfterClamp(t *testing.T) {
for _, tc := range []struct {
name string
body string
want time.Duration
}{
{"honours a sane value", `{"retry_after":2}`, 2 * time.Second},
{"clamps a hostile value", `{"retry_after":600}`, 5 * time.Second},
{"defaults when absent", `{}`, 1 * time.Second},
} {
t.Run(tc.name, func(t *testing.T) {
if got := retryAfter(http.Header{}, []byte(tc.body)); got != tc.want {
t.Fatalf("retryAfter(%s) = %v; want %v", tc.body, got, tc.want)
}
})
}
}