thermograph/backend/daemon/internal/gateway/reply.go
Emi Griffith adf824b33f
All checks were successful
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-backend (pull_request) Successful in 1m33s
PR build (required check) / gate (pull_request) Successful in 3s
daemon: move the Discord gateway and scheduler out of the web process into Go
web/app.py started two long-lived background jobs under a leader election: the
Discord gateway bot and an APScheduler. Both are stateful I/O loops -- reconnect,
RESUME, heartbeat, backoff, interval timers -- living inside an async web app
that also has to serve requests. This moves them into a single Go binary.

Go owns ONLY the stateful I/O. It owns no climate or grading logic: anything
needing data calls back into Python over a new internal-only HTTP surface
(/internal/discord/grade, /internal/jobs/warm-cities, /internal/jobs/indexnow).
Grading depends on polars and the parquet cache; reimplementing it in Go would
make the bot's grades drift from the API's, and the slash-command path
deliberately shares one grade builder so the two can never disagree. The grade
route returns gateway-ready JSON -- including the ephemeral-flag drop that
discord_bot.py used to do -- and Go relays those bytes verbatim without parsing
the embed.

Packaging: the binary is built by a golang:1.26 stage in the backend Dockerfile
and shipped in the SAME image, run as a second compose service off the SAME tag.
The daemon and backend share the /internal/* contract, so they must never skew
versions; one image makes that structural rather than a convention. Its
entrypoint bypasses entrypoint.sh -- the backend owns alembic, and two racing
migrators is a real hazard.

replicas: 1 in the Swarm stack is load-bearing. Discord permits exactly one
gateway connection per bot token; the pin replaces core/singleton.claim_leader
for this workload. update_config uses order: stop-first, since start-first would
briefly run two gateways. autoscale.sh targets ${STACK_NAME}_web only, so it
cannot scale this.

Security: the internal routes compare the token with hmac.compare_digest and the
whole router 404s when THERMOGRAPH_INTERNAL_TOKEN is unset -- fail closed, never
default open. Caddy only routes /api/*, /digest and /discord/interactions to the
backend, so /internal/* was never publicly reachable; the token is defence in
depth. The router mounts before the catch-all frontend proxy so /internal/*
cannot fall through to it. The daemon refuses to start without the token.

Behaviour preserved from the Python, with the reasoning carried into the Go
comments: non-privileged intents (no MESSAGE_CONTENT, so no portal review);
fatal close codes 4004/4010-4014 stop rather than loop; the bot-author and
self-author mention-loop guard; allowed_mentions locked to {"parse":[],
"replied_user":true} so a crafted query cannot turn a reply into an @everyone
ping; the first cron tick deferred one full interval rather than firing at boot,
since warm-cities already runs at deploy time; and no overlapping warm-cities
run, which would double-spend the archive-fetch quota.

Two deliberate improvements over the Python. A close intended for RESUME now
uses 4000 rather than 1000 -- Discord invalidates a session closed 1000/1001, so
the Python's default close silently defeated its own resume. And MESSAGE_CREATE
is handled on a bounded worker pool rather than an unbounded thread hand-off, so
a flood of mentions cannot spawn unbounded work against the backend.

A .dockerignore is added because a disposable backend/.venv was being swallowed
by COPY . /app/ and duplicated again by the chown layer, inflating the image to
1.8 GB; it builds at 578 MB.

Tests: 29 Go gateway tests covering every behaviour the deleted
test_discord_bot.py asserted, plus cron/config/apiclient suites; 10 new Python
tests for the internal routes (fail-closed, auth, flag drop, per-job 409 guard).
Full suite 359 passed / 7 skipped; go build, vet and test -race clean.
2026-07-23 15:47:07 -07:00

138 lines
4.6 KiB
Go

// The REST half of answering a message: replies go out over plain HTTP with
// the bot token, not the gateway socket (the gateway is receive-only for us).
package gateway
import (
"bytes"
"context"
"encoding/json"
"io"
"log"
"net/http"
"strconv"
"time"
)
const discordAPIBase = "https://discord.com/api/v10"
// replyBodyLimit caps how much of a failed reply's response body we echo into
// logs.
const replyBodyLimit = 4096
// injectReplyFields adds message_reference + allowed_mentions to an otherwise
// verbatim message body. Only the top level is decoded — every value that
// came from Python (embeds and all) passes through as raw bytes, unparsed, so
// the grading/embed contract stays entirely on the Python side.
//
// allowed_mentions is a SECURITY control, not decoration: the graded reply
// echoes the user's query text, so without {"parse":[]} a crafted query could
// turn our reply into an @everyone/role ping. Only the person who asked is
// pinged, via the reply reference.
func injectReplyFields(body []byte, messageID string) ([]byte, error) {
var top map[string]json.RawMessage
if err := json.Unmarshal(body, &top); err != nil {
return nil, err
}
if messageID != "" {
ref, err := json.Marshal(map[string]string{"message_id": messageID})
if err != nil {
return nil, err
}
top["message_reference"] = ref
top["allowed_mentions"] = json.RawMessage(`{"parse":[],"replied_user":true}`)
}
return json.Marshal(top)
}
// reply posts body to the triggering message's channel as a proper reply. A
// failed reply is logged and swallowed — it must never kill the read loop.
func (b *Bot) reply(ctx context.Context, channelID, messageID string, body json.RawMessage) {
if channelID == "" {
return
}
payload, err := injectReplyFields(body, messageID)
if err != nil {
log.Printf("gateway: reply body is not a JSON object: %v", err)
return
}
url := b.rest + "/channels/" + channelID + "/messages"
for attempt := 0; ; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(payload))
if err != nil {
log.Printf("gateway: reply failed: %v", err)
return
}
req.Header.Set("Authorization", "Bot "+b.token)
req.Header.Set("Content-Type", "application/json")
resp, err := b.http.Do(req)
if err != nil {
log.Printf("gateway: reply failed: %v", err)
return
}
raw, _ := io.ReadAll(io.LimitReader(resp.Body, replyBodyLimit))
resp.Body.Close()
if resp.StatusCode == http.StatusTooManyRequests && attempt == 0 {
// One retry honouring the advertised wait — parity with the
// Python REST helper's 429 handling.
if !sleepCtx(ctx, retryAfter(resp.Header, raw)) {
return
}
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
log.Printf("gateway: reply failed: status %d: %s", resp.StatusCode, raw)
}
return
}
}
// retryAfter extracts Discord's requested wait from a 429 (JSON retry_after
// in seconds, falling back to the Retry-After header), clamped so a bogus
// server value can't park the handler for minutes.
func retryAfter(h http.Header, body []byte) time.Duration {
seconds := 1.0
var d struct {
RetryAfter float64 `json:"retry_after"`
}
if err := json.Unmarshal(body, &d); err == nil && d.RetryAfter > 0 {
seconds = d.RetryAfter
} else if v, err := strconv.ParseFloat(h.Get("Retry-After"), 64); err == nil && v > 0 {
seconds = v
}
// 5s matches the Python REST helper's _MAX_BACKOFF_S. It also bounds the cost
// of a bogus/hostile retry_after: replies run on a small fixed worker pool, so
// a parked handler holds one of very few slots and mentions start being
// dropped that much sooner.
if seconds > 5 {
seconds = 5
}
return time.Duration(seconds * float64(time.Second))
}
// handleMessage triages one MESSAGE_CREATE and sends whatever reply it calls
// for. Runs in its own goroutine (see dispatch) so a slow grade lookup can
// never stall heartbeats or the read loop.
func (b *Bot) handleMessage(ctx context.Context, m *gwMessage, botID string) {
act, query := triage(m, botID)
switch act {
case actSilent:
case actHelp:
body, err := json.Marshal(map[string]string{"content": helpText})
if err != nil {
return
}
b.reply(ctx, string(m.ChannelID), string(m.ID), body)
case actGrade:
// The callback owns all grading; its JSON is relayed VERBATIM (no
// parsing, no reshaping) so the bot's grades can never drift from
// the API's. A failed callback stays silent — better no reply than a
// made-up one.
raw, err := b.api.Grade(ctx, query)
if err != nil {
log.Printf("gateway: grade callback failed for %q: %v", query, err)
return
}
b.reply(ctx, string(m.ChannelID), string(m.ID), raw)
}
}