thermograph/backend/daemon/internal/gateway/reply.go
emi 2d3f37c474
All checks were successful
Sync infra to hosts / sync-beta (push) Successful in 8s
Sync infra to hosts / sync-prod (push) Successful in 7s
secrets-guard / encrypted (push) Successful in 8s
shell-lint / shellcheck (push) Successful in 10s
Build + push backend image (Forgejo registry) / build-push (push) Successful in 1m14s
Deploy backend to beta VPS / deploy (push) Successful in 2m4s
daemon: move the Discord gateway and scheduler out of the web process into Go (#21)
The gateway bot and APScheduler were long-lived stateful I/O loops running
inside the async web app under a leader election. They move into a single Go
binary that owns ONLY that I/O -- websocket, RESUME, heartbeat, backoff, timers.

It owns no grading logic. Anything needing data calls back over a new
internal-only surface (/internal/discord/grade, /internal/jobs/*). Grading
depends on polars and the parquet cache; reimplementing it in Go would let the
bot's grades drift from the API's. The grade route returns gateway-ready JSON
and Go relays the bytes verbatim.

The binary ships in the backend image and runs as a second compose service off
the same tag, so the two ends of the /internal/* contract can never skew.
deploy.sh rolls daemon alongside backend -- without that the service would never
be created, since a single-service deploy uses --no-deps. It also probes the
image first and skips the daemon when rolling a tag that predates the binary:
infra tracks main while image tags are env-staged, so a host can legitimately be
asked to roll an older backend image, and creating the service anyway would
leave a container crash-looping on a missing binary.

replicas: 1 with order: stop-first replaces the leader election -- Discord
permits one gateway connection per bot token.

THERMOGRAPH_INTERNAL_TOKEN is optional: both ends derive it from
THERMOGRAPH_AUTH_SECRET via HMAC under a domain-separation label, so this needs
no new vault entry. The derivation is pinned to a shared cross-language test
vector asserted on both sides, so drift fails CI instead of 401ing every call.
Fail closed when neither secret is set.

Improvements over the Python: a close intended for RESUME uses 4000 rather than
1000 (Discord invalidates a session closed 1000, so the old default defeated its
own resume); MESSAGE_CREATE runs on a bounded worker pool; and a malformed HELLO
returns an error rather than a clean reconnect, which would otherwise reset
backoff and hot-loop against the gateway.

365 Python tests pass; Go build/vet/test -race clean; shellcheck 0 findings.
2026-07-23 22:49:54 +00:00

138 lines
4.6 KiB
Go

// The REST half of answering a message: replies go out over plain HTTP with
// the bot token, not the gateway socket (the gateway is receive-only for us).
package gateway
import (
"bytes"
"context"
"encoding/json"
"io"
"log"
"net/http"
"strconv"
"time"
)
const discordAPIBase = "https://discord.com/api/v10"
// replyBodyLimit caps how much of a failed reply's response body we echo into
// logs.
const replyBodyLimit = 4096
// injectReplyFields adds message_reference + allowed_mentions to an otherwise
// verbatim message body. Only the top level is decoded — every value that
// came from Python (embeds and all) passes through as raw bytes, unparsed, so
// the grading/embed contract stays entirely on the Python side.
//
// allowed_mentions is a SECURITY control, not decoration: the graded reply
// echoes the user's query text, so without {"parse":[]} a crafted query could
// turn our reply into an @everyone/role ping. Only the person who asked is
// pinged, via the reply reference.
func injectReplyFields(body []byte, messageID string) ([]byte, error) {
var top map[string]json.RawMessage
if err := json.Unmarshal(body, &top); err != nil {
return nil, err
}
if messageID != "" {
ref, err := json.Marshal(map[string]string{"message_id": messageID})
if err != nil {
return nil, err
}
top["message_reference"] = ref
top["allowed_mentions"] = json.RawMessage(`{"parse":[],"replied_user":true}`)
}
return json.Marshal(top)
}
// reply posts body to the triggering message's channel as a proper reply. A
// failed reply is logged and swallowed — it must never kill the read loop.
func (b *Bot) reply(ctx context.Context, channelID, messageID string, body json.RawMessage) {
if channelID == "" {
return
}
payload, err := injectReplyFields(body, messageID)
if err != nil {
log.Printf("gateway: reply body is not a JSON object: %v", err)
return
}
url := b.rest + "/channels/" + channelID + "/messages"
for attempt := 0; ; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(payload))
if err != nil {
log.Printf("gateway: reply failed: %v", err)
return
}
req.Header.Set("Authorization", "Bot "+b.token)
req.Header.Set("Content-Type", "application/json")
resp, err := b.http.Do(req)
if err != nil {
log.Printf("gateway: reply failed: %v", err)
return
}
raw, _ := io.ReadAll(io.LimitReader(resp.Body, replyBodyLimit))
resp.Body.Close()
if resp.StatusCode == http.StatusTooManyRequests && attempt == 0 {
// One retry honouring the advertised wait — parity with the
// Python REST helper's 429 handling.
if !sleepCtx(ctx, retryAfter(resp.Header, raw)) {
return
}
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
log.Printf("gateway: reply failed: status %d: %s", resp.StatusCode, raw)
}
return
}
}
// retryAfter extracts Discord's requested wait from a 429 (JSON retry_after
// in seconds, falling back to the Retry-After header), clamped so a bogus
// server value can't park the handler for minutes.
func retryAfter(h http.Header, body []byte) time.Duration {
seconds := 1.0
var d struct {
RetryAfter float64 `json:"retry_after"`
}
if err := json.Unmarshal(body, &d); err == nil && d.RetryAfter > 0 {
seconds = d.RetryAfter
} else if v, err := strconv.ParseFloat(h.Get("Retry-After"), 64); err == nil && v > 0 {
seconds = v
}
// 5s matches the Python REST helper's _MAX_BACKOFF_S. It also bounds the cost
// of a bogus/hostile retry_after: replies run on a small fixed worker pool, so
// a parked handler holds one of very few slots and mentions start being
// dropped that much sooner.
if seconds > 5 {
seconds = 5
}
return time.Duration(seconds * float64(time.Second))
}
// handleMessage triages one MESSAGE_CREATE and sends whatever reply it calls
// for. Runs in its own goroutine (see dispatch) so a slow grade lookup can
// never stall heartbeats or the read loop.
func (b *Bot) handleMessage(ctx context.Context, m *gwMessage, botID string) {
act, query := triage(m, botID)
switch act {
case actSilent:
case actHelp:
body, err := json.Marshal(map[string]string{"content": helpText})
if err != nil {
return
}
b.reply(ctx, string(m.ChannelID), string(m.ID), body)
case actGrade:
// The callback owns all grading; its JSON is relayed VERBATIM (no
// parsing, no reshaping) so the bot's grades can never drift from
// the API's. A failed callback stays silent — better no reply than a
// made-up one.
raw, err := b.api.Grade(ctx, query)
if err != nil {
log.Printf("gateway: grade callback failed for %q: %v", query, err)
return
}
b.reply(ctx, string(m.ChannelID), string(m.ID), raw)
}
}