All checks were successful
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / build-backend (pull_request) Successful in 1m18s
PR build (required check) / changes (pull_request) Successful in 6s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 3s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / validate-observability (pull_request) Has been skipped
web/app.py started two long-lived background jobs under a leader election: the
Discord gateway bot and an APScheduler. Both are stateful I/O loops -- reconnect,
RESUME, heartbeat, backoff, interval timers -- living inside an async web app
that also has to serve requests. This moves them into a single Go binary.
Go owns ONLY the stateful I/O. It owns no climate or grading logic: anything
needing data calls back into Python over a new internal-only HTTP surface
(/internal/discord/grade, /internal/jobs/warm-cities, /internal/jobs/indexnow).
Grading depends on polars and the parquet cache; reimplementing it in Go would
make the bot's grades drift from the API's, and the slash-command path
deliberately shares one grade builder so the two can never disagree. The grade
route returns gateway-ready JSON -- including the ephemeral-flag drop that
discord_bot.py used to do -- and Go relays those bytes verbatim without parsing
the embed.
Packaging: the binary is built by a golang:1.26 stage in the backend Dockerfile
and shipped in the SAME image, run as a second compose service off the SAME tag.
The daemon and backend share the /internal/* contract, so they must never skew
versions; one image makes that structural rather than a convention. Its
entrypoint bypasses entrypoint.sh -- the backend owns alembic, and two racing
migrators is a real hazard.
replicas: 1 in the Swarm stack is load-bearing. Discord permits exactly one
gateway connection per bot token; the pin replaces core/singleton.claim_leader
for this workload. update_config uses order: stop-first, since start-first would
briefly run two gateways. autoscale.sh targets ${STACK_NAME}_web only, so it
cannot scale this.
Security: the internal routes compare the token with hmac.compare_digest and the
whole router 404s when THERMOGRAPH_INTERNAL_TOKEN is unset -- fail closed, never
default open. Caddy only routes /api/*, /digest and /discord/interactions to the
backend, so /internal/* was never publicly reachable; the token is defence in
depth. The router mounts before the catch-all frontend proxy so /internal/*
cannot fall through to it. The daemon refuses to start without the token.
Behaviour preserved from the Python, with the reasoning carried into the Go
comments: non-privileged intents (no MESSAGE_CONTENT, so no portal review);
fatal close codes 4004/4010-4014 stop rather than loop; the bot-author and
self-author mention-loop guard; allowed_mentions locked to {"parse":[],
"replied_user":true} so a crafted query cannot turn a reply into an @everyone
ping; the first cron tick deferred one full interval rather than firing at boot,
since warm-cities already runs at deploy time; and no overlapping warm-cities
run, which would double-spend the archive-fetch quota.
Two deliberate improvements over the Python. A close intended for RESUME now
uses 4000 rather than 1000 -- Discord invalidates a session closed 1000/1001, so
the Python's default close silently defeated its own resume. And MESSAGE_CREATE
is handled on a bounded worker pool rather than an unbounded thread hand-off, so
a flood of mentions cannot spawn unbounded work against the backend.
A .dockerignore is added because a disposable backend/.venv was being swallowed
by COPY . /app/ and duplicated again by the chown layer, inflating the image to
1.8 GB; it builds at 578 MB.
Tests: 29 Go gateway tests covering every behaviour the deleted
test_discord_bot.py asserted, plus cron/config/apiclient suites; 10 new Python
tests for the internal routes (fail-closed, auth, flag drop, per-job 409 guard).
Full suite 359 passed / 7 skipped; go build, vet and test -race clean.
138 lines
4.6 KiB
Go
138 lines
4.6 KiB
Go
// The REST half of answering a message: replies go out over plain HTTP with
|
|
// the bot token, not the gateway socket (the gateway is receive-only for us).
|
|
|
|
package gateway
|
|
|
|
import (
|
|
"bytes"
|
|
"context"
|
|
"encoding/json"
|
|
"io"
|
|
"log"
|
|
"net/http"
|
|
"strconv"
|
|
"time"
|
|
)
|
|
|
|
const discordAPIBase = "https://discord.com/api/v10"
|
|
|
|
// replyBodyLimit caps how much of a failed reply's response body we echo into
|
|
// logs.
|
|
const replyBodyLimit = 4096
|
|
|
|
// injectReplyFields adds message_reference + allowed_mentions to an otherwise
|
|
// verbatim message body. Only the top level is decoded — every value that
|
|
// came from Python (embeds and all) passes through as raw bytes, unparsed, so
|
|
// the grading/embed contract stays entirely on the Python side.
|
|
//
|
|
// allowed_mentions is a SECURITY control, not decoration: the graded reply
|
|
// echoes the user's query text, so without {"parse":[]} a crafted query could
|
|
// turn our reply into an @everyone/role ping. Only the person who asked is
|
|
// pinged, via the reply reference.
|
|
func injectReplyFields(body []byte, messageID string) ([]byte, error) {
|
|
var top map[string]json.RawMessage
|
|
if err := json.Unmarshal(body, &top); err != nil {
|
|
return nil, err
|
|
}
|
|
if messageID != "" {
|
|
ref, err := json.Marshal(map[string]string{"message_id": messageID})
|
|
if err != nil {
|
|
return nil, err
|
|
}
|
|
top["message_reference"] = ref
|
|
top["allowed_mentions"] = json.RawMessage(`{"parse":[],"replied_user":true}`)
|
|
}
|
|
return json.Marshal(top)
|
|
}
|
|
|
|
// reply posts body to the triggering message's channel as a proper reply. A
|
|
// failed reply is logged and swallowed — it must never kill the read loop.
|
|
func (b *Bot) reply(ctx context.Context, channelID, messageID string, body json.RawMessage) {
|
|
if channelID == "" {
|
|
return
|
|
}
|
|
payload, err := injectReplyFields(body, messageID)
|
|
if err != nil {
|
|
log.Printf("gateway: reply body is not a JSON object: %v", err)
|
|
return
|
|
}
|
|
url := b.rest + "/channels/" + channelID + "/messages"
|
|
for attempt := 0; ; attempt++ {
|
|
req, err := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(payload))
|
|
if err != nil {
|
|
log.Printf("gateway: reply failed: %v", err)
|
|
return
|
|
}
|
|
req.Header.Set("Authorization", "Bot "+b.token)
|
|
req.Header.Set("Content-Type", "application/json")
|
|
resp, err := b.http.Do(req)
|
|
if err != nil {
|
|
log.Printf("gateway: reply failed: %v", err)
|
|
return
|
|
}
|
|
raw, _ := io.ReadAll(io.LimitReader(resp.Body, replyBodyLimit))
|
|
resp.Body.Close()
|
|
if resp.StatusCode == http.StatusTooManyRequests && attempt == 0 {
|
|
// One retry honouring the advertised wait — parity with the
|
|
// Python REST helper's 429 handling.
|
|
if !sleepCtx(ctx, retryAfter(resp.Header, raw)) {
|
|
return
|
|
}
|
|
continue
|
|
}
|
|
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
|
|
log.Printf("gateway: reply failed: status %d: %s", resp.StatusCode, raw)
|
|
}
|
|
return
|
|
}
|
|
}
|
|
|
|
// retryAfter extracts Discord's requested wait from a 429 (JSON retry_after
|
|
// in seconds, falling back to the Retry-After header), clamped so a bogus
|
|
// server value can't park the handler for minutes.
|
|
func retryAfter(h http.Header, body []byte) time.Duration {
|
|
seconds := 1.0
|
|
var d struct {
|
|
RetryAfter float64 `json:"retry_after"`
|
|
}
|
|
if err := json.Unmarshal(body, &d); err == nil && d.RetryAfter > 0 {
|
|
seconds = d.RetryAfter
|
|
} else if v, err := strconv.ParseFloat(h.Get("Retry-After"), 64); err == nil && v > 0 {
|
|
seconds = v
|
|
}
|
|
// 5s matches the Python REST helper's _MAX_BACKOFF_S. It also bounds the cost
|
|
// of a bogus/hostile retry_after: replies run on a small fixed worker pool, so
|
|
// a parked handler holds one of very few slots and mentions start being
|
|
// dropped that much sooner.
|
|
if seconds > 5 {
|
|
seconds = 5
|
|
}
|
|
return time.Duration(seconds * float64(time.Second))
|
|
}
|
|
|
|
// handleMessage triages one MESSAGE_CREATE and sends whatever reply it calls
|
|
// for. Runs in its own goroutine (see dispatch) so a slow grade lookup can
|
|
// never stall heartbeats or the read loop.
|
|
func (b *Bot) handleMessage(ctx context.Context, m *gwMessage, botID string) {
|
|
act, query := triage(m, botID)
|
|
switch act {
|
|
case actSilent:
|
|
case actHelp:
|
|
body, err := json.Marshal(map[string]string{"content": helpText})
|
|
if err != nil {
|
|
return
|
|
}
|
|
b.reply(ctx, string(m.ChannelID), string(m.ID), body)
|
|
case actGrade:
|
|
// The callback owns all grading; its JSON is relayed VERBATIM (no
|
|
// parsing, no reshaping) so the bot's grades can never drift from
|
|
// the API's. A failed callback stays silent — better no reply than a
|
|
// made-up one.
|
|
raw, err := b.api.Grade(ctx, query)
|
|
if err != nil {
|
|
log.Printf("gateway: grade callback failed for %q: %v", query, err)
|
|
return
|
|
}
|
|
b.reply(ctx, string(m.ChannelID), string(m.ID), raw)
|
|
}
|
|
}
|