thermograph/backend/daemon/internal/gateway/message.go
Emi Griffith c3b906a9ed
All checks were successful
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / build-backend (pull_request) Successful in 1m18s
PR build (required check) / changes (pull_request) Successful in 6s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 3s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / validate-observability (pull_request) Has been skipped
daemon: move the Discord gateway and scheduler out of the web process into Go
web/app.py started two long-lived background jobs under a leader election: the
Discord gateway bot and an APScheduler. Both are stateful I/O loops -- reconnect,
RESUME, heartbeat, backoff, interval timers -- living inside an async web app
that also has to serve requests. This moves them into a single Go binary.

Go owns ONLY the stateful I/O. It owns no climate or grading logic: anything
needing data calls back into Python over a new internal-only HTTP surface
(/internal/discord/grade, /internal/jobs/warm-cities, /internal/jobs/indexnow).
Grading depends on polars and the parquet cache; reimplementing it in Go would
make the bot's grades drift from the API's, and the slash-command path
deliberately shares one grade builder so the two can never disagree. The grade
route returns gateway-ready JSON -- including the ephemeral-flag drop that
discord_bot.py used to do -- and Go relays those bytes verbatim without parsing
the embed.

Packaging: the binary is built by a golang:1.26 stage in the backend Dockerfile
and shipped in the SAME image, run as a second compose service off the SAME tag.
The daemon and backend share the /internal/* contract, so they must never skew
versions; one image makes that structural rather than a convention. Its
entrypoint bypasses entrypoint.sh -- the backend owns alembic, and two racing
migrators is a real hazard.

replicas: 1 in the Swarm stack is load-bearing. Discord permits exactly one
gateway connection per bot token; the pin replaces core/singleton.claim_leader
for this workload. update_config uses order: stop-first, since start-first would
briefly run two gateways. autoscale.sh targets ${STACK_NAME}_web only, so it
cannot scale this.

Security: the internal routes compare the token with hmac.compare_digest and the
whole router 404s when THERMOGRAPH_INTERNAL_TOKEN is unset -- fail closed, never
default open. Caddy only routes /api/*, /digest and /discord/interactions to the
backend, so /internal/* was never publicly reachable; the token is defence in
depth. The router mounts before the catch-all frontend proxy so /internal/*
cannot fall through to it. The daemon refuses to start without the token.

Behaviour preserved from the Python, with the reasoning carried into the Go
comments: non-privileged intents (no MESSAGE_CONTENT, so no portal review);
fatal close codes 4004/4010-4014 stop rather than loop; the bot-author and
self-author mention-loop guard; allowed_mentions locked to {"parse":[],
"replied_user":true} so a crafted query cannot turn a reply into an @everyone
ping; the first cron tick deferred one full interval rather than firing at boot,
since warm-cities already runs at deploy time; and no overlapping warm-cities
run, which would double-spend the archive-fetch quota.

Two deliberate improvements over the Python. A close intended for RESUME now
uses 4000 rather than 1000 -- Discord invalidates a session closed 1000/1001, so
the Python's default close silently defeated its own resume. And MESSAGE_CREATE
is handled on a bounded worker pool rather than an unbounded thread hand-off, so
a flood of mentions cannot spawn unbounded work against the backend.

A .dockerignore is added because a disposable backend/.venv was being swallowed
by COPY . /app/ and duplicated again by the chown layer, inflating the image to
1.8 GB; it builds at 578 MB.

Tests: 29 Go gateway tests covering every behaviour the deleted
test_discord_bot.py asserted, plus cron/config/apiclient suites; 10 new Python
tests for the internal routes (fail-closed, auth, flag drop, per-job 409 guard).
Full suite 359 passed / 7 skipped; go build, vet and test -race clean.
2026-07-23 15:42:44 -07:00

118 lines
3.5 KiB
Go

// Message triage: the pure decision of whether — and how — to answer one
// MESSAGE_CREATE. Ported from the Python bot's _response_for_message /
// _mentions_bot / _is_dm / _strip_mentions; it is plain string/JSON logic
// with no data dependency, so it lives in Go. Anything needing climate data
// (the actual grade) goes back to Python via Grader.
package gateway
import (
"encoding/json"
"fmt"
"strings"
)
// helpText answers an empty query. Go owns this constant because it needs no
// data; the wording matches the Python original exactly.
const helpText = "Mention me with a city name — e.g. `@Thermograph Phoenix` — and I'll grade " +
"today's weather against ~45 years of local history."
// snowflake decodes a Discord id that may arrive as a JSON string (what v10
// sends) or a bare number. The Python compared ids through str() coercion;
// this keeps that tolerance instead of betting the payload shape never
// wobbles.
type snowflake string
func (s *snowflake) UnmarshalJSON(b []byte) error {
if string(b) == "null" {
*s = ""
return nil
}
var str string
if err := json.Unmarshal(b, &str); err == nil {
*s = snowflake(str)
return nil
}
var n json.Number
if err := json.Unmarshal(b, &n); err == nil {
*s = snowflake(n.String())
return nil
}
return fmt.Errorf("snowflake: cannot decode %s", b)
}
// gwMessage is the slice of a MESSAGE_CREATE payload the triage needs.
type gwMessage struct {
ID snowflake `json:"id"`
ChannelID snowflake `json:"channel_id"`
GuildID snowflake `json:"guild_id"`
Content string `json:"content"`
Author gwAuthor `json:"author"`
Mentions []gwAuthor `json:"mentions"`
}
type gwAuthor struct {
ID snowflake `json:"id"`
Bot bool `json:"bot"`
}
// action is triage's verdict for one message.
type action int
const (
actSilent action = iota // no reply
actHelp // reply with helpText
actGrade // call back into Python with the query
)
// mentionsBot reports whether the message @mentions our own user id.
func mentionsBot(m *gwMessage, botID string) bool {
for _, u := range m.Mentions {
if string(u.ID) == botID {
return true
}
}
return false
}
// isDM: a DM has no guild_id; a guild message always carries one.
func isDM(m *gwMessage) bool {
return m.GuildID == ""
}
// stripMentions drops every form of the bot's own mention (<@id> and the
// legacy <@!id>) and collapses the surrounding whitespace, leaving just the
// user's query text.
func stripMentions(content, botID string) string {
out := strings.ReplaceAll(content, "<@"+botID+">", " ")
out = strings.ReplaceAll(out, "<@!"+botID+">", " ")
return strings.Join(strings.Fields(out), " ")
}
// triage decides the reply for one MESSAGE_CREATE: silence, the help line, or
// a grade lookup for the returned query.
//
// Silent unless the message DMs the bot or @mentions it, and never for a bot
// author — which includes ourselves, the guard against a mention-loop (our
// reply quoting the mention must not trigger another reply). The text after
// the mention is treated as a city query; DMs are the query as-is (trimmed
// but not collapsed, matching the Python's .strip()).
func triage(m *gwMessage, botID string) (action, string) {
if m.Author.Bot || string(m.Author.ID) == botID {
return actSilent, ""
}
dm := isDM(m)
if !dm && !mentionsBot(m, botID) {
return actSilent, ""
}
var query string
if dm {
query = strings.TrimSpace(m.Content)
} else {
query = stripMentions(m.Content, botID)
}
if query == "" {
return actHelp, ""
}
return actGrade, query
}