thermograph/backend/daemon/main.go
Emi Griffith adf824b33f
All checks were successful
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-backend (pull_request) Successful in 1m33s
PR build (required check) / gate (pull_request) Successful in 3s
daemon: move the Discord gateway and scheduler out of the web process into Go
web/app.py started two long-lived background jobs under a leader election: the
Discord gateway bot and an APScheduler. Both are stateful I/O loops -- reconnect,
RESUME, heartbeat, backoff, interval timers -- living inside an async web app
that also has to serve requests. This moves them into a single Go binary.

Go owns ONLY the stateful I/O. It owns no climate or grading logic: anything
needing data calls back into Python over a new internal-only HTTP surface
(/internal/discord/grade, /internal/jobs/warm-cities, /internal/jobs/indexnow).
Grading depends on polars and the parquet cache; reimplementing it in Go would
make the bot's grades drift from the API's, and the slash-command path
deliberately shares one grade builder so the two can never disagree. The grade
route returns gateway-ready JSON -- including the ephemeral-flag drop that
discord_bot.py used to do -- and Go relays those bytes verbatim without parsing
the embed.

Packaging: the binary is built by a golang:1.26 stage in the backend Dockerfile
and shipped in the SAME image, run as a second compose service off the SAME tag.
The daemon and backend share the /internal/* contract, so they must never skew
versions; one image makes that structural rather than a convention. Its
entrypoint bypasses entrypoint.sh -- the backend owns alembic, and two racing
migrators is a real hazard.

replicas: 1 in the Swarm stack is load-bearing. Discord permits exactly one
gateway connection per bot token; the pin replaces core/singleton.claim_leader
for this workload. update_config uses order: stop-first, since start-first would
briefly run two gateways. autoscale.sh targets ${STACK_NAME}_web only, so it
cannot scale this.

Security: the internal routes compare the token with hmac.compare_digest and the
whole router 404s when THERMOGRAPH_INTERNAL_TOKEN is unset -- fail closed, never
default open. Caddy only routes /api/*, /digest and /discord/interactions to the
backend, so /internal/* was never publicly reachable; the token is defence in
depth. The router mounts before the catch-all frontend proxy so /internal/*
cannot fall through to it. The daemon refuses to start without the token.

Behaviour preserved from the Python, with the reasoning carried into the Go
comments: non-privileged intents (no MESSAGE_CONTENT, so no portal review);
fatal close codes 4004/4010-4014 stop rather than loop; the bot-author and
self-author mention-loop guard; allowed_mentions locked to {"parse":[],
"replied_user":true} so a crafted query cannot turn a reply into an @everyone
ping; the first cron tick deferred one full interval rather than firing at boot,
since warm-cities already runs at deploy time; and no overlapping warm-cities
run, which would double-spend the archive-fetch quota.

Two deliberate improvements over the Python. A close intended for RESUME now
uses 4000 rather than 1000 -- Discord invalidates a session closed 1000/1001, so
the Python's default close silently defeated its own resume. And MESSAGE_CREATE
is handled on a bounded worker pool rather than an unbounded thread hand-off, so
a flood of mentions cannot spawn unbounded work against the backend.

A .dockerignore is added because a disposable backend/.venv was being swallowed
by COPY . /app/ and duplicated again by the chown layer, inflating the image to
1.8 GB; it builds at 578 MB.

Tests: 29 Go gateway tests covering every behaviour the deleted
test_discord_bot.py asserted, plus cron/config/apiclient suites; 10 new Python
tests for the internal routes (fail-closed, auth, flag drop, per-job 409 guard).
Full suite 359 passed / 7 skipped; go build, vet and test -race clean.
2026-07-23 15:47:07 -07:00

130 lines
4.4 KiB
Go

// thermograph-daemon owns the backend's long-lived stateful I/O — the Discord
// gateway websocket (IDENTIFY/RESUME/heartbeat/backoff) and the recurring-job
// timers — and nothing else. Anything that needs data calls back into the
// Python app over its internal-only HTTP routes (see internal/apiclient).
//
// This replaces the two in-process background jobs web/app.py used to start
// under leader election. One daemon process replaces the election entirely:
// there is exactly one replica of this container, so the single-connection
// invariant (Discord tolerates only one gateway session per token; duplicated
// APScheduler replicas would repeat every job) holds by deployment shape
// rather than by runtime coordination.
package main
import (
"context"
"log/slog"
"os"
"os/signal"
"sync"
"syscall"
"time"
"thermograph/daemon/internal/apiclient"
"thermograph/daemon/internal/config"
"thermograph/daemon/internal/cron"
"thermograph/daemon/internal/gateway"
)
// shutdownGrace bounds how long we wait for the gateway and cron halves to
// stop after a signal. Long enough for the gateway to close its session
// cleanly (a dirty exit leaves Discord holding a stale session), short enough
// to stay inside a container runtime's stop timeout so we exit on our own
// terms rather than being SIGKILLed mid-close.
const shutdownGrace = 10 * time.Second
func main() {
os.Exit(run())
}
// run is main minus os.Exit, so deferred cleanup actually runs.
func run() int {
// Everything to stdout: the container collects stdout, nothing else.
logger := slog.New(slog.NewTextHandler(os.Stdout, nil))
// Refuse to start on bad config rather than running half-configured: a
// daemon that cannot authenticate to the backend can only fail silently.
cfg, err := config.Load()
if err != nil {
logger.Error("invalid configuration; refusing to start", "err", err)
return 1
}
// One context for everything long-lived; SIGINT/SIGTERM cancels it so the
// gateway can close its session cleanly (Discord treats an abrupt drop
// worse than a clean close for RESUME purposes).
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer stop()
api := apiclient.New(cfg.APIBase, cfg.InternalToken)
logger.Info("thermograph-daemon starting",
"api", cfg.APIBase,
"discord", cfg.DiscordEnabled,
"warm_cities_interval", cfg.WarmCitiesInterval,
"indexnow_interval", cfg.IndexNowInterval)
var wg sync.WaitGroup
// The cron half runs unconditionally — it is the floor of what this
// daemon does, gateway or not.
wg.Add(1)
go func() {
defer wg.Done()
cron.Run(ctx, logger,
cron.Job{
Name: "warm-cities",
Interval: cfg.WarmCitiesInterval,
Run: func(ctx context.Context) error { return api.WarmCities(ctx) },
},
cron.Job{
Name: "indexnow",
Interval: cfg.IndexNowInterval,
Run: func(ctx context.Context) error { return api.IndexNow(ctx) },
},
)
}()
if cfg.DiscordEnabled {
wg.Add(1)
go func() {
defer wg.Done()
// A fatal gateway error (bad token, disallowed intents) must not
// take the cron half down: the schedule is independently useful,
// and exiting here would turn a Discord misconfiguration into a
// crash-loop that also stops warm-cities/IndexNow. Log loudly and
// keep running instead.
if err := gateway.Run(ctx, cfg, api); err != nil && ctx.Err() == nil {
logger.Error("discord gateway exited with a fatal error; cron jobs keep running without it",
"err", err)
}
}()
} else {
// An unconfigured deploy must pay nothing, same as the Python side:
// no connection, no retries — just say so once and run cron-only.
logger.Info("discord gateway disabled (THERMOGRAPH_DISCORD_BOT not truthy or no bot token); running cron-only")
}
<-ctx.Done()
// Restore default signal handling immediately: a second SIGINT/SIGTERM
// during the grace period kills us the ordinary way instead of being
// swallowed by the (already-cancelled) NotifyContext.
stop()
logger.Info("shutting down", "grace", shutdownGrace)
done := make(chan struct{})
go func() {
wg.Wait()
close(done)
}()
select {
case <-done:
logger.Info("shutdown complete")
case <-time.After(shutdownGrace):
// This process is the sole holder of the gateway connection, so a
// clean close matters — but a wedged goroutine must not hold the
// container's stop hostage forever.
logger.Warn("shutdown grace period elapsed before all loops stopped; exiting anyway")
}
return 0
}