thermograph/backend/daemon/internal/cron/cron.go
Emi Griffith c3b906a9ed
All checks were successful
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / build-backend (pull_request) Successful in 1m18s
PR build (required check) / changes (pull_request) Successful in 6s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 3s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / validate-observability (pull_request) Has been skipped
daemon: move the Discord gateway and scheduler out of the web process into Go
web/app.py started two long-lived background jobs under a leader election: the
Discord gateway bot and an APScheduler. Both are stateful I/O loops -- reconnect,
RESUME, heartbeat, backoff, interval timers -- living inside an async web app
that also has to serve requests. This moves them into a single Go binary.

Go owns ONLY the stateful I/O. It owns no climate or grading logic: anything
needing data calls back into Python over a new internal-only HTTP surface
(/internal/discord/grade, /internal/jobs/warm-cities, /internal/jobs/indexnow).
Grading depends on polars and the parquet cache; reimplementing it in Go would
make the bot's grades drift from the API's, and the slash-command path
deliberately shares one grade builder so the two can never disagree. The grade
route returns gateway-ready JSON -- including the ephemeral-flag drop that
discord_bot.py used to do -- and Go relays those bytes verbatim without parsing
the embed.

Packaging: the binary is built by a golang:1.26 stage in the backend Dockerfile
and shipped in the SAME image, run as a second compose service off the SAME tag.
The daemon and backend share the /internal/* contract, so they must never skew
versions; one image makes that structural rather than a convention. Its
entrypoint bypasses entrypoint.sh -- the backend owns alembic, and two racing
migrators is a real hazard.

replicas: 1 in the Swarm stack is load-bearing. Discord permits exactly one
gateway connection per bot token; the pin replaces core/singleton.claim_leader
for this workload. update_config uses order: stop-first, since start-first would
briefly run two gateways. autoscale.sh targets ${STACK_NAME}_web only, so it
cannot scale this.

Security: the internal routes compare the token with hmac.compare_digest and the
whole router 404s when THERMOGRAPH_INTERNAL_TOKEN is unset -- fail closed, never
default open. Caddy only routes /api/*, /digest and /discord/interactions to the
backend, so /internal/* was never publicly reachable; the token is defence in
depth. The router mounts before the catch-all frontend proxy so /internal/*
cannot fall through to it. The daemon refuses to start without the token.

Behaviour preserved from the Python, with the reasoning carried into the Go
comments: non-privileged intents (no MESSAGE_CONTENT, so no portal review);
fatal close codes 4004/4010-4014 stop rather than loop; the bot-author and
self-author mention-loop guard; allowed_mentions locked to {"parse":[],
"replied_user":true} so a crafted query cannot turn a reply into an @everyone
ping; the first cron tick deferred one full interval rather than firing at boot,
since warm-cities already runs at deploy time; and no overlapping warm-cities
run, which would double-spend the archive-fetch quota.

Two deliberate improvements over the Python. A close intended for RESUME now
uses 4000 rather than 1000 -- Discord invalidates a session closed 1000/1001, so
the Python's default close silently defeated its own resume. And MESSAGE_CREATE
is handled on a bounded worker pool rather than an unbounded thread hand-off, so
a flood of mentions cannot spawn unbounded work against the backend.

A .dockerignore is added because a disposable backend/.venv was being swallowed
by COPY . /app/ and duplicated again by the chown layer, inflating the image to
1.8 GB; it builds at 578 MB.

Tests: 29 Go gateway tests covering every behaviour the deleted
test_discord_bot.py asserted, plus cron/config/apiclient suites; 10 new Python
tests for the internal routes (fail-closed, auth, flag drop, per-job 409 guard).
Full suite 359 passed / 7 skipped; go build, vet and test -race clean.
2026-07-23 15:42:44 -07:00

92 lines
3.4 KiB
Go

// Package cron is the Go replacement for notifications/scheduler.py's
// APScheduler pair (warm-cities + IndexNow). The Python side chose APScheduler
// over a Postgres-native queue because exactly one process ever runs these
// jobs — no fan-out, backpressure, or dead-letter need — and that shape
// carries over unchanged: this daemon is the sole scheduler, so plain
// per-job tickers are all the machinery required. procrastinate/pgqueuer
// remain the documented upgrade path if that ever changes — see
// docs/architecture/repo-topology-and-infrastructure.md §7.
//
// The jobs themselves stay in Python (called via the internal HTTP routes);
// this package owns only the timing.
package cron
import (
"context"
"log/slog"
"sync"
"time"
)
// Job is one recurring task: a name for logs, an interval, and the callback
// (in practice an apiclient method hitting a /internal/jobs/* route).
type Job struct {
Name string
Interval time.Duration
Run func(ctx context.Context) error
}
// Run drives every job on its own ticker until ctx is cancelled, then returns
// once all job goroutines have stopped. It never returns early: a failing job
// logs and keeps ticking — one bad tick (backend briefly down, timeout) must
// never kill the schedule, because nothing restarts it short of a container
// restart.
func Run(ctx context.Context, logger *slog.Logger, jobs ...Job) {
var wg sync.WaitGroup
for _, job := range jobs {
wg.Add(1)
go func(job Job) {
defer wg.Done()
runJob(ctx, logger, job)
}(job)
}
wg.Wait()
}
func runJob(ctx context.Context, logger *slog.Logger, job Job) {
// A plain ticker's first fire is one interval from now, NOT at startup —
// deliberately kept from the Python scheduler (its add_job had no
// next_run_time override for the same reason): warm-cities already runs at
// deploy time via the deploy hook, so an immediate run on every daemon
// boot/restart would just re-spend work (and, for warm-cities, archive-fetch
// quota) that the deploy already paid for.
ticker := time.NewTicker(job.Interval)
defer ticker.Stop()
for {
select {
case <-ctx.Done():
return
case <-ticker.C:
start := time.Now()
// Run synchronously in this loop: a job can never overlap with
// itself by construction. That matters most for warm-cities,
// which spends the archive-fetch quota — two overlapping runs would
// double-spend it.
if err := job.Run(ctx); err != nil {
if ctx.Err() != nil {
// Shutdown raced the job; the cancellation is the story,
// not the (expected) aborted call.
return
}
logger.Error("cron job failed; will retry next tick",
"job", job.Name, "err", err, "elapsed", time.Since(start))
} else {
logger.Info("cron job ok",
"job", job.Name, "elapsed", time.Since(start))
}
// If the job overran its interval, a tick may have fired mid-run.
// Skip it rather than running again back-to-back — a late tick is
// the same double-spend as an overlapping one, just serialized.
// (Go 1.23+ tickers already drop fires with no ready receiver;
// this drain pins the skip semantics rather than relying on the
// runtime's channel buffering.)
select {
case <-ticker.C:
logger.Warn("cron job overran its interval; skipping the tick that fired mid-run",
"job", job.Name, "interval", job.Interval, "elapsed", time.Since(start))
default:
}
}
}
}