All checks were successful
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-backend (pull_request) Successful in 1m33s
PR build (required check) / gate (pull_request) Successful in 3s
web/app.py started two long-lived background jobs under a leader election: the
Discord gateway bot and an APScheduler. Both are stateful I/O loops -- reconnect,
RESUME, heartbeat, backoff, interval timers -- living inside an async web app
that also has to serve requests. This moves them into a single Go binary.
Go owns ONLY the stateful I/O. It owns no climate or grading logic: anything
needing data calls back into Python over a new internal-only HTTP surface
(/internal/discord/grade, /internal/jobs/warm-cities, /internal/jobs/indexnow).
Grading depends on polars and the parquet cache; reimplementing it in Go would
make the bot's grades drift from the API's, and the slash-command path
deliberately shares one grade builder so the two can never disagree. The grade
route returns gateway-ready JSON -- including the ephemeral-flag drop that
discord_bot.py used to do -- and Go relays those bytes verbatim without parsing
the embed.
Packaging: the binary is built by a golang:1.26 stage in the backend Dockerfile
and shipped in the SAME image, run as a second compose service off the SAME tag.
The daemon and backend share the /internal/* contract, so they must never skew
versions; one image makes that structural rather than a convention. Its
entrypoint bypasses entrypoint.sh -- the backend owns alembic, and two racing
migrators is a real hazard.
replicas: 1 in the Swarm stack is load-bearing. Discord permits exactly one
gateway connection per bot token; the pin replaces core/singleton.claim_leader
for this workload. update_config uses order: stop-first, since start-first would
briefly run two gateways. autoscale.sh targets ${STACK_NAME}_web only, so it
cannot scale this.
Security: the internal routes compare the token with hmac.compare_digest and the
whole router 404s when THERMOGRAPH_INTERNAL_TOKEN is unset -- fail closed, never
default open. Caddy only routes /api/*, /digest and /discord/interactions to the
backend, so /internal/* was never publicly reachable; the token is defence in
depth. The router mounts before the catch-all frontend proxy so /internal/*
cannot fall through to it. The daemon refuses to start without the token.
Behaviour preserved from the Python, with the reasoning carried into the Go
comments: non-privileged intents (no MESSAGE_CONTENT, so no portal review);
fatal close codes 4004/4010-4014 stop rather than loop; the bot-author and
self-author mention-loop guard; allowed_mentions locked to {"parse":[],
"replied_user":true} so a crafted query cannot turn a reply into an @everyone
ping; the first cron tick deferred one full interval rather than firing at boot,
since warm-cities already runs at deploy time; and no overlapping warm-cities
run, which would double-spend the archive-fetch quota.
Two deliberate improvements over the Python. A close intended for RESUME now
uses 4000 rather than 1000 -- Discord invalidates a session closed 1000/1001, so
the Python's default close silently defeated its own resume. And MESSAGE_CREATE
is handled on a bounded worker pool rather than an unbounded thread hand-off, so
a flood of mentions cannot spawn unbounded work against the backend.
A .dockerignore is added because a disposable backend/.venv was being swallowed
by COPY . /app/ and duplicated again by the chown layer, inflating the image to
1.8 GB; it builds at 578 MB.
Tests: 29 Go gateway tests covering every behaviour the deleted
test_discord_bot.py asserted, plus cron/config/apiclient suites; 10 new Python
tests for the internal routes (fail-closed, auth, flag drop, per-job 409 guard).
Full suite 359 passed / 7 skipped; go build, vet and test -race clean.
92 lines
3.4 KiB
Go
92 lines
3.4 KiB
Go
// Package cron is the Go replacement for notifications/scheduler.py's
|
|
// APScheduler pair (warm-cities + IndexNow). The Python side chose APScheduler
|
|
// over a Postgres-native queue because exactly one process ever runs these
|
|
// jobs — no fan-out, backpressure, or dead-letter need — and that shape
|
|
// carries over unchanged: this daemon is the sole scheduler, so plain
|
|
// per-job tickers are all the machinery required. procrastinate/pgqueuer
|
|
// remain the documented upgrade path if that ever changes — see
|
|
// docs/architecture/repo-topology-and-infrastructure.md §7.
|
|
//
|
|
// The jobs themselves stay in Python (called via the internal HTTP routes);
|
|
// this package owns only the timing.
|
|
package cron
|
|
|
|
import (
|
|
"context"
|
|
"log/slog"
|
|
"sync"
|
|
"time"
|
|
)
|
|
|
|
// Job is one recurring task: a name for logs, an interval, and the callback
|
|
// (in practice an apiclient method hitting a /internal/jobs/* route).
|
|
type Job struct {
|
|
Name string
|
|
Interval time.Duration
|
|
Run func(ctx context.Context) error
|
|
}
|
|
|
|
// Run drives every job on its own ticker until ctx is cancelled, then returns
|
|
// once all job goroutines have stopped. It never returns early: a failing job
|
|
// logs and keeps ticking — one bad tick (backend briefly down, timeout) must
|
|
// never kill the schedule, because nothing restarts it short of a container
|
|
// restart.
|
|
func Run(ctx context.Context, logger *slog.Logger, jobs ...Job) {
|
|
var wg sync.WaitGroup
|
|
for _, job := range jobs {
|
|
wg.Add(1)
|
|
go func(job Job) {
|
|
defer wg.Done()
|
|
runJob(ctx, logger, job)
|
|
}(job)
|
|
}
|
|
wg.Wait()
|
|
}
|
|
|
|
func runJob(ctx context.Context, logger *slog.Logger, job Job) {
|
|
// A plain ticker's first fire is one interval from now, NOT at startup —
|
|
// deliberately kept from the Python scheduler (its add_job had no
|
|
// next_run_time override for the same reason): warm-cities already runs at
|
|
// deploy time via the deploy hook, so an immediate run on every daemon
|
|
// boot/restart would just re-spend work (and, for warm-cities, archive-fetch
|
|
// quota) that the deploy already paid for.
|
|
ticker := time.NewTicker(job.Interval)
|
|
defer ticker.Stop()
|
|
|
|
for {
|
|
select {
|
|
case <-ctx.Done():
|
|
return
|
|
case <-ticker.C:
|
|
start := time.Now()
|
|
// Run synchronously in this loop: a job can never overlap with
|
|
// itself by construction. That matters most for warm-cities,
|
|
// which spends the archive-fetch quota — two overlapping runs would
|
|
// double-spend it.
|
|
if err := job.Run(ctx); err != nil {
|
|
if ctx.Err() != nil {
|
|
// Shutdown raced the job; the cancellation is the story,
|
|
// not the (expected) aborted call.
|
|
return
|
|
}
|
|
logger.Error("cron job failed; will retry next tick",
|
|
"job", job.Name, "err", err, "elapsed", time.Since(start))
|
|
} else {
|
|
logger.Info("cron job ok",
|
|
"job", job.Name, "elapsed", time.Since(start))
|
|
}
|
|
// If the job overran its interval, a tick may have fired mid-run.
|
|
// Skip it rather than running again back-to-back — a late tick is
|
|
// the same double-spend as an overlapping one, just serialized.
|
|
// (Go 1.23+ tickers already drop fires with no ready receiver;
|
|
// this drain pins the skip semantics rather than relying on the
|
|
// runtime's channel buffering.)
|
|
select {
|
|
case <-ticker.C:
|
|
logger.Warn("cron job overran its interval; skipping the tick that fired mid-run",
|
|
"job", job.Name, "interval", job.Interval, "elapsed", time.Since(start))
|
|
default:
|
|
}
|
|
}
|
|
}
|
|
}
|