thermograph/backend/daemon/internal/cron/cron_test.go
Emi Griffith adf824b33f
All checks were successful
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / build-backend (pull_request) Successful in 1m33s
PR build (required check) / gate (pull_request) Successful in 3s
daemon: move the Discord gateway and scheduler out of the web process into Go
web/app.py started two long-lived background jobs under a leader election: the
Discord gateway bot and an APScheduler. Both are stateful I/O loops -- reconnect,
RESUME, heartbeat, backoff, interval timers -- living inside an async web app
that also has to serve requests. This moves them into a single Go binary.

Go owns ONLY the stateful I/O. It owns no climate or grading logic: anything
needing data calls back into Python over a new internal-only HTTP surface
(/internal/discord/grade, /internal/jobs/warm-cities, /internal/jobs/indexnow).
Grading depends on polars and the parquet cache; reimplementing it in Go would
make the bot's grades drift from the API's, and the slash-command path
deliberately shares one grade builder so the two can never disagree. The grade
route returns gateway-ready JSON -- including the ephemeral-flag drop that
discord_bot.py used to do -- and Go relays those bytes verbatim without parsing
the embed.

Packaging: the binary is built by a golang:1.26 stage in the backend Dockerfile
and shipped in the SAME image, run as a second compose service off the SAME tag.
The daemon and backend share the /internal/* contract, so they must never skew
versions; one image makes that structural rather than a convention. Its
entrypoint bypasses entrypoint.sh -- the backend owns alembic, and two racing
migrators is a real hazard.

replicas: 1 in the Swarm stack is load-bearing. Discord permits exactly one
gateway connection per bot token; the pin replaces core/singleton.claim_leader
for this workload. update_config uses order: stop-first, since start-first would
briefly run two gateways. autoscale.sh targets ${STACK_NAME}_web only, so it
cannot scale this.

Security: the internal routes compare the token with hmac.compare_digest and the
whole router 404s when THERMOGRAPH_INTERNAL_TOKEN is unset -- fail closed, never
default open. Caddy only routes /api/*, /digest and /discord/interactions to the
backend, so /internal/* was never publicly reachable; the token is defence in
depth. The router mounts before the catch-all frontend proxy so /internal/*
cannot fall through to it. The daemon refuses to start without the token.

Behaviour preserved from the Python, with the reasoning carried into the Go
comments: non-privileged intents (no MESSAGE_CONTENT, so no portal review);
fatal close codes 4004/4010-4014 stop rather than loop; the bot-author and
self-author mention-loop guard; allowed_mentions locked to {"parse":[],
"replied_user":true} so a crafted query cannot turn a reply into an @everyone
ping; the first cron tick deferred one full interval rather than firing at boot,
since warm-cities already runs at deploy time; and no overlapping warm-cities
run, which would double-spend the archive-fetch quota.

Two deliberate improvements over the Python. A close intended for RESUME now
uses 4000 rather than 1000 -- Discord invalidates a session closed 1000/1001, so
the Python's default close silently defeated its own resume. And MESSAGE_CREATE
is handled on a bounded worker pool rather than an unbounded thread hand-off, so
a flood of mentions cannot spawn unbounded work against the backend.

A .dockerignore is added because a disposable backend/.venv was being swallowed
by COPY . /app/ and duplicated again by the chown layer, inflating the image to
1.8 GB; it builds at 578 MB.

Tests: 29 Go gateway tests covering every behaviour the deleted
test_discord_bot.py asserted, plus cron/config/apiclient suites; 10 new Python
tests for the internal routes (fail-closed, auth, flag drop, per-job 409 guard).
Full suite 359 passed / 7 skipped; go build, vet and test -race clean.
2026-07-23 15:47:07 -07:00

204 lines
5.2 KiB
Go

package cron
import (
"context"
"errors"
"io"
"log/slog"
"sync/atomic"
"testing"
"time"
)
// Intervals here are tens of milliseconds: long enough that scheduler jitter
// cannot invert an assertion, short enough that the whole file runs in well
// under a second.
func discard() *slog.Logger {
return slog.New(slog.NewTextHandler(io.Discard, nil))
}
// The first run must be one interval after start, never at startup — the
// deploy hook already warmed the cache, so a boot-time run is pure waste.
func TestFirstRunIsDeferredOneInterval(t *testing.T) {
t.Parallel()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
var runs atomic.Int32
done := make(chan struct{})
go func() {
defer close(done)
Run(ctx, discard(), Job{
Name: "test",
Interval: 60 * time.Millisecond,
Run: func(context.Context) error {
runs.Add(1)
return nil
},
})
}()
// Well inside the first interval: nothing may have run yet.
time.Sleep(20 * time.Millisecond)
if got := runs.Load(); got != 0 {
t.Fatalf("job ran %d time(s) before the first interval elapsed; first run must be deferred", got)
}
// Well past the first interval: it must have run by now.
deadline := time.After(500 * time.Millisecond)
for runs.Load() == 0 {
select {
case <-deadline:
t.Fatal("job never ran after the first interval elapsed")
case <-time.After(5 * time.Millisecond):
}
}
cancel()
<-done
}
// A slow run must not overlap with itself, and ticks that fired mid-run must
// be skipped, not queued — an immediate back-to-back run would double-spend
// the archive-fetch quota just like a concurrent one.
func TestSlowJobNeverOverlapsAndSkipsMissedTicks(t *testing.T) {
t.Parallel()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
const interval = 30 * time.Millisecond
var inFlight, maxInFlight, runs atomic.Int32
release := make(chan struct{})
started := make(chan struct{}, 16)
done := make(chan struct{})
go func() {
defer close(done)
Run(ctx, discard(), Job{
Name: "slow",
Interval: interval,
Run: func(context.Context) error {
n := inFlight.Add(1)
for {
m := maxInFlight.Load()
if n <= m || maxInFlight.CompareAndSwap(m, n) {
break
}
}
runs.Add(1)
started <- struct{}{}
<-release // block until the test lets each run finish
inFlight.Add(-1)
return nil
},
})
}()
// First run starts; hold it across several intervals.
<-started
time.Sleep(4 * interval)
if got := maxInFlight.Load(); got != 1 {
t.Fatalf("job overlapped with itself: max in-flight = %d", got)
}
if got := runs.Load(); got != 1 {
t.Fatalf("expected exactly 1 run while the first is still blocked, got %d", got)
}
release <- struct{}{}
// After release the schedule resumes, but the ticks missed during the
// block must have been dropped: the next run arrives roughly one interval
// later, and only one more within that window (not a burst of catch-ups).
select {
case <-started:
case <-time.After(500 * time.Millisecond):
t.Fatal("job never resumed after the blocked run finished")
}
if got := runs.Load(); got != 2 {
t.Fatalf("missed ticks were replayed as a burst: %d runs total, want 2", got)
}
release <- struct{}{}
cancel()
<-done
if got := maxInFlight.Load(); got != 1 {
t.Fatalf("job overlapped with itself: max in-flight = %d", got)
}
}
// One failed tick must never kill the ticker — nothing restarts the schedule
// short of a container restart.
func TestFailedTickDoesNotStopSchedule(t *testing.T) {
t.Parallel()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
var runs atomic.Int32
done := make(chan struct{})
go func() {
defer close(done)
Run(ctx, discard(), Job{
Name: "flaky",
Interval: 20 * time.Millisecond,
Run: func(context.Context) error {
if runs.Add(1) == 1 {
return errors.New("backend briefly down")
}
return nil
},
})
}()
deadline := time.After(1 * time.Second)
for runs.Load() < 3 {
select {
case <-deadline:
t.Fatalf("schedule stalled after a failure: only %d run(s)", runs.Load())
case <-time.After(5 * time.Millisecond):
}
}
cancel()
<-done
}
// Cancellation must stop Run promptly even while a job is mid-call: the job
// callbacks are ctx-aware HTTP calls, so cancelling the context aborts the
// in-flight request and the loop must then exit instead of ticking again.
func TestCancelStopsRunWhileJobInFlight(t *testing.T) {
t.Parallel()
ctx, cancel := context.WithCancel(context.Background())
started := make(chan struct{})
done := make(chan struct{})
go func() {
defer close(done)
Run(ctx, discard(),
Job{
Name: "blocking",
Interval: 10 * time.Millisecond,
Run: func(jobCtx context.Context) error {
close(started)
<-jobCtx.Done() // behaves like an HTTP call aborted by cancellation
return jobCtx.Err()
},
},
// A second, idle job proves Run waits for ALL jobs to stop.
Job{
Name: "idle",
Interval: time.Hour,
Run: func(context.Context) error { return nil },
},
)
}()
<-started
cancel()
select {
case <-done:
case <-time.After(1 * time.Second):
t.Fatal("Run did not return promptly after context cancellation")
}
}