thermograph/infra/deploy/provision-mail.sh
Emi Griffith e4693dce58
All checks were successful
secrets-guard / encrypted (pull_request) Successful in 8s
shell-lint / shellcheck (pull_request) Successful in 10s
PR build (required check) / changes (pull_request) Successful in 16s
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
observability: add the estate's first alerting; supervise Postfix
The estate had zero alert rules. The only contact point was Grafana's factory
default pointing at the literal string <example@email.com>, and beta's
untracked compose override routed Grafana's SMTP at prod's Postfix — so
alerts, had any existed, would have been delivered by the box most likely to
be on fire. Prod served 86 5xx in 24h including five /healthz failures and
nobody was told.

Alerting (deploys to beta, which is where Grafana runs):
- 12 Loki-based rules. There is no Prometheus in this estate and Loki is the
  only datasource, so every rule is log-derived.
- Thresholds come from a 24h backtest that happens to contain a real ~20min
  prod outage at 04:20Z. Absolute counts, not ratios: prod runs ~7 req/min, so
  one bad request is 1.4% and a ratio alert would scream all night. The 5xx
  burst rule fires on the outage's 45 and 22 buckets and on nothing else in
  the day; the largest benign bucket all day was 5.
- Routed to a new private #ops-alerts channel, not to any existing channel —
  #weather-events, #announcements and #prod are product surfaces that notify
  real subscribers.
- AlertingWatchdog is a dead-man's switch; its value is its absence.
- Delivery is proven, not assumed: the rules were provisioned into a throwaway
  Grafana against beta's live Loki and a real alert arrived in Discord. This
  matters because GET /api/v1/provisioning/contact-points returns [REDACTED]
  for the URL — a contact point holding an uninterpolated env var looks
  perfectly healthy and pages nobody. The only proof is a message arriving.

CI gains a structural check, because the existing one only proves YAML parses:
an alert rule whose condition names a missing refId is valid YAML, provisions
cleanly, and never fires. It also hard-fails on a literal Discord webhook in
the repo. Verified against all three breakages deliberately introduced.

Postfix supervision:
- The 13h outage was a boot-ordering race, not a Docker renumbering: postfix
  started at 08:09:49, wg0 came up at :51, postfix fataled at :52 on a missing
  docker_gwbridge address, and dockerd did not finish starting until 08:10:16.
  Stock postfix@.service is ordered only After=network-online.target and ships
  no Restart=, so one lost race became a permanent outage.
- An ExecStartPre gate now blocks up to 60s until every inet_interfaces
  address actually exists, which absorbs the transient case inside a single
  start attempt. That makes bounded retry correct: 5 attempts in 600s, then
  failed — a genuinely broken config reaches a visible failed state in ~100s
  instead of re-fataling every 15s forever.
- A 5-minute watchdog timer retries indefinitely and runs reset-failed, so
  "failed" still self-heals. Worst case is ~5 minutes, not 13 hours.
- Wants=, not Requires=: a dockerd failure must not take down the loopback and
  mesh listeners that do not depend on Docker at all.

Health checks must read config with `postconf -c`, never postmulti/postqueue/
postfix — those three RESOLVE inet_interfaces and so fatal precisely when an
address is missing, which made the first version of this check report
status=ok bound=0/0. A health check that fails open is worse than none.

DEPLOY.md carries the monitoring contract, including that systemctl is-active
postfix is a known-false signal: postfix.service is a wrapper whose
ExecStart=/bin/true, so it reports active forever while the real postfix@-
instance is failed with zero listeners. Reproduced live.

Known gap, documented not fixed: Alloy ships Docker stdout, Caddy files and
app JSONL, not journald — so no Postfix line reaches Loki and the mail rules
cannot fire until loki.source.journal is added.

Claude-Session: https://claude.ai/code/session_0182KTMrsTHJc3TcewCatJFY
2026-07-24 13:19:19 -07:00

210 lines
9.9 KiB
Bash
Executable file

#!/usr/bin/env bash
# Outbound email for Thermograph — run once on the VPS, as root.
#
# Installs Postfix as a SEND-ONLY NULL CLIENT: it listens on 127.0.0.1:25 only,
# accepts mail from this machine, and never receives mail from the internet.
#
# Why a local MTA instead of talking to a mail provider's API from Python:
#
# * The app's only mail config becomes "SMTP on localhost". Whether delivery
# then goes direct to the recipient's MX or through a relay is a Postfix
# setting — switchable without touching, redeploying, or retesting the app.
# * Postfix queues and retries. A request handler hands the message over in
# microseconds and returns; a slow or briefly-down upstream can't stall a
# web request or lose a signup.
# * No new Python dependency: stdlib smtplib talks to it (see backend/mailer.py).
#
# DELIVERABILITY — read before pointing this at real subscribers.
#
# Mail sent straight from a VPS IP is very often junked, regardless of Postfix
# config, because the IP has no sending reputation. Two options:
#
# A. RELAY through a transactional provider (recommended for real mail).
# Set RELAYHOST + RELAY_USER + RELAY_PASSWORD below. The provider handles
# SPF/DKIM alignment and reputation; you keep the loopback-SMTP seam.
#
# B. DIRECT to MX (no third party). Then you must also set up, in DNS:
# - SPF: TXT @ "v=spf1 a mx ip4:<VPS_IP> -all"
# - DKIM: install opendkim, publish the public key as a TXT record
# - DMARC: TXT _dmarc "v=DMARC1; p=none; rua=mailto:you@domain"
# - PTR / reverse DNS on the VPS IP -> mail.thermograph.org
# The PTR record is the one people forget, and its absence alone is enough
# for Gmail and Outlook to junk everything you send.
#
# Usage:
# sudo MAIL_DOMAIN=thermograph.org bash deploy/provision-mail.sh
# sudo MAIL_DOMAIN=thermograph.org RELAYHOST='[smtp.provider.com]:587' \
# RELAY_USER=apikey RELAY_PASSWORD=secret bash deploy/provision-mail.sh
set -euo pipefail
MAIL_DOMAIN="${MAIL_DOMAIN:-thermograph.org}"
MAIL_HOSTNAME="${MAIL_HOSTNAME:-mail.${MAIL_DOMAIN}}"
RELAYHOST="${RELAYHOST:-}"
RELAY_USER="${RELAY_USER:-}"
RELAY_PASSWORD="${RELAY_PASSWORD:-}"
if [[ $EUID -ne 0 ]]; then
echo "run as root (sudo)" >&2
exit 1
fi
echo "==> installing postfix (non-interactive)"
export DEBIAN_FRONTEND=noninteractive
# Preseed so the installer doesn't open its curses dialog.
debconf-set-selections <<EOF
postfix postfix/main_mailer_type select Internet Site
postfix postfix/mailname string ${MAIL_HOSTNAME}
EOF
apt-get update -qq
apt-get install -y -qq postfix libsasl2-modules
echo "==> configuring send-only null client"
postconf -e "myhostname = ${MAIL_HOSTNAME}"
postconf -e "myorigin = ${MAIL_DOMAIN}"
# Never listen on a public interface. This box sends only. The app runs in a
# Docker container, so it can't reach the host's loopback — it hands mail to
# Postfix over the compose bridge's gateway. So Postfix also listens on that
# gateway and accepts mail from the bridge subnet (both pinned in
# docker-compose.yml). Set DOCKER_MAIL_GATEWAY="" for a pure loopback-only null
# client (app running natively on the host, not in a container).
#
# WHICH gateway depends on the host's deploy mode: plain compose (beta, LAN)
# uses the pinned compose bridge (172.19.0.1/172.19.0.0/16, the defaults);
# Swarm-stack mode (prod) uses the docker_gwbridge gateway instead --
# overlay tasks have no compose-bridge gateway -- so prod is provisioned with
# DOCKER_MAIL_GATEWAY=172.18.0.1 DOCKER_MAIL_SUBNET=172.18.0.0/16 (plus its
# MESH_MAIL_* listener below). Do NOT list an address that doesn't exist on
# the host: Postfix's master fails to bind and takes ALL listeners down --
# exactly what happened when the compose bridge (172.19.0.1) vanished at the
# stack cutover while still listed in inet_interfaces. Also note: postfix on
# this distro is an umbrella unit; restart `postfix@-`, not `postfix`, for
# inet_interfaces changes to take effect.
#
# THE SAME TRAP FIRES AT BOOT, NOT JUST ON RENUMBERING (prod outage 2026-07-24).
# A Docker bridge address does not exist until dockerd creates it, and the stock
# postfix@.service is only ordered `After=network-online.target` -- which says
# nothing about dockerd or wg-quick. At the 08:09 reboot Postfix started at
# 08:09:49 and fataled at 08:09:52 on "no local interface found for 172.18.0.1";
# dockerd did not even begin starting until 08:09:53. wg0 (10.10.0.1) won the
# same race by one second. Because postfix@.service ships no Restart=, that
# single lost race killed ALL mail for 13h -- loopback and mesh included.
# install_postfix_ordering_dropin below is what makes this survive a reboot:
# it orders postfix@ after docker.service and wg-quick@wg0.service and retries
# on failure. If you add an address here that some other daemon creates, add
# that daemon to the drop-in too.
DOCKER_MAIL_GATEWAY="${DOCKER_MAIL_GATEWAY-172.19.0.1}"
DOCKER_MAIL_SUBNET="${DOCKER_MAIL_SUBNET-172.19.0.0/16}"
# Optional WireGuard-mesh listener: other mesh nodes (e.g. beta's Forgejo, whose
# mailer posts to 10.10.0.1:25 — see deploy/forgejo/docker-stack.yml) can relay
# through this box. Prod runs with MESH_MAIL_LISTEN=10.10.0.1 and
# MESH_MAIL_PEERS=10.10.0.2/32; both default OFF so a plain run stays a strict
# null client. Without these, re-running this script on prod would silently drop
# the mesh listener and break Forgejo's outbound mail — the live config was
# originally hand-applied and this script is the source of truth for it now.
MESH_MAIL_LISTEN="${MESH_MAIL_LISTEN-}"
MESH_MAIL_PEERS="${MESH_MAIL_PEERS-}"
postconf -e "inet_protocols = ipv4"
listen="127.0.0.1"
networks="127.0.0.0/8 [::1]/128"
if [[ -n "$DOCKER_MAIL_GATEWAY" ]]; then
listen="${listen}, ${DOCKER_MAIL_GATEWAY}"
networks="${networks} ${DOCKER_MAIL_SUBNET}"
# ufw is default-deny incoming; a container connecting to the host's gateway IP
# hits the INPUT chain, so allow the bridge subnet to reach port 25.
command -v ufw >/dev/null 2>&1 && \
ufw allow from "${DOCKER_MAIL_SUBNET}" to any port 25 proto tcp \
comment 'app container -> host Postfix' || true
fi
if [[ -n "$MESH_MAIL_LISTEN" ]]; then
listen="${listen}, ${MESH_MAIL_LISTEN}"
networks="${networks} ${MESH_MAIL_PEERS}"
fi
if [[ "$listen" == "127.0.0.1" ]]; then
postconf -e "inet_interfaces = loopback-only"
else
postconf -e "inet_interfaces = ${listen}"
fi
postconf -e "mynetworks = ${networks}"
# A null client delivers nothing locally; everything is relayed out.
postconf -e "mydestination ="
postconf -e "local_transport = error:local delivery is disabled"
# Use TLS opportunistically when talking to the next hop.
postconf -e "smtp_tls_security_level = may"
postconf -e "smtp_tls_loglevel = 1"
if [[ -n "$RELAYHOST" ]]; then
echo "==> configuring relay via ${RELAYHOST}"
postconf -e "relayhost = ${RELAYHOST}"
if [[ -n "$RELAY_USER" ]]; then
postconf -e "smtp_sasl_auth_enable = yes"
postconf -e "smtp_sasl_password_maps = hash:/etc/postfix/sasl_passwd"
postconf -e "smtp_sasl_security_options = noanonymous"
printf '%s %s:%s\n' "$RELAYHOST" "$RELAY_USER" "$RELAY_PASSWORD" \
> /etc/postfix/sasl_passwd
# The credential file must not be world-readable.
chmod 600 /etc/postfix/sasl_passwd
postmap /etc/postfix/sasl_passwd
chmod 600 /etc/postfix/sasl_passwd.db
fi
else
echo "==> no RELAYHOST set: delivering direct to MX"
echo " remember SPF + DKIM + DMARC + PTR, or expect the spam folder"
postconf -e "relayhost ="
fi
# Postfix fatals if ANY inet_interfaces address is missing when it starts, and
# takes every listener down with it. Docker bridge and WireGuard addresses are
# created by other daemons, so order Postfix after them, gate the start on the
# addresses actually existing, and supervise the result. See the long comment
# above the inet_interfaces block, and the header of the script itself for the
# full incident write-up and the reasoning behind each number.
bash "$(dirname "${BASH_SOURCE[0]}")/provision-mail-supervision.sh"
systemctl enable postfix
# postfix.service is an umbrella whose ExecStart is /bin/true; the instance
# postfix@- is what actually binds. Restarting the umbrella propagates via
# PartOf=, but restart the instance directly so a failure surfaces here.
systemctl restart 'postfix@-'
# Assert, don't hope. postfix-health is the ONE definition of "mail works" on
# this estate -- the same check the watchdog, the Grafana alert and any
# mail_health tool use, so provisioning cannot pass on a laxer standard than
# monitoring. It checks the INSTANCE unit (never the active(exited) umbrella,
# which is the check that reported green through the whole 13h 2026-07-24
# outage), that every configured address is really bound, that a live 220
# greeting comes back, that no PUBLIC address is bound, and that nothing in
# inet_interfaces is missing from the host -- the latent state that stays
# invisible until the next reboot and then kills all mail.
echo "==> verifying postfix is actually up and bound"
/usr/local/sbin/postfix-health || {
echo "FATAL: postfix-health failed -- see detail above" >&2
exit 1
}
echo "==> listeners (must NOT include a public address):"
ss -lntp | grep ':25 '
cat <<'NOTE'
==> next steps
1. Point the app at it, in /etc/thermograph.env:
THERMOGRAPH_MAIL_BACKEND=smtp
THERMOGRAPH_SMTP_HOST=127.0.0.1
THERMOGRAPH_SMTP_PORT=25
THERMOGRAPH_MAIL_FROM=Thermograph <no-reply@thermograph.org>
then: sudo systemctl restart thermograph
2. Send yourself a test message:
echo "test body" | mail -s "thermograph test" you@example.com
# or, exercising the app's own path:
# python -c "import sys; sys.path.insert(0,'/opt/thermograph/backend'); \
# import mailer; print(mailer.send('you@example.com','t','body'))"
3. Watch it leave: journalctl -u postfix -f (queue: mailq)
4. Check placement with https://www.mail-tester.com — it scores SPF, DKIM,
DMARC and rDNS in one shot and tells you exactly what's missing.
NOTE