All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
Environments stop being machines. Until now each box WAS an environment --
"beta" named both a deploy target and a host, "the desktop" named both the
operator's computer and the dev server -- so every path could assume one
environment per host and hardcode /opt/thermograph, /etc/thermograph.env and
ports 8137/8080. That assumption ends here:
vps1 75.119.132.91 Forgejo, Grafana/Loki, the portfolio site, and DEV
(own Postgres, mesh-only on 10.10.0.2:8137)
vps2 169.58.46.181 PROD and BETA as two Swarm stacks sharing one
TimescaleDB instance, plus Centralis, Postfix, backups
desktop AI model hosting + flex Swarm capacity, no environment
Nothing live has moved; the ordered cutover is in
infra/deploy/RUNBOOK-vps1-vps2-cutover.md.
deploy/env-topology.sh is the single source of truth: env -> host, checkout,
branch, deploy mode, stack name, env file, LB ports, DB role/database, service
prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing
distinguishing a beta deploy from a prod one -- deploy.sh refuses to run when it
disagrees with the checkout it was invoked from. The host-wide secrets-env and
deploy-mode markers survive only as a fallback for a single-environment box.
Beta gets its own stack file rather than an overlay (a merge cannot REMOVE the
db service, and beta having no database of its own is the point). Its services
are prefixed beta-*: Swarm registers a service's short name as a DNS alias on
every network it joins, so two stacks both calling a service `web` on the shared
data network would let prod's frontend resolve a beta task. Prod's stack, env
and LB config are untouched.
One Postgres, two databases with two roles: deploy/db/provision-env-db.sh
creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT
revoked from PUBLIC) plus a <role>_ro for ad-hoc queries, and refuses to run for
the environment that owns the instance -- doing so would demote prod's bootstrap
superuser.
Fixes that co-residency would otherwise have broken silently:
- CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and
also "the Forgejo box" because those shared a machine; that conflation is what
once had the prod backup dumping beta.
- The nightly backup dumps BOTH application databases to separate off-box
prefixes, and fails loudly on a missing one instead of skipping it.
- Alloy stopped deriving the `host` label from the node -- on vps2 that would
have filed every beta line as prod, feeding prod's alert rules with beta's
traffic. It is now derived per source, with a new `node` label for the machine,
and beta's log volume is mounted via a vps2-only overlay.
- ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derived their SSH
target and env-file path from the topology instead of hardcoding beta to
75.119.132.91 -- which is vps1's address now.
- Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on
vps1) instead of 0.0.0.0: correct for a home LAN box, a public exposure of
unreviewed branches on a VPS.
Terraform's hosts variable is keyed by machine with a nested environments map,
and main.tf flattens (host, environment) pairs so two environments on one box
cannot share a checkout. Caddy's single reference config is split per host.
Docs, onboarding and the runbooks are updated throughout, including the security
rationales that co-residency makes false: separate SSH credentials no longer put
a host boundary between beta and prod, and the boundary that remains is the
database and the filesystem.
234 lines
11 KiB
Bash
Executable file
234 lines
11 KiB
Bash
Executable file
#!/usr/bin/env bash
|
|
# Outbound email for Thermograph — run once on the VPS, as root.
|
|
#
|
|
# Installs Postfix as a SEND-ONLY NULL CLIENT: it listens on 127.0.0.1:25 only,
|
|
# accepts mail from this machine, and never receives mail from the internet.
|
|
#
|
|
# Why a local MTA instead of talking to a mail provider's API from Python:
|
|
#
|
|
# * The app's only mail config becomes "SMTP on localhost". Whether delivery
|
|
# then goes direct to the recipient's MX or through a relay is a Postfix
|
|
# setting — switchable without touching, redeploying, or retesting the app.
|
|
# * Postfix queues and retries. A request handler hands the message over in
|
|
# microseconds and returns; a slow or briefly-down upstream can't stall a
|
|
# web request or lose a signup.
|
|
# * No new Python dependency: stdlib smtplib talks to it (see backend/mailer.py).
|
|
#
|
|
# DELIVERABILITY — read before pointing this at real subscribers.
|
|
#
|
|
# Mail sent straight from a VPS IP is very often junked, regardless of Postfix
|
|
# config, because the IP has no sending reputation. Two options:
|
|
#
|
|
# A. RELAY through a transactional provider (recommended for real mail).
|
|
# Set RELAYHOST + RELAY_USER + RELAY_PASSWORD below. The provider handles
|
|
# SPF/DKIM alignment and reputation; you keep the loopback-SMTP seam.
|
|
#
|
|
# B. DIRECT to MX (no third party). Then you must also set up, in DNS:
|
|
# - SPF: TXT @ "v=spf1 a mx ip4:<VPS_IP> -all"
|
|
# - DKIM: install opendkim, publish the public key as a TXT record
|
|
# - DMARC: TXT _dmarc "v=DMARC1; p=none; rua=mailto:you@domain"
|
|
# - PTR / reverse DNS on the VPS IP -> mail.thermograph.org
|
|
# The PTR record is the one people forget, and its absence alone is enough
|
|
# for Gmail and Outlook to junk everything you send.
|
|
#
|
|
# Usage:
|
|
# sudo MAIL_DOMAIN=thermograph.org bash deploy/provision-mail.sh
|
|
# sudo MAIL_DOMAIN=thermograph.org RELAYHOST='[smtp.provider.com]:587' \
|
|
# RELAY_USER=apikey RELAY_PASSWORD=secret bash deploy/provision-mail.sh
|
|
#
|
|
# MAIL_ENV picks which environment's deploy mode (env-topology.sh) sizes the
|
|
# Docker mail gateway below -- see the "WHICH gateway" comment. Defaults to
|
|
# prod; only matters if this box ever runs an environment in compose mode
|
|
# (it doesn't today -- prod and beta are both Swarm on this box; compose-mode
|
|
# dev lives on the other box entirely and doesn't run Postfix).
|
|
set -euo pipefail
|
|
|
|
MAIL_DOMAIN="${MAIL_DOMAIN:-thermograph.org}"
|
|
MAIL_HOSTNAME="${MAIL_HOSTNAME:-mail.${MAIL_DOMAIN}}"
|
|
RELAYHOST="${RELAYHOST:-}"
|
|
RELAY_USER="${RELAY_USER:-}"
|
|
RELAY_PASSWORD="${RELAY_PASSWORD:-}"
|
|
|
|
MAIL_ENV="${MAIL_ENV:-prod}"
|
|
SELF_DIR=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
|
|
# shellcheck source=infra/deploy/env-topology.sh
|
|
. "$SELF_DIR/env-topology.sh"
|
|
thermograph_topology "$MAIL_ENV"
|
|
|
|
if [[ $EUID -ne 0 ]]; then
|
|
echo "run as root (sudo)" >&2
|
|
exit 1
|
|
fi
|
|
|
|
echo "==> installing postfix (non-interactive)"
|
|
export DEBIAN_FRONTEND=noninteractive
|
|
# Preseed so the installer doesn't open its curses dialog.
|
|
debconf-set-selections <<EOF
|
|
postfix postfix/main_mailer_type select Internet Site
|
|
postfix postfix/mailname string ${MAIL_HOSTNAME}
|
|
EOF
|
|
apt-get update -qq
|
|
apt-get install -y -qq postfix libsasl2-modules
|
|
|
|
echo "==> configuring send-only null client"
|
|
postconf -e "myhostname = ${MAIL_HOSTNAME}"
|
|
postconf -e "myorigin = ${MAIL_DOMAIN}"
|
|
# Never listen on a public interface. This box sends only. The app runs in a
|
|
# Docker container, so it can't reach the host's loopback — it hands mail to
|
|
# Postfix over the compose bridge's gateway. So Postfix also listens on that
|
|
# gateway and accepts mail from the bridge subnet (both pinned in
|
|
# docker-compose.yml). Set DOCKER_MAIL_GATEWAY="" for a pure loopback-only null
|
|
# client (app running natively on the host, not in a container).
|
|
#
|
|
# WHICH gateway to listen on comes from MAIL_ENV's deploy mode (env-topology.sh,
|
|
# TG_DEPLOY_MODE) -- not a hardcoded beta/prod split. Beta moved onto this SAME
|
|
# box as prod and is Swarm too now, so "compose" no longer means "beta"; it
|
|
# means dev, which lives on the other box entirely and never runs this script.
|
|
# Compose mode uses the pinned compose bridge (172.19.0.1/172.19.0.0/16, the
|
|
# defaults); stack mode uses the docker_gwbridge gateway instead -- overlay
|
|
# tasks have no compose-bridge gateway -- so a stack-mode MAIL_ENV (prod or
|
|
# beta; same box, same gateway) gets DOCKER_MAIL_GATEWAY=172.18.0.1
|
|
# DOCKER_MAIL_SUBNET=172.18.0.0/16 (plus the MESH_MAIL_* listener below). Do
|
|
# NOT list an address that doesn't exist on the host: Postfix's master fails
|
|
# to bind and takes ALL listeners down -- exactly what happened when the
|
|
# compose bridge (172.19.0.1) vanished at the stack cutover while still listed
|
|
# in inet_interfaces. Also note: postfix on this distro is an umbrella unit;
|
|
# restart `postfix@-`, not `postfix`, for inet_interfaces changes to take
|
|
# effect.
|
|
#
|
|
# THE SAME TRAP FIRES AT BOOT, NOT JUST ON RENUMBERING (prod outage 2026-07-24).
|
|
# A Docker bridge address does not exist until dockerd creates it, and the stock
|
|
# postfix@.service is only ordered `After=network-online.target` -- which says
|
|
# nothing about dockerd or wg-quick. At the 08:09 reboot Postfix started at
|
|
# 08:09:49 and fataled at 08:09:52 on "no local interface found for 172.18.0.1";
|
|
# dockerd did not even begin starting until 08:09:53. wg0 (10.10.0.1) won the
|
|
# same race by one second. Because postfix@.service ships no Restart=, that
|
|
# single lost race killed ALL mail for 13h -- loopback and mesh included.
|
|
# install_postfix_ordering_dropin below is what makes this survive a reboot:
|
|
# it orders postfix@ after docker.service and wg-quick@wg0.service and retries
|
|
# on failure. If you add an address here that some other daemon creates, add
|
|
# that daemon to the drop-in too.
|
|
if [ "$TG_DEPLOY_MODE" = stack ]; then
|
|
DOCKER_MAIL_GATEWAY="${DOCKER_MAIL_GATEWAY-172.18.0.1}"
|
|
DOCKER_MAIL_SUBNET="${DOCKER_MAIL_SUBNET-172.18.0.0/16}"
|
|
else
|
|
DOCKER_MAIL_GATEWAY="${DOCKER_MAIL_GATEWAY-172.19.0.1}"
|
|
DOCKER_MAIL_SUBNET="${DOCKER_MAIL_SUBNET-172.19.0.0/16}"
|
|
fi
|
|
# Optional WireGuard-mesh listener: other mesh nodes (e.g. vps1's Forgejo, whose
|
|
# mailer posts to 10.10.0.1:25 — see deploy/forgejo/docker-stack.yml) can relay
|
|
# through this box. This host (vps2 -- prod AND beta) runs with
|
|
# MESH_MAIL_LISTEN=10.10.0.1 and MESH_MAIL_PEERS=10.10.0.2/32; both default OFF
|
|
# so a plain run stays a strict null client. 10.10.0.2 is vps1 (Forgejo,
|
|
# Grafana, dev) -- mesh IPs did not move in the vps1/vps2 split, only which
|
|
# environment runs where, so this is NOT "beta's" address. Without these,
|
|
# re-running this script on vps2 would silently drop the mesh listener and
|
|
# break Forgejo's outbound mail — the live config was originally hand-applied
|
|
# and this script is the source of truth for it now.
|
|
MESH_MAIL_LISTEN="${MESH_MAIL_LISTEN-}"
|
|
MESH_MAIL_PEERS="${MESH_MAIL_PEERS-}"
|
|
postconf -e "inet_protocols = ipv4"
|
|
listen="127.0.0.1"
|
|
networks="127.0.0.0/8 [::1]/128"
|
|
if [[ -n "$DOCKER_MAIL_GATEWAY" ]]; then
|
|
listen="${listen}, ${DOCKER_MAIL_GATEWAY}"
|
|
networks="${networks} ${DOCKER_MAIL_SUBNET}"
|
|
# ufw is default-deny incoming; a container connecting to the host's gateway IP
|
|
# hits the INPUT chain, so allow the bridge subnet to reach port 25.
|
|
command -v ufw >/dev/null 2>&1 && \
|
|
ufw allow from "${DOCKER_MAIL_SUBNET}" to any port 25 proto tcp \
|
|
comment 'app container -> host Postfix' || true
|
|
fi
|
|
if [[ -n "$MESH_MAIL_LISTEN" ]]; then
|
|
listen="${listen}, ${MESH_MAIL_LISTEN}"
|
|
networks="${networks} ${MESH_MAIL_PEERS}"
|
|
fi
|
|
if [[ "$listen" == "127.0.0.1" ]]; then
|
|
postconf -e "inet_interfaces = loopback-only"
|
|
else
|
|
postconf -e "inet_interfaces = ${listen}"
|
|
fi
|
|
postconf -e "mynetworks = ${networks}"
|
|
# A null client delivers nothing locally; everything is relayed out.
|
|
postconf -e "mydestination ="
|
|
postconf -e "local_transport = error:local delivery is disabled"
|
|
# Use TLS opportunistically when talking to the next hop.
|
|
postconf -e "smtp_tls_security_level = may"
|
|
postconf -e "smtp_tls_loglevel = 1"
|
|
|
|
if [[ -n "$RELAYHOST" ]]; then
|
|
echo "==> configuring relay via ${RELAYHOST}"
|
|
postconf -e "relayhost = ${RELAYHOST}"
|
|
if [[ -n "$RELAY_USER" ]]; then
|
|
postconf -e "smtp_sasl_auth_enable = yes"
|
|
postconf -e "smtp_sasl_password_maps = hash:/etc/postfix/sasl_passwd"
|
|
postconf -e "smtp_sasl_security_options = noanonymous"
|
|
printf '%s %s:%s\n' "$RELAYHOST" "$RELAY_USER" "$RELAY_PASSWORD" \
|
|
> /etc/postfix/sasl_passwd
|
|
# The credential file must not be world-readable.
|
|
chmod 600 /etc/postfix/sasl_passwd
|
|
postmap /etc/postfix/sasl_passwd
|
|
chmod 600 /etc/postfix/sasl_passwd.db
|
|
fi
|
|
else
|
|
echo "==> no RELAYHOST set: delivering direct to MX"
|
|
echo " remember SPF + DKIM + DMARC + PTR, or expect the spam folder"
|
|
postconf -e "relayhost ="
|
|
fi
|
|
|
|
# Postfix fatals if ANY inet_interfaces address is missing when it starts, and
|
|
# takes every listener down with it. Docker bridge and WireGuard addresses are
|
|
# created by other daemons, so order Postfix after them, gate the start on the
|
|
# addresses actually existing, and supervise the result. See the long comment
|
|
# above the inet_interfaces block, and the header of the script itself for the
|
|
# full incident write-up and the reasoning behind each number.
|
|
bash "$(dirname "${BASH_SOURCE[0]}")/provision-mail-supervision.sh"
|
|
|
|
systemctl enable postfix
|
|
# postfix.service is an umbrella whose ExecStart is /bin/true; the instance
|
|
# postfix@- is what actually binds. Restarting the umbrella propagates via
|
|
# PartOf=, but restart the instance directly so a failure surfaces here.
|
|
systemctl restart 'postfix@-'
|
|
|
|
# Assert, don't hope. postfix-health is the ONE definition of "mail works" on
|
|
# this estate -- the same check the watchdog, the Grafana alert and any
|
|
# mail_health tool use, so provisioning cannot pass on a laxer standard than
|
|
# monitoring. It checks the INSTANCE unit (never the active(exited) umbrella,
|
|
# which is the check that reported green through the whole 13h 2026-07-24
|
|
# outage), that every configured address is really bound, that a live 220
|
|
# greeting comes back, that no PUBLIC address is bound, and that nothing in
|
|
# inet_interfaces is missing from the host -- the latent state that stays
|
|
# invisible until the next reboot and then kills all mail.
|
|
echo "==> verifying postfix is actually up and bound"
|
|
/usr/local/sbin/postfix-health || {
|
|
echo "FATAL: postfix-health failed -- see detail above" >&2
|
|
exit 1
|
|
}
|
|
echo "==> listeners (must NOT include a public address):"
|
|
ss -lntp | grep ':25 '
|
|
|
|
cat <<'NOTE'
|
|
|
|
==> next steps
|
|
|
|
1. Point the app at it, in /etc/thermograph.env:
|
|
|
|
THERMOGRAPH_MAIL_BACKEND=smtp
|
|
THERMOGRAPH_SMTP_HOST=127.0.0.1
|
|
THERMOGRAPH_SMTP_PORT=25
|
|
THERMOGRAPH_MAIL_FROM=Thermograph <no-reply@thermograph.org>
|
|
|
|
then: sudo systemctl restart thermograph
|
|
|
|
2. Send yourself a test message:
|
|
|
|
echo "test body" | mail -s "thermograph test" you@example.com
|
|
# or, exercising the app's own path:
|
|
# python -c "import sys; sys.path.insert(0,'/opt/thermograph/backend'); \
|
|
# import mailer; print(mailer.send('you@example.com','t','body'))"
|
|
|
|
3. Watch it leave: journalctl -u postfix -f (queue: mailq)
|
|
|
|
4. Check placement with https://www.mail-tester.com — it scores SPF, DKIM,
|
|
DMARC and rDNS in one shot and tells you exactly what's missing.
|
|
NOTE
|