thermograph/infra/deploy/thermograph.env.example
Emi Griffith 4e97d8e5dc
All checks were successful
PR build (required check) / changes (pull_request) Successful in 6s
secrets-guard / encrypted (pull_request) Successful in 5s
PR build (required check) / build-backend (pull_request) Has been skipped
shell-lint / shellcheck (pull_request) Successful in 7s
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Successful in 18s
PR build (required check) / gate (pull_request) Successful in 2s
infra: split the estate into vps1/vps2, moving beta next to prod and dev onto vps1
Environments stop being machines. Until now each box WAS an environment --
"beta" named both a deploy target and a host, "the desktop" named both the
operator's computer and the dev server -- so every path could assume one
environment per host and hardcode /opt/thermograph, /etc/thermograph.env and
ports 8137/8080. That assumption ends here:

  vps1  75.119.132.91  Forgejo, Grafana/Loki, the portfolio site, and DEV
                       (own Postgres, mesh-only on 10.10.0.2:8137)
  vps2  169.58.46.181  PROD and BETA as two Swarm stacks sharing one
                       TimescaleDB instance, plus Centralis, Postfix, backups
  desktop              AI model hosting + flex Swarm capacity, no environment

Nothing live has moved; the ordered cutover is in
infra/deploy/RUNBOOK-vps1-vps2-cutover.md.

deploy/env-topology.sh is the single source of truth: env -> host, checkout,
branch, deploy mode, stack name, env file, LB ports, DB role/database, service
prefix. THERMOGRAPH_ENV is the input, and on vps2 it is the only thing
distinguishing a beta deploy from a prod one -- deploy.sh refuses to run when it
disagrees with the checkout it was invoked from. The host-wide secrets-env and
deploy-mode markers survive only as a fallback for a single-environment box.

Beta gets its own stack file rather than an overlay (a merge cannot REMOVE the
db service, and beta having no database of its own is the point). Its services
are prefixed beta-*: Swarm registers a service's short name as a DNS alias on
every network it joins, so two stacks both calling a service `web` on the shared
data network would let prod's frontend resolve a beta task. Prod's stack, env
and LB config are untouched.

One Postgres, two databases with two roles: deploy/db/provision-env-db.sh
creates thermograph_beta (NOSUPERUSER, owns only its own database, CONNECT
revoked from PUBLIC) plus a <role>_ro for ad-hoc queries, and refuses to run for
the environment that owns the instance -- doing so would demote prod's bootstrap
superuser.

Fixes that co-residency would otherwise have broken silently:

- CI secrets are keyed by host (VPS1_SSH_*, VPS2_SSH_*). SSH_* meant "beta" and
  also "the Forgejo box" because those shared a machine; that conflation is what
  once had the prod backup dumping beta.
- The nightly backup dumps BOTH application databases to separate off-box
  prefixes, and fails loudly on a missing one instead of skipping it.
- Alloy stopped deriving the `host` label from the node -- on vps2 that would
  have filed every beta line as prod, feeding prod's alert rules with beta's
  traffic. It is now derived per source, with a new `node` label for the machine,
  and beta's log volume is mounted via a vps2-only overlay.
- ops/dbq.sh, ops/iceberg.sh and the secrets seed scripts derived their SSH
  target and env-file path from the topology instead of hardcoding beta to
  75.119.132.91 -- which is vps1's address now.
- Dev's overlay binds DEV_BIND_ADDR (loopback by default, the mesh address on
  vps1) instead of 0.0.0.0: correct for a home LAN box, a public exposure of
  unreviewed branches on a VPS.

Terraform's hosts variable is keyed by machine with a nested environments map,
and main.tf flattens (host, environment) pairs so two environments on one box
cannot share a checkout. Caddy's single reference config is split per host.

Docs, onboarding and the runbooks are updated throughout, including the security
rationales that co-residency makes false: separate SSH credentials no longer put
a host boundary between beta and prod, and the boundary that remains is the
database and the filesystem.
2026-07-25 15:01:29 -07:00

224 lines
14 KiB
Text

# This is the REFERENCE for what /etc/thermograph.env contains — the full set of
# vars the app reads. On a host wired for the SOPS vault (an age key +
# /etc/thermograph/secrets-env), you do NOT hand-edit /etc/thermograph.env: it is
# rendered at deploy from the encrypted source of truth in deploy/secrets/*.yaml
# (see deploy/secrets/README.md). To change a secret there, `sops edit` + commit +
# deploy — not a host edit. (Legacy: Terraform can also render this file from tfvars,
# and a from-scratch host can start by copying this example and editing it.)
#
# Sourced by deploy/deploy.sh (and any `docker compose` invocation) so compose can
# interpolate it, AND loaded into the app container (env_file in docker-compose.yml).
# Anything secret the app needs — Postgres password, VAPID keys, auth secret —
# belongs here.
# Port uvicorn binds inside the container. The compose stack publishes it on the
# host loopback (127.0.0.1:8137) for Caddy to proxy to. Keep 8137.
PORT=8137
# --- Registry pull (deploy.sh / Terraform) ---------------------------------------
# deploy.sh logs in to git.thermograph.org's registry and `docker compose pull`s
# the image build-push.yml already pushed for the deployed commit (repo-split
# Stage 6), instead of building in place. A personal access token with at least
# read:package scope -- the same REGISTRY_TOKEN secret build-push.yml uses (that
# one needs write:package too, to push; this only needs to pull).
REGISTRY_TOKEN=
# --- PostgreSQL (docker-compose stack) ------------------------------------------
# The app and Postgres run as a docker-compose stack (see docker-compose.yml).
# POSTGRES_PASSWORD is the database password: compose uses it to initialize the
# postgres container AND to build the app's THERMOGRAPH_DATABASE_URL. It MUST be
# set here (the systemd unit sources this file so `docker compose up` can
# interpolate it). Change it from the default before the first `up`.
POSTGRES_PASSWORD=change-me
# The app's compose service already builds THERMOGRAPH_DATABASE_URL from
# POSTGRES_PASSWORD, so you normally DON'T need this. It's here for reference and
# for running the app outside compose against the same DB (keep the password in
# sync with POSTGRES_PASSWORD above).
#THERMOGRAPH_DATABASE_URL=postgresql+asyncpg://thermograph:change-me@db:5432/thermograph
# --- Historical archive (self-hosted Open-Meteo) --------------------------------
# Where the app fetches its 45-year daily history. Unset (default) → the public
# Open-Meteo archive API (rate-limited). On the self-hosting host, the
# docker-compose.openmeteo.yml overlay sets this to the internal service URL for
# you, so you normally DON'T set it here. Only pin it to run the app outside that
# overlay against a reachable Open-Meteo instance. Leave it UNSET rather than empty:
# an empty value is honored as-is and would break the fallback to the public API.
#THERMOGRAPH_ARCHIVE_URL=http://open-meteo-api:8080/v1/archive
# Mark the session cookie Secure — required behind Caddy's HTTPS. Set to 1 for
# prod and beta (both TLS-fronted); leave unset only for dev's plain-HTTP
# mesh-only URL (a Secure cookie is never sent over HTTP).
THERMOGRAPH_COOKIE_SECURE=1
# Pin these too (see their own sections below), so container restarts don't rotate
# them: THERMOGRAPH_AUTH_SECRET (else every emailed confirm/reset link breaks on
# restart) and THERMOGRAPH_VAPID_PRIVATE_KEY / _PUBLIC_KEY (else every existing push
# subscription silently stops delivering). The data dir persists on the appdata
# volume, but pinning here is the safe default.
# Number of uvicorn worker processes. More than 1 stops a single slow upstream fetch
# (e.g. a cache-miss weather lookup) from blocking every other request — the cause of
# past brief outages. Prod's Swarm `web` service defaults to 4 (WEB_WORKERS, also the
# autoscaler's per-replica count); beta's defaults to 2 (BETA_WEB_WORKERS) — a
# rehearsal environment, not a decision that beta needs less headroom per se. The
# base compose file (dev) also defaults to 4. Leave unset (defaults to 1) on a small
# box. Workers elect one leader for the subscription notifier via a lockfile
# (THERMOGRAPH_SINGLETON_LOCK, set by the compose/stack app service to
# /app/data/notifier.lock) so its timer-driven upstream sweep runs once, not once per
# worker. ~200 MB RAM per worker.
WORKERS=4
# THERMOGRAPH_SINGLETON_LOCK arbitrates workers on ONE host. Under multi-host Swarm,
# each host would independently elect its own leader — multiplying the Open-Meteo
# quota use N-fold again. Set THERMOGRAPH_SINGLETON_PG=1 (with THERMOGRAPH_DATABASE_URL
# pointing at Postgres) to switch to a cluster-wide Postgres advisory lock instead, so
# exactly one host — not one per host — runs the notifier. Leave unset on a single-host
# deploy (today's default); the flock above is sufficient there.
#THERMOGRAPH_SINGLETON_PG=1
# Which duties this process performs. Every replica runs the same image; ROLE just
# decides whether it's allowed to own the notifier once it wins the leader election
# above — this is what lets the web tier scale to N stateless replicas under Swarm
# without also scaling notifier instances, while a single worker replica owns it.
# all -> both (default; today's single-process behavior, unchanged)
# web -> never runs the notifier, even if it would win leader election
# worker -> runs the notifier if it wins leader election
# Read once at process start; changing it needs a restart, not a live toggle.
#THERMOGRAPH_ROLE=all
# The worker's own recurring jobs (city warming, IndexNow), gated by the same
# leader election as the notifier above. Hours between runs; both are cheap
# no-op skips when there's nothing to do, so the defaults rarely need changing.
#THERMOGRAPH_WARM_CITIES_INTERVAL_HOURS=24
#THERMOGRAPH_INDEXNOW_INTERVAL_HOURS=6
# Base path the app is served under.
# / -> app at the domain root (Thermograph owns the whole domain —
# this is what the Caddyfile expects: thermograph.org proxies "/")
# /thermograph -> app under a sub-path (share the host with other apps, e.g. a
# portfolio at the root)
THERMOGRAPH_BASE=/
# --- SEO: search-engine verification + IndexNow ---------------------------------
# Ownership-verification tokens, rendered as <meta> tags in every page's <head>.
# Google Search Console → add property https://thermograph.org → "HTML tag" method
# → paste just the content="…" value below. (Or verify via DNS TXT and skip this.)
#THERMOGRAPH_GOOGLE_VERIFY=
# Bing Webmaster Tools → add site → "HTML Meta Tag" (msvalidate.01) → paste the
# content value. (Bing can also import verification from Google Search Console.)
#THERMOGRAPH_BING_VERIFY=
# IndexNow key (Bing/DuckDuckGo/Yandex instant re-crawl). Auto-generated to
# data/indexnow_key.txt on first use; set here to pin a specific key.
#THERMOGRAPH_INDEXNOW_KEY=
# Public site URL for IndexNow. The deploy hook auto-pings IndexNow after a
# successful deploy, but only when the URL set changed (a new/removed city), so
# code-only deploys don't resubmit. `make indexnow` forces a full submit.
THERMOGRAPH_BASE_URL=https://thermograph.org
# --- Web Push (VAPID) -----------------------------------------------------------
# Keys that sign push notifications. If unset, the app generates a pair into
# data/vapid.json on first run — fine as long as that file PERSISTS (it lives in the
# writable data dir and survives deploys). PIN them here to be safe: if the keys ever
# change, every existing browser subscription silently stops receiving (the push
# service rejects with 401/403), and users must toggle alerts off/on to re-subscribe.
# Generate a pair: cd backend && ../.venv/bin/python -c "import push,json; k=push._generate(); print('PRIVATE=',k['private_key']); print('PUBLIC=',k['public_key'])"
#THERMOGRAPH_VAPID_PRIVATE_KEY=
#THERMOGRAPH_VAPID_PUBLIC_KEY=
# Contact (mailto: or https URL) sent to push services in the VAPID claim.
#THERMOGRAPH_VAPID_CONTACT=mailto:you@example.com
# --- Outbound email --------------------------------------------------------------
# Delivery goes through the host's Postfix null client (deploy/provision-mail.sh).
# The app runs in a container, so it can't reach the host's loopback — it speaks
# plain SMTP to the compose bridge's gateway (172.19.0.1, pinned in
# docker-compose.yml), where Postfix listens and relays out. Switching between
# "direct to MX" and "relay through a provider" is a Postfix change, no redeploy.
# Stack-mode hosts (prod) override this per-service to the docker_gwbridge
# gateway (172.18.0.1) in deploy/stack/thermograph-stack.yml -- overlay tasks
# can't reach a compose bridge gateway; leave this file's value as the compose
# default.
#
# WARNING -- this host's rendered value must name an address Postfix ACTUALLY
# LISTENS ON (`postconf -h inet_interfaces`), or mail fails at send time with
# connection-refused. prod's SOPS-rendered /etc/thermograph.env currently says
# 172.19.0.1, which is the *compose* bridge and is NOT in prod's
# inet_interfaces. It is inert today only because prod runs in stack mode, where
# thermograph-stack.yml sets THERMOGRAPH_SMTP_HOST per-service and nothing on
# the box reads this variable from here. It becomes live breakage the moment
# prod falls back to compose mode. Fix it in deploy/secrets/prod.yaml (sops)
# rather than on the box; 10.10.0.1 (the mesh address) is the durable choice,
# since it depends on wg0 rather than on Docker.
#
# Backends: console (log it, send nothing — the default, right for dev),
# smtp (actually send), disabled (drop silently).
# Leave unset until Postfix is provisioned: signups are still collected either way.
# THERMOGRAPH_MAIL_FROM is left unset here on purpose — the app default
# (Thermograph <no-reply@thermograph.org>) is correct, and its "<>" would need
# escaping in this shell-sourced file. Override only if the address differs.
#THERMOGRAPH_MAIL_BACKEND=smtp
#THERMOGRAPH_SMTP_HOST=172.19.0.1
#THERMOGRAPH_SMTP_PORT=25
# Only needed if talking to a remote SMTP server directly instead of local Postfix.
#THERMOGRAPH_SMTP_USER=
#THERMOGRAPH_SMTP_PASSWORD=
#THERMOGRAPH_SMTP_STARTTLS=1
#THERMOGRAPH_MAIL_FROM=Thermograph <no-reply@thermograph.org>
#THERMOGRAPH_MAIL_REPLY_TO=
# Signing secret for email confirmation / password-reset tokens. MUST be set to a
# fixed value before any such link is mailed: it defaults to a per-boot random
# value, which would invalidate every outstanding link on each restart.
# generate with: python -c "import secrets; print(secrets.token_urlsafe(48))"
#THERMOGRAPH_AUTH_SECRET=
# --- Discord ---------------------------------------------------------------------
# Incoming webhook URL for the daily "most unusual right now" post. Create it in the
# Discord server: Channel → Edit → Integrations → Webhooks → New Webhook → Copy URL.
# The URL IS the credential — anyone who has it can post as the webhook, so keep it
# here and never in the repo. Unset => the daily post is disabled (no-op).
# The post rides the notifier daemon (leader-only), once per day after the feed
# refresh, so it needs THERMOGRAPH_ENABLE_NOTIFIER on (the default).
#THERMOGRAPH_DISCORD_WEBHOOK=
# Slash commands (/grade) are served over Discord's HTTP interactions endpoint by
# the app itself (no bot process). Set the portal's "Interactions Endpoint URL" to
# https://thermograph.org/discord/interactions. These come from the Developer Portal:
# - PUBLIC_KEY: General Information -> Public Key (used to verify every request).
# - APP_ID / BOT_TOKEN: only needed to (re)register the commands, via
# scripts/register_discord_commands.py. The bot token is a credential.
#THERMOGRAPH_DISCORD_PUBLIC_KEY=
#THERMOGRAPH_DISCORD_APP_ID=
#THERMOGRAPH_DISCORD_BOT_TOKEN=
# Channel IDs (numeric snowflakes) for the bot to post into with the BOT_TOKEN above.
# Subscription test feed: this deployment mirrors every notification the notifier
# creates into this channel, so a run can be watched live. Set it per environment;
# leave unset to disable. In the Thermograph.org server the hidden channels are:
# dev = 1529274513009934336
# uat = 1529274514066636861
# prod = 1529274515127799940
#THERMOGRAPH_DISCORD_SUBSCRIPTION_CHANNEL=
# Notable-weather feed: prod broadcasts the daily "most unusual right now" digest
# into this channel via the bot (independent of the WEBHOOK above — either can run
# alone). Leave unset to disable. In the Thermograph.org server:
# weather-events = 1529274516746932307
#THERMOGRAPH_DISCORD_WEATHER_CHANNEL=
# Discord gateway bot (@mention/DM grading): runs INSIDE the backend leader
# process (same singleton election as the notifier — see web/app.py's lifespan),
# riding its event loop; no separate bot process. Opt-in: any truthy value here
# starts it, provided BOT_TOKEN above is also set (flag alone is a silent no-op).
# Discord allows ONE gateway connection per bot token, so enable this in exactly
# one environment — prod, the only vault with the token.
#THERMOGRAPH_DISCORD_BOT=
# Account linking (OAuth2 "identify"): lets a signed-in user connect their Discord
# account, storing their Discord user id for DM alerts. CLIENT_ID is the same App
# ID above. The CLIENT_SECRET is from OAuth2 -> Client Secret (a credential). In the
# portal, add the redirect: https://thermograph.org/api/v2/discord/link/callback
#THERMOGRAPH_DISCORD_CLIENT_SECRET=
# Gateway bot (opt-in): holds a live websocket so the bot replies to messages that
# @mention it (or DM it) with a city grade — e.g. "@Thermograph Phoenix". Reuses
# THERMOGRAPH_DISCORD_BOT_TOKEN above. Runs on the single notifier leader only, so
# it needs THERMOGRAPH_ENABLE_NOTIFIER on (the default). No privileged intent
# required — Discord delivers content for mentions/DMs. Unset/0 => no gateway
# connection. (Code: thermograph-backend notifications/discord_bot.py.)
#THERMOGRAPH_DISCORD_BOT=1