* Split web/worker duties with THERMOGRAPH_ROLE Background work (the subscription notifier) is welded to the same process that serves requests, so scaling the web tier to N replicas would also scale notifier instances unless something restricts it further than leader election alone. Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process behavior). Every replica runs the same image; ROLE only gates whether a process is allowed to own the notifier at all, layered on top of the existing leader election: web replicas never start it even if they'd win leader election, worker replicas start it if they win. The decision is pulled into _should_run_notifier() so it's unit-testable without booting the full app (DB init, places index, neighbor warmer). Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE) so a worker replica - which serves no real traffic - still has something Swarm can health-check. * Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate Three changes toward the hop-1 interim cutover, all inert until Track B stands up the platform: docker-stack.yml: the Swarm stack file for the interim cutover, distinct from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a pre-built image (IMAGE_TAG) instead of building in place; app/worker publish no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted), reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard needs a smaller MTU or large payloads silently stall); db is placement- pinned to a labelled node; app/worker skip inline migrations (RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing that ever runs Alembic; secrets are real Swarm secrets mounted at /run/secrets, read by the entrypoint shim rather than plain env vars. TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default latest-pg18, today's behavior unchanged), wired through Terraform (timescaledb_tag, default "latest-pg18") so it can actually be pinned to an exact minor without hand-editing the host - required before any host of the stack could replicate with another (a floating tag risks mismatched extension minors, which blocks a physical replica and risks compressed- chunk corruption on restore). Caddy active health-gate: both the Terraform-rendered Caddyfile and the live deploy/Caddyfile now health-check the app on the same cheap /healthz route its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage, so it's cheap enough for a tight interval and works identically for a worker replica, which serves no public traffic at all) - Caddy won't forward into a container that's still booting or unhealthy. Verified live: built and booted the real image via docker compose - both containers report healthy via the new /healthz-based HEALTHCHECK, and GET / still renders the full SSR homepage unchanged. Both Caddyfiles validated with the real caddy binary. docker-stack.yml validated with docker compose config (required-var guards fire with clear messages; secrets correctly mount at /run/secrets/<name>, matching the entrypoint shim's mapping). docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside the existing openmeteo overlay. terraform validate + fmt clean.
152 lines
9.1 KiB
Text
152 lines
9.1 KiB
Text
# Copy to /etc/thermograph.env on the VPS and edit.
|
|
# Read by the systemd unit (EnvironmentFile) so `docker compose up` can interpolate
|
|
# it, AND loaded into the app container (env_file in docker-compose.yml). Anything
|
|
# secret the app needs — Postgres password, VAPID keys, auth secret — belongs here.
|
|
|
|
# Port uvicorn binds inside the container. The compose stack publishes it on the
|
|
# host loopback (127.0.0.1:8137) for Caddy to proxy to. Keep 8137.
|
|
PORT=8137
|
|
|
|
# --- PostgreSQL (docker-compose stack) ------------------------------------------
|
|
# The app and Postgres run as a docker-compose stack (see docker-compose.yml).
|
|
# POSTGRES_PASSWORD is the database password: compose uses it to initialize the
|
|
# postgres container AND to build the app's THERMOGRAPH_DATABASE_URL. It MUST be
|
|
# set here (the systemd unit sources this file so `docker compose up` can
|
|
# interpolate it). Change it from the default before the first `up`.
|
|
POSTGRES_PASSWORD=change-me
|
|
|
|
# The app's compose service already builds THERMOGRAPH_DATABASE_URL from
|
|
# POSTGRES_PASSWORD, so you normally DON'T need this. It's here for reference and
|
|
# for running the app outside compose against the same DB (keep the password in
|
|
# sync with POSTGRES_PASSWORD above).
|
|
#THERMOGRAPH_DATABASE_URL=postgresql+asyncpg://thermograph:change-me@db:5432/thermograph
|
|
|
|
# --- Historical archive (self-hosted Open-Meteo) --------------------------------
|
|
# Where the app fetches its 45-year daily history. Unset (default) → the public
|
|
# Open-Meteo archive API (rate-limited). On the self-hosting host, the
|
|
# docker-compose.openmeteo.yml overlay sets this to the internal service URL for
|
|
# you, so you normally DON'T set it here. Only pin it to run the app outside that
|
|
# overlay against a reachable Open-Meteo instance. Leave it UNSET rather than empty:
|
|
# an empty value is honored as-is and would break the fallback to the public API.
|
|
#THERMOGRAPH_ARCHIVE_URL=http://open-meteo-api:8080/v1/archive
|
|
|
|
# Mark the session cookie Secure — required behind Caddy's HTTPS. Set to 1 in prod;
|
|
# leave unset only for plain-HTTP LAN dev (a Secure cookie is never sent over HTTP).
|
|
THERMOGRAPH_COOKIE_SECURE=1
|
|
|
|
# Pin these too (see their own sections below), so container restarts don't rotate
|
|
# them: THERMOGRAPH_AUTH_SECRET (else every emailed confirm/reset link breaks on
|
|
# restart) and THERMOGRAPH_VAPID_PRIVATE_KEY / _PUBLIC_KEY (else every existing push
|
|
# subscription silently stops delivering). The data dir persists on the appdata
|
|
# volume, but pinning here is the safe default.
|
|
|
|
# Number of uvicorn worker processes. More than 1 stops a single slow upstream fetch
|
|
# (e.g. a cache-miss weather lookup) from blocking every other request — the cause of
|
|
# past brief outages. Prod runs 4 (the compose app service also defaults to WORKERS=4);
|
|
# leave unset (defaults to 1) on a small box. Workers elect one leader for the
|
|
# subscription notifier via a lockfile (THERMOGRAPH_SINGLETON_LOCK, set by the compose
|
|
# app service to /app/data/notifier.lock) so its timer-driven upstream sweep runs once,
|
|
# not once per worker. ~200 MB RAM per worker.
|
|
WORKERS=4
|
|
|
|
# THERMOGRAPH_SINGLETON_LOCK arbitrates workers on ONE host. Under multi-host Swarm,
|
|
# each host would independently elect its own leader — multiplying the Open-Meteo
|
|
# quota use N-fold again. Set THERMOGRAPH_SINGLETON_PG=1 (with THERMOGRAPH_DATABASE_URL
|
|
# pointing at Postgres) to switch to a cluster-wide Postgres advisory lock instead, so
|
|
# exactly one host — not one per host — runs the notifier. Leave unset on a single-host
|
|
# deploy (today's default); the flock above is sufficient there.
|
|
#THERMOGRAPH_SINGLETON_PG=1
|
|
|
|
# Which duties this process performs. Every replica runs the same image; ROLE just
|
|
# decides whether it's allowed to own the notifier once it wins the leader election
|
|
# above — this is what lets the web tier scale to N stateless replicas under Swarm
|
|
# without also scaling notifier instances, while a single worker replica owns it.
|
|
# all -> both (default; today's single-process behavior, unchanged)
|
|
# web -> never runs the notifier, even if it would win leader election
|
|
# worker -> runs the notifier if it wins leader election
|
|
# Read once at process start; changing it needs a restart, not a live toggle.
|
|
#THERMOGRAPH_ROLE=all
|
|
|
|
# Base path the app is served under.
|
|
# / -> app at the domain root (Thermograph owns the whole domain —
|
|
# this is what the Caddyfile expects: thermograph.org proxies "/")
|
|
# /thermograph -> app under a sub-path (share the host with other apps, e.g. a
|
|
# portfolio at the root)
|
|
THERMOGRAPH_BASE=/
|
|
|
|
# --- SEO: search-engine verification + IndexNow ---------------------------------
|
|
# Ownership-verification tokens, rendered as <meta> tags in every page's <head>.
|
|
# Google Search Console → add property https://thermograph.org → "HTML tag" method
|
|
# → paste just the content="…" value below. (Or verify via DNS TXT and skip this.)
|
|
#THERMOGRAPH_GOOGLE_VERIFY=
|
|
# Bing Webmaster Tools → add site → "HTML Meta Tag" (msvalidate.01) → paste the
|
|
# content value. (Bing can also import verification from Google Search Console.)
|
|
#THERMOGRAPH_BING_VERIFY=
|
|
|
|
# IndexNow key (Bing/DuckDuckGo/Yandex instant re-crawl). Auto-generated to
|
|
# data/indexnow_key.txt on first use; set here to pin a specific key.
|
|
#THERMOGRAPH_INDEXNOW_KEY=
|
|
# Public site URL for IndexNow. The deploy hook auto-pings IndexNow after a
|
|
# successful deploy, but only when the URL set changed (a new/removed city), so
|
|
# code-only deploys don't resubmit. `make indexnow` forces a full submit.
|
|
THERMOGRAPH_BASE_URL=https://thermograph.org
|
|
|
|
# --- Web Push (VAPID) -----------------------------------------------------------
|
|
# Keys that sign push notifications. If unset, the app generates a pair into
|
|
# data/vapid.json on first run — fine as long as that file PERSISTS (it lives in the
|
|
# writable data dir and survives deploys). PIN them here to be safe: if the keys ever
|
|
# change, every existing browser subscription silently stops receiving (the push
|
|
# service rejects with 401/403), and users must toggle alerts off/on to re-subscribe.
|
|
# Generate a pair: cd backend && ../.venv/bin/python -c "import push,json; k=push._generate(); print('PRIVATE=',k['private_key']); print('PUBLIC=',k['public_key'])"
|
|
#THERMOGRAPH_VAPID_PRIVATE_KEY=
|
|
#THERMOGRAPH_VAPID_PUBLIC_KEY=
|
|
# Contact (mailto: or https URL) sent to push services in the VAPID claim.
|
|
#THERMOGRAPH_VAPID_CONTACT=mailto:you@example.com
|
|
|
|
# --- Outbound email --------------------------------------------------------------
|
|
# Delivery goes through a local Postfix null client on 127.0.0.1:25 — see
|
|
# deploy/provision-mail.sh. The app only ever speaks plain SMTP to loopback, so
|
|
# switching between "direct to MX" and "relay through a provider" is a Postfix
|
|
# config change and needs no code change or redeploy.
|
|
#
|
|
# Backends: console (log it, send nothing — the default, right for dev),
|
|
# smtp (actually send), disabled (drop silently).
|
|
# Leave unset until Postfix is provisioned: signups are still collected either way.
|
|
#THERMOGRAPH_MAIL_BACKEND=smtp
|
|
#THERMOGRAPH_SMTP_HOST=127.0.0.1
|
|
#THERMOGRAPH_SMTP_PORT=25
|
|
# Only needed if talking to a remote SMTP server directly instead of local Postfix.
|
|
#THERMOGRAPH_SMTP_USER=
|
|
#THERMOGRAPH_SMTP_PASSWORD=
|
|
#THERMOGRAPH_SMTP_STARTTLS=1
|
|
#THERMOGRAPH_MAIL_FROM=Thermograph <no-reply@thermograph.org>
|
|
#THERMOGRAPH_MAIL_REPLY_TO=
|
|
|
|
# Signing secret for email confirmation / password-reset tokens. MUST be set to a
|
|
# fixed value before any such link is mailed: it defaults to a per-boot random
|
|
# value, which would invalidate every outstanding link on each restart.
|
|
# generate with: python -c "import secrets; print(secrets.token_urlsafe(48))"
|
|
#THERMOGRAPH_AUTH_SECRET=
|
|
|
|
# --- Discord ---------------------------------------------------------------------
|
|
# Incoming webhook URL for the daily "most unusual right now" post. Create it in the
|
|
# Discord server: Channel → Edit → Integrations → Webhooks → New Webhook → Copy URL.
|
|
# The URL IS the credential — anyone who has it can post as the webhook, so keep it
|
|
# here and never in the repo. Unset => the daily post is disabled (no-op).
|
|
# The post rides the notifier daemon (leader-only), once per day after the feed
|
|
# refresh, so it needs THERMOGRAPH_ENABLE_NOTIFIER on (the default).
|
|
#THERMOGRAPH_DISCORD_WEBHOOK=
|
|
# Slash commands (/grade) are served over Discord's HTTP interactions endpoint by
|
|
# the app itself (no bot process). Set the portal's "Interactions Endpoint URL" to
|
|
# https://thermograph.org/discord/interactions. These come from the Developer Portal:
|
|
# - PUBLIC_KEY: General Information -> Public Key (used to verify every request).
|
|
# - APP_ID / BOT_TOKEN: only needed to (re)register the commands, via
|
|
# scripts/register_discord_commands.py. The bot token is a credential.
|
|
#THERMOGRAPH_DISCORD_PUBLIC_KEY=
|
|
#THERMOGRAPH_DISCORD_APP_ID=
|
|
#THERMOGRAPH_DISCORD_BOT_TOKEN=
|
|
# Account linking (OAuth2 "identify"): lets a signed-in user connect their Discord
|
|
# account, storing their Discord user id for DM alerts. CLIENT_ID is the same App
|
|
# ID above. The CLIENT_SECRET is from OAuth2 -> Client Secret (a credential). In the
|
|
# portal, add the redirect: https://thermograph.org/api/v2/discord/link/callback
|
|
#THERMOGRAPH_DISCORD_CLIENT_SECRET=
|