Subtree-merge thermograph-infra (origin/main) into infra/
git-subtree-dir: infra git-subtree-mainline:d6df04eab2git-subtree-split:99b4b3f78d
This commit is contained in:
commit
ae1d9bb534
77 changed files with 6829 additions and 0 deletions
78
infra/.claude/skills/key-gaps/SKILL.md
Normal file
78
infra/.claude/skills/key-gaps/SKILL.md
Normal file
|
|
@ -0,0 +1,78 @@
|
|||
---
|
||||
name: key-gaps
|
||||
description: Audit which secret keys are missing, partially configured, or drifting across Thermograph environments (prod, beta, dev) — comparing the live /etc/thermograph.env on each box and/or the SOPS vault files (deploy/secrets/*.yaml). Use when asked about "key gaps", "missing keys/secrets", "what's not set", cross-environment secret drift, or as a pre-deploy / pre-cutover safety check. Reads key NAMES only, never values.
|
||||
---
|
||||
|
||||
# Key-gap audit
|
||||
|
||||
Reports, per environment and across them, where secret **keys** are missing or
|
||||
inconsistent — with special weight on the failure modes that bite silently here.
|
||||
|
||||
Run the deterministic script; don't eyeball it:
|
||||
|
||||
```
|
||||
python3 .claude/skills/key-gaps/key_gaps.py <env>=<source> [<env>=<source> ...]
|
||||
```
|
||||
|
||||
`<source>` is either a **SOPS file** (`deploy/secrets/prod.yaml`) or a **key-list
|
||||
file** (one `KEY` per line, or `KEY=`/`KEY:` lines). Keys are readable in a SOPS file
|
||||
even encrypted, so **no age key or decryption is needed**. Exit status is non-zero if
|
||||
any CRITICAL or required gap is found (so it also works as a CI/pre-deploy gate).
|
||||
|
||||
## Audit the SOPS vault (offline, no SSH)
|
||||
|
||||
```
|
||||
python3 .claude/skills/key-gaps/key_gaps.py \
|
||||
prod=deploy/secrets/prod.yaml beta=deploy/secrets/beta.yaml
|
||||
```
|
||||
|
||||
## Audit the LIVE boxes (read-only, key names only)
|
||||
|
||||
Gather the key names over SSH (never the values), then audit. Hosts/keys per INFRA.md:
|
||||
|
||||
```sh
|
||||
K=~/.ssh/thermograph_agent_ed25519
|
||||
# sudo: /etc/thermograph.env is root-owned (0640). On prod `agent` can read it
|
||||
# directly too, but sudo works uniformly on both boxes.
|
||||
ssh -i $K agent@169.58.46.181 'sudo grep -oE "^[A-Z_]+=" /etc/thermograph.env' > /tmp/prod.keys # prod
|
||||
ssh -i $K agent@75.119.132.91 'sudo grep -oE "^[A-Z_]+=" /etc/thermograph.env' > /tmp/beta.keys # beta
|
||||
python3 .claude/skills/key-gaps/key_gaps.py prod=/tmp/prod.keys beta=/tmp/beta.keys
|
||||
```
|
||||
|
||||
## Verify a SOPS cutover matches live
|
||||
|
||||
Before flipping a host onto the vault, confirm the rendered set equals the live set —
|
||||
mix the two source types for the same box:
|
||||
|
||||
```
|
||||
python3 .claude/skills/key-gaps/key_gaps.py live=/tmp/prod.keys vault=deploy/secrets/prod.yaml
|
||||
```
|
||||
The `Cross-environment drift` section should report **none** for tracked keys.
|
||||
|
||||
## Reading the output
|
||||
|
||||
- **CRITICAL** — a required, *self-generating* secret is missing (`THERMOGRAPH_AUTH_SECRET`,
|
||||
`THERMOGRAPH_VAPID_PRIVATE_KEY/_PUBLIC_KEY`). Missing here is the worst case: the app
|
||||
silently mints a new value on boot, invalidating every login session and push
|
||||
subscription. Fix before deploying.
|
||||
- **MISSING** — a required secret is absent (`POSTGRES_PASSWORD`, `REGISTRY_TOKEN`).
|
||||
- **Feature GAP** — a feature is *partially* configured (e.g. `discord-account-linking`
|
||||
has the app id but not the client secret), so it's silently broken. A fully-unset
|
||||
feature is "off", not a gap.
|
||||
- **Dependent-key GAP** — an optional key is set but the *other* key it needs to do
|
||||
anything isn't (e.g. `THERMOGRAPH_DISCORD_WEATHER_CHANNEL` set with no
|
||||
`THERMOGRAPH_DISCORD_BOT_TOKEN` — discord.py gates the post on both). Asymmetric,
|
||||
unlike a feature group: the prerequisite key is fine set alone (the bot token alone
|
||||
already enables DMs).
|
||||
- **drift** — a tracked key is in some environments but not others; often intentional,
|
||||
but worth a glance.
|
||||
|
||||
## Keeping it current
|
||||
|
||||
The secret manifest (which keys are required / self-generating / grouped into features /
|
||||
dependent on another key) lives at the top of `key_gaps.py` in `REQUIRED`,
|
||||
`OPTIONAL_SELF_GEN`, `FEATURE_GROUPS`, `DEPENDENT_OPTIONAL`, and `STANDALONE_OPTIONAL`.
|
||||
When a new `THERMOGRAPH_*` credential is added to `deploy/thermograph.env.example` /
|
||||
`deploy/secrets/`, add it there too — check the code path that reads it (`os.environ.get`
|
||||
call site) to see whether it's truly required, self-generating, part of an all-or-nothing
|
||||
feature, or only meaningful alongside another key, rather than guessing from the name.
|
||||
216
infra/.claude/skills/key-gaps/key_gaps.py
Executable file
216
infra/.claude/skills/key-gaps/key_gaps.py
Executable file
|
|
@ -0,0 +1,216 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Audit secret-key coverage across Thermograph environments.
|
||||
|
||||
Reports, per environment and across them:
|
||||
- CRITICAL gaps: a required, *self-generating* secret is missing. Missing here is
|
||||
the worst case — the app silently mints a new value on boot (AUTH_SECRET, VAPID,
|
||||
INDEXNOW), invalidating every session / push subscription.
|
||||
- Required gaps: a required secret is missing.
|
||||
- Feature gaps: a feature is *partially* configured (some keys present, some not),
|
||||
so it's silently broken. Fully-unset features are "off", not a gap.
|
||||
- Dependent-key gaps: an optional key is set but the OTHER key it needs to do
|
||||
anything isn't (e.g. a Discord channel id with no bot token) — asymmetric,
|
||||
unlike feature gaps: the prerequisite is fine set alone.
|
||||
- Drift: a key present in one environment but absent in another.
|
||||
|
||||
It reads only KEY NAMES, never values — safe to run anywhere. Sources per environment:
|
||||
* a SOPS-encrypted YAML (deploy/secrets/<env>.yaml) — keys are plaintext even when
|
||||
encrypted, so no age key or decryption is needed; or
|
||||
* a plain key-list file (one KEY per line, or KEY=... lines) — e.g. the output of
|
||||
`ssh <box> 'grep -oE "^[A-Z_]+=" /etc/thermograph.env'` for a live audit.
|
||||
|
||||
Usage:
|
||||
key_gaps.py prod=deploy/secrets/prod.yaml beta=deploy/secrets/beta.yaml
|
||||
key_gaps.py prod=/tmp/prod.keys # a live key-list gathered over SSH
|
||||
Exit status is non-zero if any CRITICAL or required gap is found (usable in CI).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
import sys
|
||||
|
||||
# --- what counts as a secret/key, and how it behaves --------------------------
|
||||
# category: "required" (app needs it) | "optional" (fine to omit)
|
||||
# self_gen: missing => the app silently generates a replacement (data-loss hazard)
|
||||
REQUIRED = {
|
||||
# key: self_generating?
|
||||
"POSTGRES_PASSWORD": False,
|
||||
"THERMOGRAPH_AUTH_SECRET": True,
|
||||
"THERMOGRAPH_VAPID_PRIVATE_KEY": True,
|
||||
"THERMOGRAPH_VAPID_PUBLIC_KEY": True,
|
||||
"REGISTRY_TOKEN": False,
|
||||
}
|
||||
OPTIONAL_SELF_GEN = {
|
||||
"THERMOGRAPH_INDEXNOW_KEY": True, # regenerates, but persisted to a file — low risk
|
||||
}
|
||||
# Features that only work if ALL their keys are set together; some-but-not-all = broken.
|
||||
FEATURE_GROUPS = {
|
||||
"discord-account-linking": ["THERMOGRAPH_DISCORD_APP_ID", "THERMOGRAPH_DISCORD_CLIENT_SECRET"],
|
||||
"mail-smtp": ["THERMOGRAPH_SMTP_USER", "THERMOGRAPH_SMTP_PASSWORD"],
|
||||
}
|
||||
# Optional keys that only do anything when a specific OTHER key is also set. Unlike
|
||||
# FEATURE_GROUPS the relationship is asymmetric: the prerequisite is fine set alone
|
||||
# (THERMOGRAPH_DISCORD_BOT_TOKEN alone already enables DMs), but the dependent key
|
||||
# alone is a silent no-op (discord.py gates both channel posts on
|
||||
# `bool(BOT_TOKEN and <channel>)`) — flagging it as a symmetric feature group would
|
||||
# false-positive on every environment that has the bot token for DMs only.
|
||||
DEPENDENT_OPTIONAL = {
|
||||
"THERMOGRAPH_DISCORD_WEATHER_CHANNEL": "THERMOGRAPH_DISCORD_BOT_TOKEN",
|
||||
"THERMOGRAPH_DISCORD_SUBSCRIPTION_CHANNEL": "THERMOGRAPH_DISCORD_BOT_TOKEN",
|
||||
# Gateway-bot enable flag: truthy alone does nothing (web/app.py only starts the
|
||||
# bot when discord_bot.enabled() sees the token too). One gateway connection per
|
||||
# token — set in exactly one environment.
|
||||
"THERMOGRAPH_DISCORD_BOT": "THERMOGRAPH_DISCORD_BOT_TOKEN",
|
||||
}
|
||||
# Standalone optional keys: reported present/absent, never a "gap" on their own.
|
||||
STANDALONE_OPTIONAL = [
|
||||
"THERMOGRAPH_METRICS_TOKEN", # ops metrics remote access (else loopback-only, not broken)
|
||||
"THERMOGRAPH_DISCORD_WEBHOOK", # daily post
|
||||
"THERMOGRAPH_DISCORD_PUBLIC_KEY", # slash-command interactions
|
||||
"THERMOGRAPH_DISCORD_BOT_TOKEN", # bot DMs / gateway bot (prerequisite for the two below)
|
||||
]
|
||||
|
||||
C = {"red": "\033[31m", "yellow": "\033[33m", "green": "\033[32m",
|
||||
"dim": "\033[2m", "bold": "\033[1m", "off": "\033[0m"}
|
||||
|
||||
|
||||
def keys_from_source(path: str) -> set[str]:
|
||||
"""Extract KEY names from a SOPS yaml or a plain key-list file (values ignored)."""
|
||||
with open(path, encoding="utf-8", errors="replace") as fh:
|
||||
text = fh.read()
|
||||
keys: set[str] = set()
|
||||
in_sops_block = False
|
||||
for raw in text.splitlines():
|
||||
line = raw.rstrip("\n")
|
||||
if line.startswith("sops:"): # SOPS metadata block — skip its subkeys
|
||||
in_sops_block = True
|
||||
continue
|
||||
if in_sops_block:
|
||||
if line and not line[0].isspace():
|
||||
in_sops_block = False # dedented back to a real key
|
||||
else:
|
||||
continue
|
||||
s = line.strip()
|
||||
if not s or s.startswith("#"):
|
||||
continue
|
||||
# Accept "KEY: ..." (yaml), "KEY=..." (dotenv), or a bare "KEY" (key-list).
|
||||
# The whole stripped line must be just an identifier optionally followed by
|
||||
# a : or = — so prose/nested lines don't get mistaken for keys.
|
||||
m = re.match(r"^([A-Za-z_][A-Za-z0-9_]*)\s*(?:[:=].*)?$", s)
|
||||
if m:
|
||||
keys.add(m.group(1))
|
||||
return keys
|
||||
|
||||
|
||||
def audit(envs: dict[str, set[str]]) -> int:
|
||||
names = list(envs)
|
||||
print(f"{C['bold']}Key-gap audit — {', '.join(names)}{C['off']}")
|
||||
for n in names:
|
||||
print(f" {C['dim']}{n}: {len(envs[n])} keys{C['off']}")
|
||||
print()
|
||||
critical = required = feature = dependent = 0
|
||||
|
||||
def present(env, key): # noqa: ANN001
|
||||
return key in envs[env]
|
||||
|
||||
# Required + self-generating
|
||||
print(f"{C['bold']}Required secrets{C['off']}")
|
||||
for key, self_gen in {**REQUIRED, **OPTIONAL_SELF_GEN}.items():
|
||||
req = key in REQUIRED
|
||||
missing = [n for n in names if not present(n, key)]
|
||||
if not missing:
|
||||
print(f" {C['green']}OK{C['off']} {key} (all environments)")
|
||||
continue
|
||||
if req and self_gen:
|
||||
critical += 1
|
||||
tag = f"{C['red']}CRITICAL{C['off']}"
|
||||
note = " — self-generates on boot: MISSING = every session/subscription invalidated"
|
||||
elif req:
|
||||
required += 1
|
||||
tag = f"{C['red']}MISSING {C['off']}"
|
||||
note = ""
|
||||
else:
|
||||
tag = f"{C['yellow']}note {C['off']}"
|
||||
note = " — self-generates (persisted to a file; low risk)"
|
||||
print(f" {tag} {key} absent in: {', '.join(missing)}{C['dim']}{note}{C['off']}")
|
||||
print()
|
||||
|
||||
# Feature groups — partial config = broken
|
||||
print(f"{C['bold']}Feature groups (all-or-nothing){C['off']}")
|
||||
for feat, keys in FEATURE_GROUPS.items():
|
||||
for n in names:
|
||||
have = [k for k in keys if present(n, k)]
|
||||
if not have:
|
||||
print(f" {C['dim']}off {feat} @ {n} (none set){C['off']}")
|
||||
elif len(have) == len(keys):
|
||||
print(f" {C['green']}OK{C['off']} {feat} @ {n} (enabled)")
|
||||
else:
|
||||
feature += 1
|
||||
miss = [k for k in keys if not present(n, k)]
|
||||
print(f" {C['red']}GAP {C['off']} {feat} @ {n}: has {have}, MISSING {miss}")
|
||||
print()
|
||||
|
||||
# Dependent optional — the key only does anything with its prerequisite also set
|
||||
print(f"{C['bold']}Dependent optional keys{C['off']}")
|
||||
for key, prereq in DEPENDENT_OPTIONAL.items():
|
||||
for n in names:
|
||||
if not present(n, key):
|
||||
print(f" {C['dim']}off {key} @ {n} (unset){C['off']}")
|
||||
elif present(n, prereq):
|
||||
print(f" {C['green']}OK{C['off']} {key} @ {n} (enabled, {prereq} set)")
|
||||
else:
|
||||
dependent += 1
|
||||
print(f" {C['red']}GAP {C['off']} {key} @ {n}: set but {prereq} is missing — silently no-ops")
|
||||
print()
|
||||
|
||||
# Standalone optional
|
||||
print(f"{C['bold']}Optional keys{C['off']}")
|
||||
for key in STANDALONE_OPTIONAL:
|
||||
where = [n for n in names if present(n, key)]
|
||||
state = f"set in: {', '.join(where)}" if where else "unset everywhere"
|
||||
print(f" {C['dim']}{key}: {state}{C['off']}")
|
||||
print()
|
||||
|
||||
# Cross-env drift
|
||||
if len(names) > 1:
|
||||
print(f"{C['bold']}Cross-environment drift{C['off']}")
|
||||
union = set().union(*envs.values())
|
||||
drift = False
|
||||
for key in sorted(union):
|
||||
where = [n for n in names if key in envs[n]]
|
||||
if len(where) != len(names) and (key in REQUIRED or key in OPTIONAL_SELF_GEN
|
||||
or any(key in g for g in FEATURE_GROUPS.values())
|
||||
or key in DEPENDENT_OPTIONAL
|
||||
or key in STANDALONE_OPTIONAL):
|
||||
drift = True
|
||||
absent = [n for n in names if n not in where]
|
||||
print(f" {C['yellow']}drift{C['off']} {key}: in {', '.join(where)}; not in {', '.join(absent)}")
|
||||
if not drift:
|
||||
print(f" {C['green']}none{C['off']} (tracked keys are consistent)")
|
||||
print()
|
||||
|
||||
total = critical + required + feature + dependent
|
||||
color = C["red"] if total else C["green"]
|
||||
print(f"{color}{C['bold']}Summary: {critical} critical, {required} required, "
|
||||
f"{feature} feature gaps, {dependent} dependent-key gaps{C['off']}")
|
||||
return 1 if (critical or required) else 0
|
||||
|
||||
|
||||
def main(argv: list[str]) -> int:
|
||||
args = [a for a in argv[1:] if "=" in a]
|
||||
if not args:
|
||||
print(__doc__)
|
||||
return 2
|
||||
envs: dict[str, set[str]] = {}
|
||||
for a in args:
|
||||
name, _, path = a.partition("=")
|
||||
try:
|
||||
envs[name] = keys_from_source(path)
|
||||
except OSError as e:
|
||||
print(f"!! {name}: cannot read {path}: {e}", file=sys.stderr)
|
||||
return 2
|
||||
return audit(envs)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main(sys.argv))
|
||||
42
infra/.env.example
Normal file
42
infra/.env.example
Normal file
|
|
@ -0,0 +1,42 @@
|
|||
# Local docker-compose overrides. Copy to .env (gitignored) for `docker compose
|
||||
# up` / `make up` on a dev box: cp .env.example .env
|
||||
#
|
||||
# Compose auto-reads a repo-root .env for ${VAR} interpolation in
|
||||
# docker-compose.yml. In production these live in /etc/thermograph.env instead
|
||||
# (loaded by the systemd unit), so this file is only for local runs.
|
||||
|
||||
# Database password. Compose uses it to initialize the postgres container AND to
|
||||
# build the app's THERMOGRAPH_DATABASE_URL. Change it before first `up`.
|
||||
POSTGRES_PASSWORD=change-me
|
||||
|
||||
# --- Everything below is optional -- docker-compose.yml already defaults each
|
||||
# --- of these, so a plain `make up` works with none of it set. Uncomment to
|
||||
# --- override.
|
||||
|
||||
# TimescaleDB image tag. Defaults to the floating latest-pg18 tag. Pin it to an
|
||||
# exact minor (e.g. 2.17.2-pg18) before any host of this stack could ever
|
||||
# replicate with another -- see docker-compose.yml's db service comment.
|
||||
# TIMESCALEDB_TAG=latest-pg18
|
||||
|
||||
# Postgres sizing. Terraform sets these per host in prod/beta (prod DB_MEMORY
|
||||
# 16g); local/beta default to 8g / 2 CPUs.
|
||||
# DB_MEMORY=8g
|
||||
# DB_CPUS=2
|
||||
|
||||
# Backend uvicorn worker count and CPU cap. Terraform raises these on bigger
|
||||
# hosts; defaults keep a plain `docker compose up` identical to before.
|
||||
# WORKERS=4
|
||||
# APP_CPUS=4
|
||||
|
||||
# Frontend CPU cap.
|
||||
# FRONTEND_CPUS=2
|
||||
|
||||
# Registry + per-service image path/tag. Local dev normally builds each image
|
||||
# in its own app repo (thermograph-backend / thermograph-frontend) tagged
|
||||
# :local, which is the default here -- only set these to pull a specific
|
||||
# published build instead of building locally.
|
||||
# REGISTRY_HOST=git.thermograph.org
|
||||
# BACKEND_IMAGE_PATH=emi/thermograph-backend/app
|
||||
# BACKEND_IMAGE_TAG=local
|
||||
# FRONTEND_IMAGE_PATH=emi/thermograph-frontend/app
|
||||
# FRONTEND_IMAGE_TAG=local
|
||||
104
infra/.forgejo/workflows/ops-cron.yml
Normal file
104
infra/.forgejo/workflows/ops-cron.yml
Normal file
|
|
@ -0,0 +1,104 @@
|
|||
name: Ops cron (backup + IndexNow)
|
||||
|
||||
# Scheduled operational jobs that don't belong in the app's own worker-tier
|
||||
# scheduler (notifications/scheduler.py, Track A chunk 5) because they're
|
||||
# infra/ops concerns rather than app-domain background work -- see the job
|
||||
# classification table in
|
||||
# thermograph-docs/architecture/repo-topology-and-infrastructure.md §7.
|
||||
#
|
||||
# Both jobs SSH into the prod host and run inside the already-running compose
|
||||
# stack (docker compose exec), the same way deploy.sh already runs its own
|
||||
# post-deploy IndexNow ping -- no new network exposure, no separate dependency
|
||||
# install. Runs on the `docker` label (the always-on Swarm-hosted runner
|
||||
# deploy/forgejo/ stood up).
|
||||
#
|
||||
# Targets PROD via the PROD_SSH_* secrets (prod = 169.58.46.181, `agent` user,
|
||||
# in the docker group so no sudo needed; /etc/thermograph.env is agent-readable).
|
||||
# These are the same secrets deploy-prod.yml uses for the release->prod deploy --
|
||||
# NOT the SSH_* secrets, which point at BETA (deploy.yml's `main`->beta path). An
|
||||
# earlier revision reused SSH_* here, so the "prod" backup was silently dumping
|
||||
# beta; prod itself had no backup at all. The prod database is the one that must
|
||||
# be backed up, so both jobs use PROD_SSH_*.
|
||||
|
||||
on:
|
||||
schedule:
|
||||
# 03:00 UTC daily -- a low-traffic window for both jobs.
|
||||
- cron: '0 3 * * *'
|
||||
workflow_dispatch: {}
|
||||
|
||||
jobs:
|
||||
backup:
|
||||
name: pg_dump backup
|
||||
runs-on: docker
|
||||
# Guards against the schedule and a manual workflow_dispatch landing
|
||||
# close together (or two manual triggers): queue rather than cancel, so
|
||||
# a workflow_dispatch never aborts a pg_dump mid-write and leaves a
|
||||
# truncated .dump file -- same cancel-in-progress:false rationale as
|
||||
# deploy.yml/deploy-dev.yml's own SSH-script jobs.
|
||||
concurrency:
|
||||
group: ops-backup
|
||||
cancel-in-progress: false
|
||||
steps:
|
||||
- name: Dump the prod database over SSH
|
||||
uses: https://github.com/appleboy/ssh-action@v1.2.0
|
||||
with:
|
||||
host: ${{ secrets.PROD_SSH_HOST }}
|
||||
username: ${{ secrets.PROD_SSH_USER }}
|
||||
key: ${{ secrets.PROD_SSH_KEY }}
|
||||
port: ${{ secrets.PROD_SSH_PORT }}
|
||||
script: |
|
||||
set -euo pipefail
|
||||
cd /opt/thermograph
|
||||
# Source the env so `docker compose` can interpolate POSTGRES_PASSWORD;
|
||||
# without it compose fails to parse docker-compose.yml, the redirect
|
||||
# still creates the target, and the job leaves a 0-byte .dump and exits
|
||||
# non-zero (exactly how this backup was silently failing). Mirrors the
|
||||
# IndexNow job below, which already sources it.
|
||||
set -a; . /etc/thermograph.env 2>/dev/null || true; set +a
|
||||
backup_dir="$HOME/thermograph-backups"
|
||||
mkdir -p "$backup_dir"
|
||||
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
|
||||
out="$backup_dir/thermograph-$stamp.dump"
|
||||
# The db may run under plain compose OR as a Swarm stack task
|
||||
# (prod post-cutover); resolve the container either way so the
|
||||
# backup survives the deploy-mode switch.
|
||||
dbc=$(docker ps -q --filter "label=com.docker.swarm.service.name=thermograph_db" | head -1)
|
||||
[ -z "$dbc" ] && dbc=$(cd /opt/thermograph && docker compose ps -q db 2>/dev/null | head -1)
|
||||
[ -n "$dbc" ] || { echo "!! no db container found (compose or stack)"; exit 1; }
|
||||
# Write to a .partial and rename on success so a mid-dump failure can
|
||||
# never leave a truncated file that looks like a good backup.
|
||||
docker exec "$dbc" pg_dump -U thermograph -d thermograph \
|
||||
--format=custom > "$out.partial"
|
||||
mv "$out.partial" "$out"
|
||||
echo "wrote $out ($(du -h "$out" | cut -f1))"
|
||||
# The dumps are the disaster-recovery copy, not a versioned
|
||||
# archive -- keep the last 14 days and let the rest age out.
|
||||
find "$backup_dir" -name 'thermograph-*.dump' -mtime +14 -delete
|
||||
find "$backup_dir" -name 'thermograph-*.dump.partial' -mtime +1 -delete
|
||||
|
||||
indexnow:
|
||||
name: IndexNow ping
|
||||
runs-on: docker
|
||||
# Same overlap guard as the backup job above (schedule vs. manual
|
||||
# dispatch); --if-changed already makes a second ping a cheap no-op, but
|
||||
# queueing avoids two SSH sessions racing on the host regardless.
|
||||
concurrency:
|
||||
group: ops-indexnow
|
||||
cancel-in-progress: false
|
||||
steps:
|
||||
- name: Ping IndexNow if the URL set changed
|
||||
uses: https://github.com/appleboy/ssh-action@v1.2.0
|
||||
with:
|
||||
host: ${{ secrets.PROD_SSH_HOST }}
|
||||
username: ${{ secrets.PROD_SSH_USER }}
|
||||
key: ${{ secrets.PROD_SSH_KEY }}
|
||||
port: ${{ secrets.PROD_SSH_PORT }}
|
||||
script: |
|
||||
set -euo pipefail
|
||||
cd /opt/thermograph
|
||||
set -a; . /etc/thermograph.env 2>/dev/null || true; set +a
|
||||
bec=$(docker ps -q --filter "label=com.docker.swarm.service.name=thermograph_web" | head -1)
|
||||
[ -z "$bec" ] && bec=$(cd /opt/thermograph && docker compose ps -q backend 2>/dev/null | head -1)
|
||||
[ -n "$bec" ] || { echo "!! no backend/web container found"; exit 1; }
|
||||
docker exec "$bec" python indexnow.py --if-changed \
|
||||
"${THERMOGRAPH_BASE_URL:-https://thermograph.org}"
|
||||
36
infra/.forgejo/workflows/secrets-guard.yml
Normal file
36
infra/.forgejo/workflows/secrets-guard.yml
Normal file
|
|
@ -0,0 +1,36 @@
|
|||
name: secrets-guard
|
||||
|
||||
# Fail the build if any deploy/secrets/*.yaml is committed UNENCRYPTED. SOPS stores a
|
||||
# `sops:` metadata block and wraps every value in ENC[...]; a plaintext leak has
|
||||
# neither. This is the backstop against the classic "committed the decrypted file by
|
||||
# accident" incident. .sops.yaml (the plaintext config) and README.md are not *.yaml
|
||||
# under deploy/secrets/, so they're correctly ignored.
|
||||
on:
|
||||
pull_request:
|
||||
push:
|
||||
branches: [dev, main, release]
|
||||
|
||||
jobs:
|
||||
encrypted:
|
||||
runs-on: docker
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- name: Assert deploy/secrets/*.yaml are SOPS-encrypted
|
||||
run: |
|
||||
set -euo pipefail
|
||||
shopt -s nullglob
|
||||
files=(deploy/secrets/*.yaml)
|
||||
if [ ${#files[@]} -eq 0 ]; then
|
||||
echo "no deploy/secrets/*.yaml yet — nothing to check"
|
||||
exit 0
|
||||
fi
|
||||
fail=0
|
||||
for f in "${files[@]}"; do
|
||||
if grep -q '^sops:' "$f" && grep -q 'ENC\[AES256_GCM' "$f"; then
|
||||
echo "ok: $f is SOPS-encrypted"
|
||||
else
|
||||
echo "::error file=$f::NOT SOPS-encrypted — refusing to accept a plaintext secrets file"
|
||||
fail=1
|
||||
fi
|
||||
done
|
||||
exit $fail
|
||||
25
infra/.gitignore
vendored
Normal file
25
infra/.gitignore
vendored
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
# Terraform state/vars are also ignored inside terraform/.gitignore; repeated here
|
||||
# in case a stray file ever lands at repo root.
|
||||
*.tfstate
|
||||
*.tfstate.*
|
||||
.terraform/
|
||||
*.tfvars
|
||||
!*.tfvars.example
|
||||
crash.log
|
||||
crash.*.log
|
||||
|
||||
# The SOPS+age private key must never be committed (deploy/secrets/*.yaml, the
|
||||
# encrypted values, are meant to be committed).
|
||||
age.key
|
||||
*.age.key
|
||||
|
||||
.DS_Store
|
||||
|
||||
# Host-side deploy state (deploy.sh): the live per-service image tags and the
|
||||
# cross-repo deploy lock. Untracked on purpose -- they must survive the
|
||||
# `git reset --hard` at the top of every deploy (deploy.sh's comments already
|
||||
# assumed .image-tags.env was ignored; make it actually true so a stray
|
||||
# `git clean` can't destroy the record of what's running).
|
||||
deploy/.image-tags.env
|
||||
deploy/.deploy.lock
|
||||
deploy/.stack-image-tags.env
|
||||
14
infra/.sops.yaml
Normal file
14
infra/.sops.yaml
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
# SOPS creation rules — how `sops <file>` encrypts secrets in this repo.
|
||||
#
|
||||
# Every deploy/secrets/*.yaml is encrypted to the age recipient below (values only;
|
||||
# keys stay readable for clean diffs). The matching PRIVATE key is NEVER in the repo
|
||||
# — it lives at ~/.config/sops/age/keys.txt on the operator's machine and at
|
||||
# /etc/thermograph/age.key on each host that renders secrets at deploy time.
|
||||
#
|
||||
# Rotating the age identity itself: add the new recipient here, run
|
||||
# sops updatekeys deploy/secrets/*.yaml
|
||||
# to re-encrypt to both, distribute the new private key, then drop the old recipient
|
||||
# and updatekeys again. See deploy/secrets/README.md.
|
||||
creation_rules:
|
||||
- path_regex: ^deploy/secrets/.*\.yaml$
|
||||
age: age1xx4dzs0dxlwvkv9sjuqzsphl7lfrxannkfken374yu2qvvcte9sqzktqt2
|
||||
162
infra/ACCESS.md
Normal file
162
infra/ACCESS.md
Normal file
|
|
@ -0,0 +1,162 @@
|
|||
# Agent access, Docker Swarm, and Forgejo CI/CD
|
||||
|
||||
This covers three additive infrastructure layers on top of the host
|
||||
provisioning / Postgres / Terraform work described in `terraform/README.md`
|
||||
and (for the historical Track A/Track B plan this whole effort grew from,
|
||||
including the decision to split this repo out of the app monorepo)
|
||||
`thermograph-docs/runbooks/implementation-handoff.md` and
|
||||
`thermograph-docs/architecture/repo-topology-and-infrastructure.md`. Terraform
|
||||
provisions **prod** (the new 48 GB / 12-core box, `thermograph.org`) and
|
||||
**beta** (the old VPS, `75.119.132.91`) — see `terraform/README.md` /
|
||||
`terraform.tfvars.example`. **None of this touches the app repo** (its
|
||||
backend/frontend source, `Dockerfile`, or CI) — this repo owns only how and
|
||||
where the already-built app image runs; the app repo owns building it.
|
||||
|
||||
```
|
||||
Track 1: deploy/provision-agent-access.sh — a dedicated full-root login for me
|
||||
Track 2: deploy/swarm/ — Swarm cluster spanning THREE nodes
|
||||
Track 3: deploy/forgejo/ + .forgejo/ — Forgejo + Forgejo Actions, replacing GitHub
|
||||
```
|
||||
|
||||
**Three nodes, not two**: prod, beta, and **the desktop** (the LAN dev
|
||||
machine — same box that already runs the pre-Forgejo GitHub self-hosted
|
||||
runner). This matches the canonical Track B design in the handoff doc; an
|
||||
earlier revision of this whole effort covered just prod+beta and has been
|
||||
realigned.
|
||||
|
||||
## Track 1 — Agent access
|
||||
|
||||
Run `sudo bash deploy/provision-agent-access.sh` on **prod and beta** (not
|
||||
the desktop — that's wherever you're already working from, no separate
|
||||
access-provisioning step needed there). Creates a dedicated `agent` user (not
|
||||
raw root login — a distinct name gives a clean audit trail) with passwordless
|
||||
sudo, installs the agent's public key, disables SSH password auth
|
||||
fleet-wide, and turns on `auditd` logging of every root-effective command.
|
||||
See the script's header comment for the full rationale and revocation steps
|
||||
(one line to delete, or delete the whole account — your own access is never
|
||||
affected).
|
||||
|
||||
### Current access state (live)
|
||||
|
||||
Both VPS boxes are provisioned and reachable as of this writing:
|
||||
|
||||
| Host | Role | Public IP | Login |
|
||||
|------|------|-----------|-------|
|
||||
| prod | Swarm manager; Thermograph's live prod home (`release`, thermograph.org) | `169.58.46.181` | `agent` |
|
||||
| beta | Swarm worker; `main`/beta.thermograph.org + hosts Forgejo | `75.119.132.91` | `agent` |
|
||||
| desktop | Swarm worker, LAN dev machine + Forgejo Actions runner | (local) | (already have access) |
|
||||
|
||||
```
|
||||
ssh -i ~/.ssh/thermograph_agent_ed25519 agent@169.58.46.181 # prod
|
||||
ssh -i ~/.ssh/thermograph_agent_ed25519 agent@75.119.132.91 # beta
|
||||
```
|
||||
|
||||
The private half of the agent's dedicated keypair lives at
|
||||
`~/.ssh/thermograph_agent_ed25519` on the operator's own machine (never in
|
||||
this repo, never in a CI secret) — same convention `DEPLOY.md` already uses
|
||||
for the separate CI deploy key (`~/.ssh/thermograph_ci`). If you're setting
|
||||
this up fresh elsewhere, generate a new pair the same way
|
||||
(`ssh-keygen -t ed25519 -a 100 -C "claude-agent@thermograph-infra" -f
|
||||
~/.ssh/thermograph_agent_ed25519 -N ""`), embed the new public half in
|
||||
`AGENT_PUBKEY` at the top of `provision-agent-access.sh`, and re-run it on
|
||||
both boxes — that replaces the authorized key wholesale rather than
|
||||
appending, so an old key stops working the moment the new one lands.
|
||||
|
||||
**Keep the private key somewhere that survives** — not a scratch/temp
|
||||
directory that gets cleaned up, not anywhere-in-repo. Losing it just means
|
||||
regenerating and re-running the provisioning script; it does not lock either
|
||||
box, since your own login is untouched.
|
||||
|
||||
**Live as of 2026-07-21** (was: "no Docker / Terraform not applied / Swarm
|
||||
inactive" — all now done): Docker is installed on all three nodes; the
|
||||
WireGuard mesh and Docker Swarm are up with all three nodes `Ready` (prod
|
||||
manager, beta + desktop workers); Terraform has been applied to both VPS
|
||||
boxes (prod stood up fresh on `release`, beta rebuilt on `main`); Forgejo is
|
||||
serving at `git.thermograph.org`; and the desktop runs the Forgejo Actions
|
||||
runner. See the verification checklist at the end.
|
||||
|
||||
## Track 2 — Swarm
|
||||
|
||||
See `deploy/swarm/README.md` for the exact order of operations across all
|
||||
**three** nodes (WireGuard mesh first, then swarm init/join, then lock the
|
||||
Swarm ports down to the tunnel interface, then label beta for Forgejo
|
||||
placement). This cluster's only job is hosting Forgejo — it does not
|
||||
orchestrate the Terraform-managed app deploys.
|
||||
|
||||
## Track 3 — Forgejo, replacing GitHub
|
||||
|
||||
### 3a. Stand up Forgejo
|
||||
See `deploy/forgejo/README.md` — deploys the stack (pinned to beta), then
|
||||
walks through registering the Actions runner **on the desktop** as a plain
|
||||
systemd service (`register-lan-runner.sh`), not as a Swarm-scheduled
|
||||
container. Only one Swarm secret to mint now (`forgejo_db_password`) — the
|
||||
runner token is no longer a Swarm secret, since the runner isn't a Swarm
|
||||
service anymore.
|
||||
|
||||
### 3b. Migrate the repo (mirror first, verify, then cut over)
|
||||
1. In Forgejo: **+ New Migration → GitHub**. Point it at
|
||||
`https://github.com/griffemi/thermograph`, pull code + issues + PRs +
|
||||
releases + labels. This is non-destructive to GitHub — do this as a mirror
|
||||
and leave GitHub live and untouched until you're confident.
|
||||
2. Spot-check: branch list matches, a few PRs/issues render correctly, LFS (if
|
||||
any) came across.
|
||||
3. **Don't retarget any secret or disable a GitHub workflow yet** — see 3d.
|
||||
|
||||
### 3c. Workflows
|
||||
`.forgejo/workflows/{build,pr-build,deploy-dev,deploy}.yml` mirror
|
||||
`.github/workflows/*.yml`, copied with the mechanical changes Forgejo needs:
|
||||
- `runs-on: ubuntu-latest` → `runs-on: docker` (one of the two labels the
|
||||
desktop's runner registers under — see 3a; general CI/deploy jobs get a
|
||||
fresh container, the LAN-deploy job in `deploy-dev.yml` runs bare/host-native
|
||||
under the `thermograph-lan` label instead, since it needs real filesystem
|
||||
access).
|
||||
- `appleboy/ssh-action` referenced by full GitHub URL (not mirrored in
|
||||
Forgejo's default action registry, unlike `actions/checkout` /
|
||||
`actions/setup-python`, which resolve unchanged).
|
||||
- **The custom auto-merge workflow step is gone.** It existed only because
|
||||
GitHub's native branch protection/auto-merge is paywalled on private
|
||||
free-tier repos. Forgejo has no such tier: `dev` is protected with the
|
||||
`build` status check required (Settings → Branches). Forgejo does **not**
|
||||
auto-merge on green by itself — a ready, green, mergeable PR is merged
|
||||
**explicitly** (`POST .../pulls/{n}/merge`, `{"Do": "squash"}`), per
|
||||
`CLAUDE.md`. That merge is an ordinary push, so `deploy-dev.yml`'s push
|
||||
trigger fires naturally afterward.
|
||||
- **Branch/deploy mapping (settled):** `.forgejo/workflows/deploy.yml`
|
||||
("Deploy to beta VPS") triggers on `main` and SSHes to **beta**
|
||||
(`75.119.132.91`) running `deploy/deploy.sh`. **Prod is not deployed by a
|
||||
push-triggered workflow** — it's deployed with `terraform apply`
|
||||
(`terraform/README.md`) on the `release` branch. So the promotion chain is
|
||||
`dev` → `main` (this workflow deploys beta) → `release` (Terraform deploys
|
||||
prod). There is deliberately no `release`-triggered workflow.
|
||||
|
||||
A second, independent set of Forgejo workflows
|
||||
(`.forgejo/workflows/{ci,build-push,ops-cron}.yml`) was added by a parallel
|
||||
effort implementing `thermograph-docs/runbooks/implementation-handoff.md` Track A chunk
|
||||
7 — see that doc and its own reconciliation PR (which already fixed its
|
||||
runner-label and ssh-action guesses to match the real infrastructure here).
|
||||
`ci.yml` duplicated this PR's `build.yml` exactly and was dropped in that
|
||||
reconciliation; `build-push.yml` (image → Forgejo's built-in registry) and
|
||||
`ops-cron.yml` (backup + IndexNow) are novel and still present.
|
||||
|
||||
### 3d. Cut over — DONE (2026-07-21)
|
||||
The GitHub → Forgejo cutover is complete; this records what was done:
|
||||
1. Deploy secrets (`SSH_HOST`, `SSH_USER`, `SSH_KEY`, `SSH_PORT`) live in
|
||||
Forgejo's repo Secrets, pointing at beta with a dedicated CI deploy key.
|
||||
2. The LAN dev runner was re-pointed to Forgejo via
|
||||
`deploy/forgejo/register-lan-runner.sh` (systemd --user, on the desktop).
|
||||
3. The repo was migrated into Forgejo; `dev`/`main`/`release` all promoted
|
||||
through Forgejo and deploys verified live.
|
||||
4. GitHub is retired as git host and CI (PR "Retire GitHub as the git host and
|
||||
CI platform") — `origin` remains only as a read-only mirror of history.
|
||||
|
||||
## Verification checklist — all met (2026-07-21)
|
||||
- [x] `ssh agent@<prod>` and `ssh agent@<beta>` both work; `sudo whoami` → `root`
|
||||
- [x] Password SSH auth confirmed dead on both boxes (tested with an actual
|
||||
failed login attempt, not just config inspection)
|
||||
- [x] `docker node ls` (from prod) shows all three nodes `Ready`
|
||||
- [x] Swarm ports (2377/7946/4789) unreachable from outside the WireGuard tunnel
|
||||
- [x] `https://git.thermograph.org` serves Forgejo over TLS; `/v2/` registry
|
||||
API is 403 from off-mesh (firewalled to the WireGuard CIDR)
|
||||
- [x] A test PR flow verified: required `build` check runs → explicit squash
|
||||
merge → LAN dev deploy lands
|
||||
- [x] GitHub retired as git host + CI; kept only as a read-only history mirror
|
||||
42
infra/CLAUDE.md
Normal file
42
infra/CLAUDE.md
Normal file
|
|
@ -0,0 +1,42 @@
|
|||
# thermograph-infra — agent instructions
|
||||
|
||||
How and where the already-built Thermograph app images run. This repo owns
|
||||
Terraform, the SOPS+age secrets vault, compose files, deploy scripts, host
|
||||
provisioning, and the ops cron (DB backup + IndexNow). The app repos
|
||||
(`thermograph-backend`, `thermograph-frontend`) own building and testing the
|
||||
images; this repo never checks out app source. The old monorepo
|
||||
(`emi/thermograph`) is archived — do not point anything at it.
|
||||
|
||||
## The four machines
|
||||
|
||||
Same as ACCESS.md: **dev machine** (operator's box, LAN dev server, CI runner),
|
||||
**prod** (`169.58.46.181`, thermograph.org, `agent` user, passwordless sudo),
|
||||
**beta** (`75.119.132.91`, beta.thermograph.org + Forgejo, `agent` user). Hosts'
|
||||
`/opt/thermograph` (and LAN's `~/thermograph-dev`) are checkouts of THIS repo.
|
||||
|
||||
## Deploys
|
||||
|
||||
- `deploy/deploy.sh` — the one deploy path (prod/beta): resets this repo's
|
||||
checkout, renders secrets from the SOPS vault, pulls the per-service image
|
||||
tags (`BACKEND_IMAGE_TAG`/`FRONTEND_IMAGE_TAG`, persisted in untracked
|
||||
`deploy/.image-tags.env`), rolls the target service, health-checks via
|
||||
container healthchecks. Invoked over SSH by the app repos' deploy workflows.
|
||||
- `deploy/deploy-dev.sh` — thin LAN-dev wrapper around deploy.sh (dev compose
|
||||
overlay, `~/thermograph-dev`, branch `dev`).
|
||||
- Branch model: see README — infra `main` is live on prod+beta, `dev` on LAN;
|
||||
`release` is currently unused. App code is env-staged via image tags; infra
|
||||
is not.
|
||||
|
||||
## Rules
|
||||
|
||||
- **Never run `terraform apply` casually** — no tfstate exists anywhere (never
|
||||
persisted), so an apply would attempt full re-provisioning of live hosts.
|
||||
Terraform here is executable documentation until state is bootstrapped.
|
||||
- Secrets: only via the SOPS vault (`deploy/secrets/*.yaml`, `sops edit` +
|
||||
commit + deploy). Never hand-edit `/etc/thermograph.env` on a SOPS-enabled
|
||||
host — it's a rendered artifact. `secrets-guard` CI rejects plaintext.
|
||||
- The ops cron (`.forgejo/workflows/ops-cron.yml`) is THE prod backup — it uses
|
||||
this repo's `PROD_SSH_*` Actions secrets. If you touch it, verify a dump
|
||||
actually lands in `agent@prod:~/thermograph-backups/`.
|
||||
- Commits/PRs: concise and technical; never mention AI/assistants/automated
|
||||
authorship.
|
||||
172
infra/DEPLOY-DEV.md
Normal file
172
infra/DEPLOY-DEV.md
Normal file
|
|
@ -0,0 +1,172 @@
|
|||
# Dev CI/CD → LAN server on this machine
|
||||
|
||||
Parallel to the prod pipeline (`main` → VPS, see `DEPLOY.md`), the **`dev`**
|
||||
branch continuously deploys to a **LAN server running on this computer**. Git
|
||||
hosting and CI are self-hosted **Forgejo** (`git.thermograph.org`, reachable
|
||||
at `http://10.10.0.2:3080` over the WireGuard mesh) — GitHub is retired.
|
||||
|
||||
```
|
||||
open PR ──▶ CI (build + boot/health, on the LAN Forgejo runner)
|
||||
│ green
|
||||
▼
|
||||
merge into dev (explicit — Forgejo does not auto-merge on its own;
|
||||
see "Landing a PR" below)
|
||||
│
|
||||
▼
|
||||
LAN runner (thermograph-lan label) runs deploy/deploy-dev.sh
|
||||
│
|
||||
▼
|
||||
~/thermograph-dev updated ─▶ docker compose stack (backend + frontend +
|
||||
Postgres 18) brought up, backend on 0.0.0.0:8137
|
||||
```
|
||||
|
||||
Reachable at `http://<lan-ip>:8137/` from any device on your Wi-Fi.
|
||||
|
||||
The dev server runs the **same containerized stack as prod** (`docker-compose.yml`),
|
||||
overlaid with `docker-compose.dev.yml`: **uncapped** (no CPU limits, unlike prod's
|
||||
backend=4 / frontend=2 / db=2) and backend published on `0.0.0.0:8137` for the
|
||||
LAN (prod binds loopback behind Caddy; frontend has no published port here
|
||||
either way, no Caddy to reach it directly, so it's only reached through
|
||||
backend's own reverse-proxy fallback). Bring it up by hand with `make dev-up`.
|
||||
|
||||
## Why it's built this way
|
||||
|
||||
Forgejo has native branch protection + required-status-checks + auto-merge
|
||||
(unlike GitHub on a private free-tier repo, where those are paywalled). But
|
||||
"auto-merge" here still means opting a PR in — either clicking **"Auto merge
|
||||
when checks succeed"** in the web UI, or merging explicitly once CI is green
|
||||
via the API (`POST .../pulls/{n}/merge`, `{"Do": "squash"}`). Nothing merges
|
||||
a ready PR for you automatically just because checks pass, unlike GitHub's
|
||||
old in-workflow `ci-cd.yml` (retired along with the rest of `.github/`).
|
||||
|
||||
## Moving parts
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `.forgejo/workflows/pr-build.yml` | required status check for PRs into `dev` (calls `build.yml`) |
|
||||
| `.forgejo/workflows/deploy-dev.yml` | build + deploy on pushes to `dev` (direct, or a PR merge) |
|
||||
| `.forgejo/workflows/build.yml` | shared build gate: deps, backend tests, JS syntax check, boot/health check |
|
||||
| `deploy/deploy-dev.sh` | pull `dev` into `~/thermograph-dev`, `docker compose up` the uncapped LAN stack, health check |
|
||||
| `docker-compose.dev.yml` | dev overlay: no CPU limits, backend published on `0.0.0.0:8137` (frontend unpublished) |
|
||||
| `deploy/provision-dev-lan.sh` | one-time sudo-free bootstrap of the checkout + runner |
|
||||
| `deploy/forgejo/register-lan-runner.sh` | registers the Forgejo Actions runner on this machine |
|
||||
|
||||
`deploy-dev.yml`'s `deploy` job runs on the `thermograph-lan` label — **not**
|
||||
`[self-hosted, thermograph-lan]` (the array form is what the original GitHub
|
||||
version used; GitHub implicitly tags every self-hosted runner `self-hosted`,
|
||||
but Forgejo's runner has only the labels it was explicitly registered with,
|
||||
so the array form is permanently unschedulable there — a real bug found and
|
||||
fixed once, worth knowing about if a "deploy" job ever silently stops
|
||||
appearing in the Actions history again).
|
||||
|
||||
`deploy-dev.sh`'s git auth (`GH_TOKEN`/`GITHUB_TOKEN`, Forgejo's per-job
|
||||
token exposed under that name for compatibility) is scoped to `REPO_URL`'s
|
||||
own host via a per-command `http.<scheme>://<host>/.extraheader`, not a
|
||||
hardcoded one — important if the Forgejo host ever changes, since a stale
|
||||
hardcoded host would make git silently skip the header (no error) and fall
|
||||
through to whatever ambient credentials happen to be available, not fail
|
||||
loudly.
|
||||
|
||||
## The self-hosted runner
|
||||
|
||||
A Forgejo Actions **runner** (`forgejo-runner`) runs on this machine, labels
|
||||
`docker` (containerized jobs, node:20-bookworm, with the host's docker.sock
|
||||
automounted in — `container.docker_host: automount` in
|
||||
`~/forgejo-runner/config.yaml` — so `docker build`/`push` work without
|
||||
privileged Docker-in-Docker) and `thermograph-lan` (host-native jobs, for
|
||||
`deploy-dev.sh`'s systemctl/docker-compose calls). Installed under
|
||||
`~/forgejo-runner`, runs as the `forgejo-runner` systemd `--user` service.
|
||||
|
||||
```bash
|
||||
# status / logs
|
||||
systemctl --user status forgejo-runner
|
||||
journalctl --user -u forgejo-runner -f
|
||||
|
||||
# re-register if the token/host ever changes
|
||||
bash deploy/forgejo/register-lan-runner.sh <forgejo_url> <registration_token>
|
||||
```
|
||||
|
||||
`git.thermograph.org`'s public DNS resolves to beta's **public** IP, which
|
||||
the registry's mesh-only ACL (see `deploy/forgejo/README.md`) rejects for
|
||||
`/v2/` paths — `docker build-push.yml` pushes run against the *host's*
|
||||
automounted daemon, so it's this **host's** own resolver that needs to route
|
||||
git.thermograph.org over the mesh, not anything settable from a job
|
||||
container. If registry pushes ever start failing with an HTTP/TLS mismatch
|
||||
or a 403 on `/v2/`, check `getent hosts git.thermograph.org` here — it needs
|
||||
to resolve to beta's WireGuard IP (`10.10.0.2`), typically via a `/etc/hosts`
|
||||
line, not the public one.
|
||||
|
||||
## The LAN dev service (Docker Compose)
|
||||
|
||||
Runs as a Docker Compose stack, not a bare systemd unit — `docker ps --filter
|
||||
name=thermograph-dev` shows `thermograph-dev-backend-1`, `thermograph-dev-frontend-1`
|
||||
(repo-split Stage 4 split the single "app" container in two — frontend has no
|
||||
published port, reached only through backend's own reverse-proxy fallback) and
|
||||
`thermograph-dev-db-1` (TimescaleDB). A legacy pre-container
|
||||
`thermograph-dev.service` systemd --user unit may still exist from before this
|
||||
cutover; it's stale and `deploy-dev.sh` stops/disables it on every run — don't
|
||||
trust `systemctl --user status thermograph-dev` as a liveness check, use `docker
|
||||
ps` instead.
|
||||
|
||||
```bash
|
||||
docker ps --filter name=thermograph-dev # is it up?
|
||||
docker compose -f docker-compose.yml -f docker-compose.dev.yml logs -f backend frontend # app logs (from ~/thermograph-dev)
|
||||
```
|
||||
|
||||
Bootstrap (or rebuild) it by hand:
|
||||
|
||||
```bash
|
||||
bash deploy/provision-dev-lan.sh
|
||||
```
|
||||
|
||||
Config: port/base path are baked into the unit from `deploy/deploy-dev.sh`
|
||||
(`PORT`, `THERMOGRAPH_BASE`). Change them there and redeploy. The dedicated
|
||||
checkout at `~/thermograph-dev` is separate from your working tree, so a deploy
|
||||
never touches uncommitted edits in `~/Code/Thermograph`.
|
||||
|
||||
## Before monitoring a PR to land
|
||||
|
||||
A PR needs an explicit merge once CI is green — but first confirm it's
|
||||
actually set up to land, or you'll monitor a PR that can never merge:
|
||||
|
||||
1. **Not conflicting** — `GET /api/v1/repos/{owner}/{repo}/pulls/{n}` should
|
||||
show `"mergeable": true`. If not, merge `forgejo/dev` in and resolve now —
|
||||
don't monitor a PR that can't merge.
|
||||
2. **Required check actually configured** — `GET
|
||||
/api/v1/repos/{owner}/{repo}/branch_protections/dev` should list
|
||||
`status_check_contexts` containing the *exact* string Forgejo reports for
|
||||
that PR's check (workflow display name + job name + event, e.g. `"PR build
|
||||
(required check) / build (pull_request)"` — not just the job name
|
||||
`"build"`). A mismatch here means no PR can ever satisfy the check even
|
||||
with fully green CI; this bit us once already.
|
||||
3. **CI actually running** — `GET
|
||||
/api/v1/repos/{owner}/{repo}/commits/{sha}/status` shows a status, not an
|
||||
empty list. If empty, check the runner is up
|
||||
(`systemctl --user status forgejo-runner`).
|
||||
|
||||
Only start monitoring once all three pass. Once CI is green and mergeable is
|
||||
true, merge explicitly: `POST /api/v1/repos/{owner}/{repo}/pulls/{n}/merge`
|
||||
with `{"Do": "squash", "delete_branch_after_merge": true}`.
|
||||
|
||||
## Day-to-day
|
||||
|
||||
- **Ship:** open a PR into `dev`, confirm CI is green, merge it explicitly —
|
||||
the merge triggers the deploy here.
|
||||
- **Manual redeploy:** trigger `deploy-dev.yml` via `workflow_dispatch`
|
||||
(`POST /api/v1/repos/{owner}/{repo}/actions/workflows/deploy-dev.yml/dispatches`,
|
||||
`{"ref": "dev"}`), or locally `APP_DIR=~/thermograph-dev bash deploy/deploy-dev.sh`.
|
||||
- **Promote to prod:** merge/fast-forward `dev` into `main` and push — that
|
||||
triggers the VPS pipeline in `DEPLOY.md`.
|
||||
|
||||
## Monitoring
|
||||
|
||||
Logs and dashboards moved to a **separate project** — a central Grafana + Loki
|
||||
stack fed by a Grafana Alloy agent on every node (including this LAN dev box),
|
||||
over the WireGuard mesh (`thermograph-observability` on Forgejo; UI at
|
||||
`https://dashboard.thermograph.org`). The dev node's Alloy agent ships its
|
||||
container/app logs under `host="dev"`, so LAN dev shows up alongside prod and
|
||||
beta in the same dashboards. This replaced the old `scripts/dashboard.py`.
|
||||
|
||||
For a quick terminal check without Grafana, the app still serves raw counters at
|
||||
the gated `GET /api/v2/metrics` route (loopback-only; set
|
||||
`THERMOGRAPH_METRICS_TOKEN` to read it over an SSH tunnel).
|
||||
385
infra/DEPLOY.md
Normal file
385
infra/DEPLOY.md
Normal file
|
|
@ -0,0 +1,385 @@
|
|||
# Deploying Thermograph to a prod VPS
|
||||
|
||||
> **Prod and beta are now provisioned by Terraform** (`terraform/README.md`) and
|
||||
> run as a **`docker compose` stack** (backend + frontend + Postgres/TimescaleDB;
|
||||
> repo-split Stage 4 split the single "app" service in two), not the
|
||||
> manual venv + systemd model this document originally described. Current shape:
|
||||
> prod = new box `169.58.46.181`, branch `release`, deployed with `terraform
|
||||
> apply`; beta = old box `75.119.132.91`, branch `main`, deployed when `main` is
|
||||
> pushed (`.forgejo/workflows/deploy.yml` SSHes in and runs `deploy/deploy.sh`,
|
||||
> which `docker compose pull`s the image `build-push.yml` already pushed for
|
||||
> that commit and `up`s it -- repo-split Stage 6's registry-pull cutover;
|
||||
> `deploy.sh` no longer builds in place). The **manual walkthrough below
|
||||
> (`provision.sh`, `thermograph.service`, the venv) is legacy** — kept as a
|
||||
> from-scratch, no-Terraform reference; the process model is compose, so the
|
||||
> systemd/venv specifics no longer match how prod/beta actually run.
|
||||
|
||||
Pipeline (beta, on `main`): **push to `main` on Forgejo → `build-push.yml` builds
|
||||
+ pushes the image, tagged by SHA → Forgejo Actions SSHes to beta →
|
||||
`deploy/deploy.sh` (`git reset` + `docker login` + `docker compose pull && up`,
|
||||
schema migration runs in the app entrypoint) → Caddy fronts it with Let's Encrypt
|
||||
TLS.** Prod (on `release`) is deployed with `terraform apply`, not a push
|
||||
trigger — its `remote-exec` provisioner does the same pull-instead-of-build.
|
||||
(GitHub is retired/archived — Forgejo, self-hosted at `git.thermograph.org` /
|
||||
`10.10.0.2:3080` over the WireGuard mesh, is now the sole git host and CI.)
|
||||
|
||||
Files that make this work:
|
||||
|
||||
| File | Where it lives in prod | Purpose |
|
||||
|------|------------------------|---------|
|
||||
| `.forgejo/workflows/deploy.yml` | Forgejo | CI job that SSHes in and runs the deploy script (beta, on `main`) |
|
||||
| `.forgejo/workflows/build-push.yml` | Forgejo | builds + pushes the SHA-tagged image `deploy.sh`/Terraform pull |
|
||||
| `deploy/deploy.sh` | `/opt/thermograph/deploy/deploy.sh` | `git reset` + `docker login` + `docker compose pull && up` + health check + warm |
|
||||
| `docker-compose.yml` | `/opt/thermograph/docker-compose.yml` | app + Postgres/TimescaleDB services (the process model) |
|
||||
| `deploy/thermograph.env.example` | `/etc/thermograph.env` | Postgres password, VAPID/auth secrets, `WORKERS`, sizing, base path |
|
||||
| `deploy/Caddyfile` | `/etc/caddy/Caddyfile` | reverse proxy + automatic HTTPS |
|
||||
| `terraform/` | run from your machine | provisions the box + brings the stack up (replaces `provision.sh`) |
|
||||
| `deploy/provision.sh`, `deploy/thermograph.service` | *(legacy)* | pre-compose venv+systemd bootstrap — superseded by Terraform |
|
||||
|
||||
The app serves on **loopback only**; Caddy is the only thing exposed to the
|
||||
internet (ports 80/443). The parquet cache in `data/cache/` lives inside the
|
||||
`/opt/thermograph` checkout and is gitignored, so `git pull` never touches it.
|
||||
|
||||
**Any deploy host needs a `/etc/hosts` entry for the registry.**
|
||||
`git.thermograph.org`'s public DNS resolves to beta's public IP, and beta's
|
||||
own Caddy rejects `/v2/*` (the registry API) from anything outside the
|
||||
WireGuard mesh (10.10.0.0/24) — confirmed live: `docker login`/`pull` from
|
||||
prod failed with a 403 until adding `10.10.0.2 git.thermograph.org` to
|
||||
`/etc/hosts` (10.10.0.2 is beta's mesh IP; every deploy host is already on
|
||||
the mesh, this is purely a DNS-routing gap, not a connectivity one). Beta
|
||||
itself doesn't need this — it's the box Forgejo runs on, so `git.thermograph.org`
|
||||
just resolves to itself either way. `build-push.yml`'s own header comment
|
||||
documents the same requirement for the CI runner host.
|
||||
|
||||
**Current production layout** (`deploy/Caddyfile`): the app owns
|
||||
**`thermograph.org`** at its root (`THERMOGRAPH_BASE=/`, so uvicorn serves `/`,
|
||||
`/calendar`, `/api/v2/…` with no prefix), and **`emigriffith.dev`** serves a
|
||||
static portfolio at its root with `emigriffith.dev/thermograph*` permanently
|
||||
redirecting to `thermograph.org` (prefix stripped). Both domains' `A` records
|
||||
point at the same VPS; Caddy provisions a separate cert for each.
|
||||
|
||||
> **Note:** `deploy.sh` (the CI deploy) only pulls code, reinstalls deps, and
|
||||
> restarts the app service — it does **not** touch Caddy or `/etc/thermograph.env`.
|
||||
> After changing `deploy/Caddyfile` or the env, copy it onto the VPS and reload
|
||||
> Caddy / restart the service by hand (see Part 2 step 4 and Part 3).
|
||||
|
||||
---
|
||||
|
||||
## Part 1 — Keys: how to manage them
|
||||
|
||||
You need **two** separate SSH keypairs. Don't reuse one for both.
|
||||
|
||||
### Key A — your login key (you ↔ VPS), you probably already have this
|
||||
This is how *you* SSH in to run `provision.sh`. Nothing to set up here beyond
|
||||
your existing access.
|
||||
|
||||
### Key B — the CI deploy key (Forgejo Actions ↔ VPS), you create this
|
||||
A dedicated key whose **private half** lives in a Forgejo Actions secret and
|
||||
whose **public half** is authorized on the VPS's deploy user. It should only
|
||||
be able to log in as `deploy` and run the deploy script.
|
||||
|
||||
Generate it on your machine (no passphrase — CI can't type one):
|
||||
|
||||
```bash
|
||||
ssh-keygen -t ed25519 -C "thermograph-ci" -f ~/.ssh/thermograph_ci -N ""
|
||||
```
|
||||
|
||||
That makes two files:
|
||||
- `~/.ssh/thermograph_ci` → **private** key → goes into the Forgejo Actions secret `SSH_KEY`
|
||||
- `~/.ssh/thermograph_ci.pub` → **public** key → goes on the VPS
|
||||
|
||||
Install the **public** key on the VPS (as the `deploy` user):
|
||||
|
||||
```bash
|
||||
ssh-copy-id -i ~/.ssh/thermograph_ci.pub deploy@YOUR_VPS_IP
|
||||
# or manually: append the .pub line to /home/deploy/.ssh/authorized_keys
|
||||
```
|
||||
|
||||
Put the **private** key into Forgejo → repo **Settings → Actions → Secrets →
|
||||
Add Secret**. Set all of these:
|
||||
|
||||
| Secret | Value |
|
||||
|--------|-------|
|
||||
| `SSH_KEY` | full contents of `~/.ssh/thermograph_ci` (the private key, incl. BEGIN/END lines) |
|
||||
| `SSH_HOST` | VPS IP or hostname |
|
||||
| `SSH_USER` | `deploy` |
|
||||
| `SSH_PORT` | `22` (or your custom SSH port) |
|
||||
|
||||
No `tea` CLI is installed for a command-line alternative; use the web UI, or
|
||||
the REST API with a personal access token:
|
||||
`POST /api/v1/repos/{owner}/{repo}/actions/secrets/{name}` (body `{"data": "..."}`).
|
||||
|
||||
**Why I can't do this for you:** Forgejo secrets are write-only and require
|
||||
your Forgejo auth; the VPS's `authorized_keys` requires your SSH access. Both
|
||||
are credentials only you should hold.
|
||||
|
||||
### Deploy key for cloning a private repo (optional)
|
||||
If the Forgejo repo is **private**, the VPS also needs to *pull* from it. Two
|
||||
options:
|
||||
- Make `provision.sh`'s `REPO_URL` an **HTTPS** URL with an embedded personal
|
||||
access token (what the live VPS checkout actually uses today — see its
|
||||
`origin` remote), or
|
||||
- Add the `deploy` user's own key (`ssh-keygen` on the VPS) as a **read-only
|
||||
Deploy Key** in the repo (Settings → Deploy keys). This is separate from Key B.
|
||||
|
||||
If the repo is **public**, skip this — `git pull` needs no auth.
|
||||
|
||||
---
|
||||
|
||||
## Part 2 — Install (once, on the VPS)
|
||||
|
||||
1. **Create the Forgejo repo and push** (this project isn't versioned yet):
|
||||
|
||||
```bash
|
||||
cd ~/Code/Thermograph
|
||||
git init && git add -A && git commit -m "Initial commit"
|
||||
# create an empty private repo on the Forgejo instance first (web UI or
|
||||
# POST /api/v1/user/repos), then:
|
||||
git remote add origin https://<forgejo-host>/OWNER/thermograph.git
|
||||
git push -u origin main
|
||||
```
|
||||
|
||||
2. **Point DNS** for each domain you serve (`thermograph.org` for the app,
|
||||
`emigriffith.dev` for the portfolio) at the VPS IP with an `A` record. Let's
|
||||
Encrypt won't issue a cert until the name resolves to the box.
|
||||
|
||||
3. **Provision the box.** Copy `deploy/provision.sh` over (or clone the repo)
|
||||
and edit `REPO_URL` at the top, then:
|
||||
|
||||
```bash
|
||||
REPO_URL="https://<forgejo-host>/OWNER/thermograph.git" bash deploy/provision.sh
|
||||
```
|
||||
|
||||
It installs Python + git + Caddy, creates the `deploy` user, clones to
|
||||
`/opt/thermograph`, installs the systemd unit + env file, grants the deploy
|
||||
user a password-less `systemctl restart thermograph` (so CI can restart it),
|
||||
and starts the service.
|
||||
|
||||
4. **Install the Caddy config** and reload. `deploy/Caddyfile` already carries the
|
||||
real domains (`thermograph.org` + `emigriffith.dev`); edit them if yours differ,
|
||||
then copy it into place, validate, and reload:
|
||||
|
||||
```bash
|
||||
sudo cp /opt/thermograph/deploy/Caddyfile /etc/caddy/Caddyfile
|
||||
sudo caddy validate --config /etc/caddy/Caddyfile # syntax check before reload
|
||||
sudo systemctl reload caddy # Caddy fetches each domain's TLS cert automatically
|
||||
```
|
||||
|
||||
Do this again any time `deploy/Caddyfile` changes — the CI deploy does not ship
|
||||
it. Set the app's base path to match in `/etc/thermograph.env`
|
||||
(`THERMOGRAPH_BASE=/` for the app-at-root layout), then `sudo systemctl restart
|
||||
thermograph`.
|
||||
|
||||
5. **Install Key B** (Part 1) so CI can log in.
|
||||
|
||||
6. **Verify:** open `https://YOUR_DOMAIN/` — you should get a clean padlock.
|
||||
|
||||
---
|
||||
|
||||
## Part 3 — Day-to-day
|
||||
|
||||
- **Deploy:** just `git push` to `main`. Actions runs `deploy.sh` and
|
||||
health-checks the service. Watch it under the repo's **Actions** tab on
|
||||
Forgejo (or query `/api/v1/repos/{owner}/{repo}/actions/tasks`).
|
||||
- **Manual deploy / rollback:** `ssh deploy@vps '/opt/thermograph/deploy/deploy.sh'`
|
||||
(set `BRANCH=` or check out a tag first to roll back).
|
||||
- **Logs:** `journalctl -u thermograph -f` (app), `/var/log/caddy/thermograph.log` (proxy).
|
||||
- **Restart:** `sudo systemctl restart thermograph`.
|
||||
- **Change port / base path:** edit `/etc/thermograph.env`, then `sudo systemctl
|
||||
restart thermograph`. Changing the base path usually pairs with a Caddy change
|
||||
(proxy target / redirects) — update `/etc/caddy/Caddyfile` and reload Caddy too.
|
||||
- **Change routing / domains:** edit `/etc/caddy/Caddyfile` (or copy the repo's),
|
||||
`sudo caddy validate --config /etc/caddy/Caddyfile`, then `sudo systemctl reload
|
||||
caddy`. The CI deploy never ships the Caddyfile, so this is always by hand.
|
||||
|
||||
## Security notes
|
||||
- App is bound to `127.0.0.1` — not reachable except through Caddy.
|
||||
- Firewall: allow only `22`, `80`, `443` (e.g. `ufw allow 22,80,443/tcp`).
|
||||
- The CI sudoers rule is scoped to exactly `systemctl restart/status thermograph`.
|
||||
- Rotate Key B by regenerating it and updating the `SSH_KEY` secret +
|
||||
`authorized_keys`.
|
||||
|
||||
## Accounts & notifications
|
||||
|
||||
Accounts, subscriptions, and notifications are **authoritative and not
|
||||
regenerable** — back them up. On **prod/beta they live in Postgres/TimescaleDB**
|
||||
(the compose `db` service, on the `pgdata` volume) alongside the climate record;
|
||||
on **LAN dev** they fall back to SQLite (`data/accounts.sqlite`) when
|
||||
`THERMOGRAPH_DATABASE_URL` isn't a Postgres URL.
|
||||
|
||||
- **Back up the Postgres volume** (e.g. `pg_dump` the `thermograph` database) as
|
||||
part of the VPS backup routine — losing it wipes all accounts and alerts. (On
|
||||
a SQLite LAN dev box, back up `data/accounts.sqlite` and its `-wal`/`-shm`
|
||||
sidecars instead.)
|
||||
- **Multiple workers, one notifier (leader election).** Prod runs several uvicorn
|
||||
workers (`WORKERS`, currently 8 on prod / 4 on beta). The subscription
|
||||
evaluator and the recurring scheduler must run in exactly one, so the workers
|
||||
elect a single leader — a host-local `flock` (`THERMOGRAPH_SINGLETON_LOCK`,
|
||||
set by the compose app service to `/app/data/notifier.lock`), or a cluster-wide
|
||||
Postgres advisory lock (`THERMOGRAPH_SINGLETON_PG=1`) if the notifier must be
|
||||
single across *hosts*. See `backend/core/singleton.py`. `THERMOGRAPH_ROLE`
|
||||
(`web`/`worker`/`all`, default `all`) further restricts which processes may own
|
||||
it, so a stateless web tier can scale without scaling notifiers.
|
||||
- **Env vars** (all optional, sensible defaults):
|
||||
- `THERMOGRAPH_COOKIE_SECURE` — set to `1` when serving over HTTPS (the VPS/TLS
|
||||
deploy) so the session cookie is `Secure`. Leave unset for plain-HTTP LAN, where
|
||||
a `Secure` cookie would never be sent.
|
||||
- `THERMOGRAPH_SESSION_TTL_DAYS` — login/cookie lifetime (default 30).
|
||||
- `THERMOGRAPH_NOTIFY_INTERVAL` — seconds between evaluation passes (default 900).
|
||||
- `THERMOGRAPH_NOTIFY_ARCHIVE_FETCHES` — max missing-archive fetches per pass
|
||||
(default 4). A newly-subscribed cell with no cached ~45-year archive is fetched
|
||||
once (then only read); this caps how many such fetches one pass may do so a burst
|
||||
of new subscriptions can't exhaust the archive quota.
|
||||
- `THERMOGRAPH_ENABLE_NOTIFIER` — set to `0` to disable the background evaluator.
|
||||
- `THERMOGRAPH_AUTH_SECRET` — signing secret for (future) email reset/verify
|
||||
tokens; unused today, a per-boot random value is used if unset.
|
||||
|
||||
### Web Push (VAPID)
|
||||
|
||||
Push notifications are signed with a VAPID keypair. If `THERMOGRAPH_VAPID_PRIVATE_KEY`
|
||||
/ `THERMOGRAPH_VAPID_PUBLIC_KEY` are unset, the app generates a pair into
|
||||
**`data/vapid.json`** on first run. That's fine **only if the file persists** — the
|
||||
key the browser subscribed with must keep matching the key the server signs with. If
|
||||
the keys ever change (regenerated `data/vapid.json`, a non-persistent data dir), every
|
||||
existing subscription silently stops delivering: the push service returns 401/403 and
|
||||
the app records it (`logs/errors/*.jsonl`, tagged `"phase":"push"`) but the `/push/test`
|
||||
call still returns 202 with `sent:0, failed:N`. To be safe, **pin the keys** in
|
||||
`/etc/thermograph.env` (see `deploy/thermograph.env.example` for the generate command).
|
||||
After changing keys, users must toggle alerts off/on to re-subscribe.
|
||||
|
||||
Diagnose a "test says sent but nothing arrives": check the `/push/test` response
|
||||
(`failed:N` ⇒ delivery rejected) or `tail logs/errors/*.jsonl | grep push`, and
|
||||
`docker compose logs backend | grep -i push` for the exact status code.
|
||||
|
||||
## Homepage "unusual right now" feed
|
||||
|
||||
`backend/api/homepage.py` sweeps the tracked cities, grades each one's latest
|
||||
observation from the parquet cache, and writes the ranked result to
|
||||
`data/homepage.json` for the homepage template. It runs on the notifier's timer
|
||||
(hourly guard) and at the tail of `warm_cities.py`, so a deploy repopulates it.
|
||||
|
||||
**It spends a small amount of forecast quota, and has to.** Nothing else keeps
|
||||
the recent/forecast cache current for the tracked cities: `warm_cities.py` skips
|
||||
any city whose archive is already cached, so it never re-fetches their forecast
|
||||
bundle, and the notifier only touches cells somebody has subscribed to. Reading
|
||||
that cache without refreshing it made the strip present days-old readings as
|
||||
"right now". Each pass therefore tops up the **stalest** cells first, capped:
|
||||
|
||||
- `THERMOGRAPH_HOMEPAGE_REFRESH` — recent/forecast fetches per pass (default 40,
|
||||
one per cell). At the hourly cadence that is ~960/day. Set `0` to disable, in
|
||||
which case the strip only ranks cities something else happened to refresh.
|
||||
- `THERMOGRAPH_HOMEPAGE_CITIES` — how many cities the strip ranks over (default
|
||||
250). The strip shows ~12 cards; ranking over a smaller set that is genuinely
|
||||
fresh beats ranking over every cached city when most of those are stale.
|
||||
|
||||
Readings older than a day are dropped rather than ranked, so an empty strip means
|
||||
the cache went cold — not that the weather is unremarkable.
|
||||
|
||||
## Outbound email
|
||||
|
||||
The app never talks to a mail provider directly. It speaks plain SMTP to
|
||||
**Postfix running as a send-only null client on `127.0.0.1:25`**
|
||||
(`deploy/provision-mail.sh`, run once as root). `backend/mailer.py` is the only
|
||||
code that sends, over stdlib `smtplib` — no new dependency.
|
||||
|
||||
Why route through a local MTA rather than a provider's API:
|
||||
|
||||
- **Delivery policy stays swappable.** Direct-to-MX or relayed through a
|
||||
transactional provider is a Postfix setting; the app's config is just
|
||||
"localhost:25" either way, so switching needs no code change or redeploy.
|
||||
- **Postfix queues and retries.** A handler hands the message off in
|
||||
microseconds; a slow or briefly-down upstream can't stall a request or lose a
|
||||
signup.
|
||||
- **Nothing new is exposed.** `inet_interfaces = loopback-only`, so the box
|
||||
accepts mail from itself and nothing else. A loopback socket is not a
|
||||
filesystem write, so the hardened unit (`ProtectSystem=full`,
|
||||
`ReadWritePaths=…`) needs no change.
|
||||
|
||||
**Default is safe:** `THERMOGRAPH_MAIL_BACKEND` defaults to `console` — it logs
|
||||
the message and sends nothing — so LAN dev and the test suite exercise the whole
|
||||
signup path with no mail server and no risk of mailing a real person. Production
|
||||
opts in with `=smtp`.
|
||||
|
||||
### Deliverability (do this before mailing real subscribers)
|
||||
|
||||
Mail from a bare VPS IP is very often junked no matter how Postfix is configured,
|
||||
because the IP has no sending reputation. Either:
|
||||
|
||||
- **Relay through a provider** (recommended) — `RELAYHOST` + credentials in
|
||||
`provision-mail.sh`. They handle SPF/DKIM alignment and reputation.
|
||||
- **Direct to MX** — then you own the DNS work: an **SPF** TXT record, **DKIM**
|
||||
(opendkim) with the public key published, a **DMARC** TXT record, and a
|
||||
**PTR / reverse-DNS** record on the VPS IP pointing at `mail.thermograph.org`.
|
||||
Missing rDNS alone is enough for Gmail and Outlook to junk everything.
|
||||
|
||||
Check placement with <https://www.mail-tester.com>, which scores all four at once.
|
||||
|
||||
- **Env vars** (all in `deploy/thermograph.env.example`):
|
||||
`THERMOGRAPH_MAIL_BACKEND` (`console` | `smtp` | `disabled`),
|
||||
`THERMOGRAPH_SMTP_HOST` / `_PORT` / `_USER` / `_PASSWORD` / `_STARTTLS`,
|
||||
`THERMOGRAPH_MAIL_FROM`, `THERMOGRAPH_MAIL_REPLY_TO`.
|
||||
- **`THERMOGRAPH_AUTH_SECRET` must be set to a fixed value** before any
|
||||
confirmation or password-reset link is mailed. It currently defaults to a
|
||||
per-boot random value, so every outstanding link would break on restart.
|
||||
- **Digest signups** land in the `pending_digest` table in the accounts DB
|
||||
(Postgres on prod/beta; authoritative, back it up). The form ships ahead of
|
||||
delivery on purpose:
|
||||
addresses are collected now and confirmed once SMTP is live.
|
||||
- **Logs:** `journalctl -u postfix -f`; queue with `mailq`.
|
||||
|
||||
## SEO content pages
|
||||
|
||||
Crawlable, server-rendered pages (Jinja2) sit alongside the interactive tool and
|
||||
link into it: `/climate` (hub), `/climate/<slug>` (per-city), `/climate/<slug>/<month>`,
|
||||
`/climate/<slug>/records`, `/glossary`, `/about`, plus `/robots.txt` and
|
||||
`/sitemap.xml`. The routable city set is `backend/cities.json` (~1000 metros: top
|
||||
~500 global + ~250 core-English + ~250 extended-English/high-proficiency),
|
||||
regenerated with `python gen_cities.py [n_global] [n_english] [n_extended]`.
|
||||
|
||||
- A **values filter** in `gen_cities.py` drops cities in countries that criminalize
|
||||
LGBTQ+ people (plus China/Pakistan/Indonesia by choice) via `EXCLUDE_CC`, except a
|
||||
hand-kept list of notable/tourist/high-English hubs in `KEEP_SLUGS` (Lagos, Nairobi,
|
||||
Cairo, Dubai, KL, Shanghai×5-from-China, Karachi, Jakarta, …). Excluded slots are
|
||||
backfilled from the next-ranked non-excluded cities, so the total stays ~1000. Edit
|
||||
either set and regenerate to change the policy.
|
||||
|
||||
- **Per-city blurbs** live in `backend/cities_flavor.json` (a short Wikipedia summary
|
||||
per city, CC BY-SA, attributed on the page). `python gen_flavor.py` fetches
|
||||
Wikipedia's free REST summary API and validates each match by coordinates; it's
|
||||
**incremental** by default (only fetches cities missing a blurb, and prunes ones no
|
||||
longer in cities.json) — pass `--full` to rebuild. ~94% of cities get a blurb; the
|
||||
rest render without one.
|
||||
- **Notable weather events** are a small **hand-curated** list in `backend/city_events.py`
|
||||
(`{slug: {text, url}}`) — one verified iconic event per city (Katrina/New Orleans,
|
||||
Harvey/Houston, the 1952 Great Smog/London, …). Auto-sourcing these from search
|
||||
proved unreliable (wrong matches), so they're edited by hand; a city not in the list
|
||||
simply shows its flavor blurb. Add entries freely — a test checks the slugs are valid.
|
||||
|
||||
- **Archive warming runs automatically on every deploy.** `deploy.sh` (prod) and
|
||||
`deploy-dev.sh` (dev) launch `warm_cities.py` detached in the background after the
|
||||
health check, so the ~750 city pages serve from cache and a crawl can't burst the
|
||||
archive API quota. It's idempotent (skips already-cached cells), so only the first
|
||||
deploy does the full ~25-min warm; later deploys just top up new cities. A page
|
||||
also self-heals (fetches its archive once) if hit before warming finishes. To run
|
||||
it by hand: `cd backend && python warm_cities.py --pace 2`. Logs: prod
|
||||
`logs/warm-cities.log`; dev `journalctl --user -u thermograph-warm-cities`.
|
||||
- **Submit the sitemap** in Google Search Console: `https://thermograph.org/sitemap.xml`.
|
||||
- `jinja2` is a new dependency (already in `requirements.txt`).
|
||||
|
||||
## Monitoring
|
||||
|
||||
Fleet-wide logs and dashboards live in a **separate project** — a central
|
||||
Grafana + Loki stack fed by a Grafana Alloy agent on every node, over the
|
||||
WireGuard mesh (`thermograph-observability` on Forgejo; UI at
|
||||
**`https://dashboard.thermograph.org`**, Google SSO). It ingests every
|
||||
container's stdout, Caddy's access logs, and the app's structured JSON logs
|
||||
(`logs/{errors,access,audit}/*.jsonl`). This replaced the old SSH-tailed
|
||||
`scripts/dashboard.py`.
|
||||
|
||||
The app still exposes raw counters at the gated `GET /api/v2/metrics` route
|
||||
(loopback-only; refuses any request carrying a proxy `X-Forwarded-*` header, so
|
||||
it's never reachable through Caddy). It reports inbound requests per endpoint,
|
||||
outbound calls per upstream source, and background-daemon heartbeats
|
||||
(e.g. the subscription notifier). Read it over an SSH tunnel with
|
||||
`THERMOGRAPH_METRICS_TOKEN` (set in `/etc/thermograph.env`) as the `token=`
|
||||
query param — handy for a quick check without opening Grafana.
|
||||
61
infra/Makefile
Normal file
61
infra/Makefile
Normal file
|
|
@ -0,0 +1,61 @@
|
|||
# docker-compose orchestration only. The app-level targets that used to live
|
||||
# here (run, lan-run, stop, migrate, migrate-cache, test, shots, indexnow,
|
||||
# venv, clean) moved to the backend/frontend app repos along with the code
|
||||
# they operate on.
|
||||
|
||||
.PHONY: up down db-up db-down dev-up dev-down om-up om-down om-backfill
|
||||
|
||||
# --- docker-compose stack (backend + frontend + PostgreSQL) ---------------------
|
||||
# Requires a POSTGRES_PASSWORD (copy .env.example -> .env). See docker-compose.yml.
|
||||
|
||||
# Pull each service's published image and start the whole stack (backend +
|
||||
# frontend + db) in the background. docker-compose.yml no longer has a
|
||||
# top-level `build:` for backend/frontend -- each ships as its OWN image
|
||||
# (BACKEND_IMAGE_TAG / FRONTEND_IMAGE_TAG), built in its own app repo (a plain
|
||||
# `docker build` there, or that repo's own build-push.yml in CI) -- so `up`
|
||||
# here only pulls + runs; it can no longer build in place.
|
||||
up:
|
||||
docker compose pull
|
||||
docker compose up -d
|
||||
|
||||
# Stop and remove the stack's containers (named volumes persist).
|
||||
down:
|
||||
docker compose down
|
||||
|
||||
# Start just the Postgres service (e.g. to run the app against it from a venv).
|
||||
db-up:
|
||||
docker compose up -d db
|
||||
|
||||
# Stop the Postgres service.
|
||||
db-down:
|
||||
docker compose stop db
|
||||
|
||||
# The LAN dev stack: uncapped (no CPU limits) and published on 0.0.0.0:8137 so
|
||||
# phones on the Wi-Fi can reach it (the base file is prod: capped + loopback).
|
||||
DEV_COMPOSE = docker compose -f docker-compose.yml -f docker-compose.dev.yml
|
||||
dev-up:
|
||||
POSTGRES_PASSWORD=$${POSTGRES_PASSWORD:-thermograph-dev} $(DEV_COMPOSE) up -d --build
|
||||
|
||||
dev-down:
|
||||
$(DEV_COMPOSE) down
|
||||
|
||||
# Self-hosted Open-Meteo (prod-only overlay). Serves the ERA5 archive from object
|
||||
# storage so the app is off the public archive API. See deploy/openmeteo/README.md
|
||||
# for the object-storage bucket + rclone mount this expects on the host.
|
||||
OM_COMPOSE = docker compose -f docker-compose.yml -f docker-compose.openmeteo.yml
|
||||
om-up:
|
||||
$(OM_COMPOSE) up -d --build
|
||||
|
||||
om-down:
|
||||
$(OM_COMPOSE) down
|
||||
|
||||
# One-time historical backfill: write the full 1980→present archive (~1–1.5 TB of
|
||||
# .om) to object storage. Runs each dataset once (no --repeat-interval), independent
|
||||
# of the app/db, so it can run before the stack is flipped over. Hours-long.
|
||||
OM_LAND_VARS = temperature_2m,dew_point_2m,precipitation,shortwave_radiation,wind_u_component_10m,wind_v_component_10m
|
||||
OM_ERA5_VARS = wind_gusts_10m,temperature_2m,dew_point_2m,precipitation,shortwave_radiation,wind_u_component_10m,wind_v_component_10m
|
||||
om-backfill:
|
||||
$(OM_COMPOSE) run --rm --no-deps open-meteo-sync-land \
|
||||
sync copernicus_era5_land $(OM_LAND_VARS) --past-days $${OM_BACKFILL_DAYS:-17000}
|
||||
$(OM_COMPOSE) run --rm --no-deps open-meteo-sync-era5 \
|
||||
sync copernicus_era5 $(OM_ERA5_VARS) --past-days $${OM_BACKFILL_DAYS:-17000}
|
||||
46
infra/README.md
Normal file
46
infra/README.md
Normal file
|
|
@ -0,0 +1,46 @@
|
|||
# thermograph-infra
|
||||
|
||||
Infrastructure for [Thermograph](https://thermograph.org): Terraform host
|
||||
provisioning, the SOPS+age secrets vault, Docker Swarm/WireGuard networking,
|
||||
Forgejo, Caddy, and the deploy scripts that run the already-built app image on
|
||||
each host. Extracted from the app monorepo (`emi/thermograph`) — the app repo
|
||||
owns building and testing the app; this repo owns running it.
|
||||
|
||||
- **`terraform/`** — provisions/configures hosts (SSH-driven by default; an
|
||||
optional GCP-creating module is scaffolded, no live resources yet) and
|
||||
triggers each deploy. See `terraform/README.md`.
|
||||
- **`deploy/secrets/`** — the git-native SOPS+age secrets vault (every app
|
||||
secret, encrypted at rest, rendered at deploy time). See
|
||||
`deploy/secrets/README.md`.
|
||||
- **`deploy/swarm/`, `deploy/forgejo/`** — the 3-node WireGuard/Swarm cluster
|
||||
that hosts Forgejo (git + CI + registry); does not run the app itself. See
|
||||
`ACCESS.md` and the READMEs under each directory.
|
||||
- **`deploy/deploy.sh`** — pulls the pinned app image (`IMAGE_TAG`) and rolls
|
||||
the compose stack; invoked by Terraform and by the app repo's
|
||||
`.forgejo/workflows/deploy.yml` over SSH.
|
||||
- **`docker-compose*.yml`, `docker-stack.yml`** — how the app image runs
|
||||
(compose in production today; `docker-stack.yml` is a design record for a
|
||||
possible future Swarm-based app deploy, not currently live).
|
||||
|
||||
The app's own source, `Dockerfile`, and build/test CI stay in the app repo —
|
||||
this repo never checks out app source; hosts only pull tagged images from the
|
||||
registry. See `ACCESS.md` for host access and the Swarm/Forgejo topology, and
|
||||
`terraform/README.md` for the day-to-day `plan`/`apply` workflow.
|
||||
|
||||
## Branches & how changes reach each environment
|
||||
|
||||
- **`main`** — what **prod and beta** run: their `/opt/thermograph` checkouts
|
||||
`git reset --hard origin/main` at the start of every deploy (`deploy/deploy.sh`).
|
||||
A merge to `main` reaches those hosts on the next app deploy (or a by-hand
|
||||
`deploy.sh` run); there is no separate infra deploy trigger.
|
||||
- **`dev`** — what **LAN dev** runs: `~/thermograph-dev` resets to it via
|
||||
`deploy/deploy-dev.sh`. Keep it fast-forwarded to `main` (infra changes are not
|
||||
environment-staged today; the branches exist so LAN dev *can* trail or lead
|
||||
when needed).
|
||||
- **`release`** — currently consumed by nothing (prod tracks `main`, not
|
||||
`release`). It exists to mirror the app repos' dev→main→release promotion
|
||||
shape if per-environment infra staging is ever wanted; until then, treat
|
||||
`main` as live-everywhere.
|
||||
|
||||
Note the asymmetry with the app repos: app code IS environment-staged
|
||||
(dev→main→release maps to LAN→beta→prod via image tags), infra is not.
|
||||
97
infra/deploy/Caddyfile
Normal file
97
infra/deploy/Caddyfile
Normal file
|
|
@ -0,0 +1,97 @@
|
|||
# /etc/caddy/Caddyfile on the VPS.
|
||||
# Each domain's A/AAAA record must already point at this VPS — Caddy provisions a
|
||||
# Let's Encrypt cert on first request and auto-renews. Nothing else to do for TLS
|
||||
# (just make sure ports 80 and 443 are open).
|
||||
#
|
||||
# Layout:
|
||||
# thermograph.org/* -> the Thermograph app (path-split across
|
||||
# backend/frontend -- see below)
|
||||
# emigriffith.dev/ -> static portfolio site (served straight from disk)
|
||||
# emigriffith.dev/thermograph* -> permanent redirect to thermograph.org (the app moved)
|
||||
#
|
||||
# Thermograph now owns thermograph.org's root, so both services run with
|
||||
# THERMOGRAPH_BASE=/ (see /etc/thermograph.env) — pages, assets and API all sit
|
||||
# at "/" with no sub-path prefix. Repo-split Stage 4: backend and frontend are
|
||||
# two containers now (docker-compose.yml), each on its own loopback port --
|
||||
# Caddy path-splits directly to whichever owns a given path, so the browser
|
||||
# still sees one apparent origin. Repo-split Stage 7a flipped which side owns
|
||||
# the enumerated list: frontend now owns everything (content pages, the
|
||||
# interactive tool's SPA shells, every static asset, the dynamic IndexNow key
|
||||
# file) except the short, stable set below, which mirrors backend/web/app.py's
|
||||
# own routing exactly (a single catch-all proxy to frontend for everything
|
||||
# else) -- unlike frontend's paths, backend's don't grow every time a new
|
||||
# static asset filename is added. A gap in this list still just degrades to
|
||||
# "one extra hop" through backend's own proxy fallback, never a 404.
|
||||
|
||||
thermograph.org {
|
||||
encode zstd gzip
|
||||
|
||||
@backend_paths path /api/* /digest /discord/interactions
|
||||
|
||||
# Active health check on the same cheap /healthz route each container's own
|
||||
# HEALTHCHECK uses (Dockerfile) — so a deploy that's still restarting/booting
|
||||
# never gets proxied into (a reload alone has no gate, hop-1 runbook hazard #10).
|
||||
handle @backend_paths {
|
||||
reverse_proxy 127.0.0.1:8137 {
|
||||
health_uri /healthz
|
||||
health_interval 5s
|
||||
health_timeout 3s
|
||||
health_status 2xx
|
||||
}
|
||||
}
|
||||
|
||||
handle {
|
||||
reverse_proxy 127.0.0.1:8080 {
|
||||
health_uri /healthz
|
||||
health_interval 5s
|
||||
health_timeout 3s
|
||||
health_status 2xx
|
||||
}
|
||||
}
|
||||
|
||||
log {
|
||||
output file /var/log/caddy/thermograph.log
|
||||
}
|
||||
}
|
||||
|
||||
emigriffith.dev {
|
||||
encode zstd gzip
|
||||
|
||||
# Thermograph moved to its own domain. Send the old sub-path there with a
|
||||
# permanent redirect, stripping the /thermograph prefix so deep links map
|
||||
# straight across (…/thermograph/calendar -> thermograph.org/calendar). The
|
||||
# bare /thermograph (no trailing slash) goes to the new root.
|
||||
handle_path /thermograph/* {
|
||||
redir https://thermograph.org{uri} permanent
|
||||
}
|
||||
handle /thermograph {
|
||||
redir https://thermograph.org/ permanent
|
||||
}
|
||||
|
||||
# Portfolio at the root. Point `root` at the built static site (for the Astro
|
||||
# portfolio that's its `dist/` output). file_server serves index.html for
|
||||
# directories and returns a real 404 for missing paths.
|
||||
handle {
|
||||
root * /var/www/emigriffith
|
||||
file_server
|
||||
}
|
||||
|
||||
log {
|
||||
output file /var/log/caddy/emigriffith.log
|
||||
}
|
||||
}
|
||||
|
||||
# Old bookmarks to the raw IP (the pre-domain URL) would otherwise get bounced to
|
||||
# HTTPS-on-the-IP, which has no cert and fails. Redirect them to the portfolio domain.
|
||||
http://75.119.132.91 {
|
||||
redir https://emigriffith.dev{uri} permanent
|
||||
}
|
||||
|
||||
# Optional: redirect www -> apex for either domain. Add the www CNAME/A record
|
||||
# first, then uncomment the matching block.
|
||||
# www.emigriffith.dev {
|
||||
# redir https://emigriffith.dev{uri} permanent
|
||||
# }
|
||||
# www.thermograph.org {
|
||||
# redir https://thermograph.org{uri} permanent
|
||||
# }
|
||||
120
infra/deploy/POSTGRES-MIGRATION.md
Normal file
120
infra/deploy/POSTGRES-MIGRATION.md
Normal file
|
|
@ -0,0 +1,120 @@
|
|||
# Cutover: SQLite → PostgreSQL 18 (containerized stack)
|
||||
|
||||
The app now runs as a docker-compose stack (`app` + `db`) and standardizes on
|
||||
PostgreSQL. Locally you need nothing but Docker; `make up` builds and starts both.
|
||||
This doc is the **one-time production cutover** from the old bare-systemd/SQLite
|
||||
deploy to the compose stack, migrating the authoritative accounts data.
|
||||
|
||||
## What moved
|
||||
|
||||
- **accounts** (users/subscriptions/notifications/…): SQLite → Postgres **durable**
|
||||
tables. Managed by Alembic (`backend/alembic/`); `create_all` on a fresh DB.
|
||||
- **derived cache** (`store.py`) and **metrics** (`metrics.py`): now Postgres
|
||||
**UNLOGGED** tables (fast, non-durable — same throwaway semantics). No data to
|
||||
migrate — the cache rebuilds, metrics start empty.
|
||||
- **climate record** (`climate.py`): the raw daily archive + recent/forecast bundle
|
||||
moved from per-cell **parquet** files into **TimescaleDB hypertables**
|
||||
(`climate_history` / `climate_recent` / `climate_sync`, managed by Alembic). The
|
||||
app reads/writes them via `data/climate_store.py`. Backfilled from the existing
|
||||
parquet cache with `migrate_cache_to_pg.py` (see below).
|
||||
- The app selects Postgres when `THERMOGRAPH_DATABASE_URL` is a `postgresql+asyncpg`
|
||||
URL; unset ⇒ the old SQLite/parquet behavior (this is how the **test suite** stays
|
||||
Postgres-free — no Postgres needed in CI, and climate falls back to parquet).
|
||||
- `db` runs the stock **`timescale/timescaledb:latest-pg18`** image (genuine PG18 +
|
||||
TimescaleDB). The earlier pg_duckdb / `read_parquet('/parquet/…')` capability is
|
||||
gone — the climate record is real tables now, queryable directly (see
|
||||
`deploy/db/README.md`).
|
||||
|
||||
## Connection model
|
||||
|
||||
Per worker (4 workers): a read-write asyncpg engine + a read-only asyncpg engine
|
||||
(`default_transaction_read_only=on`, used by the pure-GET endpoints), plus one sync
|
||||
psycopg engine for the notifier thread. Single Postgres primary — the RO engine is
|
||||
a guardrail, not a replica.
|
||||
|
||||
## Prod cutover steps (maintenance window)
|
||||
|
||||
Prereq: the deploy user can run `docker`. Secrets in `/etc/thermograph.env` must
|
||||
include `POSTGRES_PASSWORD`, and pinned `THERMOGRAPH_AUTH_SECRET`,
|
||||
`THERMOGRAPH_VAPID_PRIVATE_KEY`/`_PUBLIC_KEY`, `THERMOGRAPH_COOKIE_SECURE=1`.
|
||||
|
||||
1. **Ship the stack, don't cut over yet.** Deploy the branch; `docker compose build`.
|
||||
Bring up only the DB: `docker compose up -d db`. Add a nightly `pg_dump` backup.
|
||||
2. **Freeze + back up.** Stop the old app (hard stop — no writes). Back up the live
|
||||
`data/accounts.sqlite` + `-wal`/`-shm` (this is the rollback artifact).
|
||||
3. **Build the schema.** `docker compose run --rm backend alembic upgrade head`
|
||||
(or let the `app` entrypoint do it on first start — but do the data copy before
|
||||
real traffic).
|
||||
4. **Copy the accounts data:**
|
||||
```
|
||||
docker compose run --rm \
|
||||
-e THERMOGRAPH_DATABASE_URL="postgresql+asyncpg://thermograph:$POSTGRES_PASSWORD@db:5432/thermograph" \
|
||||
-v /opt/thermograph/data/accounts.sqlite:/src.sqlite:ro \
|
||||
backend python migrate_accounts_to_pg.py --sqlite /src.sqlite
|
||||
```
|
||||
It copies user → subscription/push_subscription/pending_digest → notification,
|
||||
**preserving PKs**, **skips access_token** (everyone re-logins once), and
|
||||
**resets the integer-PK sequences** (skipping that = a runtime PK collision).
|
||||
Verify the printed per-table counts against the old DB.
|
||||
5. **Start serving:** `docker compose up -d` (Caddy already proxies `127.0.0.1:8137`).
|
||||
6. **Smoke test:** login, list/create/delete a subscription, one notifier tick, a
|
||||
cached calendar/day request (derived store on PG), the metrics dashboard.
|
||||
7. Keep the frozen `accounts.sqlite` as rollback for a few days; keep PG backed up.
|
||||
Enable pg_duckdb on the (already-initialized) volume once, by hand:
|
||||
`docker compose exec db psql -U thermograph -d thermograph -c 'CREATE EXTENSION IF NOT EXISTS pg_duckdb;'`
|
||||
|
||||
## Rollback
|
||||
|
||||
Redeploy the previous SQLite release (or point `THERMOGRAPH_DATABASE_URL` back to
|
||||
unset/SQLite and restart). **Clean only during/immediately after the window** — once
|
||||
real users write to Postgres, those writes are lost on rollback (the copy is
|
||||
one-way), so keep the window short and writes frozen during the copy.
|
||||
|
||||
## TimescaleDB cutover (from the earlier pg_duckdb PG18 stack)
|
||||
|
||||
Swapping the `db` image from `pgduckdb/pgduckdb:18` to
|
||||
`timescale/timescaledb:latest-pg18` means a **fresh PGDATA volume** — and that
|
||||
volume also holds the durable **accounts** data. So this cutover is not a plain
|
||||
`docker compose pull`; it re-seeds both accounts and the climate record.
|
||||
|
||||
1. **Back up first.** `pg_dump` the accounts tables from the *old* db container
|
||||
(`docker compose exec db pg_dump -U thermograph -t user -t subscription \
|
||||
-t notification -t push_subscription -t pending_digest thermograph > accounts.sql`),
|
||||
or keep the pre-Postgres `accounts.sqlite` as the source of truth.
|
||||
2. **Replace the DB.** Deploy this branch (compose now points at the stock
|
||||
TimescaleDB image, no `build:`). Bring down the stack and remove the old volume
|
||||
so PGDATA re-inits on the new image: `docker compose down && docker volume rm
|
||||
<project>_pgdata`.
|
||||
3. **Build the schema.** `docker compose up -d db`, then
|
||||
`docker compose run --rm backend alembic upgrade head` — this creates the accounts
|
||||
tables *and* the climate hypertables, and `CREATE EXTENSION timescaledb`.
|
||||
4. **Restore accounts.** Either `psql < accounts.sql`, or re-run
|
||||
`migrate_accounts_to_pg.py --sqlite <accounts.sqlite>` (§ the SQLite cutover
|
||||
above). Verify per-table counts.
|
||||
5. **Backfill the climate record** from the parquet cache (mount it and run the
|
||||
backfill; idempotent, resumable):
|
||||
```
|
||||
docker compose run --rm \
|
||||
-v /opt/thermograph/data/cache:/app/data/cache:ro \
|
||||
backend python migrate_cache_to_pg.py
|
||||
```
|
||||
It loads `climate_history` + `climate_recent` and sets each cell's
|
||||
`climate_sync` to the parquet **file mtime**, so `recent_stamp` (hence every
|
||||
derived-payload token) is unchanged across the cutover. Cells without a
|
||||
schema-complete parquet record are skipped and refetch lazily on first request.
|
||||
6. **Start serving:** `docker compose up -d`. Smoke test: login, a subscription,
|
||||
a cached calendar/day request (served from `climate_history`), the metrics
|
||||
dashboard.
|
||||
7. Keep the pre-cutover DB backup (accounts dump + parquet cache) for a few days.
|
||||
The climate record is now **durable** — make sure `pg_dump`/PITR covers it.
|
||||
|
||||
## Notes
|
||||
|
||||
- The two `deploy/migrations/*.sql` files are superseded by Alembic on Postgres
|
||||
(they were SQLite-specific `ALTER TABLE` column-adds for the old prod DB).
|
||||
- The TimescaleDB image's data dir is `/var/lib/postgresql/data`; the compose
|
||||
volume is mounted at the parent `/var/lib/postgresql` so it persists without the
|
||||
initdb permission issue that pinning `PGDATA` to the mountpoint can trigger.
|
||||
- Climate falls back to the parquet cache whenever `THERMOGRAPH_DATABASE_URL` is
|
||||
not a Postgres URL (dev, tests, offline tooling) — the same dialect switch as the
|
||||
accounts DB, so a DB outage in prod degrades to upstream refetch, never a 500.
|
||||
87
infra/deploy/db/README.md
Normal file
87
infra/deploy/db/README.md
Normal file
|
|
@ -0,0 +1,87 @@
|
|||
# Thermograph DB — TimescaleDB (PostgreSQL 18)
|
||||
|
||||
The `db` service runs the stock **`timescale/timescaledb:latest-pg18`** image
|
||||
(TimescaleDB 2.24+, genuine PostgreSQL 18). The app's climate record lives here in
|
||||
**hypertables** — the DB, not the filesystem, is the source of truth:
|
||||
|
||||
- **`climate_history`** — a hypertable of the full daily archive per grid cell
|
||||
(`cell_id, date, tmax, tmin, precip, wind, gust, humid, fmax, fmin, feels`), back
|
||||
to 1980. Durable/LOGGED (a 45-year, rate-limited refetch is expensive),
|
||||
range-partitioned on `date` (5-year chunks), compressed for chunks older than a
|
||||
year.
|
||||
- **`climate_recent`** — the recent-observations + forward-forecast bundle (a plain
|
||||
table: it holds future dates and is rewritten hourly).
|
||||
- **`climate_sync`** — per-cell freshness (epoch seconds) that replaces the old
|
||||
parquet file mtimes: it drives the hourly history top-up, the 1-hour forecast
|
||||
TTL, and the `recent_stamp` token embedded in derived-payload validity.
|
||||
|
||||
The schema is created by Alembic (`backend/alembic/versions/0002_climate_hypertables.py`,
|
||||
run at app boot via `deploy/entrypoint.sh`). The app reads/writes it through
|
||||
`backend/data/climate_store.py` (psycopg + polars). See
|
||||
`deploy/POSTGRES-MIGRATION.md` for the parquet→hypertable cutover.
|
||||
|
||||
## Why the stock image (no custom Dockerfile)
|
||||
|
||||
The previous DB image was a custom `pgduckdb/pgduckdb:18` build whose only purpose
|
||||
was ad-hoc `read_parquet()` over the parquet cache. Now the climate record is in
|
||||
real tables, so that capability is gone and the DB is the **stock TimescaleDB
|
||||
image** — no build step. The image already sets
|
||||
`shared_preload_libraries=timescaledb`; never `ALTER SYSTEM SET
|
||||
shared_preload_libraries` (it would land in `postgresql.auto.conf` and override the
|
||||
image's preload).
|
||||
|
||||
## Files here
|
||||
|
||||
- **`init/10-timescaledb.sql`** — `CREATE EXTENSION IF NOT EXISTS timescaledb;`
|
||||
(runs from `/docker-entrypoint-initdb.d` on first cluster init; Alembic also does
|
||||
this idempotently at boot).
|
||||
- **`init/20-tuning.sh`** — scales `shared_buffers` (25%), `effective_cache_size`
|
||||
(75%), `work_mem`, and `maintenance_work_mem` from `DB_MEMORY` via `ALTER SYSTEM`.
|
||||
|
||||
## Compose `db` service
|
||||
|
||||
```yaml
|
||||
db:
|
||||
image: timescale/timescaledb:latest-pg18
|
||||
environment:
|
||||
POSTGRES_USER: thermograph
|
||||
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD}
|
||||
POSTGRES_DB: thermograph
|
||||
DB_MEMORY: ${DB_MEMORY:-8g}
|
||||
volumes:
|
||||
- pgdata:/var/lib/postgresql
|
||||
- ./deploy/db/init:/docker-entrypoint-initdb.d
|
||||
# healthcheck / cpus / mem_limit / shm_size: unchanged
|
||||
```
|
||||
|
||||
Notes:
|
||||
- The volume is mounted at the **parent** of the data dir; the image picks its own
|
||||
PGDATA subdir (`/var/lib/postgresql/data`) under it. The whole tree persists on
|
||||
the named volume.
|
||||
- No host port on purpose: the app reaches Postgres as `db:5432` on the compose
|
||||
network. Nothing outside the stack should touch the database.
|
||||
- Init scripts only run when PGDATA is empty. On an **existing** database, enable
|
||||
the extension once by hand:
|
||||
```
|
||||
docker compose exec db psql -U thermograph -d thermograph \
|
||||
-c 'CREATE EXTENSION IF NOT EXISTS timescaledb;'
|
||||
```
|
||||
|
||||
## Inspecting the hypertable
|
||||
|
||||
```sql
|
||||
-- Chunk / compression overview
|
||||
SELECT hypertable_name, num_chunks, compression_enabled
|
||||
FROM timescaledb_information.hypertables;
|
||||
|
||||
-- One cell, most recent archived days
|
||||
SELECT date, tmax, tmin, precip
|
||||
FROM climate_history WHERE cell_id = '1026_-2857'
|
||||
ORDER BY date DESC LIMIT 5;
|
||||
|
||||
-- Per-year highs for one cell
|
||||
SELECT EXTRACT(YEAR FROM date) AS yr,
|
||||
ROUND(AVG(tmax)::numeric, 1) AS avg_tmax, MAX(tmax) AS record_high
|
||||
FROM climate_history WHERE cell_id = '1026_-2857' AND date >= '2020-01-01'
|
||||
GROUP BY yr ORDER BY yr;
|
||||
```
|
||||
14
infra/deploy/db/init/10-timescaledb.sql
Normal file
14
infra/deploy/db/init/10-timescaledb.sql
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
-- Enable TimescaleDB so the app's climate record lives in hypertables
|
||||
-- (climate_history) instead of parquet files. The stock timescale/timescaledb
|
||||
-- image already preloads the extension (shared_preload_libraries=timescaledb);
|
||||
-- this just runs CREATE EXTENSION on first cluster init.
|
||||
--
|
||||
-- This runs once, on first cluster init (empty PGDATA), from
|
||||
-- /docker-entrypoint-initdb.d. Because the compose `db` service keeps its data on
|
||||
-- a persistent named volume, init scripts do NOT re-run on an existing database.
|
||||
-- Alembic (backend/alembic) also runs `CREATE EXTENSION IF NOT EXISTS timescaledb`
|
||||
-- at app boot, so an existing volume gets it there; to enable it by hand instead:
|
||||
--
|
||||
-- docker compose exec db psql -U thermograph -d thermograph \
|
||||
-- -c 'CREATE EXTENSION IF NOT EXISTS timescaledb;'
|
||||
CREATE EXTENSION IF NOT EXISTS timescaledb;
|
||||
62
infra/deploy/db/init/20-tuning.sh
Executable file
62
infra/deploy/db/init/20-tuning.sh
Executable file
|
|
@ -0,0 +1,62 @@
|
|||
#!/bin/bash
|
||||
# Postgres memory / performance tuning, scaled to the container's DB_MEMORY budget so
|
||||
# the same init serves every host (beta 8g; prod 16g on the 48 GB box) with no
|
||||
# hardcoding. Runs once on a fresh data volume from /docker-entrypoint-initdb.d, after
|
||||
# 10-timescaledb.sql enables timescaledb. Settings are written via ALTER SYSTEM
|
||||
# (persisted to postgresql.auto.conf, which the timescaledb image's own
|
||||
# timescaledb-tune postgresql.conf defers to); the container's post-init restart brings
|
||||
# restart-only settings (shared_buffers, …) into effect. NB: never ALTER SYSTEM SET
|
||||
# shared_preload_libraries here — that would land in auto.conf and override the image's
|
||||
# `timescaledb` preload. Init scripts do NOT re-run on an existing volume — to
|
||||
# re-tune later, set DB_MEMORY and run this by hand, then restart:
|
||||
# docker compose exec -e DB_MEMORY=16g db bash /docker-entrypoint-initdb.d/20-tuning.sh
|
||||
# docker compose restart db
|
||||
set -euo pipefail
|
||||
|
||||
# Parse DB_MEMORY ("16g" / "8192m" / plain MB) into whole MB; default + floor at 8 GB.
|
||||
budget="${DB_MEMORY:-8g}"
|
||||
num="${budget//[!0-9]/}"
|
||||
num="${num:-8}"
|
||||
unit="$(printf '%s' "$budget" | tr -dc '[:alpha:]' | tr '[:upper:]' '[:lower:]')"
|
||||
case "$unit" in
|
||||
g | gb) mb=$((num * 1024)) ;;
|
||||
m | mb | "") mb="$num" ;;
|
||||
*) mb=8192 ;;
|
||||
esac
|
||||
if [ "$mb" -lt 1024 ]; then mb=8192; fi
|
||||
|
||||
# Derive settings from the budget. The ratios reproduce the historical 8 GB tuning
|
||||
# (shared_buffers 2 GB, effective_cache_size 6 GB, work_mem 64 MB,
|
||||
# maintenance_work_mem 512 MB) and scale linearly on a bigger box.
|
||||
shared_buffers=$((mb / 4)) # 25% — the shared page cache
|
||||
effective_cache=$((mb * 3 / 4)) # 75% — planner's view of total cache (PG + OS)
|
||||
work_mem=$((mb / 128)) # ~64 MB at 8 GB (per-operation; kept modest)
|
||||
if [ "$work_mem" -lt 16 ]; then work_mem=16; fi
|
||||
maint_mem=$((mb / 16)) # 512 MB at 8 GB — index builds / VACUUM
|
||||
|
||||
echo "[tuning] DB_MEMORY=${budget} -> ${mb}MB: shared_buffers=${shared_buffers}MB" \
|
||||
"effective_cache_size=${effective_cache}MB work_mem=${work_mem}MB" \
|
||||
"maintenance_work_mem=${maint_mem}MB"
|
||||
|
||||
psql -v ON_ERROR_STOP=1 --username "$POSTGRES_USER" --dbname "$POSTGRES_DB" <<SQL
|
||||
-- Caching: the shared page cache, and the planner's view of total cache (PG + OS).
|
||||
ALTER SYSTEM SET shared_buffers = '${shared_buffers}MB';
|
||||
ALTER SYSTEM SET effective_cache_size = '${effective_cache}MB';
|
||||
|
||||
-- Processing: per-operation sort/hash memory, and maintenance (index builds, VACUUM).
|
||||
ALTER SYSTEM SET work_mem = '${work_mem}MB';
|
||||
ALTER SYSTEM SET maintenance_work_mem = '${maint_mem}MB';
|
||||
|
||||
-- Write throughput: fewer, larger checkpoints.
|
||||
ALTER SYSTEM SET wal_buffers = '16MB';
|
||||
ALTER SYSTEM SET min_wal_size = '1GB';
|
||||
ALTER SYSTEM SET max_wal_size = '4GB';
|
||||
ALTER SYSTEM SET checkpoint_completion_target = 0.9;
|
||||
|
||||
-- SSD-friendly planner + IO concurrency.
|
||||
ALTER SYSTEM SET random_page_cost = 1.1;
|
||||
ALTER SYSTEM SET effective_io_concurrency = 200;
|
||||
|
||||
-- Room for parallel scans/aggregates on the bigger analytic queries.
|
||||
ALTER SYSTEM SET max_parallel_workers_per_gather = 2;
|
||||
SQL
|
||||
68
infra/deploy/deploy-dev.sh
Executable file
68
infra/deploy/deploy-dev.sh
Executable file
|
|
@ -0,0 +1,68 @@
|
|||
#!/usr/bin/env bash
|
||||
# LAN-dev deploy: roll the per-service registry-pull stack (see deploy.sh) onto
|
||||
# the ~/thermograph-dev overlay instead of prod/beta's loopback-only stack.
|
||||
#
|
||||
# This script assumes $APP_DIR is a checkout of THIS repo (thermograph-infra),
|
||||
# the same way /opt/thermograph is on prod/beta. That is the live state: the LAN
|
||||
# box's ~/thermograph-dev was reprovisioned as an infra checkout during the
|
||||
# 2026-07-22 cutover (the app monorepo is archived), provision-dev-lan.sh's
|
||||
# REPO_URL defaults to this repo, and the app repos' deploy-dev.yml workflows
|
||||
# invoke this script on the thermograph-lan runner. This IS the live LAN-dev
|
||||
# deploy path.
|
||||
#
|
||||
# Design: a thin wrapper around deploy.sh, not a duplicate. deploy.sh already owns
|
||||
# the entire per-service registry-pull mechanism -- secrets sourcing, docker login,
|
||||
# the retry-pull loop, --no-deps single-service rolls vs. --remove-orphans `all`,
|
||||
# .image-tags.env persistence, and the 8137/8080 health checks. None of that is
|
||||
# dev-specific; the only things LAN dev actually changes are WHERE it deploys
|
||||
# (a separate checkout + branch) and WHICH compose files are in play (the base
|
||||
# file plus docker-compose.dev.yml's uncapped/LAN-published overrides). Both are
|
||||
# expressible as environment (APP_DIR/BRANCH that deploy.sh already reads, and
|
||||
# docker compose's own COMPOSE_FILE variable), so re-exec'ing deploy.sh with that
|
||||
# environment set covers it with no forked copy of the pull/roll/health logic to
|
||||
# drift out of sync. If LAN dev ever needs deploy logic that genuinely diverges
|
||||
# from prod/beta (not just "different files/directory"), promote this to a
|
||||
# standalone script at that point rather than growing special cases into deploy.sh.
|
||||
#
|
||||
# Usage, mirrors deploy.sh directly:
|
||||
# # roll just the backend onto a dev-tagged image:
|
||||
# ssh dev-box 'SERVICE=backend BACKEND_IMAGE_TAG=sha-<12hex> deploy/deploy-dev.sh'
|
||||
# # bring the whole dev stack up (both tags required, same as deploy.sh):
|
||||
# ssh dev-box 'SERVICE=all BACKEND_IMAGE_TAG=sha-<a> FRONTEND_IMAGE_TAG=sha-<b> deploy/deploy-dev.sh'
|
||||
set -euo pipefail
|
||||
|
||||
# Dev context: a separate checkout + branch from prod/beta's /opt/thermograph
|
||||
# (main), so a dev deploy never touches or is touched by the prod/beta one.
|
||||
APP_DIR="${APP_DIR:-$HOME/thermograph-dev}"
|
||||
BRANCH="${BRANCH:-dev}"
|
||||
export APP_DIR BRANCH
|
||||
|
||||
# Point every `docker compose` invocation inside deploy.sh at the LAN-dev overlay
|
||||
# (uncapped CPU, backend published on 0.0.0.0:8137, frontend unpublished -- see
|
||||
# docker-compose.dev.yml) without deploy.sh needing to know dev exists at all.
|
||||
# COMPOSE_FILE is docker compose's own env var for this; ':' is the Linux/macOS
|
||||
# path-list separator it expects (';' only on Windows).
|
||||
export COMPOSE_FILE="docker-compose.yml:docker-compose.dev.yml"
|
||||
|
||||
# deploy.sh needs POSTGRES_PASSWORD to interpolate the compose file (both to init
|
||||
# the db container and to build backend's THERMOGRAPH_DATABASE_URL) -- normally
|
||||
# supplied by /etc/thermograph.env via render-secrets.sh's SOPS render, but dev has
|
||||
# no entry in deploy/secrets/ (only common/beta/prod.yaml exist), so
|
||||
# render_thermograph_secrets degrades to a no-op here (see render-secrets.sh) and
|
||||
# /etc/thermograph.env likely doesn't exist on a fresh dev box either. Postgres is
|
||||
# loopback-only inside the compose network on dev (never published, no real
|
||||
# credentials on this box), so a fixed default is fine -- same value the monorepo's
|
||||
# deploy-dev.sh used historically. Override via the environment if you ever want a
|
||||
# different one; this only sets it when nothing else already has.
|
||||
export POSTGRES_PASSWORD="${POSTGRES_PASSWORD:-thermograph-dev}"
|
||||
|
||||
# REGISTRY_TOKEN (for `docker login` in deploy.sh) is a REAL credential -- read
|
||||
# access to git.thermograph.org's registry -- and deliberately has NO default here.
|
||||
# Until dev gets its own deploy/secrets/dev.yaml entry (see deploy/secrets/README.md)
|
||||
# it must come from the calling environment: exported by hand for a manual run, or
|
||||
# injected by whatever eventually SSHes in to trigger a dev deploy (analogous to
|
||||
# how deploy.yml supplies it for prod/beta today). deploy.sh's own `docker login`
|
||||
# step fails loudly if it's unset or wrong -- nothing silently no-ops.
|
||||
|
||||
echo "==> LAN-dev deploy: APP_DIR=$APP_DIR BRANCH=$BRANCH COMPOSE_FILE=$COMPOSE_FILE"
|
||||
exec "$(dirname "$0")/deploy.sh" "$@"
|
||||
293
infra/deploy/deploy.sh
Executable file
293
infra/deploy/deploy.sh
Executable file
|
|
@ -0,0 +1,293 @@
|
|||
#!/usr/bin/env bash
|
||||
# Pull this INFRA repo's checkout up to date, then roll ONE (or all) of the
|
||||
# docker-compose app services onto its separately-published image. Run on the
|
||||
# VPS — each app repo's Forgejo Actions workflow invokes this over SSH (see
|
||||
# .forgejo/workflows/deploy.yml in thermograph-backend / thermograph-frontend),
|
||||
# and you can run it by hand too.
|
||||
#
|
||||
# # roll just the backend onto a specific image:
|
||||
# ssh deploy@vps 'SERVICE=backend BACKEND_IMAGE_TAG=sha-<12hex> /opt/thermograph/deploy/deploy.sh'
|
||||
# # roll just the frontend:
|
||||
# ssh deploy@vps 'SERVICE=frontend FRONTEND_IMAGE_TAG=sha-<12hex> /opt/thermograph/deploy/deploy.sh'
|
||||
# # bring the whole stack up (both tags required):
|
||||
# ssh deploy@vps 'SERVICE=all BACKEND_IMAGE_TAG=sha-<a> FRONTEND_IMAGE_TAG=sha-<b> /opt/thermograph/deploy/deploy.sh'
|
||||
#
|
||||
# FE/BE CI-CD split: backend and frontend are published from separate repos as
|
||||
# separate images (emi/thermograph-backend/app, emi/thermograph-frontend/app),
|
||||
# so a deploy targets ONE service and leaves the other's running container +
|
||||
# tag untouched. Each service's live tag is persisted host-side in
|
||||
# deploy/.image-tags.env (untracked -- survives the git reset below) so a
|
||||
# single-service roll re-renders compose with BOTH services' real tags and
|
||||
# never accidentally recreates or downgrades the sibling.
|
||||
#
|
||||
# This checkout is thermograph-infra, not an app repo: BRANCH is this repo's
|
||||
# branch (compose files, db init, the secrets vault, this script itself);
|
||||
# the *_IMAGE_TAG values are the separately-published app images to run -- the
|
||||
# two axes are independent and rarely change together.
|
||||
set -euo pipefail
|
||||
|
||||
APP_DIR="${APP_DIR:-/opt/thermograph}"
|
||||
BRANCH="${BRANCH:-main}"
|
||||
# Which service this deploy rolls: backend | frontend | all. Defaults to `all`
|
||||
# (a full-stack bring-up) so a by-hand run with both tags still works; the
|
||||
# per-repo deploy.yml workflows always pass an explicit single service.
|
||||
SERVICE="${SERVICE:-all}"
|
||||
HEALTH_PORT="${HEALTH_PORT:-8137}"
|
||||
cd "$APP_DIR"
|
||||
|
||||
case "$SERVICE" in
|
||||
backend|frontend|all) ;;
|
||||
*) echo "!! SERVICE must be backend|frontend|all, got '$SERVICE'" >&2; exit 2 ;;
|
||||
esac
|
||||
|
||||
# Serialize deploys on this host. Backend and frontend deploy from SEPARATE
|
||||
# repos whose workflows can fire for the same push within seconds of each
|
||||
# other, and both SSH into this one checkout: concurrent runs race on the
|
||||
# `git reset` below, the shared compose project, and the .image-tags.env
|
||||
# read/modify/write (a lost update there re-rolls the sibling onto a stale
|
||||
# tag). flock makes the second deploy wait its turn instead. The lock fd is
|
||||
# inherited across the self re-exec below, so the lock spans the whole run;
|
||||
# -w 600 bounds the wait (a deploy holding the lock >10 min is already
|
||||
# broken), after which this exits non-zero and the CI job fails loudly.
|
||||
DEPLOY_LOCK="$APP_DIR/deploy/.deploy.lock"
|
||||
if [ -z "${DEPLOY_SH_FLOCKED:-}" ]; then
|
||||
export DEPLOY_SH_FLOCKED=1
|
||||
exec flock -w 600 "$DEPLOY_LOCK" "$0" "$@"
|
||||
fi
|
||||
|
||||
# Secrets (POSTGRES_PASSWORD, VAPID keys, AUTH_SECRET, ...) drive compose
|
||||
# interpolation and are also loaded into the backend container via env_file.
|
||||
#
|
||||
# When this host is configured for SOPS (an age key + /etc/thermograph/secrets-env),
|
||||
# first render /etc/thermograph.env from the committed encrypted source of truth
|
||||
# (deploy/secrets/*.yaml) so a key rotation is just an edit+commit+deploy. The guard
|
||||
# on the helper's existence keeps the very deploy that INTRODUCES this file safe: on
|
||||
# the first pass the checkout may predate it (it arrives with the git reset below,
|
||||
# after which deploy.sh re-execs), so a missing helper simply falls back to the
|
||||
# existing /etc/thermograph.env. Then source it so a by-hand run interpolates the
|
||||
# same as the systemd unit does. See deploy/render-secrets.sh + deploy/secrets/.
|
||||
if [ -f "$APP_DIR/deploy/render-secrets.sh" ]; then
|
||||
# shellcheck source=deploy/render-secrets.sh
|
||||
. "$APP_DIR/deploy/render-secrets.sh"
|
||||
render_thermograph_secrets "$APP_DIR"
|
||||
fi
|
||||
set -a; . /etc/thermograph.env 2>/dev/null || true; set +a
|
||||
|
||||
# Pre-warm the ~750 city-page archives so /climate pages serve from cache and a
|
||||
# search-engine crawl never bursts the archive API quota. Detached inside the
|
||||
# backend container (compose exec -d), idempotent (skips already-cached cells),
|
||||
# so it never blocks the deploy or health check and is cheap on every deploy
|
||||
# after the first full warm.
|
||||
warm_city_archives() {
|
||||
echo "==> Warming city-page archives in the background (backend:/app/logs/warm-cities.log)"
|
||||
docker compose exec -d backend sh -c \
|
||||
'python warm_cities.py --pace 2 >> /app/logs/warm-cities.log 2>&1' || true
|
||||
}
|
||||
|
||||
# Notify IndexNow (Bing / DuckDuckGo / Yandex) of the site's URLs, but only when
|
||||
# the set of pages actually changed (a new/removed city) — code-only deploys skip.
|
||||
# Best-effort: never fails the deploy.
|
||||
ping_indexnow() {
|
||||
echo "==> Pinging IndexNow (only if the URL set changed)"
|
||||
local base="${THERMOGRAPH_BASE_URL:-https://thermograph.org}"
|
||||
docker compose exec -T backend python indexnow.py --if-changed "$base" \
|
||||
|| echo "!! IndexNow ping failed (non-fatal)" >&2
|
||||
}
|
||||
|
||||
echo "==> Fetching $BRANCH"
|
||||
git fetch --prune origin "$BRANCH"
|
||||
git reset --hard "origin/$BRANCH"
|
||||
|
||||
# Re-exec: git reset --hard just rewrote this very file's bytes on disk while
|
||||
# it's still running. bash reads a script via buffered, byte-offset I/O, so
|
||||
# anything AFTER this point in the OLD execution can read from the wrong
|
||||
# offset once the file's size/content changed underneath it -- a classic
|
||||
# self-modifying-script footgun. Confirmed live: after this PR added ~15
|
||||
# lines above, one deploy ran with the OLD "Building images" log lines even
|
||||
# though `git status` showed the checkout correctly at the NEW commit --
|
||||
# the file changed under a running interpreter, not the checkout. Restart
|
||||
# fresh from the now-updated file so everything after this line is
|
||||
# guaranteed self-consistent. Guarded so the second invocation doesn't
|
||||
# fetch+reset+re-exec forever.
|
||||
if [ -z "${DEPLOY_SH_REEXECED:-}" ]; then
|
||||
export DEPLOY_SH_REEXECED=1
|
||||
exec "$0" "$@"
|
||||
fi
|
||||
|
||||
# Stack-mode routing: a host whose /etc/thermograph/deploy-mode says "stack"
|
||||
# (prod, after the Swarm cutover) deploys via the Swarm stack path instead of
|
||||
# compose. Checked AFTER the reset+re-exec so the stack script is always the
|
||||
# freshly-pulled one, and the SERVICE/tag contract passes through unchanged --
|
||||
# the app repos' workflows never need to know which mode a host runs.
|
||||
if [ "$(cat /etc/thermograph/deploy-mode 2>/dev/null || true)" = "stack" ]; then
|
||||
exec bash "$APP_DIR/deploy/stack/deploy-stack.sh"
|
||||
fi
|
||||
|
||||
# Registry-pull cutover: pull the image each app repo's build-push.yml already
|
||||
# built and pushed, instead of building in place. This checkout is
|
||||
# thermograph-infra, not an app repo, so there's no "current commit" to derive
|
||||
# a tag from -- the caller (the deploying repo's deploy.yml) exports its own
|
||||
# service's tag (BACKEND_IMAGE_TAG or FRONTEND_IMAGE_TAG = sha-<12 hex> of the
|
||||
# app commit, or a semver tag).
|
||||
REGISTRY_HOST="${REGISTRY_HOST:-git.thermograph.org}"
|
||||
export REGISTRY_HOST BACKEND_IMAGE_PATH FRONTEND_IMAGE_PATH
|
||||
|
||||
# Load the last-deployed tag for BOTH services first, so a single-service roll
|
||||
# still renders compose with the sibling's real, currently-running tag (never a
|
||||
# bare `local` that would recreate/downgrade it). The incoming env for the
|
||||
# service being deployed then overrides its line below. This file is untracked
|
||||
# (see .gitignore), so `git reset --hard` above leaves it in place.
|
||||
TAGS_FILE="$APP_DIR/deploy/.image-tags.env"
|
||||
# Capture the tags the caller explicitly passed BEFORE sourcing -- they must
|
||||
# WIN. The file only supplies the *sibling's* last-known tag; sourcing it
|
||||
# unconditionally would clobber an incoming tag (e.g. a frontend deploy whose
|
||||
# FRONTEND_IMAGE_TAG got overwritten by the stale `local` the first backend-only
|
||||
# deploy persisted for the not-yet-known sibling -> pull `:local` -> "manifest
|
||||
# unknown"). So source for the sibling, then re-apply the caller's own value.
|
||||
_incoming_backend="${BACKEND_IMAGE_TAG:-}"
|
||||
_incoming_frontend="${FRONTEND_IMAGE_TAG:-}"
|
||||
if [ -f "$TAGS_FILE" ]; then
|
||||
set -a; . "$TAGS_FILE"; set +a
|
||||
fi
|
||||
[ -n "$_incoming_backend" ] && BACKEND_IMAGE_TAG="$_incoming_backend"
|
||||
[ -n "$_incoming_frontend" ] && FRONTEND_IMAGE_TAG="$_incoming_frontend"
|
||||
|
||||
# Guard: the service(s) being rolled MUST have a concrete tag supplied now (the
|
||||
# sibling's may come from the persisted file). `all` needs both.
|
||||
case "$SERVICE" in
|
||||
backend) : "${BACKEND_IMAGE_TAG:?set BACKEND_IMAGE_TAG=sha-<12-hex> for a backend deploy}" ;;
|
||||
frontend) : "${FRONTEND_IMAGE_TAG:?set FRONTEND_IMAGE_TAG=sha-<12-hex> for a frontend deploy}" ;;
|
||||
all)
|
||||
: "${BACKEND_IMAGE_TAG:?set BACKEND_IMAGE_TAG=sha-<12-hex> (SERVICE=all needs both)}"
|
||||
: "${FRONTEND_IMAGE_TAG:?set FRONTEND_IMAGE_TAG=sha-<12-hex> (SERVICE=all needs both)}"
|
||||
;;
|
||||
esac
|
||||
# Compose interpolates both vars for the whole file even when we act on one
|
||||
# service; default the not-yet-known sibling (first-ever deploy) to `local` so
|
||||
# interpolation doesn't warn -- harmless since --no-deps never touches it.
|
||||
export BACKEND_IMAGE_TAG="${BACKEND_IMAGE_TAG:-local}"
|
||||
export FRONTEND_IMAGE_TAG="${FRONTEND_IMAGE_TAG:-local}"
|
||||
|
||||
# Which compose services this run pulls/rolls.
|
||||
case "$SERVICE" in
|
||||
backend) TARGETS=(backend) ;;
|
||||
frontend) TARGETS=(frontend) ;;
|
||||
all) TARGETS=(backend frontend) ;;
|
||||
esac
|
||||
|
||||
# Login only when a token is supplied. The SSH/CI deploy paths (deploy.yml,
|
||||
# deploy-prod.yml, deploy-dev on the LAN runner) don't pass REGISTRY_TOKEN --
|
||||
# the host is already `docker login`ed to the registry (persistent cred in
|
||||
# ~/.docker/config.json), so an unconditional login with an empty token would
|
||||
# abort the deploy under `set -e`. Use the token if present, else trust the
|
||||
# host's existing cred; a genuine auth problem then fails loudly at `pull`.
|
||||
if [ -n "${REGISTRY_TOKEN:-}" ]; then
|
||||
echo "==> Logging in to the registry ($REGISTRY_HOST)"
|
||||
echo "$REGISTRY_TOKEN" | docker login "$REGISTRY_HOST" --username emi --password-stdin
|
||||
else
|
||||
echo "==> No REGISTRY_TOKEN in env; relying on the host's existing docker login to $REGISTRY_HOST"
|
||||
fi
|
||||
|
||||
echo "==> Pulling images (backend=$BACKEND_IMAGE_TAG frontend=$FRONTEND_IMAGE_TAG; rolling: ${TARGETS[*]})"
|
||||
# Retry: build-push.yml (triggered by the same push) has no ordering guarantee
|
||||
# against this deploy -- Forgejo Actions `needs:` only works between jobs in ONE
|
||||
# workflow file, not across the separate build-push.yml triggered by the same
|
||||
# event. Confirmed live: a deploy raced ahead of the push and failed with "not
|
||||
# found". A bounded retry (~5 min) covers a normal build; a genuine problem
|
||||
# (bad tag, registry down) still fails loudly after that.
|
||||
pull_ok=0
|
||||
for i in $(seq 1 30); do
|
||||
if docker compose pull "${TARGETS[@]}"; then
|
||||
pull_ok=1
|
||||
break
|
||||
fi
|
||||
echo " pull attempt $i/30 failed (image may not be pushed yet); retrying in 10s..." >&2
|
||||
sleep 10
|
||||
done
|
||||
if [ "$pull_ok" != 1 ]; then
|
||||
echo "!! docker compose pull failed after 30 attempts" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Roll only the target service(s). Backend schema migrations run inside its own
|
||||
# entrypoint (alembic upgrade head) before uvicorn, so there's no separate
|
||||
# migrate step; frontend is stateless.
|
||||
#
|
||||
# Single-service rolls use --no-deps so recreating backend doesn't also bounce
|
||||
# db, and recreating frontend doesn't touch backend -- that independence is the
|
||||
# whole point of the split. A full `all` deploy instead uses --remove-orphans,
|
||||
# which matters when the service topology itself changes (a renamed-away
|
||||
# service's old container would otherwise keep running and squat its port --
|
||||
# confirmed live, this is what blocked beta's first dual-service deploy with
|
||||
# "port is already allocated").
|
||||
echo "==> Rolling ${TARGETS[*]}"
|
||||
if [ "$SERVICE" = all ]; then
|
||||
docker compose up -d --remove-orphans
|
||||
else
|
||||
docker compose up -d --no-deps "${TARGETS[@]}"
|
||||
fi
|
||||
|
||||
# Persist the now-live tags so the next single-service deploy knows the
|
||||
# sibling's real tag. Written after `up` so a failed pull never records a tag
|
||||
# that isn't actually running.
|
||||
mkdir -p "$(dirname "$TAGS_FILE")"
|
||||
cat > "$TAGS_FILE" <<EOF
|
||||
# Written by deploy.sh -- the image tag each service is currently running.
|
||||
# Untracked (see .gitignore); lets a single-service deploy leave the other alone.
|
||||
BACKEND_IMAGE_TAG=$BACKEND_IMAGE_TAG
|
||||
FRONTEND_IMAGE_TAG=$FRONTEND_IMAGE_TAG
|
||||
EOF
|
||||
|
||||
# Health check the service(s) we rolled: backend on 8137, frontend on 8080.
|
||||
# (For `all`, backend's `/` serving is the readiness signal the old script used
|
||||
# and the frontend depends_on backend anyway.)
|
||||
# Health via each container's own HEALTHCHECK (docker inspect), NOT a host-port
|
||||
# curl: the dev overlay leaves the frontend port UNpublished (reached through the
|
||||
# backend's _proxy_to_frontend), so a localhost:8080 curl spuriously fails there.
|
||||
# Both images HEALTHCHECK-curl /healthz internally, so this works published or not.
|
||||
health_ok=1
|
||||
for svc in "${TARGETS[@]}"; do
|
||||
cid=$(docker compose ps -q "$svc" 2>/dev/null)
|
||||
echo "==> Health check: $svc (container health)"
|
||||
ok=0; st=unknown
|
||||
for i in $(seq 1 40); do
|
||||
st=$(docker inspect --format '{{if .State.Health}}{{.State.Health.Status}}{{else}}none{{end}}' "$cid" 2>/dev/null || echo gone)
|
||||
[ "$st" = healthy ] && { ok=1; break; }
|
||||
if [ "$st" = none ] && [ "$(docker inspect --format '{{.State.Status}}' "$cid" 2>/dev/null)" = running ]; then ok=1; break; fi
|
||||
sleep 2
|
||||
done
|
||||
if [ "$ok" = 1 ]; then
|
||||
echo "==> OK: $svc is healthy"
|
||||
else
|
||||
echo "!! Health check failed for $svc (status=$st)" >&2
|
||||
health_ok=0
|
||||
fi
|
||||
done
|
||||
|
||||
if [ "$health_ok" != 1 ]; then
|
||||
docker compose ps || true
|
||||
for svc in "${TARGETS[@]}"; do docker compose logs --tail=50 "$svc" || true; done
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Every deploy pulls a new sha-tagged image and nothing ever removed the old
|
||||
# ones -- hosts accumulate gigabytes of dead tags at one per app commit. Keep
|
||||
# only the tags recorded as now-live (both services) and delete other tags of
|
||||
# the two app-image repos. Best-effort and after the health gate, so a failed
|
||||
# roll never garbage-collects the image a rollback would need; docker also
|
||||
# refuses to remove an image any container still uses.
|
||||
echo "==> Pruning old app-image tags"
|
||||
_be_repo="${REGISTRY_HOST}/${BACKEND_IMAGE_PATH:-emi/thermograph-backend/app}"
|
||||
_fe_repo="${REGISTRY_HOST}/${FRONTEND_IMAGE_PATH:-emi/thermograph-frontend/app}"
|
||||
docker images --format '{{.Repository}}:{{.Tag}}' \
|
||||
| grep -E "^(${_be_repo}|${_fe_repo}):" \
|
||||
| grep -v -e "^${_be_repo}:${BACKEND_IMAGE_TAG}$" -e "^${_fe_repo}:${FRONTEND_IMAGE_TAG}$" \
|
||||
| xargs -r docker rmi 2>/dev/null || true
|
||||
|
||||
# Post-deploy warm/IndexNow only make sense once the backend is (re)deployed --
|
||||
# they exec inside the backend container. Skip them on a frontend-only roll.
|
||||
if [ "$SERVICE" = backend ] || [ "$SERVICE" = all ]; then
|
||||
warm_city_archives
|
||||
ping_indexnow
|
||||
fi
|
||||
exit 0
|
||||
125
infra/deploy/forgejo/README.md
Normal file
125
infra/deploy/forgejo/README.md
Normal file
|
|
@ -0,0 +1,125 @@
|
|||
# Forgejo on the Swarm cluster
|
||||
|
||||
Runs as `deploy/forgejo/docker-stack.yml` — the only Swarm-scheduled workload
|
||||
this cluster carries (the Thermograph app itself stays on the
|
||||
Terraform-managed `docker compose` deploys; see `terraform/README.md`).
|
||||
Pinned to the **beta** node (old VPS) via the `role=forge` label from
|
||||
`deploy/swarm/label-forge-node.sh`.
|
||||
|
||||
The Actions **runner** is deliberately *not* part of this stack — it runs on
|
||||
the **desktop** as a plain systemd service (`register-lan-runner.sh` below),
|
||||
per `thermograph-docs/runbooks/implementation-handoff.md` Track B step 5. That's the
|
||||
canonical placement; an earlier revision of this stack ran the runner as a
|
||||
Swarm-scheduled Docker-in-Docker sidecar pinned to beta, which is gone now.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
1. All three nodes have joined the swarm (`deploy/swarm/`) and beta is
|
||||
labeled `role=forge`.
|
||||
2. `docker node ls` (from the manager) shows all three `Ready`.
|
||||
|
||||
## One-time setup: Swarm secret
|
||||
|
||||
One secret the stack expects to already exist (a Swarm secret, not a file —
|
||||
`external: true` in the stack file, so `docker stack deploy` never creates or
|
||||
sees the value, only references it):
|
||||
|
||||
```bash
|
||||
# A strong random password for Forgejo's own Postgres (NOT related to
|
||||
# Thermograph's app database — entirely separate instance/network).
|
||||
openssl rand -base64 32 | docker secret create forgejo_db_password -
|
||||
```
|
||||
|
||||
## Deploy / update
|
||||
|
||||
```bash
|
||||
docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo
|
||||
```
|
||||
|
||||
Re-running is safe — Swarm only touches services whose spec actually changed.
|
||||
Do this **before** the DNS + Caddy step below — Caddy's reverse_proxy target
|
||||
(`127.0.0.1:3080`) needs the `forgejo` service actually listening first, or
|
||||
its first health check just fails harmlessly until it is.
|
||||
|
||||
## DNS + TLS: reusing beta's existing Caddy, not a second reverse proxy
|
||||
|
||||
Forgejo is pinned to beta (`role=forge`) — but beta is also **today's live
|
||||
thermograph.org host**, and its Caddy already owns ports 80/443
|
||||
(`/etc/caddy/Caddyfile` on that box). A second ingress (Traefik) trying to
|
||||
bind the same ports would collide with it. So there's no Traefik in this
|
||||
stack: `forgejo`'s web port publishes to `127.0.0.1:3080` only (host-local),
|
||||
and beta's *existing* Caddy gets one more site block reverse-proxying to it —
|
||||
same pattern as its `thermograph.org` block, same automatic-HTTPS.
|
||||
|
||||
1. Point the Forgejo domain (default `git.thermograph.org`; override with
|
||||
`FORGEJO_DOMAIN=...` before `docker stack deploy`) at **beta's** public IP
|
||||
— that's where the task actually runs, not prod's or the desktop's.
|
||||
2. Append `deploy/forgejo/caddy-git.conf` to beta's `/etc/caddy/Caddyfile`,
|
||||
adjusting the domain if you didn't use the default, then `systemctl reload
|
||||
caddy`.
|
||||
3. That file also resolves the registry-exposure hazard (#15 in
|
||||
`thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md`): `/v2/*` (the registry
|
||||
API) is blocked to everything except the WireGuard mesh CIDR
|
||||
(`10.10.0.0/24`); the git/web UI stays public. CI runners and Swarm nodes
|
||||
reach the registry over the mesh, not the public internet — see "Registry
|
||||
access from mesh clients" below.
|
||||
|
||||
## Registry access from mesh clients
|
||||
|
||||
Any node that needs `docker login`/push/pull against the registry (the
|
||||
desktop's CI runner building/pushing images, later any Swarm node pulling
|
||||
them) must reach `git.thermograph.org` **over the WireGuard tunnel**, not
|
||||
beta's public IP — otherwise Caddy's `/v2/*` block above refuses the
|
||||
connection. Public DNS resolves the domain to beta's public IP, so add a
|
||||
`/etc/hosts` override on each such node pinning it to beta's WireGuard
|
||||
address instead:
|
||||
|
||||
```
|
||||
echo "10.10.0.2 git.thermograph.org" | sudo tee -a /etc/hosts
|
||||
```
|
||||
|
||||
(`10.10.0.2` is beta's WG address per `deploy/swarm/README.md`'s peer
|
||||
numbering — adjust if you assigned it differently.) The git/web UI keeps
|
||||
working normally for everyone else since only `/v2/*` is restricted.
|
||||
|
||||
## Register the Actions runner (on the desktop, not through Swarm)
|
||||
|
||||
Once Forgejo answers at its domain:
|
||||
|
||||
```bash
|
||||
# On the desktop:
|
||||
# Forgejo web UI -> Site Administration -> Actions -> Runners -> Create new Runner
|
||||
# (or, for a repo-scoped runner: <repo> -> Settings -> Actions -> Runners)
|
||||
# copy the registration token, then:
|
||||
bash deploy/forgejo/register-lan-runner.sh https://<forgejo-domain> <token>
|
||||
```
|
||||
|
||||
See that script's header for exactly what it replaces (the pre-Forgejo GitHub
|
||||
self-hosted runner on this same machine) and why it registers with two
|
||||
labels where there used to be two separate runners.
|
||||
|
||||
## Why Postgres here and not the Thermograph app's TimescaleDB
|
||||
|
||||
Separate instance, separate network (`forgejo_net`, not the app's compose
|
||||
network), separate volume. Forgejo is a distinct product with its own schema
|
||||
and its own backup/restore lifecycle — sharing a database with the app would
|
||||
couple two things that should be able to fail, migrate, and restore
|
||||
independently.
|
||||
|
||||
## Verifying
|
||||
|
||||
```bash
|
||||
docker service ls # forgejo_db, forgejo_forgejo both Running, 1/1
|
||||
curl -I https://git.thermograph.org/ # 200, valid cert (Caddy's, not a new one)
|
||||
curl -I https://git.thermograph.org/v2/ # 403 from anywhere off the WireGuard mesh
|
||||
# On the desktop, after registering the runner:
|
||||
systemctl --user status forgejo-runner # active, both labels registered
|
||||
```
|
||||
|
||||
## Rollback / removal
|
||||
|
||||
```bash
|
||||
docker stack rm forgejo
|
||||
# volumes (forgejo_data, forgejo_db, ...) survive a stack rm — remove them
|
||||
# explicitly only if you actually want to destroy the Forgejo instance's data.
|
||||
```
|
||||
37
infra/deploy/forgejo/caddy-git.conf
Normal file
37
infra/deploy/forgejo/caddy-git.conf
Normal file
|
|
@ -0,0 +1,37 @@
|
|||
# Only the block below (from "git.thermograph.org {") goes into the live
|
||||
# Caddyfile — append just that part, not this instructional header, or it
|
||||
# ends up as a confusing orphaned comment in production config with no
|
||||
# surrounding context (docker-stack.yml isn't visible from there).
|
||||
#
|
||||
# Target: /etc/caddy/Caddyfile on beta (the box `forgejo` is pinned to — see
|
||||
# docker-stack.yml). Reuses beta's existing Caddy instead of a second reverse
|
||||
# proxy/ACME flow: Forgejo publishes its web UI to 127.0.0.1:3080 only
|
||||
# (host-local), and this block is the only thing that ever talks to it.
|
||||
#
|
||||
# Registry exposure (hazard #15, see docker-stack.yml's header comment):
|
||||
# git.thermograph.org/v2/* is the built-in OCI registry API. It's blocked
|
||||
# from anywhere except the WireGuard mesh (10.10.0.0/24) — reachable to CI
|
||||
# runners and Swarm nodes over wg0, not from the public internet. Everything
|
||||
# else (web UI, git-over-HTTP, PR pages) stays public like the rest of the
|
||||
# site.
|
||||
#
|
||||
# Update the domain below to match whatever you actually pointed at beta's
|
||||
# IP, if not git.thermograph.org.
|
||||
|
||||
# --- copy from here down into /etc/caddy/Caddyfile ---
|
||||
|
||||
git.thermograph.org {
|
||||
encode zstd gzip
|
||||
|
||||
@registry_external {
|
||||
path /v2/*
|
||||
not remote_ip 10.10.0.0/24
|
||||
}
|
||||
respond @registry_external 403
|
||||
|
||||
reverse_proxy 127.0.0.1:3080
|
||||
|
||||
log {
|
||||
output file /var/log/caddy/git.log
|
||||
}
|
||||
}
|
||||
148
infra/deploy/forgejo/docker-stack.yml
Normal file
148
infra/deploy/forgejo/docker-stack.yml
Normal file
|
|
@ -0,0 +1,148 @@
|
|||
# Forgejo (self-hosted Git + CI/CD) as a Docker Swarm stack — the only workload
|
||||
# this Swarm cluster runs (see deploy/swarm/README.md). Deliberately separate
|
||||
# from the Terraform-managed docker-compose.yml that runs the Thermograph app
|
||||
# itself: this stack's only job is Forgejo and its container registry.
|
||||
#
|
||||
# The Actions RUNNER is deliberately NOT a service in this stack. Per
|
||||
# thermograph-docs/runbooks/implementation-handoff.md (Track B step 5), the canonical
|
||||
# design registers the runner on the desktop node as a plain systemd service
|
||||
# (deploy/forgejo/register-lan-runner.sh) — the same place the pre-Forgejo
|
||||
# GitHub self-hosted runner already lived, not a Swarm-scheduled container.
|
||||
#
|
||||
# No Traefik here. Forgejo is pinned to beta (node.labels.role == forge)
|
||||
# because that's what was chosen, but beta is ALSO today's live thermograph.org
|
||||
# host — Caddy already owns its ports 80/443 (see /etc/caddy/Caddyfile on that
|
||||
# box). A second reverse proxy binding those same ports would either fail to
|
||||
# start or fight Caddy. Instead: forgejo's web port publishes to
|
||||
# 127.0.0.1:3080 only (host-local, mode: host), and beta's existing Caddy gets
|
||||
# a new site block reverse-proxying git.thermograph.org -> 127.0.0.1:3080,
|
||||
# same pattern as its thermograph.org block. TLS is Caddy's existing
|
||||
# automatic-HTTPS (HTTP-01), not a second ACME flow.
|
||||
#
|
||||
# Deploy from the manager node (prod), after all three nodes (prod, beta,
|
||||
# desktop) have joined the swarm and beta is labeled role=forge:
|
||||
#
|
||||
# docker stack deploy -c deploy/forgejo/docker-stack.yml forgejo
|
||||
#
|
||||
# Requires this Swarm secret to exist first (see deploy/forgejo/README.md):
|
||||
# forgejo_db_password
|
||||
#
|
||||
# Registry exposure (hazard #15 in thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md):
|
||||
# resolved as the runbook's first listed option — serve Git/UI publicly but
|
||||
# firewall the /v2/ registry API paths so only WireGuard-mesh clients can
|
||||
# reach them. That's implemented in the Caddy site block (see
|
||||
# deploy/forgejo/README.md), not here — nothing in this stack file is
|
||||
# registry-specific, the restriction lives entirely in Caddy's config on beta.
|
||||
|
||||
services:
|
||||
db:
|
||||
image: postgres:16-alpine
|
||||
environment:
|
||||
POSTGRES_USER: forgejo
|
||||
POSTGRES_DB: forgejo
|
||||
POSTGRES_PASSWORD_FILE: /run/secrets/forgejo_db_password
|
||||
secrets: [forgejo_db_password]
|
||||
volumes:
|
||||
- forgejo_db:/var/lib/postgresql/data
|
||||
networks: [forgejo_net]
|
||||
deploy:
|
||||
placement:
|
||||
constraints: [node.labels.role == forge]
|
||||
restart_policy:
|
||||
condition: on-failure
|
||||
|
||||
forgejo:
|
||||
image: codeberg.org/forgejo/forgejo:9-rootless
|
||||
depends_on: [db]
|
||||
environment:
|
||||
FORGEJO__database__DB_TYPE: postgres
|
||||
FORGEJO__database__HOST: db:5432
|
||||
FORGEJO__database__NAME: forgejo
|
||||
FORGEJO__database__USER: forgejo
|
||||
FORGEJO__server__DOMAIN: "${FORGEJO_DOMAIN:-git.thermograph.org}"
|
||||
FORGEJO__server__ROOT_URL: "https://${FORGEJO_DOMAIN:-git.thermograph.org}/"
|
||||
FORGEJO__server__SSH_PORT: "2222"
|
||||
# Without this, Forgejo starts no SSH server at all despite SSH_PORT
|
||||
# being set — git@ clones get "connection refused", not a slow failure.
|
||||
FORGEJO__server__START_SSH_SERVER: "true"
|
||||
# Config is fully supplied via env, so lock the install wizard rather
|
||||
# than leave it open on the public internet waiting for someone to
|
||||
# complete it first (DB config alone does NOT imply this).
|
||||
FORGEJO__security__INSTALL_LOCK: "true"
|
||||
# Actions on: repo/org/user-level runners register against this instance.
|
||||
FORGEJO__actions__ENABLED: "true"
|
||||
# --- Access model: OAuth-only login, multi-user via an approval gate ---
|
||||
# Login is Google SSO only. Forgejo 9.0.3 has no setting to hide the
|
||||
# password sign-in form, so "OAuth-only" is enforced operationally: every
|
||||
# account is given an unusable (random) password and no NEW local account
|
||||
# can be created (ALLOW_ONLY_EXTERNAL_REGISTRATION). ENABLE_PASSWORD_SIGNIN_FORM
|
||||
# is set anyway — ignored on 9.0.3, effective if the instance is upgraded.
|
||||
# Auto-registration stays ON so a Google login auto-links (by verified email)
|
||||
# to a PRE-CREATED account (`forgejo admin user create --username X --email
|
||||
# X@gmail.com --random-password`). A single required-claim-value can't
|
||||
# allowlist more than one email, so instead of pinning one address the
|
||||
# required-claim-value is CLEARED (a DB change on the auth source, persisted
|
||||
# in the forgejo_data volume — not re-applied from here) and every NEW signup
|
||||
# is gated: a stranger who signs in with Google lands INACTIVE
|
||||
# (REGISTER_MANUAL_CONFIRM) and RESTRICTED (DEFAULT_USER_IS_RESTRICTED)
|
||||
# pending admin approval, so an un-provisioned Google user gets zero access.
|
||||
FORGEJO__oauth2_client__ACCOUNT_LINKING: "auto"
|
||||
FORGEJO__oauth2_client__ENABLE_AUTO_REGISTRATION: "true"
|
||||
FORGEJO__service__ALLOW_ONLY_EXTERNAL_REGISTRATION: "true"
|
||||
FORGEJO__service__ENABLE_PASSWORD_SIGNIN_FORM: "false"
|
||||
FORGEJO__service__REGISTER_MANUAL_CONFIRM: "true"
|
||||
FORGEJO__service__DEFAULT_USER_IS_RESTRICTED: "true"
|
||||
# Outbound mail via prod's Postfix null client over the WireGuard mesh
|
||||
# (10.10.0.1:25 — mynetworks permits beta); self-signed cert on :25, so
|
||||
# trust it. Send-only; used for admin/approval and notification mail.
|
||||
FORGEJO__mailer__ENABLED: "true"
|
||||
FORGEJO__mailer__PROTOCOL: "smtp"
|
||||
FORGEJO__mailer__SMTP_ADDR: "10.10.0.1"
|
||||
FORGEJO__mailer__SMTP_PORT: "25"
|
||||
FORGEJO__mailer__FORCE_TRUST_SERVER_CERT: "true"
|
||||
FORGEJO__mailer__FROM: "Thermograph Git <no-reply@thermograph.org>"
|
||||
# Gitea/Forgejo's app.ini env-mapping honors a __FILE suffix to read a
|
||||
# value from a file instead of the literal env var — same convention as
|
||||
# the official Postgres image's POSTGRES_PASSWORD_FILE above.
|
||||
FORGEJO__database__PASSWD__FILE: /run/secrets/forgejo_db_password
|
||||
secrets:
|
||||
- source: forgejo_db_password
|
||||
target: forgejo_db_password
|
||||
volumes:
|
||||
- forgejo_data:/var/lib/gitea
|
||||
networks: [forgejo_net]
|
||||
ports:
|
||||
# SSH for git@ clones — published on whichever node the task lands on
|
||||
# (pinned to beta by the placement constraint below, so effectively
|
||||
# always beta's public IP, port 2222).
|
||||
- target: 2222
|
||||
published: 2222
|
||||
protocol: tcp
|
||||
mode: host
|
||||
# Web UI/API. Swarm's port schema has no host_ip scoping, so this binds
|
||||
# 0.0.0.0:3080 on beta — NOT actually localhost-only by itself. A
|
||||
# DOCKER-USER iptables rule (see deploy/forgejo/README.md) is what
|
||||
# actually restricts it, since Docker's own iptables rules bypass ufw
|
||||
# for published ports. beta's Caddy reverse-proxies to 127.0.0.1:3080.
|
||||
- target: 3000
|
||||
published: 3080
|
||||
protocol: tcp
|
||||
mode: host
|
||||
deploy:
|
||||
placement:
|
||||
constraints: [node.labels.role == forge]
|
||||
restart_policy:
|
||||
condition: on-failure
|
||||
|
||||
networks:
|
||||
forgejo_net:
|
||||
driver: overlay
|
||||
attachable: false
|
||||
|
||||
volumes:
|
||||
forgejo_db:
|
||||
forgejo_data:
|
||||
|
||||
secrets:
|
||||
forgejo_db_password:
|
||||
external: true
|
||||
104
infra/deploy/forgejo/register-lan-runner.sh
Executable file
104
infra/deploy/forgejo/register-lan-runner.sh
Executable file
|
|
@ -0,0 +1,104 @@
|
|||
#!/usr/bin/env bash
|
||||
# Re-points the existing LAN dev self-hosted runner from GitHub Actions to
|
||||
# Forgejo Actions. Run on the SAME machine that already runs the GitHub
|
||||
# runner (see DEPLOY-DEV.md) — this replaces that runner, it doesn't add a
|
||||
# second one. Sudo-free, systemd --user, same pattern as the app service.
|
||||
#
|
||||
# This is THE runner (docs/runbooks/implementation-handoff.md Track B step 5
|
||||
# — canonical placement is the desktop, not a Swarm-hosted container), so it
|
||||
# registers with BOTH labels the workflows need, where one runner used to be
|
||||
# two: general CI/build/deploy.yml jobs (`docker`, containerized via this
|
||||
# machine's own already-installed Docker — no Docker-in-Docker sidecar needed,
|
||||
# unlike the Swarm-hosted design this replaces, since a real host with a real
|
||||
# Docker install needs no such indirection) and the LAN-deploy job
|
||||
# (`thermograph-lan`, bare/host-native — it writes to ~/thermograph-dev and
|
||||
# restarts a systemd --user service, which only works running directly on the
|
||||
# host, not inside a container).
|
||||
#
|
||||
# bash deploy/forgejo/register-lan-runner.sh <forgejo_url> <registration_token>
|
||||
#
|
||||
# Get <registration_token> from the Forgejo web UI:
|
||||
# repo -> Settings -> Actions -> Runners -> Create new Runner
|
||||
# (or an org/instance-level runner page, if you want it to serve more than
|
||||
# this one repo — same as the GitHub runner did).
|
||||
set -euo pipefail
|
||||
|
||||
FORGEJO_URL="${1:?usage: $0 <forgejo_url> <registration_token>}"
|
||||
TOKEN="${2:?}"
|
||||
RUNNER_DIR="${RUNNER_DIR:-$HOME/forgejo-runner}"
|
||||
LABELS="${LABELS:-docker:docker://node:20-bookworm,thermograph-lan}"
|
||||
|
||||
echo "==> Stopping and disabling the old GitHub Actions runner service, if present"
|
||||
systemctl --user stop github-actions-runner 2>/dev/null || true
|
||||
systemctl --user disable github-actions-runner 2>/dev/null || true
|
||||
|
||||
if ! id -nG "$USER" 2>/dev/null | grep -qw docker; then
|
||||
echo "WARNING: $USER is not in the 'docker' group — the 'docker:docker://...'"
|
||||
echo "labeled jobs (general CI) will fail to start a container until this is"
|
||||
echo "fixed: sudo usermod -aG docker $USER && (log out and back in)."
|
||||
fi
|
||||
|
||||
echo "==> Installing forgejo-runner into $RUNNER_DIR"
|
||||
mkdir -p "$RUNNER_DIR"
|
||||
cd "$RUNNER_DIR"
|
||||
if [ ! -x ./forgejo-runner ]; then
|
||||
ARCH="$(uname -m)"
|
||||
case "$ARCH" in
|
||||
x86_64) BIN_ARCH=amd64 ;;
|
||||
aarch64) BIN_ARCH=arm64 ;;
|
||||
*) echo "Unsupported arch: $ARCH — download the right binary by hand from" \
|
||||
"https://code.forgejo.org/forgejo/runner/releases" >&2; exit 1 ;;
|
||||
esac
|
||||
VER="${FORGEJO_RUNNER_VERSION:-6.3.1}"
|
||||
curl -fsSL -o forgejo-runner \
|
||||
"https://code.forgejo.org/forgejo/runner/releases/download/v${VER}/forgejo-runner-${VER}-linux-${BIN_ARCH}"
|
||||
chmod +x forgejo-runner
|
||||
fi
|
||||
|
||||
echo "==> Registering with $FORGEJO_URL (label: $LABELS)"
|
||||
./forgejo-runner register --no-interactive \
|
||||
--instance "$FORGEJO_URL" \
|
||||
--token "$TOKEN" \
|
||||
--name "thermograph-lan-$(hostname -s)" \
|
||||
--labels "$LABELS"
|
||||
|
||||
echo "==> Runner config (docker.sock automount, so docker:-labeled jobs like"
|
||||
echo " build-push.yml can actually run 'docker build/push' -- generated"
|
||||
echo " config only overridden on the one setting that matters here)"
|
||||
./forgejo-runner generate-config \
|
||||
| sed 's/docker_host: "-"/docker_host: "automount"/' \
|
||||
> "${RUNNER_DIR}/config.yaml"
|
||||
|
||||
echo "==> systemd --user unit"
|
||||
mkdir -p "$HOME/.config/systemd/user"
|
||||
cat > "$HOME/.config/systemd/user/forgejo-runner.service" <<EOF
|
||||
[Unit]
|
||||
Description=Forgejo Actions runner (docker + thermograph-lan)
|
||||
After=network-online.target
|
||||
|
||||
[Service]
|
||||
WorkingDirectory=${RUNNER_DIR}
|
||||
ExecStart=${RUNNER_DIR}/forgejo-runner daemon -c config.yaml
|
||||
Restart=on-failure
|
||||
RestartSec=5
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
EOF
|
||||
|
||||
systemctl --user daemon-reload
|
||||
systemctl --user enable --now forgejo-runner
|
||||
loginctl enable-linger "$USER" 2>/dev/null || true
|
||||
|
||||
cat <<EOF
|
||||
|
||||
Done. The runner now serves Forgejo, not GitHub — one runner, both labels
|
||||
($LABELS), replacing what used to be a separate Swarm-hosted runner for
|
||||
general CI plus this machine's own GitHub runner for LAN deploys.
|
||||
status: systemctl --user status forgejo-runner
|
||||
logs: journalctl --user -u forgejo-runner -f
|
||||
restart: systemctl --user restart forgejo-runner
|
||||
|
||||
The old github-actions-runner unit was stopped and disabled but not deleted —
|
||||
remove ~/actions-runner by hand once you've confirmed Forgejo deploys work.
|
||||
EOF
|
||||
118
infra/deploy/migrate-db.py
Normal file
118
infra/deploy/migrate-db.py
Normal file
|
|
@ -0,0 +1,118 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Apply pending SQL migrations to the accounts database.
|
||||
|
||||
This project has no Alembic. SQLAlchemy ``create_all()`` only creates *missing
|
||||
tables*, so column additions to the long-lived accounts DB are done with the
|
||||
hand-written ``.sql`` files in ``deploy/migrations/``. This runner applies the
|
||||
ones a given database hasn't seen yet and records them in a ``schema_migrations``
|
||||
table, so it is safe (and cheap) to run on every deploy.
|
||||
|
||||
It resolves the database exactly as the app does (``db.py``):
|
||||
``THERMOGRAPH_ACCOUNTS_DB`` if set, otherwise ``<repo>/data/accounts.sqlite``.
|
||||
|
||||
Fresh database: the app's ``create_all()`` builds the *current* schema on first
|
||||
start, so there is nothing to ALTER. This runner detects that (no file yet, or no
|
||||
``user`` table) and records the existing migrations as a baseline instead of
|
||||
replaying them, so they never run against a schema that already has the columns.
|
||||
|
||||
Usage (wired into deploy.sh / deploy-dev.sh, but runnable by hand too):
|
||||
|
||||
/opt/thermograph/.venv/bin/python deploy/migrate-db.py
|
||||
"""
|
||||
import glob
|
||||
import os
|
||||
import shutil
|
||||
import sqlite3
|
||||
import sys
|
||||
import time
|
||||
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
REPO = os.path.dirname(HERE)
|
||||
MIGRATIONS_DIR = os.path.join(HERE, "migrations")
|
||||
DB_PATH = os.environ.get("THERMOGRAPH_ACCOUNTS_DB") or os.path.join(REPO, "data", "accounts.sqlite")
|
||||
|
||||
|
||||
def _applied(conn: sqlite3.Connection) -> set[str]:
|
||||
conn.execute(
|
||||
"CREATE TABLE IF NOT EXISTS schema_migrations ("
|
||||
"name TEXT PRIMARY KEY, applied_at TEXT NOT NULL)"
|
||||
)
|
||||
return {row[0] for row in conn.execute("SELECT name FROM schema_migrations")}
|
||||
|
||||
|
||||
def _record(conn: sqlite3.Connection, name: str) -> None:
|
||||
conn.execute(
|
||||
"INSERT OR IGNORE INTO schema_migrations(name, applied_at) "
|
||||
"VALUES (?, datetime('now'))",
|
||||
(name,),
|
||||
)
|
||||
|
||||
|
||||
def _has_user_table(conn: sqlite3.Connection) -> bool:
|
||||
row = conn.execute(
|
||||
"SELECT 1 FROM sqlite_master WHERE type='table' AND name='user'"
|
||||
).fetchone()
|
||||
return row is not None
|
||||
|
||||
|
||||
def main() -> None:
|
||||
migrations = sorted(glob.glob(os.path.join(MIGRATIONS_DIR, "*.sql")))
|
||||
if not migrations:
|
||||
print("migrate-db: no migrations found")
|
||||
return
|
||||
|
||||
os.makedirs(os.path.dirname(DB_PATH), exist_ok=True)
|
||||
fresh = not os.path.exists(DB_PATH)
|
||||
conn = sqlite3.connect(DB_PATH)
|
||||
try:
|
||||
applied = _applied(conn)
|
||||
pending = [m for m in migrations if os.path.basename(m) not in applied]
|
||||
if not pending:
|
||||
print(f"migrate-db: up to date ({len(applied)} applied) -> {DB_PATH}")
|
||||
return
|
||||
|
||||
# Fresh DB (or one whose accounts schema hasn't been built yet): the app's
|
||||
# create_all() will produce the current schema on next start, so record the
|
||||
# existing migrations as a baseline rather than ALTERing a table that isn't
|
||||
# there. Only genuinely older databases need the ALTERs replayed.
|
||||
if fresh or not _has_user_table(conn):
|
||||
for m in pending:
|
||||
_record(conn, os.path.basename(m))
|
||||
conn.commit()
|
||||
print(f"migrate-db: fresh database, baselined {len(pending)} migration(s) -> {DB_PATH}")
|
||||
return
|
||||
|
||||
# These are the authoritative, non-regenerable accounts. Back the DB up
|
||||
# before touching it, next to the DB with a timestamp.
|
||||
backup = f"{DB_PATH}.bak-{time.strftime('%Y%m%d-%H%M%S')}"
|
||||
shutil.copy2(DB_PATH, backup)
|
||||
print(f"migrate-db: backed up -> {backup}")
|
||||
|
||||
for m in pending:
|
||||
name = os.path.basename(m)
|
||||
with open(m, encoding="utf-8") as f:
|
||||
sql = f.read()
|
||||
try:
|
||||
conn.executescript(sql)
|
||||
except sqlite3.OperationalError as e:
|
||||
msg = str(e).lower()
|
||||
# Idempotent: a column/index this migration adds already exists
|
||||
# (applied by hand before, or created fresh from the model). Treat
|
||||
# as already-applied and record it; re-raise anything unexpected.
|
||||
if "duplicate column" in msg or "already exists" in msg:
|
||||
print(f"migrate-db: {name} already present in schema, recording")
|
||||
else:
|
||||
conn.rollback()
|
||||
print(f"migrate-db: FAILED on {name}: {e}", file=sys.stderr)
|
||||
raise
|
||||
_record(conn, name)
|
||||
conn.commit()
|
||||
print(f"migrate-db: applied {name}")
|
||||
|
||||
print(f"migrate-db: done ({len(pending)} newly recorded) -> {DB_PATH}")
|
||||
finally:
|
||||
conn.close()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
17
infra/deploy/migrations/001-user-discord-id.sql
Normal file
17
infra/deploy/migrations/001-user-discord-id.sql
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
-- Adds User.discord_id for Discord account linking (backend/discord_link.py).
|
||||
--
|
||||
-- This project has no Alembic: SQLAlchemy create_all() only creates *missing
|
||||
-- tables*, so it will NOT add this column to an existing accounts DB. deploy.sh
|
||||
-- and deploy-dev.sh run deploy/migrate-db.py, which applies the pending files
|
||||
-- here on every deploy (tracked in a schema_migrations table, backing the DB up
|
||||
-- first). A fresh DB gets the column + uniqueness from the model and is baselined
|
||||
-- rather than ALTERed. To apply by hand instead (back the DB up first):
|
||||
--
|
||||
-- sqlite3 /opt/thermograph/data/accounts.sqlite < deploy/migrations/001-user-discord-id.sql
|
||||
|
||||
ALTER TABLE user ADD COLUMN discord_id VARCHAR(32);
|
||||
|
||||
-- SQLite can't add a UNIQUE column via ALTER, so enforce it with a partial unique
|
||||
-- index (NULLs are allowed to repeat; a real id links to at most one account).
|
||||
CREATE UNIQUE INDEX IF NOT EXISTS ix_user_discord_id
|
||||
ON user (discord_id) WHERE discord_id IS NOT NULL;
|
||||
9
infra/deploy/migrations/002-user-discord-dm.sql
Normal file
9
infra/deploy/migrations/002-user-discord-dm.sql
Normal file
|
|
@ -0,0 +1,9 @@
|
|||
-- Adds User.discord_dm for Discord DM alert opt-in (notify.py DM channel).
|
||||
-- Same no-Alembic caveat as 001: create_all won't add a column to an existing
|
||||
-- accounts DB. deploy/migrate-db.py applies this on every deploy; existing rows
|
||||
-- default to 0 (DMs off) until the user opts in by linking Discord. To apply by
|
||||
-- hand instead (back the DB up first):
|
||||
--
|
||||
-- sqlite3 /opt/thermograph/data/accounts.sqlite < deploy/migrations/002-user-discord-dm.sql
|
||||
|
||||
ALTER TABLE user ADD COLUMN discord_dm BOOLEAN NOT NULL DEFAULT 0;
|
||||
189
infra/deploy/openmeteo/README.md
Normal file
189
infra/deploy/openmeteo/README.md
Normal file
|
|
@ -0,0 +1,189 @@
|
|||
# Self-hosted Open-Meteo (ERA5 archive)
|
||||
|
||||
Operator runbook for running a private Open-Meteo instance that serves the
|
||||
ERA5 historical archive to Thermograph, with the `.om` data held in object
|
||||
storage and surfaced on the host through an rclone FUSE mount.
|
||||
|
||||
## 1. What this is and why
|
||||
|
||||
Thermograph reads daily historical weather from an ERA5 archive. Off the
|
||||
shelf that means the public Open-Meteo archive API, which is rate-limited and
|
||||
not something to lean on for a production workload. This overlay runs our own
|
||||
Open-Meteo instance instead:
|
||||
|
||||
- `open-meteo-api` serves `era5_seamless` locally (internal to the compose
|
||||
network). The app points at it via `THERMOGRAPH_ARCHIVE_URL`.
|
||||
- Two sync workers (`open-meteo-sync-land`, `open-meteo-sync-era5`) pull `.om`
|
||||
files from Open-Meteo's free AWS Open-Data bucket (no API key, no rate
|
||||
limit) and write them into the archive.
|
||||
|
||||
`era5_seamless` is a blend: 0.1° ERA5-Land for temperature, precipitation,
|
||||
humidity, and wind, plus 0.25° ERA5 for wind gusts (which ERA5-Land does not
|
||||
carry) and as the over-water fallback. The 0.1° resolution is a hard
|
||||
requirement for city-level accuracy.
|
||||
|
||||
The archive is ~1–1.5 TB of `.om`. That does not fit on the host's 400 GB
|
||||
disk, so it lives in an object-storage bucket and is mounted read/write via
|
||||
rclone. The host disk holds only the bounded rclone VFS cache, the app's own
|
||||
parquet cache, and Postgres — never a full copy. The app never reads object
|
||||
storage directly; it only talks to `open-meteo-api`, which reads the mount.
|
||||
|
||||
## 2. Object storage prerequisites
|
||||
|
||||
Provision a bucket of ~2 TB (holds the ~1–1.5 TB archive with headroom):
|
||||
|
||||
- **Co-located with the VPS** and **low- or no-egress** — e.g. Cloudflare R2,
|
||||
or same-provider object storage in the VPS's region.
|
||||
- Co-location and low egress matter because the mount serves per-request
|
||||
**range reads**: every archive query pulls byte ranges out of many `.om`
|
||||
files. Cross-region or metered egress turns each read into latency and cost.
|
||||
Keep the bucket next to the compute and on a plan that does not bill egress.
|
||||
|
||||
You'll need S3-compatible credentials (access key id + secret) and the
|
||||
bucket's S3 endpoint.
|
||||
|
||||
## 3. Host rclone mount setup
|
||||
|
||||
Install rclone:
|
||||
|
||||
```sh
|
||||
curl https://rclone.org/install.sh | sudo bash
|
||||
```
|
||||
|
||||
Create `/etc/rclone/rclone.conf` with an S3-compatible remote. Use real
|
||||
values for your provider; **never commit real secrets**:
|
||||
|
||||
```ini
|
||||
[om-archive]
|
||||
type = s3
|
||||
provider = Cloudflare
|
||||
endpoint = https://<ACCOUNT_ID>.r2.cloudflarestorage.com
|
||||
access_key_id = REPLACE_WITH_ACCESS_KEY_ID
|
||||
secret_access_key = REPLACE_WITH_SECRET_ACCESS_KEY
|
||||
```
|
||||
|
||||
Install and enable the mount unit (see `rclone-mount.service.example`):
|
||||
|
||||
```sh
|
||||
sudo install -m0644 rclone-mount.service.example /etc/systemd/system/rclone-om.service
|
||||
# edit BUCKET_NAME in the unit first; see the unit's header comments
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl enable --now rclone-om
|
||||
```
|
||||
|
||||
Verify the mount:
|
||||
|
||||
```sh
|
||||
mountpoint -q /mnt/om-archive && echo mounted
|
||||
ls /mnt/om-archive
|
||||
```
|
||||
|
||||
The unit runs with `--vfs-cache-mode full` and a bounded
|
||||
`--vfs-cache-max-size` (e.g. `80G`). Full VFS cache mode keeps hot cells on
|
||||
local disk after first read so repeat range reads don't go back to the bucket,
|
||||
and the size cap keeps that cache inside the 400 GB disk budget by evicting
|
||||
cold data.
|
||||
|
||||
**Order Docker after the mount (reboots).** So the `restart: unless-stopped`
|
||||
containers never start against an empty mount point, make Docker wait for the
|
||||
mount (`rclone-om` is `Type=notify`, so this waits until the mount is actually
|
||||
ready). Terraform installs this automatically; for a manual setup:
|
||||
|
||||
```sh
|
||||
sudo install -d /etc/systemd/system/docker.service.d
|
||||
printf '[Unit]\nWants=rclone-om.service\nAfter=rclone-om.service\n' \
|
||||
| sudo tee /etc/systemd/system/docker.service.d/10-wait-rclone.conf
|
||||
sudo systemctl daemon-reload
|
||||
```
|
||||
|
||||
(If the mount ever drops and remounts *while* the containers are running, the
|
||||
existing bind won't see the new mount — restart the Open-Meteo containers to
|
||||
re-bind. The app stays safe either way: it rejects a short/empty archive and
|
||||
falls back to NASA rather than caching a gap. See `make om-up`.)
|
||||
|
||||
## 4. Point the overlay at the mount
|
||||
|
||||
`OM_DATA_DIR` is read from the environment at `docker compose` time; in prod
|
||||
it lives in `/etc/thermograph.env` (which the systemd unit sources). Set it to
|
||||
the mount:
|
||||
|
||||
```sh
|
||||
# /etc/thermograph.env
|
||||
OM_DATA_DIR=/mnt/om-archive
|
||||
```
|
||||
|
||||
All three services bind-mount `${OM_DATA_DIR}` to `/app/data`, so with this
|
||||
set the archive reads and writes go to object storage.
|
||||
|
||||
## 5. One-time backfill
|
||||
|
||||
The sync workers only maintain a rolling recent window. To populate full
|
||||
history, run the backfill once:
|
||||
|
||||
```sh
|
||||
make om-backfill
|
||||
```
|
||||
|
||||
This runs each dataset's `sync ... --past-days 17000` once via
|
||||
`docker compose run --rm --no-deps`. It writes the full ~1–1.5 TB of `.om`
|
||||
**to object storage**, is **hours-long**, and should be watched against the
|
||||
2 TB budget. Run it **before** flipping the app over — if the app is pointed
|
||||
at an empty instance it falls back to NASA POWER, so bring the archive up to
|
||||
full history first.
|
||||
|
||||
For a smoke test, shorten the window with `OM_BACKFILL_DAYS`:
|
||||
|
||||
```sh
|
||||
make om-backfill OM_BACKFILL_DAYS=30
|
||||
```
|
||||
|
||||
## 6. Bring the overlay up
|
||||
|
||||
```sh
|
||||
make om-up
|
||||
```
|
||||
|
||||
This is `docker compose -f docker-compose.yml -f docker-compose.openmeteo.yml
|
||||
up -d --build`. The overlay sets `THERMOGRAPH_ARCHIVE_URL=http://open-meteo-api:8080/v1/archive`
|
||||
automatically.
|
||||
|
||||
Smoke test the internal API and confirm every daily field is present and
|
||||
non-null. `open-meteo-api` has **no published host port**, so either run the
|
||||
curl from inside the compose network, or temporarily publish the port:
|
||||
|
||||
```sh
|
||||
docker compose -f docker-compose.yml -f docker-compose.openmeteo.yml \
|
||||
exec app curl "http://open-meteo-api:8080/v1/archive?latitude=47.6&longitude=-122.3&start_date=2026-06-01&end_date=2026-06-10&daily=temperature_2m_max,temperature_2m_min,precipitation_sum,wind_speed_10m_max,wind_gusts_10m_max,apparent_temperature_max,apparent_temperature_min,relative_humidity_2m_mean&models=era5_seamless&temperature_unit=fahrenheit&wind_speed_unit=mph&precipitation_unit=inch"
|
||||
```
|
||||
|
||||
If you've temporarily published the port instead, the same query works
|
||||
against `http://127.0.0.1:8080/...`. Check that each `daily` array is present
|
||||
and free of nulls across the date range.
|
||||
|
||||
## 7. Keeping current
|
||||
|
||||
The two sync workers re-sync `--past-days 14` every 1440 minutes (daily). If a
|
||||
worker stalls, the recent tail of history goes stale — the last couple of
|
||||
weeks stop updating. (The separate forecast path is unaffected; this only
|
||||
touches the historical archive.)
|
||||
|
||||
Check the workers:
|
||||
|
||||
```sh
|
||||
docker compose -f docker-compose.yml -f docker-compose.openmeteo.yml \
|
||||
logs open-meteo-sync-land
|
||||
docker compose -f docker-compose.yml -f docker-compose.openmeteo.yml \
|
||||
logs open-meteo-sync-era5
|
||||
```
|
||||
|
||||
## 8. Attribution
|
||||
|
||||
The ERA5 and ERA5-Land data is CC-BY-4.0 (Copernicus/ECMWF), sourced via
|
||||
Open-Meteo. The Open-Meteo software is AGPLv3. The app already surfaces this
|
||||
credit; keep it in place.
|
||||
|
||||
## 9. Not on dev/beta
|
||||
|
||||
This overlay runs only on the self-hosting host (prod). Dev and beta leave
|
||||
`THERMOGRAPH_ARCHIVE_URL` unset and use the public Open-Meteo archive API — do
|
||||
not bring the overlay up there.
|
||||
32
infra/deploy/openmeteo/rclone-mount.service.example
Normal file
32
infra/deploy/openmeteo/rclone-mount.service.example
Normal file
|
|
@ -0,0 +1,32 @@
|
|||
# rclone FUSE mount for the Open-Meteo (ERA5) archive bucket.
|
||||
#
|
||||
# Before installing:
|
||||
# - Replace BUCKET_NAME below with the object-storage bucket name.
|
||||
# - `--allow-other` requires `user_allow_other` to be set in /etc/fuse.conf.
|
||||
# - The `om-archive` remote must exist in /etc/rclone/rclone.conf (S3 remote).
|
||||
#
|
||||
# Install:
|
||||
# sudo install -m0644 rclone-mount.service.example /etc/systemd/system/rclone-om.service
|
||||
# sudo systemctl daemon-reload
|
||||
# sudo systemctl enable --now rclone-om
|
||||
|
||||
[Unit]
|
||||
Description=rclone mount for Open-Meteo ERA5 archive
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=notify
|
||||
ExecStartPre=/bin/mkdir -p /mnt/om-archive
|
||||
ExecStart=/usr/bin/rclone mount om-archive:BUCKET_NAME /mnt/om-archive \
|
||||
--config /etc/rclone/rclone.conf \
|
||||
--vfs-cache-mode full \
|
||||
--vfs-cache-max-size 80G \
|
||||
--dir-cache-time 12h \
|
||||
--allow-other \
|
||||
--umask 000
|
||||
ExecStop=/bin/fusermount -u /mnt/om-archive
|
||||
Restart=on-failure
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
90
infra/deploy/provision-agent-access.sh
Executable file
90
infra/deploy/provision-agent-access.sh
Executable file
|
|
@ -0,0 +1,90 @@
|
|||
#!/usr/bin/env bash
|
||||
# Provisions a dedicated, full-sudo login for an AI coding agent on this VPS.
|
||||
# Idempotent — safe to re-run (e.g. to rotate the key or re-apply hardening
|
||||
# after an OS upgrade). Run as root (or via an existing sudo-capable user):
|
||||
#
|
||||
# sudo bash /opt/thermograph/deploy/provision-agent-access.sh
|
||||
#
|
||||
# The agent's public key is embedded below (AGENT_PUBKEY) rather than passed
|
||||
# as an argument, so this file alone is the whole bootstrap — copy the current
|
||||
# one from the repo and run it; there's nothing else to fetch or fill in.
|
||||
#
|
||||
# What this does, and why it's a SEPARATE user rather than direct root login:
|
||||
# - A distinct login name ("agent") gives a clean "who did this" audit trail
|
||||
# that shared/root login destroys. It is still full-root-capable via
|
||||
# passwordless sudo — this is a full-trust grant, just an attributable and
|
||||
# independently revocable one.
|
||||
# - Revocation is one line: delete AGENT_PUBKEY's line from
|
||||
# /home/agent/.ssh/authorized_keys, or `deluser --remove-home agent` +
|
||||
# `rm /etc/sudoers.d/agent`. Your own access is untouched either way.
|
||||
#
|
||||
# This script does NOT touch the app, Docker, Postgres, or Terraform-managed
|
||||
# state — it only manages the "agent" login and sshd/auditd hardening.
|
||||
set -euo pipefail
|
||||
|
||||
# --- fill this in when rotating the key; the current value is the one
|
||||
# generated for this engagement --------------------------------------------
|
||||
AGENT_PUBKEY="${AGENT_PUBKEY:-ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAICUwrp/neW5UJVmsK9SRveghJ+ayX3r59qgkL9Zo0mqO claude-agent@thermograph-infra}"
|
||||
|
||||
AGENT_USER="${AGENT_USER:-agent}"
|
||||
|
||||
if [ "$(id -u)" -ne 0 ]; then
|
||||
echo "Run as root (sudo bash $0)." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "==> User: $AGENT_USER"
|
||||
if ! id "$AGENT_USER" >/dev/null 2>&1; then
|
||||
adduser --disabled-password --gecos "" "$AGENT_USER"
|
||||
fi
|
||||
|
||||
echo "==> Passwordless sudo (full, by design — see header comment)"
|
||||
echo "$AGENT_USER ALL=(ALL) NOPASSWD:ALL" > /etc/sudoers.d/"$AGENT_USER"
|
||||
chmod 440 /etc/sudoers.d/"$AGENT_USER"
|
||||
visudo -cf /etc/sudoers.d/"$AGENT_USER" # fail loudly on a syntax error rather than silently locking sudo
|
||||
|
||||
echo "==> Authorized key"
|
||||
install -d -m 700 -o "$AGENT_USER" -g "$AGENT_USER" /home/"$AGENT_USER"/.ssh
|
||||
# Replace wholesale rather than append, so re-running after a key rotation
|
||||
# doesn't accumulate stale keys.
|
||||
echo "$AGENT_PUBKEY" > /home/"$AGENT_USER"/.ssh/authorized_keys
|
||||
chown "$AGENT_USER":"$AGENT_USER" /home/"$AGENT_USER"/.ssh/authorized_keys
|
||||
chmod 600 /home/"$AGENT_USER"/.ssh/authorized_keys
|
||||
|
||||
echo "==> sshd hardening (fleet-wide, not just for this account)"
|
||||
SSHD_DROPIN=/etc/ssh/sshd_config.d/99-agent-hardening.conf
|
||||
cat > "$SSHD_DROPIN" <<'EOF'
|
||||
# Managed by deploy/provision-agent-access.sh — key-only login, no passwords.
|
||||
# Root itself still can't password in; PermitRootLogin prohibit-password
|
||||
# leaves direct root key-login available for a human who already has a key
|
||||
# there, without opening a password path for anyone.
|
||||
PasswordAuthentication no
|
||||
KbdInteractiveAuthentication no
|
||||
PermitRootLogin prohibit-password
|
||||
EOF
|
||||
sshd -t # validate before reload; a bad config here can lock out the whole box
|
||||
systemctl reload sshd 2>/dev/null || systemctl reload ssh 2>/dev/null || true
|
||||
|
||||
echo "==> auditd: log root-effective commands (survives shell history -c)"
|
||||
if ! command -v auditctl >/dev/null 2>&1; then
|
||||
apt-get update -y -q
|
||||
apt-get install -y -q auditd audispd-plugins
|
||||
fi
|
||||
AUDIT_RULES=/etc/audit/rules.d/agent-root-cmds.rules
|
||||
cat > "$AUDIT_RULES" <<'EOF'
|
||||
# Every root-effective exec, tagged for `ausearch -k agentcmd`.
|
||||
-a always,exit -F arch=b64 -S execve -F euid=0 -k agentcmd
|
||||
-a always,exit -F arch=b32 -S execve -F euid=0 -k agentcmd
|
||||
EOF
|
||||
augenrules --load 2>/dev/null || true
|
||||
systemctl enable --now auditd 2>/dev/null || true
|
||||
|
||||
echo
|
||||
echo "Done. Verify from your own machine (not this box):"
|
||||
echo " ssh -i <agent-private-key> $AGENT_USER@<this-host> sudo whoami # -> root"
|
||||
echo " ssh -o PasswordAuthentication=no <any-user>@<this-host> # password path confirmed dead"
|
||||
echo
|
||||
echo "Revoke any time: delete the key line in"
|
||||
echo " /home/$AGENT_USER/.ssh/authorized_keys"
|
||||
echo "or remove the account entirely:"
|
||||
echo " deluser --remove-home $AGENT_USER && rm -f /etc/sudoers.d/$AGENT_USER"
|
||||
30
infra/deploy/provision-dev-lan.sh
Executable file
30
infra/deploy/provision-dev-lan.sh
Executable file
|
|
@ -0,0 +1,30 @@
|
|||
#!/usr/bin/env bash
|
||||
# One-time bootstrap for the Thermograph LAN dev server on THIS machine.
|
||||
#
|
||||
# Sudo-free: the app runs as a Docker Compose stack (see deploy-dev.sh), and
|
||||
# `linger` keeps the runner/services running across logout/reboot. Re-runnable
|
||||
# (idempotent).
|
||||
#
|
||||
# bash deploy/provision-dev-lan.sh
|
||||
set -euo pipefail
|
||||
|
||||
APP_DIR="${APP_DIR:-$HOME/thermograph-dev}"
|
||||
REPO_URL="${REPO_URL:-http://10.10.0.2:3080/emi/thermograph-infra.git}"
|
||||
BRANCH="${BRANCH:-dev}"
|
||||
here="$(cd "$(dirname "$0")" && pwd)"
|
||||
|
||||
echo "==> Enabling linger so the service survives logout/reboot"
|
||||
loginctl enable-linger "$USER" \
|
||||
|| echo " (couldn't enable linger; the service still runs while you're logged in)"
|
||||
|
||||
echo "==> Cloning/refreshing $APP_DIR and starting the service"
|
||||
APP_DIR="$APP_DIR" REPO_URL="$REPO_URL" BRANCH="$BRANCH" bash "$here/deploy-dev.sh"
|
||||
|
||||
cat <<EOF
|
||||
|
||||
Done. The dev server runs as the 'thermograph-dev' systemd --user service.
|
||||
status: systemctl --user status thermograph-dev
|
||||
logs: journalctl --user -u thermograph-dev -f
|
||||
restart: systemctl --user restart thermograph-dev
|
||||
stop: systemctl --user stop thermograph-dev
|
||||
EOF
|
||||
172
infra/deploy/provision-mail.sh
Executable file
172
infra/deploy/provision-mail.sh
Executable file
|
|
@ -0,0 +1,172 @@
|
|||
#!/usr/bin/env bash
|
||||
# Outbound email for Thermograph — run once on the VPS, as root.
|
||||
#
|
||||
# Installs Postfix as a SEND-ONLY NULL CLIENT: it listens on 127.0.0.1:25 only,
|
||||
# accepts mail from this machine, and never receives mail from the internet.
|
||||
#
|
||||
# Why a local MTA instead of talking to a mail provider's API from Python:
|
||||
#
|
||||
# * The app's only mail config becomes "SMTP on localhost". Whether delivery
|
||||
# then goes direct to the recipient's MX or through a relay is a Postfix
|
||||
# setting — switchable without touching, redeploying, or retesting the app.
|
||||
# * Postfix queues and retries. A request handler hands the message over in
|
||||
# microseconds and returns; a slow or briefly-down upstream can't stall a
|
||||
# web request or lose a signup.
|
||||
# * No new Python dependency: stdlib smtplib talks to it (see backend/mailer.py).
|
||||
#
|
||||
# DELIVERABILITY — read before pointing this at real subscribers.
|
||||
#
|
||||
# Mail sent straight from a VPS IP is very often junked, regardless of Postfix
|
||||
# config, because the IP has no sending reputation. Two options:
|
||||
#
|
||||
# A. RELAY through a transactional provider (recommended for real mail).
|
||||
# Set RELAYHOST + RELAY_USER + RELAY_PASSWORD below. The provider handles
|
||||
# SPF/DKIM alignment and reputation; you keep the loopback-SMTP seam.
|
||||
#
|
||||
# B. DIRECT to MX (no third party). Then you must also set up, in DNS:
|
||||
# - SPF: TXT @ "v=spf1 a mx ip4:<VPS_IP> -all"
|
||||
# - DKIM: install opendkim, publish the public key as a TXT record
|
||||
# - DMARC: TXT _dmarc "v=DMARC1; p=none; rua=mailto:you@domain"
|
||||
# - PTR / reverse DNS on the VPS IP -> mail.thermograph.org
|
||||
# The PTR record is the one people forget, and its absence alone is enough
|
||||
# for Gmail and Outlook to junk everything you send.
|
||||
#
|
||||
# Usage:
|
||||
# sudo MAIL_DOMAIN=thermograph.org bash deploy/provision-mail.sh
|
||||
# sudo MAIL_DOMAIN=thermograph.org RELAYHOST='[smtp.provider.com]:587' \
|
||||
# RELAY_USER=apikey RELAY_PASSWORD=secret bash deploy/provision-mail.sh
|
||||
set -euo pipefail
|
||||
|
||||
MAIL_DOMAIN="${MAIL_DOMAIN:-thermograph.org}"
|
||||
MAIL_HOSTNAME="${MAIL_HOSTNAME:-mail.${MAIL_DOMAIN}}"
|
||||
RELAYHOST="${RELAYHOST:-}"
|
||||
RELAY_USER="${RELAY_USER:-}"
|
||||
RELAY_PASSWORD="${RELAY_PASSWORD:-}"
|
||||
|
||||
if [[ $EUID -ne 0 ]]; then
|
||||
echo "run as root (sudo)" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "==> installing postfix (non-interactive)"
|
||||
export DEBIAN_FRONTEND=noninteractive
|
||||
# Preseed so the installer doesn't open its curses dialog.
|
||||
debconf-set-selections <<EOF
|
||||
postfix postfix/main_mailer_type select Internet Site
|
||||
postfix postfix/mailname string ${MAIL_HOSTNAME}
|
||||
EOF
|
||||
apt-get update -qq
|
||||
apt-get install -y -qq postfix libsasl2-modules
|
||||
|
||||
echo "==> configuring send-only null client"
|
||||
postconf -e "myhostname = ${MAIL_HOSTNAME}"
|
||||
postconf -e "myorigin = ${MAIL_DOMAIN}"
|
||||
# Never listen on a public interface. This box sends only. The app runs in a
|
||||
# Docker container, so it can't reach the host's loopback — it hands mail to
|
||||
# Postfix over the compose bridge's gateway. So Postfix also listens on that
|
||||
# gateway and accepts mail from the bridge subnet (both pinned in
|
||||
# docker-compose.yml). Set DOCKER_MAIL_GATEWAY="" for a pure loopback-only null
|
||||
# client (app running natively on the host, not in a container).
|
||||
#
|
||||
# WHICH gateway depends on the host's deploy mode: plain compose (beta, LAN)
|
||||
# uses the pinned compose bridge (172.19.0.1/172.19.0.0/16, the defaults);
|
||||
# Swarm-stack mode (prod) uses the docker_gwbridge gateway instead --
|
||||
# overlay tasks have no compose-bridge gateway -- so prod is provisioned with
|
||||
# DOCKER_MAIL_GATEWAY=172.18.0.1 DOCKER_MAIL_SUBNET=172.18.0.0/16 (plus its
|
||||
# MESH_MAIL_* listener below). Do NOT list an address that doesn't exist on
|
||||
# the host: Postfix's master fails to bind and takes ALL listeners down --
|
||||
# exactly what happened when the compose bridge (172.19.0.1) vanished at the
|
||||
# stack cutover while still listed in inet_interfaces. Also note: postfix on
|
||||
# this distro is an umbrella unit; restart `postfix@-`, not `postfix`, for
|
||||
# inet_interfaces changes to take effect.
|
||||
DOCKER_MAIL_GATEWAY="${DOCKER_MAIL_GATEWAY-172.19.0.1}"
|
||||
DOCKER_MAIL_SUBNET="${DOCKER_MAIL_SUBNET-172.19.0.0/16}"
|
||||
# Optional WireGuard-mesh listener: other mesh nodes (e.g. beta's Forgejo, whose
|
||||
# mailer posts to 10.10.0.1:25 — see deploy/forgejo/docker-stack.yml) can relay
|
||||
# through this box. Prod runs with MESH_MAIL_LISTEN=10.10.0.1 and
|
||||
# MESH_MAIL_PEERS=10.10.0.2/32; both default OFF so a plain run stays a strict
|
||||
# null client. Without these, re-running this script on prod would silently drop
|
||||
# the mesh listener and break Forgejo's outbound mail — the live config was
|
||||
# originally hand-applied and this script is the source of truth for it now.
|
||||
MESH_MAIL_LISTEN="${MESH_MAIL_LISTEN-}"
|
||||
MESH_MAIL_PEERS="${MESH_MAIL_PEERS-}"
|
||||
postconf -e "inet_protocols = ipv4"
|
||||
listen="127.0.0.1"
|
||||
networks="127.0.0.0/8 [::1]/128"
|
||||
if [[ -n "$DOCKER_MAIL_GATEWAY" ]]; then
|
||||
listen="${listen}, ${DOCKER_MAIL_GATEWAY}"
|
||||
networks="${networks} ${DOCKER_MAIL_SUBNET}"
|
||||
# ufw is default-deny incoming; a container connecting to the host's gateway IP
|
||||
# hits the INPUT chain, so allow the bridge subnet to reach port 25.
|
||||
command -v ufw >/dev/null 2>&1 && \
|
||||
ufw allow from "${DOCKER_MAIL_SUBNET}" to any port 25 proto tcp \
|
||||
comment 'app container -> host Postfix' || true
|
||||
fi
|
||||
if [[ -n "$MESH_MAIL_LISTEN" ]]; then
|
||||
listen="${listen}, ${MESH_MAIL_LISTEN}"
|
||||
networks="${networks} ${MESH_MAIL_PEERS}"
|
||||
fi
|
||||
if [[ "$listen" == "127.0.0.1" ]]; then
|
||||
postconf -e "inet_interfaces = loopback-only"
|
||||
else
|
||||
postconf -e "inet_interfaces = ${listen}"
|
||||
fi
|
||||
postconf -e "mynetworks = ${networks}"
|
||||
# A null client delivers nothing locally; everything is relayed out.
|
||||
postconf -e "mydestination ="
|
||||
postconf -e "local_transport = error:local delivery is disabled"
|
||||
# Use TLS opportunistically when talking to the next hop.
|
||||
postconf -e "smtp_tls_security_level = may"
|
||||
postconf -e "smtp_tls_loglevel = 1"
|
||||
|
||||
if [[ -n "$RELAYHOST" ]]; then
|
||||
echo "==> configuring relay via ${RELAYHOST}"
|
||||
postconf -e "relayhost = ${RELAYHOST}"
|
||||
if [[ -n "$RELAY_USER" ]]; then
|
||||
postconf -e "smtp_sasl_auth_enable = yes"
|
||||
postconf -e "smtp_sasl_password_maps = hash:/etc/postfix/sasl_passwd"
|
||||
postconf -e "smtp_sasl_security_options = noanonymous"
|
||||
printf '%s %s:%s\n' "$RELAYHOST" "$RELAY_USER" "$RELAY_PASSWORD" \
|
||||
> /etc/postfix/sasl_passwd
|
||||
# The credential file must not be world-readable.
|
||||
chmod 600 /etc/postfix/sasl_passwd
|
||||
postmap /etc/postfix/sasl_passwd
|
||||
chmod 600 /etc/postfix/sasl_passwd.db
|
||||
fi
|
||||
else
|
||||
echo "==> no RELAYHOST set: delivering direct to MX"
|
||||
echo " remember SPF + DKIM + DMARC + PTR, or expect the spam folder"
|
||||
postconf -e "relayhost ="
|
||||
fi
|
||||
|
||||
systemctl enable postfix
|
||||
systemctl restart postfix
|
||||
|
||||
echo "==> verifying it listens on loopback only"
|
||||
ss -lntp | grep ':25 ' || true
|
||||
|
||||
cat <<'NOTE'
|
||||
|
||||
==> next steps
|
||||
|
||||
1. Point the app at it, in /etc/thermograph.env:
|
||||
|
||||
THERMOGRAPH_MAIL_BACKEND=smtp
|
||||
THERMOGRAPH_SMTP_HOST=127.0.0.1
|
||||
THERMOGRAPH_SMTP_PORT=25
|
||||
THERMOGRAPH_MAIL_FROM=Thermograph <no-reply@thermograph.org>
|
||||
|
||||
then: sudo systemctl restart thermograph
|
||||
|
||||
2. Send yourself a test message:
|
||||
|
||||
echo "test body" | mail -s "thermograph test" you@example.com
|
||||
# or, exercising the app's own path:
|
||||
# python -c "import sys; sys.path.insert(0,'/opt/thermograph/backend'); \
|
||||
# import mailer; print(mailer.send('you@example.com','t','body'))"
|
||||
|
||||
3. Watch it leave: journalctl -u postfix -f (queue: mailq)
|
||||
|
||||
4. Check placement with https://www.mail-tester.com — it scores SPF, DKIM,
|
||||
DMARC and rDNS in one shot and tells you exactly what's missing.
|
||||
NOTE
|
||||
45
infra/deploy/provision-secrets.sh
Executable file
45
infra/deploy/provision-secrets.sh
Executable file
|
|
@ -0,0 +1,45 @@
|
|||
#!/usr/bin/env bash
|
||||
# Provision a host to render secrets from the SOPS vault at deploy time. Idempotent;
|
||||
# run once per box (prod/beta) as a sudo-capable user.
|
||||
#
|
||||
# SECRETS_ENV=prod AGE_KEY_SRC=/path/to/age.key bash deploy/provision-secrets.sh
|
||||
#
|
||||
# Installs `sops` + `age`, then installs the age PRIVATE key at /etc/thermograph/age.key
|
||||
# (0400) and writes /etc/thermograph/secrets-env. After this, deploy.sh renders
|
||||
# /etc/thermograph.env from deploy/secrets/*.yaml. See deploy/secrets/README.md.
|
||||
set -euo pipefail
|
||||
|
||||
SECRETS_ENV="${SECRETS_ENV:?set SECRETS_ENV=prod|beta}"
|
||||
AGE_KEY_SRC="${AGE_KEY_SRC:?set AGE_KEY_SRC=/path/to/the/age/private/key}"
|
||||
SOPS_VERSION="${SOPS_VERSION:-v3.13.2}"
|
||||
AGE_VERSION="${AGE_VERSION:-v1.3.1}"
|
||||
BIN="${BIN:-/usr/local/bin}"
|
||||
ARCH="$(uname -m)"; case "$ARCH" in x86_64) ARCH=amd64;; aarch64|arm64) ARCH=arm64;; esac
|
||||
|
||||
install_sops() {
|
||||
command -v sops >/dev/null 2>&1 && { echo "==> sops already installed"; return; }
|
||||
echo "==> Installing sops ${SOPS_VERSION}"
|
||||
sudo curl -fsSL "https://github.com/getsops/sops/releases/download/${SOPS_VERSION}/sops-${SOPS_VERSION}.linux.${ARCH}" -o "${BIN}/sops"
|
||||
sudo chmod +x "${BIN}/sops"
|
||||
}
|
||||
|
||||
install_age() {
|
||||
command -v age >/dev/null 2>&1 && { echo "==> age already installed"; return; }
|
||||
echo "==> Installing age ${AGE_VERSION}"
|
||||
tmp="$(mktemp -d)"
|
||||
curl -fsSL "https://github.com/FiloSottile/age/releases/download/${AGE_VERSION}/age-${AGE_VERSION}-linux-${ARCH}.tar.gz" | tar -xz -C "$tmp"
|
||||
sudo install -m 0755 "$tmp/age/age" "$tmp/age/age-keygen" "$BIN/"
|
||||
rm -rf "$tmp"
|
||||
}
|
||||
|
||||
install_sops
|
||||
install_age
|
||||
|
||||
echo "==> Installing age key -> /etc/thermograph/age.key (0400) and marker (${SECRETS_ENV})"
|
||||
sudo mkdir -p /etc/thermograph
|
||||
sudo install -m 0400 "$AGE_KEY_SRC" /etc/thermograph/age.key
|
||||
printf '%s\n' "$SECRETS_ENV" | sudo tee /etc/thermograph/secrets-env >/dev/null
|
||||
sudo chmod 0644 /etc/thermograph/secrets-env
|
||||
|
||||
echo "==> Done. Verify: sops --version && ls -l /etc/thermograph/"
|
||||
echo " Next: seed deploy/secrets/*.yaml, dry-run render, then deploy (see README)."
|
||||
80
infra/deploy/render-secrets.sh
Executable file
80
infra/deploy/render-secrets.sh
Executable file
|
|
@ -0,0 +1,80 @@
|
|||
#!/usr/bin/env bash
|
||||
# Render /etc/thermograph.env from the SOPS-encrypted source of truth
|
||||
# (deploy/secrets/<env>.yaml, committed encrypted) — sourced by deploy.sh and
|
||||
# deploy-dev.sh, not run directly.
|
||||
#
|
||||
# The encrypted files are the single source of truth; /etc/thermograph.env becomes a
|
||||
# rendered artifact that the existing seams (docker-compose `env_file:`, the systemd
|
||||
# unit's EnvironmentFile, and entrypoint.sh's /run/secrets shim) keep consuming
|
||||
# unchanged. Common secrets + the per-host overrides are concatenated with the host
|
||||
# file LAST, so a host value wins (env_file and `source` both take the last
|
||||
# occurrence of a duplicate key).
|
||||
#
|
||||
# Presence-detected: a host is "configured for SOPS" only when it has an age key at
|
||||
# /etc/thermograph/age.key AND an environment name in /etc/thermograph/secrets-env
|
||||
# (prod|beta|dev). A host without them keeps whatever /etc/thermograph.env it already
|
||||
# has — so merging this is a safe no-op until a host is deliberately migrated. See
|
||||
# deploy/secrets/README.md.
|
||||
render_thermograph_secrets() {
|
||||
local repo="${1:-.}" # repo root holding deploy/secrets
|
||||
local key="${THERMOGRAPH_AGE_KEY:-/etc/thermograph/age.key}"
|
||||
local marker="${THERMOGRAPH_SECRETS_ENV_FILE:-/etc/thermograph/secrets-env}"
|
||||
local env_name
|
||||
env_name=$(cat "$marker" 2>/dev/null || true)
|
||||
|
||||
# The per-host file is required; common.yaml is optional so the initial cutover can
|
||||
# seed each host as an exact copy of its live env (byte-identical render, no value
|
||||
# changes) and factor out shared secrets into common.yaml later.
|
||||
if [ -z "$env_name" ] || [ ! -f "$key" ] \
|
||||
|| [ ! -f "$repo/deploy/secrets/${env_name}.yaml" ]; then
|
||||
echo "==> SOPS secrets not configured here; using existing /etc/thermograph.env"
|
||||
return 0
|
||||
fi
|
||||
if ! command -v sops >/dev/null 2>&1; then
|
||||
echo "!! sops not installed but this host is configured for SOPS secrets" >&2
|
||||
return 1
|
||||
fi
|
||||
|
||||
echo "==> Rendering /etc/thermograph.env from deploy/secrets (common + ${env_name})"
|
||||
# The age private key is root-owned (0400). Read it directly if we can, else via
|
||||
# sudo into SOPS_AGE_KEY — so the key never has to be readable by the deploy user.
|
||||
# (The render needs sudo to write /etc/thermograph.env below anyway.)
|
||||
local key_env=()
|
||||
if [ -r "$key" ]; then
|
||||
key_env=("SOPS_AGE_KEY_FILE=$key")
|
||||
else
|
||||
local keymat; keymat=$(sudo cat "$key" 2>/dev/null | grep '^AGE-SECRET-KEY-' || true)
|
||||
[ -n "$keymat" ] || { echo "!! cannot read age key at $key (need sudo)" >&2; return 1; }
|
||||
key_env=("SOPS_AGE_KEY=$keymat")
|
||||
fi
|
||||
local tmp; tmp=$(mktemp)
|
||||
: > "$tmp"
|
||||
# set -e in the caller makes a decrypt failure fatal here (no partial env). common
|
||||
# first, host second, so a host value overrides a shared one (last-wins).
|
||||
if [ -f "$repo/deploy/secrets/common.yaml" ]; then
|
||||
env "${key_env[@]}" sops -d --input-type yaml --output-type dotenv \
|
||||
"$repo/deploy/secrets/common.yaml" >> "$tmp"
|
||||
fi
|
||||
env "${key_env[@]}" sops -d --input-type yaml --output-type dotenv \
|
||||
"$repo/deploy/secrets/${env_name}.yaml" >> "$tmp"
|
||||
|
||||
# Write /etc/thermograph.env. Prefer an in-place write when the existing file is
|
||||
# writable by us (e.g. a group-writable 0660 root:<deploygroup> on a box whose CI
|
||||
# deploy user isn't root and has no broad sudo — beta's `deploy`), since that needs
|
||||
# only file write, not /etc dir write or sudo. Else install; else sudo install (a
|
||||
# root/agent deploy) -- and on that sudo path CHOWN the result to the invoking
|
||||
# deploy user (-o/-g $(id -un/-gn)), or the file lands root:root 0640 and the very
|
||||
# next line of deploy.sh (`. /etc/thermograph.env` as that non-root user) can't read
|
||||
# it, so POSTGRES_PASSWORD never enters the env and `docker compose` dies on
|
||||
# interpolation. Fail loudly rather than deploy against stale secrets.
|
||||
if [ -f /etc/thermograph.env ] && [ -w /etc/thermograph.env ]; then
|
||||
cat "$tmp" > /etc/thermograph.env
|
||||
elif install -m 0640 "$tmp" /etc/thermograph.env 2>/dev/null; then :
|
||||
elif sudo install -m 0640 -o "$(id -un)" -g "$(id -gn)" "$tmp" /etc/thermograph.env 2>/dev/null; then :
|
||||
else
|
||||
rm -f "$tmp"
|
||||
echo "!! cannot write /etc/thermograph.env (need file write access or passwordless sudo)" >&2
|
||||
return 1
|
||||
fi
|
||||
rm -f "$tmp"
|
||||
}
|
||||
107
infra/deploy/secrets/README.md
Normal file
107
infra/deploy/secrets/README.md
Normal file
|
|
@ -0,0 +1,107 @@
|
|||
# Secrets — the git-native vault (SOPS + age)
|
||||
|
||||
The **single source of truth** for Thermograph's secrets is the set of
|
||||
SOPS-encrypted YAML files in this directory. They're committed to the repo
|
||||
**encrypted** (values only — keys stay readable so diffs are meaningful) and
|
||||
**rendered into `/etc/thermograph.env` at deploy time** by
|
||||
[`deploy/render-secrets.sh`](../render-secrets.sh), which the prod/beta
|
||||
`deploy.sh` invokes.
|
||||
|
||||
Cycling a key is: **edit → commit → deploy**. No SSHing in to hand-edit a root file,
|
||||
no per-host duplication.
|
||||
|
||||
## Files
|
||||
|
||||
| File | Holds | Encrypted? |
|
||||
|------|-------|------------|
|
||||
| `common.yaml` | Secrets shared by every host — Discord tokens/secret, VAPID keypair, `THERMOGRAPH_AUTH_SECRET`, metrics token, indexnow key, SMTP creds | ✅ |
|
||||
| `prod.yaml` | Prod-only overrides — `POSTGRES_PASSWORD`, `REGISTRY_TOKEN` | ✅ |
|
||||
| `beta.yaml` | Beta-only overrides | ✅ |
|
||||
| `example.yaml` | Format reference / CI fixture (fake values) | ✅ |
|
||||
| `../../.sops.yaml` | Which age recipient files are encrypted to (plaintext config) | — |
|
||||
|
||||
At deploy the renderer concatenates `common.yaml` then `<env>.yaml`, so a **host
|
||||
value wins** (env_file / `source` take the last occurrence of a duplicate key).
|
||||
`<env>` comes from `/etc/thermograph/secrets-env` on the box (`prod`/`beta`/`dev`).
|
||||
`dev` has no real secrets and is intentionally not wired in — see `deploy-dev.sh`.
|
||||
|
||||
## The age key (the one thing that matters)
|
||||
|
||||
- The **private key is the single recovery root.** Lose it and every secret here is
|
||||
unrecoverable; leak it and every secret is compromised.
|
||||
- It is **never in the repo.** It lives at:
|
||||
- `~/.config/sops/age/keys.txt` on the operator's machine (for editing), and
|
||||
- `/etc/thermograph/age.key` (`0400`) on each host that renders at deploy time.
|
||||
- **Back it up** in the password manager. The public recipient is in `.sops.yaml`.
|
||||
|
||||
## Prerequisites (once per machine)
|
||||
|
||||
Install `sops` + `age` (single static binaries). On a host, use
|
||||
[`deploy/provision-secrets.sh`](../provision-secrets.sh); on your laptop:
|
||||
|
||||
```sh
|
||||
# see the pinned URLs in provision-secrets.sh, or:
|
||||
go install github.com/getsops/sops/v3/cmd/sops@latest # if you have Go
|
||||
```
|
||||
|
||||
## Everyday: rotate a key
|
||||
|
||||
```sh
|
||||
sops deploy/secrets/common.yaml # opens DECRYPTED in $EDITOR; re-encrypts on save
|
||||
git commit -am "rotate <thing>" && git push forgejo
|
||||
# beta auto-deploys on push to main. prod (no CI) — one command:
|
||||
ssh agent@169.58.46.181 'cd /opt/thermograph && git pull && deploy/deploy.sh'
|
||||
```
|
||||
|
||||
A host-specific value (e.g. `POSTGRES_PASSWORD`) is the same, editing `prod.yaml` /
|
||||
`beta.yaml` instead. **`POSTGRES_PASSWORD` is special**: the database only reads it on
|
||||
a *fresh* volume, so also `ALTER ROLE thermograph PASSWORD '…'` inside the running
|
||||
Postgres to match — the env change alone won't re-key an initialized DB.
|
||||
|
||||
## Add a new secret
|
||||
|
||||
1. `sops deploy/secrets/common.yaml`, add `THERMOGRAPH_NEW_THING: value`.
|
||||
2. Add the app reader (`os.environ.get("THERMOGRAPH_NEW_THING")`).
|
||||
3. Commit + deploy. It flows into `/etc/thermograph.env` automatically.
|
||||
|
||||
## Rotate the age identity itself (re-key)
|
||||
|
||||
```sh
|
||||
age-keygen -o new.key # new identity
|
||||
# add the new PUBLIC recipient alongside the old in .sops.yaml, then:
|
||||
sops updatekeys deploy/secrets/*.yaml # re-encrypt to BOTH keys
|
||||
# distribute new.key to /etc/thermograph/age.key on every host + ~/.config/sops/age,
|
||||
# deploy once so every host can decrypt with the new key, then remove the OLD
|
||||
# recipient from .sops.yaml and:
|
||||
sops updatekeys deploy/secrets/*.yaml # drop the old key
|
||||
git commit -am "rotate age identity" && push + deploy
|
||||
```
|
||||
|
||||
## First-time bootstrap (seed from the live boxes)
|
||||
|
||||
Order matters — values must match the live env exactly so `AUTH_SECRET` / VAPID /
|
||||
`POSTGRES_PASSWORD` don't rotate unintentionally:
|
||||
|
||||
1. Install `sops`+`age` and drop `/etc/thermograph/age.key` (0400) +
|
||||
`/etc/thermograph/secrets-env` on prod & beta (`provision-secrets.sh`).
|
||||
2. Seed each host's file from its live env with the helper (pulls over SSH, encrypts,
|
||||
and verifies the render round-trips — run it on your own machine, it reads live
|
||||
secrets):
|
||||
```sh
|
||||
deploy/secrets/seed-from-live.sh prod agent@169.58.46.181
|
||||
deploy/secrets/seed-from-live.sh beta agent@75.119.132.91
|
||||
```
|
||||
That writes exact per-host copies (`prod.yaml`/`beta.yaml`); factor shared values
|
||||
into `common.yaml` later. Commit.
|
||||
3. **Dry-run:** on the box, render to a temp file and `diff` against the live
|
||||
`/etc/thermograph.env` — must be byte-identical before cutover.
|
||||
4. Merge the `deploy.sh` wiring; run one deploy; confirm the app is healthy.
|
||||
|
||||
## Safety
|
||||
|
||||
- CI (`.forgejo/workflows/secrets-guard.yml`) fails if any `*.yaml` here is committed
|
||||
unencrypted.
|
||||
- Rendering is **presence-detected**: a host without the age key + marker keeps its
|
||||
existing `/etc/thermograph.env`, so nothing breaks before a host is migrated.
|
||||
- Rendering happens **on the target box**, so CI/runners never see the age key or the
|
||||
plaintext.
|
||||
33
infra/deploy/secrets/beta.yaml
Normal file
33
infra/deploy/secrets/beta.yaml
Normal file
|
|
@ -0,0 +1,33 @@
|
|||
PORT: ENC[AES256_GCM,data:/knJSQ==,iv:g8vSX20vD9Hz16lMdj46nToKYzLKzNskxzG/VHVVPbM=,tag:+X7H/JnC2n2Ie7RGzsaM8Q==,type:str]
|
||||
POSTGRES_PASSWORD: ENC[AES256_GCM,data:Vz/qCb50A3BLoQw+sDn7J7lvqld/urPj96qtKziR57c=,iv:ZjZi34GZZmMOTUKbSKdDV9KDNTrpSZTWww70LMV75ns=,tag:GuISiND97ThKgPl0wCywQg==,type:str]
|
||||
THERMOGRAPH_DATABASE_URL: ENC[AES256_GCM,data:o7rHapWWNbgMCNPz15NrpUbTbo7mzXgc4a2gGhXrcy7oNTAqA2ExV/+kvKv8CIvIKtgaMUrEmeh38XSd/ARVYauLRTjzCiAJIWtkXeN8fD8E4eGpfw==,iv:d8uZDon4Ub9g+I8MlgXMT+dWYrJUGCGrWdYfRNJ7kro=,tag:r/vR1zzGjZa7/pTOjBwk6g==,type:str]
|
||||
WORKERS: ENC[AES256_GCM,data:/A==,iv:kvFqDjAGq50aXhNrXmCxwYQXT24Tczh5GNPagGj28jc=,tag:oz0r4maN94LcBbzWDVOc1g==,type:str]
|
||||
APP_CPUS: ENC[AES256_GCM,data:Og==,iv:7UGacuCWIJlXu2B3WDjT6Z5BxFxUKxOlpUH8Xi+Xsgc=,tag:GOV5xS6B8u766kM6UBuwOQ==,type:str]
|
||||
DB_CPUS: ENC[AES256_GCM,data:Eg==,iv:rlZ5BjkagQR+T9BmOMIZz6R51x4boGYu12cxdG2ztvc=,tag:MQXbDXRCYlm8FFvXWwEzpg==,type:str]
|
||||
DB_MEMORY: ENC[AES256_GCM,data:5FQ=,iv:7FHgihO5DQ1se1Ng9KjKJHnY2NyoKJ+yWHt2I/jbTmU=,tag:CcsNGnpBFELh3IOrGrVjPA==,type:str]
|
||||
TIMESCALEDB_TAG: ENC[AES256_GCM,data:CWzVuxA7bbcAOb0=,iv:hRRNh11ETDTbIx7kkhumAc0QBaTV9SW8CLiq6McRqKA=,tag:3hBURxt2/BJpAUSgZyZ99g==,type:str]
|
||||
THERMOGRAPH_BASE: ENC[AES256_GCM,data:lw==,iv:kZOXeO841rEIQEnVnstm+u9BHKolnM0vvhNukx7RC8o=,tag:EMxBP63Zn1uIC8M5L7fmsg==,type:str]
|
||||
THERMOGRAPH_BASE_URL: ENC[AES256_GCM,data:DO2KCmfqs2zdxVuJivPVNn1YeFF8DED0tzET/g==,iv:Mrsrb6ToITqyatu5heC5s55SL8/Y4GMXqoYREMkOVd4=,tag:net8qhJ4ptXo3QcWYlbv6g==,type:str]
|
||||
THERMOGRAPH_COOKIE_SECURE: ENC[AES256_GCM,data:ow==,iv:U/I93B7gjYSSZjqudHd6QDQIRATx+l5QZ1MRWYD3aGk=,tag:f6Bs47hby1WQxAaDQ/fZ1w==,type:str]
|
||||
THERMOGRAPH_AUTH_SECRET: ENC[AES256_GCM,data:KsQO8ZxRp8n9AGL/z9kLxpO7vsmTBtM5H6ZpXzhm1VLVtc4CV+evEs4sAA==,iv:pMC8QruPjTRc6cXIwPBWHLY4eVomNkf4MYfWSeG/6qM=,tag:4W6Px9Xz8uvew95DIHDWVg==,type:str]
|
||||
THERMOGRAPH_METRICS_TOKEN: ENC[AES256_GCM,data:Ba1kxhSUNxq+oVHCskmQHhPuIcZPQQXTx2ht57D8Lyg=,iv:vT+aerexKF5ZXddQoXtBuzqAlzYQ8jT2wbef7HfXy9U=,tag:VkZb+s0C8UjnQnJKMC4HHQ==,type:str]
|
||||
THERMOGRAPH_INDEXNOW_KEY: ENC[AES256_GCM,data:GPytN7OsoKVO0n3VsiKERr0V5CwfXId3dCr4rMHNuHg=,iv:DlFEG3BiSMPh0jlB5YgSn6GrYgT0nxmoQGCDt1jUt28=,tag:/jclaZE6hd4EKpL0/7/WbA==,type:str]
|
||||
THERMOGRAPH_VAPID_PRIVATE_KEY: ENC[AES256_GCM,data:sTtkAbMVUtwvRssCZJa2PWUjlfnYSahciq9okgZj+DaqU4/mzNO3EDinMA==,iv:+sS0E+OP8qlBw+BEaSggEaaS2SCSBTrjaMzks8tdrJQ=,tag:pRN32rNjN4dp9OCzTHNkcQ==,type:str]
|
||||
THERMOGRAPH_VAPID_PUBLIC_KEY: ENC[AES256_GCM,data:SnFw+pwNgZlxMEhMUH/1daMZlRB2QFNJHtL+4C3S1eFcJJdT2uNGkY7Ym9FPfx1V7cwt/3urHg1oVgdDzYaIal2wye4l5HhKBP5PjtQkDW+Pck24kxul,iv:Ku4MhsrAfyA2Y3CqhGIsYKrjhKQ4cqAYjCqjFeEvBWE=,tag:hw42iS+79D7gNf+BPTYM0g==,type:str]
|
||||
THERMOGRAPH_VAPID_CONTACT: ENC[AES256_GCM,data:ox+hHiFMTPZqYwICDSwPwcT13ZDSKiGcAVHfVVMNYQ==,iv:DssV/olPXERpdmXNMmzjVYvhn/0zmuSp2XlTs42pIi0=,tag:4YHXI1BFbGZ6JsdXqJuf3w==,type:str]
|
||||
REGISTRY_TOKEN: ENC[AES256_GCM,data:pC1Zzs56la8qjzWcbkuKyeVHwS5WfIe+mrLm5Sdi/LKYziAvRU5cOA==,iv:57PlJEBdgwNJsfFdNOGqAnKk8l6zpmMSWYFEKS6pn+4=,tag:pPftNHUVU6+CY2nWMaXO6w==,type:str]
|
||||
sops:
|
||||
age:
|
||||
- enc: |
|
||||
-----BEGIN AGE ENCRYPTED FILE-----
|
||||
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSBUbjg4ZmQ2bXRzc1FoSjRj
|
||||
RzYwTXlzYkwyYjNBbUZoU1hlY0ludjNlamlJClBQK0NRRHFzcGpqZmIvTkNBMXNZ
|
||||
VmFRaXVwbTJPTHY1dzI1UGdOb0QvSzgKLS0tIGs1Q2NQUHJXZFd4WnpUaFVpY0NP
|
||||
TmhhMEN6ekg4Q0F1aUtuaDdOTVJEZHcKeCAW+Gqt6IRHizb0cOpKinyJzTkWORX5
|
||||
0+PQdtDxom72BqrAoj5lhepxW8YWPqQIRXsVS/XYQEQfPJuuuiRMKw==
|
||||
-----END AGE ENCRYPTED FILE-----
|
||||
recipient: age1xx4dzs0dxlwvkv9sjuqzsphl7lfrxannkfken374yu2qvvcte9sqzktqt2
|
||||
lastmodified: "2026-07-22T03:17:28Z"
|
||||
mac: ENC[AES256_GCM,data:npF+ngF4Khj3NJ2PNFcQjQYIrZVpdPM0drk43pAI6HfFaynX8JOqRFZRsH0sCfMoPkXB/ffCJxla1Tt3Mu3zgKTNSWcn+mjzl+dYvlruY1VylHLS+GnTO9zKeQcwbrYzTfZNWzQhm6/mw8KCL3791U/xNHwJc6g5VxRUWl6TBdo=,iv:JLzY8rm33NLG8CqyAlU6vBhlODUiRipKu9daoRvgyzg=,tag:yIsxVZAnw8FD7Bt/x55wMQ==,type:str]
|
||||
unencrypted_suffix: _unencrypted
|
||||
version: 3.13.2
|
||||
42
infra/deploy/secrets/dry-run-render.sh
Executable file
42
infra/deploy/secrets/dry-run-render.sh
Executable file
|
|
@ -0,0 +1,42 @@
|
|||
#!/usr/bin/env bash
|
||||
# Run ON a prod/beta box (via `ssh <box> 'bash -s' < this`), after the encrypted vault
|
||||
# file has been scp'd to /tmp/tg-<env>.yaml. NON-DESTRUCTIVE go/no-go for the cutover:
|
||||
# renders the vault and diffs the KEY=VALUE set against the live /etc/thermograph.env.
|
||||
# Prints only PASS/FAIL and any differing key NAMES — never secret values. No git, no
|
||||
# writes; touches nothing. (The encrypted file is safe to ship; only ciphertext.)
|
||||
set -euo pipefail
|
||||
key=/etc/thermograph/age.key
|
||||
env_name="$(sudo cat /etc/thermograph/secrets-env)"
|
||||
host="/tmp/tg-${env_name}.yaml"
|
||||
common="/tmp/tg-common.yaml"
|
||||
|
||||
command -v sops >/dev/null || { echo "FAIL: sops not installed"; exit 1; }
|
||||
[ -f "$host" ] || { echo "FAIL: $host not found — scp the encrypted vault file first"; exit 1; }
|
||||
|
||||
keymat="$(sudo grep '^AGE-SECRET-KEY-' "$key")"
|
||||
tmp="$(mktemp)"; live="$(mktemp)"; trap 'rm -f "$tmp" "$live"' EXIT
|
||||
: > "$tmp"
|
||||
[ -f "$common" ] && SOPS_AGE_KEY="$keymat" sops -d --input-type yaml --output-type dotenv "$common" >> "$tmp"
|
||||
SOPS_AGE_KEY="$keymat" sops -d --input-type yaml --output-type dotenv "$host" >> "$tmp"
|
||||
|
||||
sudo cat /etc/thermograph.env > "$live"
|
||||
python3 - "$live" "$tmp" <<'PY'
|
||||
import sys
|
||||
def load(p):
|
||||
out={}
|
||||
for line in open(p, encoding="utf-8"):
|
||||
s=line.rstrip("\n").strip()
|
||||
if not s or s.startswith("#") or "=" not in s: continue
|
||||
k,_,v=s.partition("="); out[k.strip()]=v
|
||||
return out
|
||||
live,rend=load(sys.argv[1]),load(sys.argv[2])
|
||||
lost=sorted(set(live)-set(rend)); added=sorted(set(rend)-set(live))
|
||||
changed=sorted(k for k in live if k in rend and live[k]!=rend[k])
|
||||
if lost or added or changed:
|
||||
print("FAIL — the render would NOT match the live env")
|
||||
if lost: print(" lost keys:", lost)
|
||||
if added: print(" extra keys:", added)
|
||||
if changed: print(" value-changed keys:", changed)
|
||||
sys.exit(1)
|
||||
print(f"PASS — all {len(live)} keys render byte-identical to the live /etc/thermograph.env")
|
||||
PY
|
||||
21
infra/deploy/secrets/example.yaml
Normal file
21
infra/deploy/secrets/example.yaml
Normal file
|
|
@ -0,0 +1,21 @@
|
|||
#ENC[AES256_GCM,data:mVl42QkTGozrOS8g1ufphjX65wDrrQ6CwB2ZkeliOUmvQ9TvaGe7M9i702gDBOkrqmwYRYdcPDLiLzJXt/mqpi8n0t0rnhduZxQG9WDO8ny0E+OHSO4=,iv:mr6GaRYJ8BKtrUz04MRevaPwfevwEFInKXs3MTFtY3w=,tag:gO/T8sy8Q9VLpYl2G9lXhg==,type:comment]
|
||||
#ENC[AES256_GCM,data:r+UY39rGucBlaPg9l2ce9KMzVXbZ6pv/GeuMmkfdUjiYNE9b+oH0NUNP/7kQAUrLG/ggjFAZ3sIGkIGTqi3iPuMDWW1fYv4Xs5fbe7ZfYoOI+A4=,iv:RfpoTanyXjCl03zMX4eZP/6BUhcEUyjQIPRnuCt9Bm4=,tag:AsFmRmjvnJ8eMRXrqLO78g==,type:comment]
|
||||
#ENC[AES256_GCM,data:bRsSXptMUyQ0i0JohyBhxUgn17Ozd4qltPf+fVF7vcP8IcuIln5DieCzW2TcsYQjTlhgXPPzjQA8gq4bXw8p3h2QUR9g4YLm2THRg85dKqdobQDnbIU=,iv:YbPFoIZXxpC4gpRh3spoQfYdEOpsKszqlNanPuPCXVI=,tag:VW2sCR74BlUosbxUj5P9tQ==,type:comment]
|
||||
#ENC[AES256_GCM,data:EMLogDZpYWRvdI2t7N95d8UP4a5v0qt+eS6FwHt/xRU2hUO008Ybm9z61bULqNm9iL/j6p4rPye7SNbckFU7oT/U4CvC0QitKNYf,iv:TAaCF/0f3Iv1cRbkKT7ZwnDamMOS1ec6jWiHOm/3Rgk=,tag:wC5mHKw5Q032zS2k4V6DvA==,type:comment]
|
||||
THERMOGRAPH_EXAMPLE_TOKEN: ENC[AES256_GCM,data:azyOkty43gxbFZvFfOK92aBcEwgsNAa0JuxOoQ==,iv:JI+M8AJ1XXyviaB7mz68ocv/VKIoxFyC7mD/k335LSQ=,tag:Tbglfpg6nw3EAzU2Cv4aUw==,type:str]
|
||||
THERMOGRAPH_EXAMPLE_API_KEY: ENC[AES256_GCM,data:LNDu2eHHIHykWBetNMn7wZM=,iv:dD+XmzW2xI2LkZ3DbJ6AVqOm/bmQ9e3Qrvls9dZo1T8=,tag:v1JHnPNRHXuvYlF+xHlR7w==,type:str]
|
||||
sops:
|
||||
age:
|
||||
- enc: |
|
||||
-----BEGIN AGE ENCRYPTED FILE-----
|
||||
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSA1VGlocmFpZFVqUzRsRVRR
|
||||
LzRHUWxydkFBdHhoQ0xnWnhYQ0h3Y1dSY1JFCm9ESnUwKzVxNk9PUzgwUzdpWFk0
|
||||
a1dtTVNlNkJhZFJxOVBHZDMrTFZUYzQKLS0tIGl5SzMxdE90bm9sNThFWkxvSXFl
|
||||
eWlDUFY5RlMvd1BqMExwaGU0ekE5bjQKdQV93BjiY+9aMRzWF+Fce9qMQokEaj8g
|
||||
V7fziLeEW1mXJFXiq7KhL2ryV+chZOtgGFn2sfzQlWOTu2LxO/+0ug==
|
||||
-----END AGE ENCRYPTED FILE-----
|
||||
recipient: age1xx4dzs0dxlwvkv9sjuqzsphl7lfrxannkfken374yu2qvvcte9sqzktqt2
|
||||
lastmodified: "2026-07-22T02:20:08Z"
|
||||
mac: ENC[AES256_GCM,data:pObgYq9Tu7EcfHmLIq2Q5hPF7RA3yR6L2xEIOI+d2zTU+t0gLKdOE3l2kpiHf568PLcbkdZbzF057H7MW1AQd192y5Ph2OidA05D0eZmWaykmFHHY9ZyjbTondXSbTWmbVMYDXZGkZnvPnaJaivEmfFM7n0xeOf6ecI6C/tBmG8=,iv:fWN+C1MryDXzCpMcuKddsRqEt7StW6dc+K9K3+JIGxs=,tag:+wGtLFw+yBNa8HQltUrCkg==,type:str]
|
||||
unencrypted_suffix: _unencrypted
|
||||
version: 3.13.2
|
||||
41
infra/deploy/secrets/prod.yaml
Normal file
41
infra/deploy/secrets/prod.yaml
Normal file
|
|
@ -0,0 +1,41 @@
|
|||
PORT: ENC[AES256_GCM,data:3VXjXA==,iv:9tU5b/Ept1xDFVi5oRq/rimdXYxQj5qCQkbtYKQLT5Y=,tag:1gPFP4oUwp66aNP2bbDOLg==,type:str]
|
||||
POSTGRES_PASSWORD: ENC[AES256_GCM,data:Hct/SZ4mFqb3Vm4a3OgCSbKhdLLTxOOSb/kHBiAZ9pM=,iv:fURYQkJ4biEDruWhfMqjEESF8Y64JR+Tjf429Rrt7J4=,tag:h1ZAoUImYINVzLcvTbs4dQ==,type:str]
|
||||
THERMOGRAPH_DATABASE_URL: ENC[AES256_GCM,data:gfmGcH19GDNjKCffMADnKrViei2LuGJae06MsNS7biOFcBr23oY76gCEZVRa4YAaoUfNjr5HyyS3wiIUB+qAz/XvYXilzl9cYONVY2wXWplmQ0sW4Q==,iv:4h6LxDucrJ6Ma13z23KJlzz4lDFXaz9fg134kwpLeX0=,tag:BqMQbndCngN6I0UIhjsE6w==,type:str]
|
||||
WORKERS: ENC[AES256_GCM,data:6g==,iv:R0vRawz9uTPFL/GoVQTOFybau/QHVSmXU2N4x/xPjcQ=,tag:rZ2GnEYo4EBISmgSMQdjXA==,type:str]
|
||||
APP_CPUS: ENC[AES256_GCM,data:AQ==,iv:DILyG0LefEnnVHJF7PUshbAX+wYmPqTlduOnec7lMWY=,tag:U7vQJ5l1C4vTWvf1bVSglg==,type:str]
|
||||
DB_CPUS: ENC[AES256_GCM,data:7Q==,iv:6Tc5DYEPSBeJMpdPj+zS8B8S5DEWl28/3YqnFlFHTRs=,tag:9AOgfaEkcvs4jRTo6mKYLA==,type:str]
|
||||
DB_MEMORY: ENC[AES256_GCM,data:/FIY,iv:jvVm7C9CTnCFudGDOcpCn3VgZlvHmSq27eFV/1/PMuA=,tag:xfljFaZ6ImG3c03gtjC32A==,type:str]
|
||||
TIMESCALEDB_TAG: ENC[AES256_GCM,data:HQ3xemefK2/1bPE=,iv:yJHJac2ZTzw0lCESoqywogLhUs+EAHIQKPSCXTa+kzs=,tag:POiiKXCP2CEHAatQy9T7sg==,type:str]
|
||||
THERMOGRAPH_BASE: ENC[AES256_GCM,data:Gw==,iv:2VJRKggVu41cjnq3zdl9gBYhD+bWU1HXSpTCTwXDsPM=,tag:Sosggs9Fg4YN9eBTYSgzxw==,type:str]
|
||||
THERMOGRAPH_BASE_URL: ENC[AES256_GCM,data:dbONWMu5QUsbmS3Z27v9v5w9Q9w7V18=,iv:mywyTNJLbHueL2jMeOMQBu9os3CvldaAuVP+Irg3xOE=,tag:FB/86crdYna2TyGpys/spw==,type:str]
|
||||
THERMOGRAPH_COOKIE_SECURE: ENC[AES256_GCM,data:NQ==,iv:j6BUIBWP9ebS7xFw/2RrKo+fvmhOn9mYTR6+K7lizvw=,tag:B/PEjByAT4XsXcECOLSoPw==,type:str]
|
||||
THERMOGRAPH_AUTH_SECRET: ENC[AES256_GCM,data:8+s0NGrtJBV6nOgHfmoODukS8hSij8BoUjxYTH4s398LGABwro+MzB4Xsg==,iv:tmxMUkHND2x9RZquWKesDiPePMJA1qNLD+TMQLYR35M=,tag:iSHOu3En5rZpY+MOH4Us/A==,type:str]
|
||||
THERMOGRAPH_METRICS_TOKEN: ENC[AES256_GCM,data:PcEsEp+bAMsiXxAFHDX4bS3dnwQz5q23gMb1C5tVopI=,iv:VkbuTVkOBKlCyq6mNFAIM26VSzUbJ/Nyvg1vEWuUWJg=,tag:CR4ALs4UCXZ1ybwfkwO+1Q==,type:str]
|
||||
THERMOGRAPH_INDEXNOW_KEY: ENC[AES256_GCM,data:4NzyzEyy8khcWVBJj7Rs/lP4fuAn13aNzrQRTHGgVDw=,iv:fc9VwStsWO2gbbgqK43vFdSR3zl4+vmKUQsFY9Fu1YQ=,tag:D/aEkNBlzE1K5ldLt5SGGw==,type:str]
|
||||
THERMOGRAPH_VAPID_PRIVATE_KEY: ENC[AES256_GCM,data:v2U/CVUMOAkZG8a9IadkaSPWHhg9nOQnsIHoziE7A5GxiNTwm946+Uy2UQ==,iv:ZokatpIp58zj+KSPvA5/D7Xh6n6NSiX9z7PDmuHPrLs=,tag:FDDVU1KcBd1newEpd9PKWw==,type:str]
|
||||
THERMOGRAPH_VAPID_PUBLIC_KEY: ENC[AES256_GCM,data:0Noa0XlR3NJz9SxY/IOfBei+DbeWnv+yLSb7SkKY+hmrOnNwb6Hcz5AOVOP9nVIgHZaT9cpl0e7LVJyO4QofLzZqms9Uy8HtQh3J58vTXAh24sKX/ZBR,iv:Aa3766JRR9OLeyrkmN6/yi5UqexITXObpQwDIH6BuBI=,tag:MjIVwQoxXpEwIbKjSgXYsQ==,type:str]
|
||||
THERMOGRAPH_VAPID_CONTACT: ENC[AES256_GCM,data:4UM7/dB8BDPa/VodQkEuIi5ku2JlgJkSqUVdojTR4g==,iv:09t1CpbnOIunKSegqlKKXkEubmjgidACOmXuQNMLzfk=,tag:pPz2MUzgCTZVXX0TnlgKwg==,type:str]
|
||||
THERMOGRAPH_MAIL_BACKEND: ENC[AES256_GCM,data:qucAxA==,iv:XTBitLKgpOrWGePxe09K/76hUCmPADZpKT2l/Kp0HXc=,tag:WAwk3kW/b2aXBXjau3lMyQ==,type:str]
|
||||
THERMOGRAPH_SMTP_HOST: ENC[AES256_GCM,data:GzxyI/CRiwx4gw==,iv:VleW+mP3BQii9axQVm3zHXTZsLvS89DOzdt0rFd81z0=,tag:13+I5LEasTKYkYDu2eGGYA==,type:str]
|
||||
THERMOGRAPH_SMTP_PORT: ENC[AES256_GCM,data:yTM=,iv:ja9tBwmTgSJLcUXzrvPYZKUAuFdzn753wZ5FuQJUrNM=,tag:5t0/G9/TuHl8LJNTCW5Q5w==,type:str]
|
||||
THERMOGRAPH_DISCORD_CLIENT_SECRET: ENC[AES256_GCM,data:CrhrDfFmk7F2lWQeY9kd7yxzofz6D2beXf8Iluqf3OA=,iv:0cLNwuuCqThcp7PanztQWjzNjt/QE3Q0p+R+2sd/dvI=,tag:EM2+VmwSaDj0/IItHJA5Uw==,type:str]
|
||||
THERMOGRAPH_DISCORD_APP_ID: ENC[AES256_GCM,data:lxwB08B143dCFsobG9JOkmScIQ==,iv:TFjzSDZWptqWgswiPHxJAzG1hlVI5bUA87ZKjBTsCGs=,tag:eRgvfiSS1Ujd5T3AeZCwUw==,type:str]
|
||||
THERMOGRAPH_DISCORD_BOT_TOKEN: ENC[AES256_GCM,data:zfVJLH8ddct2Jvstl/olZLQ9SM15ccWLCIpYe/k5BUZiV50atzpQfE13FifB1jWF6C+5P95odxVFtne7YJyaDxryfIX8Gbea,iv:wDM4qFpGYGtUJfEk5IbhMg9J9gr8ambEOkSWWBA2CSU=,tag:FhtUOe5WKMH2h7qIi4yKgQ==,type:str]
|
||||
REGISTRY_TOKEN: ENC[AES256_GCM,data:3VETZR026YXvWA7DLly4/mz+hqEK7vkEMwKlDKU6CQZeZiFpbQXWig==,iv:ixZJi2VWAZak6dFPvVyErVRanp9bzqXyV3MVa4QsBuU=,tag:Pc2Mdw58S0tSdkgJ5yuVNw==,type:str]
|
||||
THERMOGRAPH_DISCORD_BOT: ENC[AES256_GCM,data:Kg==,iv:CPYN8aCu79PHaz3Zj6OgqyJhXdHLt0tLodgML2NuwQM=,tag:1B3IhlfjCwSlPq7FnKGeRg==,type:str]
|
||||
THERMOGRAPH_DISCORD_WEATHER_CHANNEL: ENC[AES256_GCM,data:WIxYcoOnBDdDcNj8uJ5eUqXjVQ==,iv:l6kU5PGVcW2Nv7EqBUGCRxBzaJOJ3AW1nEEH33aAbmI=,tag:/KgdCKmrlQioIWdz21PcTg==,type:str]
|
||||
sops:
|
||||
age:
|
||||
- enc: |
|
||||
-----BEGIN AGE ENCRYPTED FILE-----
|
||||
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSA0THIwYmxSRnIveGlCWHZS
|
||||
NDR1bHZ6MGNGdFJZbFRNdEJXZk8zdHBMdkMwCnNiTlVFRGpKNHB3TzYycEY5dUd1
|
||||
RktTaWZ3dFNVQzV1OXhCSmp2VUJtaXMKLS0tIEJ0dGJXajFYcDFKaXhKZ0FaOUEz
|
||||
eUh2VFlYeW5CQ1hnTVQvN2lBMU9TclUKC3riyd2eEFsej7ItoQ8rpsAiRms8yJgo
|
||||
PDXCB8EqKwdC3BnKlpD9ptYcuPHVUSgve1wIZYJ7iHIEhW9Kc6RXOw==
|
||||
-----END AGE ENCRYPTED FILE-----
|
||||
recipient: age1xx4dzs0dxlwvkv9sjuqzsphl7lfrxannkfken374yu2qvvcte9sqzktqt2
|
||||
lastmodified: "2026-07-23T04:33:06Z"
|
||||
mac: ENC[AES256_GCM,data:BjNMEudsS9ZViR55OjSXu5qrOJ9yu2kCGx61qu4vE1ZZJIZtnnyZ3CfpsPygvLLhiFRj76xcqFnuSqTdr17BCEHz3UzsHFv0vBi17lZ5ksN5Oz+9wr+3Pg3TvxC0VF2VHJVP84VH61CmB5+G4WOms8WqLtmIGSLYcTVtcIFoUhk=,iv:jTEB8Mxg8qIyNfhpSgJtUvULVx4+XFrPLyQKhNcbkTg=,tag:SIyvNltdGCMfDGvi4kD3MA==,type:str]
|
||||
unencrypted_suffix: _unencrypted
|
||||
version: 3.13.2
|
||||
56
infra/deploy/secrets/seed-encrypt-on-host.sh
Executable file
56
infra/deploy/secrets/seed-encrypt-on-host.sh
Executable file
|
|
@ -0,0 +1,56 @@
|
|||
#!/usr/bin/env bash
|
||||
# Seed deploy/secrets/<env>.yaml by encrypting a box's live /etc/thermograph.env
|
||||
# ENTIRELY ON THE BOX. The plaintext never leaves the box; only the already-encrypted
|
||||
# ciphertext is written back here. Encryption uses the age PUBLIC key (safe to hardcode)
|
||||
# so the machine you run this from needs nothing but SSH access + the agent key — no
|
||||
# sops, no age, no private key. sops is installed on the box on demand.
|
||||
#
|
||||
# Run from any checkout of this branch that can SSH to the target with the agent key:
|
||||
# deploy/secrets/seed-encrypt-on-host.sh prod agent@169.58.46.181
|
||||
# deploy/secrets/seed-encrypt-on-host.sh beta agent@75.119.132.91
|
||||
# Then commit + push the resulting deploy/secrets/<env>.yaml. Verify faithfulness with
|
||||
# the key-gaps skill after. See README.md.
|
||||
set -euo pipefail
|
||||
cd "$(dirname "$(readlink -f "$0")")/../.."
|
||||
|
||||
ENV_NAME="${1:?usage: seed-encrypt-on-host.sh <env> <ssh-target> [ssh-key]}"
|
||||
SSH_TARGET="${2:?ssh target, e.g. agent@169.58.46.181}"
|
||||
SSH_KEY="${3:-$HOME/.ssh/thermograph_agent_ed25519}"
|
||||
OUT="deploy/secrets/${ENV_NAME}.yaml"
|
||||
|
||||
# Public age recipient — NOT a secret (matches .sops.yaml). The private key never
|
||||
# leaves the operator's machine / the hosts' /etc/thermograph/age.key.
|
||||
AGE_PUB="age1xx4dzs0dxlwvkv9sjuqzsphl7lfrxannkfken374yu2qvvcte9sqzktqt2"
|
||||
SOPS_URL="https://github.com/getsops/sops/releases/download/v3.13.2/sops-v3.13.2.linux.amd64"
|
||||
|
||||
echo "==> Encrypting ${SSH_TARGET}:/etc/thermograph.env on-host -> ${OUT}"
|
||||
ssh -i "$SSH_KEY" "$SSH_TARGET" "AGE_PUB='$AGE_PUB' SOPS_URL='$SOPS_URL' bash -s" > "$OUT" <<'REMOTE'
|
||||
set -euo pipefail
|
||||
exec 3>&1 # real stdout carries ONLY the ciphertext
|
||||
{ # setup noise -> stderr, so it can't corrupt the file
|
||||
if ! command -v sops >/dev/null 2>&1; then
|
||||
sudo curl -fsSL "$SOPS_URL" -o /usr/local/bin/sops
|
||||
sudo chmod +x /usr/local/bin/sops
|
||||
fi
|
||||
} >&2
|
||||
tmp="$(mktemp)"; tmy="$(mktemp)"; trap 'rm -f "$tmp" "$tmy"' EXIT
|
||||
sudo cat /etc/thermograph.env > "$tmp" # plaintext stays on this box only
|
||||
python3 -c '
|
||||
import json, sys
|
||||
for line in open(sys.argv[1], encoding="utf-8"):
|
||||
s = line.rstrip("\n").strip()
|
||||
if not s or s.startswith("#") or "=" not in s:
|
||||
continue
|
||||
k, _, v = s.partition("=")
|
||||
print(f"{k.strip()}: {json.dumps(v)}") # JSON-escaped scalar = valid YAML
|
||||
' "$tmp" > "$tmy"
|
||||
cd /tmp # no .sops.yaml here, so --age is honored
|
||||
sops -e --age "$AGE_PUB" --input-type yaml --output-type yaml "$tmy" >&3
|
||||
REMOTE
|
||||
|
||||
if grep -q 'ENC\[AES256_GCM' "$OUT"; then
|
||||
echo "==> OK: ${OUT} written and encrypted ($(grep -cE '^[A-Za-z_][A-Za-z0-9_]*: ' "$OUT") keys)"
|
||||
else
|
||||
echo "!! ${OUT} is not encrypted (the on-host step failed) — removing" >&2
|
||||
rm -f "$OUT"; exit 1
|
||||
fi
|
||||
83
infra/deploy/secrets/seed-from-live.sh
Executable file
83
infra/deploy/secrets/seed-from-live.sh
Executable file
|
|
@ -0,0 +1,83 @@
|
|||
#!/usr/bin/env bash
|
||||
# Seed (or re-seed) an encrypted secrets file from a box's live /etc/thermograph.env.
|
||||
# Run on the operator's machine — it reads production secrets, so it is deliberately
|
||||
# NOT something the agent runs for you. Requires: sops + age on PATH, the age private
|
||||
# key at ~/.config/sops/age/keys.txt (or SOPS_AGE_KEY_FILE), and SSH access to the box.
|
||||
#
|
||||
# deploy/secrets/seed-from-live.sh prod agent@169.58.46.181 ~/.ssh/thermograph_agent_ed25519
|
||||
# deploy/secrets/seed-from-live.sh beta agent@75.119.132.91 ~/.ssh/thermograph_agent_ed25519
|
||||
#
|
||||
# Writes deploy/secrets/<env>.yaml ENCRYPTED (exact copy of the live env, so the
|
||||
# deploy render is faithful), then verifies the render round-trips to the same
|
||||
# KEY=VALUE set. Never prints secret values. See README.md.
|
||||
set -euo pipefail
|
||||
|
||||
# Anchor to the repo that holds this script, so it works when invoked by absolute path
|
||||
# from any cwd (a different checkout, your home dir, etc.) — OUT and .sops.yaml resolve
|
||||
# relative to here, not to wherever you happened to run it.
|
||||
cd "$(dirname "$(readlink -f "$0")")/../.."
|
||||
|
||||
ENV_NAME="${1:?usage: seed-from-live.sh <env> <ssh-target> [ssh-key]}"
|
||||
SSH_TARGET="${2:?ssh target, e.g. agent@169.58.46.181}"
|
||||
SSH_KEY="${3:-$HOME/.ssh/thermograph_agent_ed25519}"
|
||||
OUT="deploy/secrets/${ENV_NAME}.yaml"
|
||||
command -v sops >/dev/null || { echo "!! sops not on PATH" >&2; exit 1; }
|
||||
|
||||
tmp="$(mktemp)"; trap 'rm -f "$tmp" "$tmp.env"' EXIT
|
||||
umask 077
|
||||
|
||||
echo "==> Pulling ${SSH_TARGET}:/etc/thermograph.env (sudo)"
|
||||
ssh -i "$SSH_KEY" "$SSH_TARGET" 'sudo cat /etc/thermograph.env' > "$tmp"
|
||||
n=$(grep -cE '^[A-Za-z_][A-Za-z0-9_]*=' "$tmp" || true)
|
||||
[ "$n" -gt 0 ] || { echo "!! no KEY=VALUE lines read (permission? path?)" >&2; exit 1; }
|
||||
echo " ${n} vars"
|
||||
|
||||
echo "==> Writing ${OUT} (flat YAML) and encrypting"
|
||||
python3 - "$tmp" "$OUT" "$ENV_NAME" <<'PY'
|
||||
import json, sys
|
||||
src, dst, env = sys.argv[1:4]
|
||||
pairs = {}
|
||||
for line in open(src, encoding="utf-8"):
|
||||
s = line.rstrip("\n").strip()
|
||||
if not s or s.startswith("#") or "=" not in s:
|
||||
continue
|
||||
k, _, v = s.partition("=")
|
||||
k = k.strip()
|
||||
if not all(c.isalnum() or c == "_" for c in k) or not k[0].isalpha() and k[0] != "_":
|
||||
continue
|
||||
pairs[k] = v
|
||||
with open(dst, "w", encoding="utf-8") as fh:
|
||||
fh.write(f"# SOPS-encrypted — {env} host secrets, seeded from the live /etc/thermograph.env.\n")
|
||||
for k, v in pairs.items():
|
||||
fh.write(f"{k}: {json.dumps(v)}\n") # JSON-escaped scalar is valid YAML
|
||||
print(f" {len(pairs)} keys")
|
||||
PY
|
||||
sops -e -i "$OUT"
|
||||
grep -q 'ENC\[AES256_GCM' "$OUT" || { echo "!! $OUT did not encrypt — refusing to keep it" >&2; rm -f "$OUT"; exit 1; }
|
||||
|
||||
echo "==> Verifying the render round-trips to the same KEY=VALUE set"
|
||||
sops -d --input-type yaml --output-type dotenv "$OUT" > "$tmp.env"
|
||||
python3 - "$tmp" "$tmp.env" <<'PY'
|
||||
import sys
|
||||
def load(p):
|
||||
d = {}
|
||||
for line in open(p, encoding="utf-8"):
|
||||
s = line.rstrip("\n").strip()
|
||||
if not s or s.startswith("#") or "=" not in s:
|
||||
continue
|
||||
k, _, v = s.partition("=")
|
||||
d[k.strip()] = v
|
||||
return d
|
||||
live, rendered = load(sys.argv[1]), load(sys.argv[2])
|
||||
missing = sorted(set(live) - set(rendered))
|
||||
extra = sorted(set(rendered) - set(live))
|
||||
diff = sorted(k for k in live if k in rendered and live[k] != rendered[k])
|
||||
if missing or extra or diff:
|
||||
if missing: print(" !! keys lost in render:", missing)
|
||||
if extra: print(" !! keys added in render:", extra)
|
||||
if diff: print(" !! values changed (keys only):", diff)
|
||||
sys.exit(1)
|
||||
print(f" OK — {len(live)} keys render identically")
|
||||
PY
|
||||
echo "==> ${OUT} is ready. Review with: sops ${OUT} (or git diff --stat)"
|
||||
echo " Then commit + push; the agent can take it from there."
|
||||
93
infra/deploy/stack/autoscale.sh
Executable file
93
infra/deploy/stack/autoscale.sh
Executable file
|
|
@ -0,0 +1,93 @@
|
|||
#!/bin/sh
|
||||
# Autoscaler for the stack's `web` service: scale replicas between
|
||||
# MIN_REPLICAS and MAX_REPLICAS on sustained per-task CPU.
|
||||
#
|
||||
# Runs as a Swarm service on the manager with the docker socket mounted (see
|
||||
# thermograph-stack.yml). Every web task is placed on this node today, so
|
||||
# node-local `docker stats` sees them all — when a second app node exists,
|
||||
# this needs a per-node reader or a metrics-based signal instead; that's the
|
||||
# documented upgrade path, not a today problem.
|
||||
#
|
||||
# Semantics (deliberately boring):
|
||||
# - Sample avg CPU% per web task every POLL_SECONDS (docker stats CPUPerc:
|
||||
# 100 = one full host core).
|
||||
# - UP_SAMPLES consecutive samples above SCALE_UP_CPU -> scale +1.
|
||||
# - DOWN_SAMPLES consecutive samples below SCALE_DOWN_CPU -> scale -1.
|
||||
# (Down is ~7x slower than up on defaults: flap-averse by construction.)
|
||||
# - COOLDOWN_SECONDS after any change: samples are ignored entirely.
|
||||
# - Clamped to [MIN_REPLICAS, MAX_REPLICAS]; scaling waits for convergence
|
||||
# (--detach=false), so a stuck rollout blocks further changes rather than
|
||||
# stacking them.
|
||||
set -eu
|
||||
|
||||
STACK_NAME="${STACK_NAME:-thermograph}"
|
||||
SERVICE="${STACK_NAME}_web"
|
||||
MIN="${MIN_REPLICAS:-1}"
|
||||
MAX="${MAX_REPLICAS:-3}"
|
||||
UP_AT="${SCALE_UP_CPU:-220}"
|
||||
DOWN_AT="${SCALE_DOWN_CPU:-60}"
|
||||
POLL="${POLL_SECONDS:-15}"
|
||||
UP_N="${UP_SAMPLES:-3}"
|
||||
DOWN_N="${DOWN_SAMPLES:-20}"
|
||||
COOLDOWN="${COOLDOWN_SECONDS:-180}"
|
||||
|
||||
up_hits=0
|
||||
down_hits=0
|
||||
last_change=0
|
||||
|
||||
log() { echo "[autoscale] $(date -u +%H:%M:%S) $*"; }
|
||||
|
||||
replicas() {
|
||||
docker service inspect "$SERVICE" \
|
||||
--format '{{.Spec.Mode.Replicated.Replicas}}' 2>/dev/null || echo ""
|
||||
}
|
||||
|
||||
avg_cpu() {
|
||||
# Mean CPUPerc across this node's web tasks, as an integer percent.
|
||||
docker stats --no-stream --format '{{.Name}} {{.CPUPerc}}' 2>/dev/null \
|
||||
| awk -v svc="$SERVICE" '
|
||||
index($1, svc".") == 1 {
|
||||
gsub(/%/, "", $2); sum += $2; n++
|
||||
}
|
||||
END { if (n > 0) printf "%d", sum / n; else print "" }'
|
||||
}
|
||||
|
||||
log "watching $SERVICE: min=$MIN max=$MAX up>@${UP_AT}%x${UP_N} down<@${DOWN_AT}%x${DOWN_N} poll=${POLL}s cooldown=${COOLDOWN}s"
|
||||
|
||||
while :; do
|
||||
sleep "$POLL"
|
||||
|
||||
now=$(date +%s)
|
||||
if [ $((now - last_change)) -lt "$COOLDOWN" ]; then
|
||||
continue
|
||||
fi
|
||||
|
||||
cur=$(replicas)
|
||||
[ -n "$cur" ] || { log "service $SERVICE not found; waiting"; continue; }
|
||||
cpu=$(avg_cpu)
|
||||
[ -n "$cpu" ] || continue # no running tasks visible this sample
|
||||
|
||||
if [ "$cpu" -gt "$UP_AT" ]; then
|
||||
up_hits=$((up_hits + 1)); down_hits=0
|
||||
elif [ "$cpu" -lt "$DOWN_AT" ]; then
|
||||
down_hits=$((down_hits + 1)); up_hits=0
|
||||
else
|
||||
up_hits=0; down_hits=0
|
||||
fi
|
||||
|
||||
if [ "$up_hits" -ge "$UP_N" ] && [ "$cur" -lt "$MAX" ]; then
|
||||
target=$((cur + 1))
|
||||
log "avg cpu ${cpu}% > ${UP_AT}% x${UP_N}: scaling $cur -> $target"
|
||||
if docker service scale --detach=false "$SERVICE=$target"; then
|
||||
last_change=$(date +%s)
|
||||
fi
|
||||
up_hits=0; down_hits=0
|
||||
elif [ "$down_hits" -ge "$DOWN_N" ] && [ "$cur" -gt "$MIN" ]; then
|
||||
target=$((cur - 1))
|
||||
log "avg cpu ${cpu}% < ${DOWN_AT}% x${DOWN_N}: scaling $cur -> $target"
|
||||
if docker service scale --detach=false "$SERVICE=$target"; then
|
||||
last_change=$(date +%s)
|
||||
fi
|
||||
up_hits=0; down_hits=0
|
||||
fi
|
||||
done
|
||||
226
infra/deploy/stack/deploy-stack.sh
Executable file
226
infra/deploy/stack/deploy-stack.sh
Executable file
|
|
@ -0,0 +1,226 @@
|
|||
#!/usr/bin/env bash
|
||||
# Swarm-stack deploy for prod — the stack-mode counterpart of deploy/deploy.sh,
|
||||
# speaking the SAME contract the app repos' workflows already use
|
||||
# (SERVICE=backend|frontend|all + BACKEND_IMAGE_TAG/FRONTEND_IMAGE_TAG), so
|
||||
# switching a host to stack mode needs no workflow changes: deploy.sh execs
|
||||
# this when /etc/thermograph/deploy-mode contains "stack".
|
||||
#
|
||||
# What a roll does here vs compose:
|
||||
# backend -> one-shot migrate, then `docker service update --image` on
|
||||
# web AND worker (same image; start-first, health-gated,
|
||||
# auto-rollback on failure).
|
||||
# frontend -> `docker service update --image` on frontend.
|
||||
# all -> full `docker stack deploy` (+ migrate first), which also
|
||||
# applies stack-file changes (new services, env, limits).
|
||||
#
|
||||
# TEST MODE (STACK_TEST=1): deploys under stack name thermograph-test with
|
||||
# throwaway volumes and the LB on 127.0.0.1:18137/18080 — a full parallel
|
||||
# rehearsal on the same host that cannot touch live data or ports.
|
||||
set -euo pipefail
|
||||
|
||||
APP_DIR="${APP_DIR:-/opt/thermograph}"
|
||||
SERVICE="${SERVICE:-all}"
|
||||
cd "$APP_DIR"
|
||||
|
||||
case "$SERVICE" in
|
||||
backend|frontend|all) ;;
|
||||
*) echo "!! SERVICE must be backend|frontend|all, got '$SERVICE'" >&2; exit 2 ;;
|
||||
esac
|
||||
|
||||
if [ "${STACK_TEST:-0}" = "1" ]; then
|
||||
STACK_NAME="thermograph-test"
|
||||
LB_NAME="thermograph-test-lb"
|
||||
LB_HTTP_PORT=18137; LB_FE_PORT=18080
|
||||
export PGDATA_VOLUME="thermograph-test_pgdata"
|
||||
export APPDATA_VOLUME="thermograph-test_appdata"
|
||||
export APPLOGS_VOLUME="thermograph-test_applogs"
|
||||
docker volume create "$PGDATA_VOLUME" >/dev/null
|
||||
docker volume create "$APPDATA_VOLUME" >/dev/null
|
||||
docker volume create "$APPLOGS_VOLUME" >/dev/null
|
||||
else
|
||||
STACK_NAME="${STACK_NAME:-thermograph}"
|
||||
LB_NAME="thermograph-lb"
|
||||
LB_HTTP_PORT=8137; LB_FE_PORT=8080
|
||||
fi
|
||||
export STACK_NAME
|
||||
|
||||
# --- secrets ------------------------------------------------------------------
|
||||
# Render (SOPS) + source /etc/thermograph.env exactly like deploy.sh, then
|
||||
# install the uid-10001-readable copy the tasks' env-entrypoint shim sources.
|
||||
# 10001 = the app images' `thermograph` user; the file is 0400 to that uid.
|
||||
if [ -f "$APP_DIR/deploy/render-secrets.sh" ]; then
|
||||
# shellcheck source=deploy/render-secrets.sh
|
||||
. "$APP_DIR/deploy/render-secrets.sh"
|
||||
render_thermograph_secrets "$APP_DIR"
|
||||
fi
|
||||
set -a; . /etc/thermograph.env 2>/dev/null || true; set +a
|
||||
sudo install -o 10001 -g 0 -m 0400 /etc/thermograph.env /etc/thermograph/stack.env \
|
||||
|| install -o 10001 -g 0 -m 0400 /etc/thermograph.env /etc/thermograph/stack.env
|
||||
|
||||
# --- image tags -----------------------------------------------------------------
|
||||
# Same persisted-tags contract as deploy.sh: incoming env wins, the file
|
||||
# supplies the sibling. Stack mode keeps its own file so test/real never mix.
|
||||
REGISTRY_HOST="${REGISTRY_HOST:-git.thermograph.org}"
|
||||
export REGISTRY_HOST
|
||||
TAGS_FILE="$APP_DIR/deploy/.stack-image-tags.env"
|
||||
_incoming_backend="${BACKEND_IMAGE_TAG:-}"
|
||||
_incoming_frontend="${FRONTEND_IMAGE_TAG:-}"
|
||||
if [ -f "$TAGS_FILE" ]; then set -a; . "$TAGS_FILE"; set +a; fi
|
||||
[ -n "$_incoming_backend" ] && BACKEND_IMAGE_TAG="$_incoming_backend"
|
||||
[ -n "$_incoming_frontend" ] && FRONTEND_IMAGE_TAG="$_incoming_frontend"
|
||||
case "$SERVICE" in
|
||||
backend) : "${BACKEND_IMAGE_TAG:?set BACKEND_IMAGE_TAG=sha-<12hex>}" ;;
|
||||
frontend) : "${FRONTEND_IMAGE_TAG:?set FRONTEND_IMAGE_TAG=sha-<12hex>}" ;;
|
||||
all)
|
||||
: "${BACKEND_IMAGE_TAG:?set BACKEND_IMAGE_TAG (SERVICE=all needs both)}"
|
||||
: "${FRONTEND_IMAGE_TAG:?set FRONTEND_IMAGE_TAG (SERVICE=all needs both)}" ;;
|
||||
esac
|
||||
export BACKEND_IMAGE_TAG="${BACKEND_IMAGE_TAG:-local}"
|
||||
export FRONTEND_IMAGE_TAG="${FRONTEND_IMAGE_TAG:-local}"
|
||||
BACKEND_IMAGE="$REGISTRY_HOST/${BACKEND_IMAGE_PATH:-emi/thermograph-backend/app}:$BACKEND_IMAGE_TAG"
|
||||
FRONTEND_IMAGE="$REGISTRY_HOST/${FRONTEND_IMAGE_PATH:-emi/thermograph-frontend/app}:$FRONTEND_IMAGE_TAG"
|
||||
|
||||
# --- timescale image pin ---------------------------------------------------------
|
||||
# Hazard #7: the db image under an existing volume must never drift. Resolve
|
||||
# the digest-pinned ref from whatever is running (stack task or compose
|
||||
# container), falling back to the local latest-pg18's digest on first bring-up.
|
||||
if [ -z "${TIMESCALEDB_IMAGE:-}" ]; then
|
||||
cid=$(docker ps -q --filter "label=com.docker.swarm.service.name=${STACK_NAME}_db" | head -1)
|
||||
[ -z "$cid" ] && cid=$(docker ps -q --filter "name=thermograph-db-1" | head -1)
|
||||
if [ -n "$cid" ]; then
|
||||
img=$(docker inspect --format '{{.Image}}' "$cid")
|
||||
else
|
||||
img="timescale/timescaledb:${TIMESCALEDB_TAG:-latest-pg18}"
|
||||
fi
|
||||
TIMESCALEDB_IMAGE=$(docker image inspect --format '{{index .RepoDigests 0}}' "$img" 2>/dev/null | head -1)
|
||||
[ -n "$TIMESCALEDB_IMAGE" ] || TIMESCALEDB_IMAGE="$img"
|
||||
fi
|
||||
export TIMESCALEDB_IMAGE
|
||||
echo "==> Images: web/worker=$BACKEND_IMAGE frontend=$FRONTEND_IMAGE db=$TIMESCALEDB_IMAGE"
|
||||
|
||||
# --- registry ------------------------------------------------------------------
|
||||
if [ -n "${REGISTRY_TOKEN:-}" ]; then
|
||||
echo "$REGISTRY_TOKEN" | docker login "$REGISTRY_HOST" --username emi --password-stdin
|
||||
fi
|
||||
echo "==> Pulling images"
|
||||
pull_ok=0
|
||||
for i in $(seq 1 30); do
|
||||
if docker pull -q "$BACKEND_IMAGE" >/dev/null && docker pull -q "$FRONTEND_IMAGE" >/dev/null; then
|
||||
pull_ok=1; break
|
||||
fi
|
||||
echo " pull attempt $i/30 failed (image may not be pushed yet); retrying in 10s..." >&2
|
||||
sleep 10
|
||||
done
|
||||
[ "$pull_ok" = 1 ] || { echo "!! image pull failed after 30 attempts" >&2; exit 1; }
|
||||
|
||||
# --- one-shot migrations ---------------------------------------------------------
|
||||
# Before any backend roll: N replicas must never race Alembic (RUN_MIGRATIONS=0
|
||||
# in the stack). Runs on the stack's overlay so `db` resolves. First-ever
|
||||
# deploy: the network doesn't exist yet — create it exactly as the stack will
|
||||
# (attachable overlay) so the name is adopted, then migrate, then deploy.
|
||||
NET="${STACK_NAME}_internal"
|
||||
# Never pre-create $NET: docker stack deploy must own it (a pre-existing
|
||||
# unlabeled network makes it fail with "already exists"). On first deploy the
|
||||
# migrate runs AFTER stack deploy instead (FIRST_DEPLOY_MIGRATE below).
|
||||
if [ "$SERVICE" = "backend" ] || [ "$SERVICE" = "all" ]; then
|
||||
if docker service inspect "${STACK_NAME}_db" >/dev/null 2>&1; then
|
||||
echo "==> One-shot migrate ($BACKEND_IMAGE)"
|
||||
docker run --rm --network "$NET" \
|
||||
-e THERMOGRAPH_DATABASE_URL="postgresql+asyncpg://thermograph:${POSTGRES_PASSWORD}@db:5432/thermograph" \
|
||||
--entrypoint /app/deploy/entrypoint.sh "$BACKEND_IMAGE" migrate
|
||||
else
|
||||
echo "==> First deploy: db not up yet; replicas will be rolled after stack deploy runs migrate below"
|
||||
fi
|
||||
fi
|
||||
|
||||
# --- deploy ----------------------------------------------------------------------
|
||||
FIRST_DEPLOY_MIGRATE=0
|
||||
docker service inspect "${STACK_NAME}_db" >/dev/null 2>&1 || FIRST_DEPLOY_MIGRATE=1
|
||||
if [ "$SERVICE" = "all" ] || ! docker service inspect "${STACK_NAME}_web" >/dev/null 2>&1; then
|
||||
echo "==> docker stack deploy ($STACK_NAME)"
|
||||
docker stack deploy --with-registry-auth -c "$APP_DIR/deploy/stack/thermograph-stack.yml" "$STACK_NAME"
|
||||
# First-ever deploy ran no migrate above (db didn't exist): wait for db,
|
||||
# migrate, then force web/worker to restart cleanly against the schema.
|
||||
if [ "${FIRST_DEPLOY_MIGRATE:-0}" = "1" ]; then
|
||||
echo "==> Waiting for db, then first-boot migrate"
|
||||
for i in $(seq 1 60); do
|
||||
cid=$(docker ps -q --filter "label=com.docker.swarm.service.name=${STACK_NAME}_db" | head -1)
|
||||
[ -n "$cid" ] && docker exec "$cid" pg_isready -U thermograph -d thermograph >/dev/null 2>&1 && break
|
||||
sleep 5
|
||||
done
|
||||
docker run --rm --network "$NET" \
|
||||
-e THERMOGRAPH_DATABASE_URL="postgresql+asyncpg://thermograph:${POSTGRES_PASSWORD}@db:5432/thermograph" \
|
||||
--entrypoint /app/deploy/entrypoint.sh "$BACKEND_IMAGE" migrate
|
||||
docker service update --force --detach=false "${STACK_NAME}_web"
|
||||
docker service update --force --detach=false "${STACK_NAME}_worker"
|
||||
fi
|
||||
else
|
||||
case "$SERVICE" in
|
||||
backend)
|
||||
echo "==> Rolling web + worker to $BACKEND_IMAGE"
|
||||
docker service update --with-registry-auth --detach=false --image "$BACKEND_IMAGE" "${STACK_NAME}_web"
|
||||
docker service update --with-registry-auth --detach=false --image "$BACKEND_IMAGE" "${STACK_NAME}_worker"
|
||||
;;
|
||||
frontend)
|
||||
echo "==> Rolling frontend to $FRONTEND_IMAGE"
|
||||
docker service update --with-registry-auth --detach=false --image "$FRONTEND_IMAGE" "${STACK_NAME}_frontend"
|
||||
;;
|
||||
esac
|
||||
fi
|
||||
|
||||
# --- loopback LB bridge -----------------------------------------------------------
|
||||
# A PLAIN container (only plain containers can bind 127.0.0.1; Swarm publishes
|
||||
# 0.0.0.0) on the attachable overlay, proxying to the service VIPs. Recreated
|
||||
# only when missing/dead — its config rarely changes; `docker rm -f $LB_NAME`
|
||||
# to force a refresh after editing lb/Caddyfile.
|
||||
if ! docker ps --format '{{.Names}}' | grep -qx "$LB_NAME"; then
|
||||
docker rm -f "$LB_NAME" >/dev/null 2>&1 || true
|
||||
echo "==> Starting loopback LB bridge $LB_NAME (127.0.0.1:$LB_HTTP_PORT, :$LB_FE_PORT)"
|
||||
docker run -d --name "$LB_NAME" --restart unless-stopped \
|
||||
--network "$NET" \
|
||||
-p "127.0.0.1:${LB_HTTP_PORT}:8137" -p "127.0.0.1:${LB_FE_PORT}:8080" \
|
||||
-v "$APP_DIR/deploy/stack/lb/Caddyfile:/etc/caddy/Caddyfile:ro" \
|
||||
caddy:2-alpine >/dev/null
|
||||
fi
|
||||
|
||||
# --- verify -----------------------------------------------------------------------
|
||||
echo "==> Health check via the LB"
|
||||
ok=0
|
||||
for i in $(seq 1 60); do
|
||||
if curl -fsS -m 3 -o /dev/null "http://127.0.0.1:${LB_HTTP_PORT}/healthz"; then ok=1; break; fi
|
||||
sleep 2
|
||||
done
|
||||
if [ "$ok" != 1 ]; then
|
||||
echo "!! stack health check failed" >&2
|
||||
docker stack ps "$STACK_NAME" --no-trunc | head -20
|
||||
exit 1
|
||||
fi
|
||||
echo "==> OK: $STACK_NAME serving on 127.0.0.1:${LB_HTTP_PORT}"
|
||||
|
||||
# Persist now-live tags (after health, like deploy.sh).
|
||||
cat > "$TAGS_FILE" <<EOF
|
||||
# Written by deploy-stack.sh -- the image tag each service is currently running.
|
||||
BACKEND_IMAGE_TAG=$BACKEND_IMAGE_TAG
|
||||
FRONTEND_IMAGE_TAG=$FRONTEND_IMAGE_TAG
|
||||
EOF
|
||||
|
||||
# GC superseded app-image tags (keep the running pair), same as deploy.sh.
|
||||
_be_repo="$REGISTRY_HOST/${BACKEND_IMAGE_PATH:-emi/thermograph-backend/app}"
|
||||
_fe_repo="$REGISTRY_HOST/${FRONTEND_IMAGE_PATH:-emi/thermograph-frontend/app}"
|
||||
docker images --format '{{.Repository}}:{{.Tag}}' \
|
||||
| grep -E "^(${_be_repo}|${_fe_repo}):" \
|
||||
| grep -v -e "^${_be_repo}:${BACKEND_IMAGE_TAG}$" -e "^${_fe_repo}:${FRONTEND_IMAGE_TAG}$" \
|
||||
| xargs -r docker rmi 2>/dev/null || true
|
||||
|
||||
# Post-deploy warm + IndexNow, via any web task (skip in test mode: no data,
|
||||
# and the warmer would burn upstream quota against an empty cache).
|
||||
if [ "${STACK_TEST:-0}" != "1" ] && { [ "$SERVICE" = backend ] || [ "$SERVICE" = all ]; }; then
|
||||
wcid=$(docker ps -q --filter "label=com.docker.swarm.service.name=${STACK_NAME}_web" | head -1)
|
||||
if [ -n "$wcid" ]; then
|
||||
echo "==> Warming city archives (detached) + IndexNow"
|
||||
docker exec -d "$wcid" sh -c 'python warm_cities.py --pace 2 >> /app/logs/warm-cities.log 2>&1' || true
|
||||
docker exec "$wcid" python indexnow.py --if-changed "${THERMOGRAPH_BASE_URL:-https://thermograph.org}" \
|
||||
|| echo "!! IndexNow ping failed (non-fatal)" >&2
|
||||
fi
|
||||
fi
|
||||
exit 0
|
||||
46
infra/deploy/stack/env-entrypoint.sh
Executable file
46
infra/deploy/stack/env-entrypoint.sh
Executable file
|
|
@ -0,0 +1,46 @@
|
|||
#!/usr/bin/env bash
|
||||
# Stack-task entrypoint shim: source the host-rendered secrets env, then hand
|
||||
# off to the image's real entrypoint.
|
||||
#
|
||||
# Why: `docker stack deploy` does not support compose's `env_file:`, and
|
||||
# enumerating every vault key in the stack yml's `environment:` blocks would
|
||||
# drift the moment a key is added to deploy/secrets/. Instead deploy-stack.sh
|
||||
# installs a uid-10001-readable copy of the rendered env at
|
||||
# /etc/thermograph/stack.env, the stack bind-mounts it (with this script) into
|
||||
# every app task, and this shim exports each KEY=value — but ONLY for keys not
|
||||
# already set, so the yml's `environment:` blocks keep compose's env_file
|
||||
# precedence (environment always wins). Same set-if-unset contract as the
|
||||
# image's own /run/secrets shim in deploy/entrypoint.sh, which still runs
|
||||
# after this and stays a no-op here.
|
||||
set -euo pipefail
|
||||
|
||||
ENV_FILE="${THERMOGRAPH_HOST_ENV:-/host/thermograph.env}"
|
||||
|
||||
if [ -f "$ENV_FILE" ]; then
|
||||
while IFS= read -r line || [ -n "$line" ]; do
|
||||
case "$line" in
|
||||
''|'#'*) continue ;;
|
||||
*=*)
|
||||
key="${line%%=*}"
|
||||
# Only sane identifiers; only if not already set by `environment:`.
|
||||
case "$key" in
|
||||
*[!A-Za-z0-9_]*|'') continue ;;
|
||||
esac
|
||||
if [ -z "${!key:-}" ]; then
|
||||
export "$key=${line#*=}"
|
||||
fi
|
||||
;;
|
||||
esac
|
||||
done < "$ENV_FILE"
|
||||
fi
|
||||
|
||||
# Hand off: the backend image has a real entrypoint script; the frontend
|
||||
# image is plain-CMD (docker passes that CMD to us as $@ when the stack
|
||||
# overrides the entrypoint) -- exec whichever this image actually is.
|
||||
if [ -x /app/deploy/entrypoint.sh ]; then
|
||||
exec /app/deploy/entrypoint.sh "$@"
|
||||
elif [ "$#" -gt 0 ]; then
|
||||
exec "$@"
|
||||
else
|
||||
exec uvicorn app:app --host 0.0.0.0 --port "${PORT:-8080}"
|
||||
fi
|
||||
30
infra/deploy/stack/lb/Caddyfile
Normal file
30
infra/deploy/stack/lb/Caddyfile
Normal file
|
|
@ -0,0 +1,30 @@
|
|||
# The loopback LB bridge's own config — NOT the host Caddy (that one still
|
||||
# terminates TLS for thermograph.org and proxies to 127.0.0.1:8137 exactly as
|
||||
# before; it needs no change for the stack cutover).
|
||||
#
|
||||
# This Caddy runs as a PLAIN container (deploy-stack.sh manages it) because
|
||||
# only plain containers can bind a specific host IP: Swarm port configs
|
||||
# publish on 0.0.0.0 (routing mesh or host mode alike), which would expose
|
||||
# the plaintext app un-fronted — hazard #6. It joins the stack's attachable
|
||||
# overlay and proxies to the service VIPs; Swarm's VIP round-robins across
|
||||
# however many web replicas the autoscaler is running, so this bridge never
|
||||
# needs to know the replica count.
|
||||
|
||||
{
|
||||
auto_https off
|
||||
admin off
|
||||
}
|
||||
|
||||
:8137 {
|
||||
reverse_proxy web:8137 {
|
||||
# Fail fast to the client if the VIP has no healthy task; Swarm's own
|
||||
# task healthchecks handle ejecting dead replicas from the VIP.
|
||||
lb_try_duration 5s
|
||||
}
|
||||
}
|
||||
|
||||
:8080 {
|
||||
reverse_proxy frontend:8080 {
|
||||
lb_try_duration 5s
|
||||
}
|
||||
}
|
||||
223
infra/deploy/stack/thermograph-stack.yml
Normal file
223
infra/deploy/stack/thermograph-stack.yml
Normal file
|
|
@ -0,0 +1,223 @@
|
|||
# Docker Swarm stack for prod — the autoscaling successor to docker-compose.yml
|
||||
# on that host (beta and LAN dev stay on plain compose; deploy.sh routes by the
|
||||
# /etc/thermograph/deploy-mode marker). Two-image world: web/worker run the
|
||||
# backend image, frontend its own — per-service tags, same contract as compose.
|
||||
#
|
||||
# Deployed by deploy/stack/deploy-stack.sh, which:
|
||||
# - sources /etc/thermograph.env (SOPS-rendered) so ${VARS} here interpolate,
|
||||
# - installs a uid-10001-readable copy at /etc/thermograph/stack.env that
|
||||
# deploy/stack/env-entrypoint.sh sources inside each app task (set-if-unset,
|
||||
# so `environment:` blocks below always win) — this replaces compose's
|
||||
# env_file:, which `docker stack deploy` does not support,
|
||||
# - runs migrations as a ONE-SHOT task before rolling (RUN_MIGRATIONS=0 in
|
||||
# every replica — N replicas must never race Alembic),
|
||||
# - manages the loopback LB bridge (see lb/README note below): Swarm's mesh
|
||||
# can only publish on 0.0.0.0 (would expose the plaintext app un-fronted),
|
||||
# so nothing here has `ports:`. A plain container on this attachable
|
||||
# overlay binds 127.0.0.1:8137/8080 for the host Caddy and proxies to the
|
||||
# service VIPs — Swarm's VIP does the actual load balancing across
|
||||
# replicas.
|
||||
#
|
||||
# Scaling model: `web` is stateless (ROLE=web never runs the notifier) and
|
||||
# scales 1..N — the autoscaler service adjusts replicas between
|
||||
# WEB_MIN_REPLICAS/WEB_MAX_REPLICAS on task CPU. `worker` owns the notifier +
|
||||
# scheduler: exactly 1 replica, with the cluster-wide Postgres advisory lock
|
||||
# (THERMOGRAPH_SINGLETON_PG) as belt-and-suspenders. `db` is exactly 1 — a
|
||||
# database does not scale by container count on one host; its levers are
|
||||
# DB_CPUS/DB_MEMORY (and, multi-host later, Patroni replicas per the topology
|
||||
# doc). Everything is pinned to the manager node: all volumes are local to
|
||||
# prod today. That constraint is the ONLY thing to relax when a second app
|
||||
# node joins.
|
||||
#
|
||||
# TIMESCALEDB_IMAGE must be the exact image (digest-pinned) the compose stack
|
||||
# was running — see hazard #7 in the hop-1 runbook: a floating tag can change
|
||||
# the extension minor under an existing volume. deploy-stack.sh resolves it
|
||||
# from the running/last-known container automatically.
|
||||
|
||||
services:
|
||||
db:
|
||||
image: ${TIMESCALEDB_IMAGE:?required}
|
||||
environment:
|
||||
POSTGRES_USER: thermograph
|
||||
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?required}
|
||||
POSTGRES_DB: thermograph
|
||||
DB_MEMORY: ${DB_MEMORY:-16g}
|
||||
volumes:
|
||||
- pgdata:/var/lib/postgresql
|
||||
- /opt/thermograph/deploy/db/init:/docker-entrypoint-initdb.d:ro
|
||||
networks:
|
||||
- internal
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "pg_isready -U thermograph -d thermograph"]
|
||||
interval: 5s
|
||||
timeout: 5s
|
||||
retries: 10
|
||||
deploy:
|
||||
replicas: 1
|
||||
placement:
|
||||
constraints: ["node.role == manager"]
|
||||
resources:
|
||||
limits:
|
||||
cpus: "${DB_CPUS:-4}"
|
||||
memory: ${DB_MEMORY:-16g}
|
||||
restart_policy:
|
||||
condition: on-failure
|
||||
|
||||
web:
|
||||
image: ${REGISTRY_HOST:-git.thermograph.org}/${BACKEND_IMAGE_PATH:-emi/thermograph-backend/app}:${BACKEND_IMAGE_TAG:?required}
|
||||
entrypoint: ["/host/env-entrypoint.sh"]
|
||||
environment:
|
||||
THERMOGRAPH_DATABASE_URL: postgresql+asyncpg://thermograph:${POSTGRES_PASSWORD}@db:5432/thermograph
|
||||
THERMOGRAPH_BASE: /
|
||||
PORT: 8137
|
||||
THERMOGRAPH_SERVICE_ROLE: backend
|
||||
THERMOGRAPH_FRONTEND_BASE_INTERNAL: http://frontend:8080
|
||||
WORKERS: ${WEB_WORKERS:-4}
|
||||
THERMOGRAPH_DATA_DIR: /state
|
||||
# web NEVER runs the notifier/scheduler, even if it would win election —
|
||||
# that's the worker service's job. This is what makes web replicas safe.
|
||||
THERMOGRAPH_ROLE: web
|
||||
RUN_MIGRATIONS: "0"
|
||||
# Overlay tasks reach the HOST's Postfix via the docker_gwbridge gateway,
|
||||
# not the compose bridge's 172.19.0.1 (an overlay has no host gateway).
|
||||
# provision-mail.sh's DOCKER_MAIL_GATEWAY/SUBNET cover this listener.
|
||||
THERMOGRAPH_SMTP_HOST: ${STACK_SMTP_HOST:-172.18.0.1}
|
||||
volumes:
|
||||
- appdata:/state
|
||||
- applogs:/app/logs
|
||||
- /opt/thermograph/deploy/stack/env-entrypoint.sh:/host/env-entrypoint.sh:ro
|
||||
- /etc/thermograph/stack.env:/host/thermograph.env:ro
|
||||
networks:
|
||||
- internal
|
||||
deploy:
|
||||
replicas: ${WEB_MIN_REPLICAS:-1}
|
||||
placement:
|
||||
constraints: ["node.role == manager"]
|
||||
resources:
|
||||
limits:
|
||||
cpus: "${WEB_CPUS:-4}"
|
||||
restart_policy:
|
||||
condition: on-failure
|
||||
update_config:
|
||||
# New task must pass the image's own HEALTHCHECK before the old one
|
||||
# stops — zero-downtime single-service rolls.
|
||||
order: start-first
|
||||
failure_action: rollback
|
||||
|
||||
worker:
|
||||
image: ${REGISTRY_HOST:-git.thermograph.org}/${BACKEND_IMAGE_PATH:-emi/thermograph-backend/app}:${BACKEND_IMAGE_TAG:?required}
|
||||
entrypoint: ["/host/env-entrypoint.sh"]
|
||||
environment:
|
||||
THERMOGRAPH_DATABASE_URL: postgresql+asyncpg://thermograph:${POSTGRES_PASSWORD}@db:5432/thermograph
|
||||
THERMOGRAPH_BASE: /
|
||||
PORT: 8137
|
||||
THERMOGRAPH_SERVICE_ROLE: backend
|
||||
THERMOGRAPH_FRONTEND_BASE_INTERNAL: http://frontend:8080
|
||||
WORKERS: "1"
|
||||
THERMOGRAPH_DATA_DIR: /state
|
||||
THERMOGRAPH_ROLE: worker
|
||||
# Cluster-wide Postgres advisory lock, not the host flock: correct at
|
||||
# replicas=1 today and stays correct if a second worker ever appears.
|
||||
THERMOGRAPH_SINGLETON_PG: "1"
|
||||
RUN_MIGRATIONS: "0"
|
||||
THERMOGRAPH_SMTP_HOST: ${STACK_SMTP_HOST:-172.18.0.1}
|
||||
volumes:
|
||||
- appdata:/state
|
||||
- applogs:/app/logs
|
||||
- /opt/thermograph/deploy/stack/env-entrypoint.sh:/host/env-entrypoint.sh:ro
|
||||
- /etc/thermograph/stack.env:/host/thermograph.env:ro
|
||||
networks:
|
||||
- internal
|
||||
deploy:
|
||||
replicas: 1
|
||||
placement:
|
||||
constraints: ["node.role == manager"]
|
||||
resources:
|
||||
limits:
|
||||
cpus: "${WORKER_CPUS:-2}"
|
||||
restart_policy:
|
||||
condition: on-failure
|
||||
|
||||
frontend:
|
||||
image: ${REGISTRY_HOST:-git.thermograph.org}/${FRONTEND_IMAGE_PATH:-emi/thermograph-frontend/app}:${FRONTEND_IMAGE_TAG:?required}
|
||||
entrypoint: ["/host/env-entrypoint.sh"]
|
||||
environment:
|
||||
THERMOGRAPH_BASE: /
|
||||
PORT: 8080
|
||||
THERMOGRAPH_SERVICE_ROLE: frontend
|
||||
THERMOGRAPH_API_BASE_INTERNAL: http://web:8137
|
||||
volumes:
|
||||
- /opt/thermograph/deploy/stack/env-entrypoint.sh:/host/env-entrypoint.sh:ro
|
||||
- /etc/thermograph/stack.env:/host/thermograph.env:ro
|
||||
networks:
|
||||
- internal
|
||||
deploy:
|
||||
replicas: 1
|
||||
placement:
|
||||
constraints: ["node.role == manager"]
|
||||
resources:
|
||||
limits:
|
||||
cpus: "${FRONTEND_CPUS:-2}"
|
||||
restart_policy:
|
||||
condition: on-failure
|
||||
update_config:
|
||||
order: start-first
|
||||
failure_action: rollback
|
||||
|
||||
# Scales `web` between WEB_MIN_REPLICAS and WEB_MAX_REPLICAS on sustained
|
||||
# task CPU (docker stats on this node — every task is pinned here today).
|
||||
# Deliberately a dumb shell loop with hysteresis + cooldown, not an
|
||||
# autoscaling framework: Swarm has no native autoscaler and this workload
|
||||
# doesn't justify one (topology doc §6 — declarative scaling as a feature).
|
||||
autoscaler:
|
||||
image: docker:27-cli
|
||||
entrypoint: ["/bin/sh", "/host/autoscale.sh"]
|
||||
environment:
|
||||
STACK_NAME: ${STACK_NAME:-thermograph}
|
||||
MIN_REPLICAS: ${WEB_MIN_REPLICAS:-1}
|
||||
MAX_REPLICAS: ${WEB_MAX_REPLICAS:-3}
|
||||
# Thresholds are avg docker-stats CPU% PER TASK (host-core-relative: 100
|
||||
# = one full core). Up fast, down slow.
|
||||
SCALE_UP_CPU: ${SCALE_UP_CPU:-220}
|
||||
SCALE_DOWN_CPU: ${SCALE_DOWN_CPU:-60}
|
||||
POLL_SECONDS: ${POLL_SECONDS:-15}
|
||||
UP_SAMPLES: ${UP_SAMPLES:-3}
|
||||
DOWN_SAMPLES: ${DOWN_SAMPLES:-20}
|
||||
COOLDOWN_SECONDS: ${COOLDOWN_SECONDS:-180}
|
||||
volumes:
|
||||
- /var/run/docker.sock:/var/run/docker.sock
|
||||
- /opt/thermograph/deploy/stack/autoscale.sh:/host/autoscale.sh:ro
|
||||
networks:
|
||||
- internal
|
||||
deploy:
|
||||
replicas: 1
|
||||
placement:
|
||||
constraints: ["node.role == manager"]
|
||||
resources:
|
||||
limits:
|
||||
cpus: "0.2"
|
||||
memory: 64m
|
||||
restart_policy:
|
||||
condition: any
|
||||
|
||||
networks:
|
||||
internal:
|
||||
driver: overlay
|
||||
# Attachable so the loopback LB bridge (a PLAIN container — only plain
|
||||
# containers can bind 127.0.0.1; Swarm port configs cannot) can join and
|
||||
# reach the service VIPs.
|
||||
attachable: true
|
||||
|
||||
volumes:
|
||||
# External and explicitly named: the real stack REUSES the volumes the
|
||||
# compose stack created (same data, zero migration). deploy-stack.sh's test
|
||||
# mode points these at throwaway names instead.
|
||||
pgdata:
|
||||
external: true
|
||||
name: ${PGDATA_VOLUME:-thermograph_pgdata}
|
||||
appdata:
|
||||
external: true
|
||||
name: ${APPDATA_VOLUME:-thermograph_appdata}
|
||||
applogs:
|
||||
external: true
|
||||
name: ${APPLOGS_VOLUME:-thermograph_applogs}
|
||||
83
infra/deploy/swarm/README.md
Normal file
83
infra/deploy/swarm/README.md
Normal file
|
|
@ -0,0 +1,83 @@
|
|||
# 3-node Docker Swarm (prod + beta + desktop), for hosting Forgejo
|
||||
|
||||
This Swarm cluster's only job is to run Forgejo (`deploy/forgejo/`) — it does
|
||||
**not** orchestrate the Thermograph app itself, which stays on the
|
||||
Terraform-managed `docker compose` deploys on prod/beta independently (see
|
||||
`terraform/README.md`). Keeping those separate means nothing here can strand
|
||||
or interfere with the app's already-working, single-writer Postgres/TimescaleDB
|
||||
deploys.
|
||||
|
||||
This is the canonical topology from
|
||||
`thermograph-docs/runbooks/implementation-handoff.md` (Track B steps 2-3) — three nodes,
|
||||
not two. An earlier revision of this doc/scripts covered just prod+beta;
|
||||
the desktop (this LAN dev machine) joins too.
|
||||
|
||||
**Nodes:**
|
||||
- **manager** — prod, the new 48 GB / 12-core box (more headroom).
|
||||
- **worker** — beta, the old VPS (`75.119.132.91`).
|
||||
- **worker** — desktop, this LAN dev machine (also runs the Forgejo Actions
|
||||
runner as a plain systemd service — see `deploy/forgejo/README.md` — not as
|
||||
a Swarm-scheduled container).
|
||||
|
||||
One manager, not more: Raft needs 3 nodes for real quorum-based HA, and this
|
||||
cluster only has 3 nodes total, so making even one more of them a manager
|
||||
would still fall short of real HA while adding split-brain risk. If the
|
||||
manager (prod) goes down, the workers keep running whatever was already
|
||||
scheduled on them (Forgejo, pinned to beta) but the cluster can't reschedule
|
||||
anything until prod's back — acceptable for a small cluster whose only job is
|
||||
CI/CD.
|
||||
|
||||
## Order of operations
|
||||
|
||||
1. **Agent access first** (`deploy/provision-agent-access.sh`) on prod and
|
||||
beta — everything below on those two boxes is run through that access. The
|
||||
desktop is wherever you're already working from; no separate access step
|
||||
needed there.
|
||||
2. **WireGuard mesh** (`setup-wireguard.sh <my_wg_ip> <peers_file>`) — run on
|
||||
**all three** nodes. See the script's header for the peer-list format and
|
||||
the two-pass key-exchange dance (pubkeys aren't known until every node has
|
||||
run it once). Verify with `ping <peer_wg_ip>` to each of the other two
|
||||
before continuing.
|
||||
3. **Swarm init** (`init-swarm.sh <manager_wg_ip>`) on the manager (prod) only.
|
||||
4. **Swarm join** (`join-swarm.sh <manager_wg_ip> <token>`) on **each** of the
|
||||
two workers (beta, desktop) — same token for both.
|
||||
5. **Firewall lockdown** (`firewall-swarm.sh`) on **all three** nodes — closes
|
||||
2377/7946/4789 to everything except the WireGuard interface. Do this
|
||||
*after* joining is confirmed working on all three, not before (locking the
|
||||
ports first would make the join itself fail).
|
||||
6. **Label beta** (`label-forge-node.sh <beta-node-name>`) on the manager —
|
||||
`docker node ls` shows each node's name/ID. Only beta gets `role=forge`;
|
||||
the desktop and prod don't need a Swarm label for anything in this setup.
|
||||
7. Deploy Forgejo: see `deploy/forgejo/README.md`.
|
||||
8. Register the Actions runner **on the desktop** (not through Swarm):
|
||||
`deploy/forgejo/register-lan-runner.sh`.
|
||||
|
||||
## Why WireGuard instead of relying on Swarm's built-in TLS alone
|
||||
|
||||
Swarm's control plane (port 2377) is TLS-encrypted and mutually authenticated
|
||||
by default. Its overlay data plane (VXLAN, port 4789) is **not** encrypted by
|
||||
default, and Docker's own guidance is that port must never face the public
|
||||
internet — these nodes are on different networks (two separate providers'
|
||||
public IPs, plus a home/LAN connection for the desktop), not one private LAN,
|
||||
so the tunnel is the network boundary the Swarm ports advertise into, rather
|
||||
than trusting the public internet (or the desktop's home network) directly.
|
||||
|
||||
## Verifying
|
||||
|
||||
```bash
|
||||
# On the manager:
|
||||
docker node ls # all three nodes Ready
|
||||
docker node inspect <beta-node> --format '{{.Spec.Labels}}' # role:forge
|
||||
|
||||
# From a FOURTH machine outside the mesh entirely, confirm the Swarm ports
|
||||
# are NOT reachable on either VPS's public IP (the desktop has no public IP
|
||||
# to check this way):
|
||||
nc -zv -w2 <prod_or_beta_public_ip> 2377 # should fail/timeout
|
||||
nc -zvu -w2 <prod_or_beta_public_ip> 4789 # should fail/timeout
|
||||
```
|
||||
|
||||
## Adding a node label back out (undo)
|
||||
|
||||
```bash
|
||||
docker node update --label-rm role <beta-node>
|
||||
```
|
||||
41
infra/deploy/swarm/firewall-swarm.sh
Executable file
41
infra/deploy/swarm/firewall-swarm.sh
Executable file
|
|
@ -0,0 +1,41 @@
|
|||
#!/usr/bin/env bash
|
||||
# Locks the Swarm ports (2377 control, 7946 gossip, 4789 overlay VXLAN) to the
|
||||
# WireGuard interface only — they must never be reachable from the public
|
||||
# internet. Run on BOTH boxes after joining the swarm. Existing app-facing
|
||||
# rules (80/443, SSH, etc.) are untouched.
|
||||
set -euo pipefail
|
||||
|
||||
WG_IFACE="${WG_IFACE:-wg0}"
|
||||
|
||||
if ! command -v ufw >/dev/null 2>&1; then
|
||||
echo "ufw not found — apply the equivalent iptables/nftables rules by hand:" >&2
|
||||
echo " allow 2377/tcp, 7946/tcp, 7946/udp, 4789/udp only on interface ${WG_IFACE}" >&2
|
||||
echo " deny 2377/tcp, 7946/tcp, 7946/udp, 4789/udp on every other interface" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "==> Allowing Swarm ports on ${WG_IFACE} only"
|
||||
ufw allow in on "$WG_IFACE" to any port 2377 proto tcp
|
||||
ufw allow in on "$WG_IFACE" to any port 7946 proto tcp
|
||||
ufw allow in on "$WG_IFACE" to any port 7946 proto udp
|
||||
ufw allow in on "$WG_IFACE" to any port 4789 proto udp
|
||||
|
||||
echo "==> Explicitly denying the same ports on every other interface"
|
||||
ufw deny 2377/tcp
|
||||
ufw deny 7946/tcp
|
||||
ufw deny 7946/udp
|
||||
ufw deny 4789/udp
|
||||
|
||||
echo
|
||||
echo "Rules added (not yet necessarily active — check ufw status):"
|
||||
ufw status numbered | grep -E "2377|7946|4789|Status" || true
|
||||
echo
|
||||
if ! ufw status | grep -q "^Status: active"; then
|
||||
echo "ufw is currently INACTIVE on this node — 'ufw enable' switches its"
|
||||
echo "default policy to deny-incoming for EVERYTHING, not just the Swarm"
|
||||
echo "ports above. Before enabling, explicitly allow every port this node"
|
||||
echo "already serves publicly (SSH at minimum; on beta specifically, also"
|
||||
echo "80/tcp and 443/tcp for the live thermograph.org Caddy) — check"
|
||||
echo "'ss -tlnp' for what's actually listening first. Enabling ufw without"
|
||||
echo "doing this WILL drop live traffic the moment it activates."
|
||||
fi
|
||||
29
infra/deploy/swarm/init-swarm.sh
Executable file
29
infra/deploy/swarm/init-swarm.sh
Executable file
|
|
@ -0,0 +1,29 @@
|
|||
#!/usr/bin/env bash
|
||||
# Run ONCE, on the manager node only (prod — the new 48 GB box). Initializes
|
||||
# the Swarm advertising the WireGuard address, so cluster traffic never
|
||||
# touches the public interface. Run setup-wireguard.sh on all THREE nodes
|
||||
# first (prod, beta, and the desktop — see
|
||||
# docs/runbooks/implementation-handoff.md Track B steps 2-3).
|
||||
set -euo pipefail
|
||||
|
||||
MY_WG_IP="${1:?usage: $0 <my_wg_ip>}"
|
||||
|
||||
if docker info 2>/dev/null | grep -q "Swarm: active"; then
|
||||
echo "Swarm is already active on this node. Current state:"
|
||||
docker node ls
|
||||
echo
|
||||
echo "Worker join token:"
|
||||
docker swarm join-token -q worker
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "==> docker swarm init, advertising ${MY_WG_IP} (the WireGuard address, not the public IP)"
|
||||
docker swarm init --advertise-addr "$MY_WG_IP" --listen-addr "${MY_WG_IP}:2377"
|
||||
|
||||
echo
|
||||
echo "==> Worker join command (run this on EACH of the two workers — beta and"
|
||||
echo " the desktop; the same token works for both):"
|
||||
TOKEN="$(docker swarm join-token -q worker)"
|
||||
echo " docker swarm join --token ${TOKEN} ${MY_WG_IP}:2377"
|
||||
echo
|
||||
echo "(Or run deploy/swarm/join-swarm.sh <manager_wg_ip> <token> on each worker.)"
|
||||
19
infra/deploy/swarm/join-swarm.sh
Executable file
19
infra/deploy/swarm/join-swarm.sh
Executable file
|
|
@ -0,0 +1,19 @@
|
|||
#!/usr/bin/env bash
|
||||
# Run ONCE on EACH worker node (beta, and the desktop — not the manager,
|
||||
# prod). Joins the Swarm initialized by init-swarm.sh, over the WireGuard
|
||||
# tunnel. The join token is the same for both workers.
|
||||
set -euo pipefail
|
||||
|
||||
MANAGER_WG_IP="${1:?usage: $0 <manager_wg_ip> <join_token>}"
|
||||
JOIN_TOKEN="${2:?}"
|
||||
|
||||
if docker info 2>/dev/null | grep -q "Swarm: active"; then
|
||||
echo "This node is already part of a swarm:"
|
||||
docker info 2>/dev/null | grep -A2 "Swarm:"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
docker swarm join --token "$JOIN_TOKEN" "${MANAGER_WG_IP}:2377"
|
||||
|
||||
echo
|
||||
echo "Joined. Verify from the manager: docker node ls"
|
||||
13
infra/deploy/swarm/label-forge-node.sh
Executable file
13
infra/deploy/swarm/label-forge-node.sh
Executable file
|
|
@ -0,0 +1,13 @@
|
|||
#!/usr/bin/env bash
|
||||
# Run ONCE, on the manager node, after both boxes have joined the swarm.
|
||||
# Labels the worker (beta) so the Forgejo stack's placement constraint
|
||||
# (node.labels.role == forge) schedules there and nowhere else.
|
||||
set -euo pipefail
|
||||
|
||||
WORKER_HOSTNAME="${1:?usage: $0 <worker_node_hostname_or_id> (see: docker node ls)}"
|
||||
|
||||
docker node update --label-add role=forge "$WORKER_HOSTNAME"
|
||||
|
||||
echo
|
||||
echo "Labels on ${WORKER_HOSTNAME}:"
|
||||
docker node inspect "$WORKER_HOSTNAME" --format '{{.Spec.Labels}}'
|
||||
115
infra/deploy/swarm/setup-wireguard.sh
Executable file
115
infra/deploy/swarm/setup-wireguard.sh
Executable file
|
|
@ -0,0 +1,115 @@
|
|||
#!/usr/bin/env bash
|
||||
# Sets up this node's side of a full-mesh WireGuard tunnel across all of
|
||||
# prod, beta, and the desktop (three nodes, not two — see
|
||||
# docs/runbooks/implementation-handoff.md Track B step 2). Docker Swarm's
|
||||
# control plane (2377/tcp) is TLS-encrypted by default, but the overlay data
|
||||
# plane (VXLAN, 4789/udp) is NOT — and it should never face the public
|
||||
# internet. Swarm joins over these private WireGuard IPs instead.
|
||||
#
|
||||
# A full mesh means every node peers directly with every other node (not a
|
||||
# hub-and-spoke through one box) — with three nodes that's 2 [Peer] blocks per
|
||||
# node. Run this on EACH of the three nodes, in any order, using a small peer
|
||||
# list file so the script isn't hardcoded to any particular node count.
|
||||
#
|
||||
# --- Peer list format --------------------------------------------------
|
||||
# A text file, one line per OTHER node (not including the one you're running
|
||||
# on), each line: <wg_ip> <public_ip> <pubkey-or-dash>
|
||||
# 10.10.0.1 169.58.46.181 <prod's pubkey, or - if not yet known>
|
||||
# 10.10.0.2 75.119.132.91 <beta's pubkey, or ->
|
||||
# 10.10.0.3 <desktop's public/reachable IP or a DDNS name> <desktop's pubkey, or ->
|
||||
#
|
||||
# First pass on any node: peer pubkeys you don't have yet are "-". Run this
|
||||
# script on all three nodes once (prints each node's own pubkey), fill those
|
||||
# into everyone's peer-list files, then re-run on all three once more to
|
||||
# actually establish the tunnels. Re-running is safe (idempotent) at any point.
|
||||
#
|
||||
# bash setup-wireguard.sh <my_wg_ip> <peers_file>
|
||||
set -euo pipefail
|
||||
|
||||
MY_WG_IP="${1:?usage: $0 <my_wg_ip> <peers_file>}"
|
||||
PEERS_FILE="${2:?}"
|
||||
WG_PORT="${WG_PORT:-51820}"
|
||||
WG_DIR=/etc/wireguard
|
||||
|
||||
if [ "$(id -u)" -ne 0 ]; then
|
||||
echo "Run as root (or via the agent user's sudo)." >&2
|
||||
exit 1
|
||||
fi
|
||||
if [ ! -f "$PEERS_FILE" ]; then
|
||||
echo "Peer list not found: $PEERS_FILE (see this script's header for the format)" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "==> Installing WireGuard if needed"
|
||||
command -v wg >/dev/null 2>&1 || { apt-get update -y -q && apt-get install -y -q wireguard; }
|
||||
|
||||
install -d -m 700 "$WG_DIR"
|
||||
if [ ! -f "$WG_DIR/privatekey" ]; then
|
||||
echo "==> Generating this node's WireGuard keypair"
|
||||
umask 077
|
||||
wg genkey | tee "$WG_DIR/privatekey" | wg pubkey > "$WG_DIR/publickey"
|
||||
fi
|
||||
MY_PRIVKEY="$(cat "$WG_DIR/privatekey")"
|
||||
MY_PUBKEY="$(cat "$WG_DIR/publickey")"
|
||||
|
||||
echo
|
||||
echo "==> This node's WireGuard public key (put this in the OTHER two nodes'"
|
||||
echo " peer-list files, replacing the '-' for this node's line):"
|
||||
echo " $MY_PUBKEY"
|
||||
echo
|
||||
|
||||
missing=0
|
||||
{
|
||||
echo "[Interface]"
|
||||
echo "Address = ${MY_WG_IP}/24"
|
||||
echo "PrivateKey = ${MY_PRIVKEY}"
|
||||
echo "ListenPort = ${WG_PORT}"
|
||||
echo
|
||||
|
||||
while read -r peer_wg_ip peer_public_ip peer_pubkey; do
|
||||
[ -z "${peer_wg_ip:-}" ] && continue
|
||||
case "$peer_wg_ip" in \#*) continue ;; esac
|
||||
if [ "$peer_pubkey" = "-" ] || [ -z "$peer_pubkey" ]; then
|
||||
echo "# SKIPPED (no pubkey yet) — peer at ${peer_public_ip}" >&2
|
||||
missing=$((missing + 1))
|
||||
continue
|
||||
fi
|
||||
echo "[Peer]"
|
||||
echo "PublicKey = ${peer_pubkey}"
|
||||
echo "Endpoint = ${peer_public_ip}:${WG_PORT}"
|
||||
echo "AllowedIPs = ${peer_wg_ip}/32"
|
||||
echo "PersistentKeepalive = 25"
|
||||
echo
|
||||
done < "$PEERS_FILE"
|
||||
} > "$WG_DIR/wg0.conf"
|
||||
chmod 600 "$WG_DIR/wg0.conf"
|
||||
|
||||
echo "==> Bringing up wg0"
|
||||
systemctl enable --now wg-quick@wg0 2>/dev/null || (wg-quick down wg0 2>/dev/null; wg-quick up wg0)
|
||||
wg syncconf wg0 <(wg-quick strip wg0) 2>/dev/null || true
|
||||
|
||||
echo "==> Firewall: only let listed peers' public IPs reach the WireGuard port"
|
||||
if command -v ufw >/dev/null 2>&1; then
|
||||
while read -r _wg_ip peer_public_ip _pubkey; do
|
||||
[ -z "${peer_public_ip:-}" ] && continue
|
||||
case "$peer_public_ip" in \#*) continue ;; esac
|
||||
ufw allow from "$peer_public_ip" to any port "$WG_PORT" proto udp comment "wireguard peer"
|
||||
done < "$PEERS_FILE"
|
||||
fi
|
||||
|
||||
echo
|
||||
echo "wg0 status:"
|
||||
wg show wg0 || true
|
||||
echo
|
||||
if [ "$missing" -gt 0 ]; then
|
||||
echo "NOTE: $missing peer(s) skipped (no pubkey yet in $PEERS_FILE). Fill them in"
|
||||
echo "and re-run on this node once all three nodes have printed their pubkey."
|
||||
else
|
||||
echo "==> All peers configured. Verify once every node has run this:"
|
||||
while read -r peer_wg_ip _peer_public_ip peer_pubkey; do
|
||||
[ -z "${peer_wg_ip:-}" ] && continue
|
||||
case "$peer_wg_ip" in \#*) continue ;; esac
|
||||
[ "$peer_pubkey" = "-" ] && continue
|
||||
echo " ping -c2 ${peer_wg_ip}"
|
||||
done < "$PEERS_FILE"
|
||||
fi
|
||||
20
infra/deploy/thermograph-dev.service
Normal file
20
infra/deploy/thermograph-dev.service
Normal file
|
|
@ -0,0 +1,20 @@
|
|||
# systemd --user unit for the Thermograph LAN dev server.
|
||||
# Installed by deploy/deploy-dev.sh into ~/.config/systemd/user/ with the
|
||||
# @PLACEHOLDER@ tokens substituted for this machine's paths/port.
|
||||
# Runs sudo-free as your own user; linger keeps it up across logout/reboot.
|
||||
[Unit]
|
||||
Description=Thermograph (dev, LAN) - FastAPI/uvicorn
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=exec
|
||||
WorkingDirectory=@APP_DIR@/backend
|
||||
Environment=THERMOGRAPH_BASE=@BASE@
|
||||
# Bind all interfaces so other devices on the LAN (e.g. your phone) can reach it.
|
||||
ExecStart=@APP_DIR@/.venv/bin/uvicorn app:app --host 0.0.0.0 --port @PORT@
|
||||
Restart=on-failure
|
||||
RestartSec=2
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
209
infra/deploy/thermograph.env.example
Normal file
209
infra/deploy/thermograph.env.example
Normal file
|
|
@ -0,0 +1,209 @@
|
|||
# This is the REFERENCE for what /etc/thermograph.env contains — the full set of
|
||||
# vars the app reads. On a host wired for the SOPS vault (an age key +
|
||||
# /etc/thermograph/secrets-env), you do NOT hand-edit /etc/thermograph.env: it is
|
||||
# rendered at deploy from the encrypted source of truth in deploy/secrets/*.yaml
|
||||
# (see deploy/secrets/README.md). To change a secret there, `sops edit` + commit +
|
||||
# deploy — not a host edit. (Legacy: Terraform can also render this file from tfvars,
|
||||
# and a from-scratch host can start by copying this example and editing it.)
|
||||
#
|
||||
# Sourced by deploy/deploy.sh (and any `docker compose` invocation) so compose can
|
||||
# interpolate it, AND loaded into the app container (env_file in docker-compose.yml).
|
||||
# Anything secret the app needs — Postgres password, VAPID keys, auth secret —
|
||||
# belongs here.
|
||||
|
||||
# Port uvicorn binds inside the container. The compose stack publishes it on the
|
||||
# host loopback (127.0.0.1:8137) for Caddy to proxy to. Keep 8137.
|
||||
PORT=8137
|
||||
|
||||
# --- Registry pull (deploy.sh / Terraform) ---------------------------------------
|
||||
# deploy.sh logs in to git.thermograph.org's registry and `docker compose pull`s
|
||||
# the image build-push.yml already pushed for the deployed commit (repo-split
|
||||
# Stage 6), instead of building in place. A personal access token with at least
|
||||
# read:package scope -- the same REGISTRY_TOKEN secret build-push.yml uses (that
|
||||
# one needs write:package too, to push; this only needs to pull).
|
||||
REGISTRY_TOKEN=
|
||||
|
||||
# --- PostgreSQL (docker-compose stack) ------------------------------------------
|
||||
# The app and Postgres run as a docker-compose stack (see docker-compose.yml).
|
||||
# POSTGRES_PASSWORD is the database password: compose uses it to initialize the
|
||||
# postgres container AND to build the app's THERMOGRAPH_DATABASE_URL. It MUST be
|
||||
# set here (the systemd unit sources this file so `docker compose up` can
|
||||
# interpolate it). Change it from the default before the first `up`.
|
||||
POSTGRES_PASSWORD=change-me
|
||||
|
||||
# The app's compose service already builds THERMOGRAPH_DATABASE_URL from
|
||||
# POSTGRES_PASSWORD, so you normally DON'T need this. It's here for reference and
|
||||
# for running the app outside compose against the same DB (keep the password in
|
||||
# sync with POSTGRES_PASSWORD above).
|
||||
#THERMOGRAPH_DATABASE_URL=postgresql+asyncpg://thermograph:change-me@db:5432/thermograph
|
||||
|
||||
# --- Historical archive (self-hosted Open-Meteo) --------------------------------
|
||||
# Where the app fetches its 45-year daily history. Unset (default) → the public
|
||||
# Open-Meteo archive API (rate-limited). On the self-hosting host, the
|
||||
# docker-compose.openmeteo.yml overlay sets this to the internal service URL for
|
||||
# you, so you normally DON'T set it here. Only pin it to run the app outside that
|
||||
# overlay against a reachable Open-Meteo instance. Leave it UNSET rather than empty:
|
||||
# an empty value is honored as-is and would break the fallback to the public API.
|
||||
#THERMOGRAPH_ARCHIVE_URL=http://open-meteo-api:8080/v1/archive
|
||||
|
||||
# Mark the session cookie Secure — required behind Caddy's HTTPS. Set to 1 in prod;
|
||||
# leave unset only for plain-HTTP LAN dev (a Secure cookie is never sent over HTTP).
|
||||
THERMOGRAPH_COOKIE_SECURE=1
|
||||
|
||||
# Pin these too (see their own sections below), so container restarts don't rotate
|
||||
# them: THERMOGRAPH_AUTH_SECRET (else every emailed confirm/reset link breaks on
|
||||
# restart) and THERMOGRAPH_VAPID_PRIVATE_KEY / _PUBLIC_KEY (else every existing push
|
||||
# subscription silently stops delivering). The data dir persists on the appdata
|
||||
# volume, but pinning here is the safe default.
|
||||
|
||||
# Number of uvicorn worker processes. More than 1 stops a single slow upstream fetch
|
||||
# (e.g. a cache-miss weather lookup) from blocking every other request — the cause of
|
||||
# past brief outages. Prod runs 4 (the compose app service also defaults to WORKERS=4);
|
||||
# leave unset (defaults to 1) on a small box. Workers elect one leader for the
|
||||
# subscription notifier via a lockfile (THERMOGRAPH_SINGLETON_LOCK, set by the compose
|
||||
# app service to /app/data/notifier.lock) so its timer-driven upstream sweep runs once,
|
||||
# not once per worker. ~200 MB RAM per worker.
|
||||
WORKERS=4
|
||||
|
||||
# THERMOGRAPH_SINGLETON_LOCK arbitrates workers on ONE host. Under multi-host Swarm,
|
||||
# each host would independently elect its own leader — multiplying the Open-Meteo
|
||||
# quota use N-fold again. Set THERMOGRAPH_SINGLETON_PG=1 (with THERMOGRAPH_DATABASE_URL
|
||||
# pointing at Postgres) to switch to a cluster-wide Postgres advisory lock instead, so
|
||||
# exactly one host — not one per host — runs the notifier. Leave unset on a single-host
|
||||
# deploy (today's default); the flock above is sufficient there.
|
||||
#THERMOGRAPH_SINGLETON_PG=1
|
||||
|
||||
# Which duties this process performs. Every replica runs the same image; ROLE just
|
||||
# decides whether it's allowed to own the notifier once it wins the leader election
|
||||
# above — this is what lets the web tier scale to N stateless replicas under Swarm
|
||||
# without also scaling notifier instances, while a single worker replica owns it.
|
||||
# all -> both (default; today's single-process behavior, unchanged)
|
||||
# web -> never runs the notifier, even if it would win leader election
|
||||
# worker -> runs the notifier if it wins leader election
|
||||
# Read once at process start; changing it needs a restart, not a live toggle.
|
||||
#THERMOGRAPH_ROLE=all
|
||||
|
||||
# The worker's own recurring jobs (city warming, IndexNow), gated by the same
|
||||
# leader election as the notifier above. Hours between runs; both are cheap
|
||||
# no-op skips when there's nothing to do, so the defaults rarely need changing.
|
||||
#THERMOGRAPH_WARM_CITIES_INTERVAL_HOURS=24
|
||||
#THERMOGRAPH_INDEXNOW_INTERVAL_HOURS=6
|
||||
|
||||
# Base path the app is served under.
|
||||
# / -> app at the domain root (Thermograph owns the whole domain —
|
||||
# this is what the Caddyfile expects: thermograph.org proxies "/")
|
||||
# /thermograph -> app under a sub-path (share the host with other apps, e.g. a
|
||||
# portfolio at the root)
|
||||
THERMOGRAPH_BASE=/
|
||||
|
||||
# --- SEO: search-engine verification + IndexNow ---------------------------------
|
||||
# Ownership-verification tokens, rendered as <meta> tags in every page's <head>.
|
||||
# Google Search Console → add property https://thermograph.org → "HTML tag" method
|
||||
# → paste just the content="…" value below. (Or verify via DNS TXT and skip this.)
|
||||
#THERMOGRAPH_GOOGLE_VERIFY=
|
||||
# Bing Webmaster Tools → add site → "HTML Meta Tag" (msvalidate.01) → paste the
|
||||
# content value. (Bing can also import verification from Google Search Console.)
|
||||
#THERMOGRAPH_BING_VERIFY=
|
||||
|
||||
# IndexNow key (Bing/DuckDuckGo/Yandex instant re-crawl). Auto-generated to
|
||||
# data/indexnow_key.txt on first use; set here to pin a specific key.
|
||||
#THERMOGRAPH_INDEXNOW_KEY=
|
||||
# Public site URL for IndexNow. The deploy hook auto-pings IndexNow after a
|
||||
# successful deploy, but only when the URL set changed (a new/removed city), so
|
||||
# code-only deploys don't resubmit. `make indexnow` forces a full submit.
|
||||
THERMOGRAPH_BASE_URL=https://thermograph.org
|
||||
|
||||
# --- Web Push (VAPID) -----------------------------------------------------------
|
||||
# Keys that sign push notifications. If unset, the app generates a pair into
|
||||
# data/vapid.json on first run — fine as long as that file PERSISTS (it lives in the
|
||||
# writable data dir and survives deploys). PIN them here to be safe: if the keys ever
|
||||
# change, every existing browser subscription silently stops receiving (the push
|
||||
# service rejects with 401/403), and users must toggle alerts off/on to re-subscribe.
|
||||
# Generate a pair: cd backend && ../.venv/bin/python -c "import push,json; k=push._generate(); print('PRIVATE=',k['private_key']); print('PUBLIC=',k['public_key'])"
|
||||
#THERMOGRAPH_VAPID_PRIVATE_KEY=
|
||||
#THERMOGRAPH_VAPID_PUBLIC_KEY=
|
||||
# Contact (mailto: or https URL) sent to push services in the VAPID claim.
|
||||
#THERMOGRAPH_VAPID_CONTACT=mailto:you@example.com
|
||||
|
||||
# --- Outbound email --------------------------------------------------------------
|
||||
# Delivery goes through the host's Postfix null client (deploy/provision-mail.sh).
|
||||
# The app runs in a container, so it can't reach the host's loopback — it speaks
|
||||
# plain SMTP to the compose bridge's gateway (172.19.0.1, pinned in
|
||||
# docker-compose.yml), where Postfix listens and relays out. Switching between
|
||||
# "direct to MX" and "relay through a provider" is a Postfix change, no redeploy.
|
||||
# Stack-mode hosts (prod) override this per-service to the docker_gwbridge
|
||||
# gateway (172.18.0.1) in deploy/stack/thermograph-stack.yml -- overlay tasks
|
||||
# can't reach a compose bridge gateway; leave this file's value as the compose
|
||||
# default.
|
||||
#
|
||||
# Backends: console (log it, send nothing — the default, right for dev),
|
||||
# smtp (actually send), disabled (drop silently).
|
||||
# Leave unset until Postfix is provisioned: signups are still collected either way.
|
||||
# THERMOGRAPH_MAIL_FROM is left unset here on purpose — the app default
|
||||
# (Thermograph <no-reply@thermograph.org>) is correct, and its "<>" would need
|
||||
# escaping in this shell-sourced file. Override only if the address differs.
|
||||
#THERMOGRAPH_MAIL_BACKEND=smtp
|
||||
#THERMOGRAPH_SMTP_HOST=172.19.0.1
|
||||
#THERMOGRAPH_SMTP_PORT=25
|
||||
# Only needed if talking to a remote SMTP server directly instead of local Postfix.
|
||||
#THERMOGRAPH_SMTP_USER=
|
||||
#THERMOGRAPH_SMTP_PASSWORD=
|
||||
#THERMOGRAPH_SMTP_STARTTLS=1
|
||||
#THERMOGRAPH_MAIL_FROM=Thermograph <no-reply@thermograph.org>
|
||||
#THERMOGRAPH_MAIL_REPLY_TO=
|
||||
|
||||
# Signing secret for email confirmation / password-reset tokens. MUST be set to a
|
||||
# fixed value before any such link is mailed: it defaults to a per-boot random
|
||||
# value, which would invalidate every outstanding link on each restart.
|
||||
# generate with: python -c "import secrets; print(secrets.token_urlsafe(48))"
|
||||
#THERMOGRAPH_AUTH_SECRET=
|
||||
|
||||
# --- Discord ---------------------------------------------------------------------
|
||||
# Incoming webhook URL for the daily "most unusual right now" post. Create it in the
|
||||
# Discord server: Channel → Edit → Integrations → Webhooks → New Webhook → Copy URL.
|
||||
# The URL IS the credential — anyone who has it can post as the webhook, so keep it
|
||||
# here and never in the repo. Unset => the daily post is disabled (no-op).
|
||||
# The post rides the notifier daemon (leader-only), once per day after the feed
|
||||
# refresh, so it needs THERMOGRAPH_ENABLE_NOTIFIER on (the default).
|
||||
#THERMOGRAPH_DISCORD_WEBHOOK=
|
||||
# Slash commands (/grade) are served over Discord's HTTP interactions endpoint by
|
||||
# the app itself (no bot process). Set the portal's "Interactions Endpoint URL" to
|
||||
# https://thermograph.org/discord/interactions. These come from the Developer Portal:
|
||||
# - PUBLIC_KEY: General Information -> Public Key (used to verify every request).
|
||||
# - APP_ID / BOT_TOKEN: only needed to (re)register the commands, via
|
||||
# scripts/register_discord_commands.py. The bot token is a credential.
|
||||
#THERMOGRAPH_DISCORD_PUBLIC_KEY=
|
||||
#THERMOGRAPH_DISCORD_APP_ID=
|
||||
#THERMOGRAPH_DISCORD_BOT_TOKEN=
|
||||
# Channel IDs (numeric snowflakes) for the bot to post into with the BOT_TOKEN above.
|
||||
# Subscription test feed: this deployment mirrors every notification the notifier
|
||||
# creates into this channel, so a run can be watched live. Set it per environment;
|
||||
# leave unset to disable. In the Thermograph.org server the hidden channels are:
|
||||
# dev = 1529274513009934336
|
||||
# uat = 1529274514066636861
|
||||
# prod = 1529274515127799940
|
||||
#THERMOGRAPH_DISCORD_SUBSCRIPTION_CHANNEL=
|
||||
# Notable-weather feed: prod broadcasts the daily "most unusual right now" digest
|
||||
# into this channel via the bot (independent of the WEBHOOK above — either can run
|
||||
# alone). Leave unset to disable. In the Thermograph.org server:
|
||||
# weather-events = 1529274516746932307
|
||||
#THERMOGRAPH_DISCORD_WEATHER_CHANNEL=
|
||||
# Discord gateway bot (@mention/DM grading): runs INSIDE the backend leader
|
||||
# process (same singleton election as the notifier — see web/app.py's lifespan),
|
||||
# riding its event loop; no separate bot process. Opt-in: any truthy value here
|
||||
# starts it, provided BOT_TOKEN above is also set (flag alone is a silent no-op).
|
||||
# Discord allows ONE gateway connection per bot token, so enable this in exactly
|
||||
# one environment — prod, the only vault with the token.
|
||||
#THERMOGRAPH_DISCORD_BOT=
|
||||
# Account linking (OAuth2 "identify"): lets a signed-in user connect their Discord
|
||||
# account, storing their Discord user id for DM alerts. CLIENT_ID is the same App
|
||||
# ID above. The CLIENT_SECRET is from OAuth2 -> Client Secret (a credential). In the
|
||||
# portal, add the redirect: https://thermograph.org/api/v2/discord/link/callback
|
||||
#THERMOGRAPH_DISCORD_CLIENT_SECRET=
|
||||
# Gateway bot (opt-in): holds a live websocket so the bot replies to messages that
|
||||
# @mention it (or DM it) with a city grade — e.g. "@Thermograph Phoenix". Reuses
|
||||
# THERMOGRAPH_DISCORD_BOT_TOKEN above. Runs on the single notifier leader only, so
|
||||
# it needs THERMOGRAPH_ENABLE_NOTIFIER on (the default). No privileged intent
|
||||
# required — Discord delivers content for mentions/DMs. Unset/0 => no gateway
|
||||
# connection. (Code: thermograph-backend notifications/discord_bot.py.)
|
||||
#THERMOGRAPH_DISCORD_BOT=1
|
||||
22
infra/deploy/thermograph.service
Normal file
22
infra/deploy/thermograph.service
Normal file
|
|
@ -0,0 +1,22 @@
|
|||
[Unit]
|
||||
Description=Thermograph docker-compose stack (app + PostgreSQL)
|
||||
# Docker must be up, and the network online, before compose can pull/interpolate.
|
||||
After=docker.service network-online.target
|
||||
Wants=network-online.target
|
||||
Requires=docker.service
|
||||
|
||||
[Service]
|
||||
# A thin manager for the compose stack: `up -d` starts it and returns, the unit
|
||||
# stays "active" (RemainAfterExit) so `systemctl restart thermograph` re-runs it.
|
||||
# Compose owns container ordering and restarts via depends_on + restart policies.
|
||||
Type=oneshot
|
||||
RemainAfterExit=yes
|
||||
WorkingDirectory=/opt/thermograph
|
||||
# POSTGRES_PASSWORD (and the other secrets) are needed at `docker compose`
|
||||
# interpolation time. The leading `-` makes a missing file non-fatal.
|
||||
EnvironmentFile=-/etc/thermograph.env
|
||||
ExecStart=/usr/bin/docker compose up -d
|
||||
ExecStop=/usr/bin/docker compose down
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
85
infra/deploy/twa/README.md
Normal file
85
infra/deploy/twa/README.md
Normal file
|
|
@ -0,0 +1,85 @@
|
|||
# Thermograph Android app (Trusted Web Activity)
|
||||
|
||||
The Android app is a **Trusted Web Activity (TWA)**: a thin native shell that opens
|
||||
`https://thermograph.org` full-screen in the user's Chrome engine. It *is* the live
|
||||
PWA — no content is duplicated, and the existing service worker + VAPID **web push
|
||||
keeps working** inside the app via Android notification delegation (no Firebase, no
|
||||
backend change). `enableNotifications: true` in `twa-manifest.json` turns that on
|
||||
and wires the Android 13+ `POST_NOTIFICATIONS` runtime permission.
|
||||
|
||||
This folder holds the build config. **Building the APK is a local/manual step** —
|
||||
it needs a JDK, the Android SDK, and a signing keystore that only you should hold.
|
||||
The current goal is a **signed APK you sideload for testing**; Play Store
|
||||
submission ($25 one-time) is deferred.
|
||||
|
||||
## What's in the repo
|
||||
- `twa-manifest.json` — Bubblewrap config (reference values; see comment inside).
|
||||
- `../../frontend/.well-known/assetlinks.json` — Digital Asset Links, served at
|
||||
`https://thermograph.org/.well-known/assetlinks.json`. Its fingerprint is a
|
||||
**placeholder** until you generate the signing key (step 4).
|
||||
|
||||
## One-time build (you)
|
||||
|
||||
Prereqs: Node 18+, JDK 17. Bubblewrap can download the Android SDK for you.
|
||||
|
||||
```bash
|
||||
npm install -g @bubblewrap/cli
|
||||
|
||||
# 1. Scaffold from the live manifest (creates ./twa-manifest.json + a keystore).
|
||||
# Accept defaults, then reconcile with this repo's twa-manifest.json — in
|
||||
# particular set: packageId = org.thermograph.twa, enableNotifications = true.
|
||||
bubblewrap init --manifest https://thermograph.org/manifest.webmanifest
|
||||
|
||||
# 2. (Or) copy this repo's config and let Bubblewrap fill the SDK bits:
|
||||
# cp deploy/twa/twa-manifest.json ./twa-manifest.json && bubblewrap update
|
||||
|
||||
# 3. Build the signed APK (and .aab for the store, for later).
|
||||
bubblewrap build
|
||||
# -> ./app-release-signed.apk and ./app-release-bundle.aab
|
||||
```
|
||||
|
||||
### 4. Wire Digital Asset Links (required — or the app opens with a URL bar)
|
||||
|
||||
```bash
|
||||
# Print the signing key's SHA-256 fingerprint:
|
||||
bubblewrap fingerprint list # or: keytool -list -v -keystore android-keystore.jks
|
||||
```
|
||||
|
||||
Copy the `SHA256` value into `frontend/.well-known/assetlinks.json`, replacing
|
||||
`REPLACE_WITH_TWA_SIGNING_KEY_SHA256_FINGERPRINT`, then **deploy the site** so the
|
||||
new file is live. Confirm:
|
||||
|
||||
```bash
|
||||
curl -s https://thermograph.org/.well-known/assetlinks.json # must show your fingerprint, Content-Type: application/json
|
||||
```
|
||||
|
||||
> If you later publish to Play Store with **Play App Signing**, add *Google's* app
|
||||
> signing certificate fingerprint (from the Play Console) to `assetlinks.json` too —
|
||||
> the upload-key fingerprint alone won't verify the store build.
|
||||
|
||||
### 5. Sideload and test
|
||||
|
||||
```bash
|
||||
adb install app-release-signed.apk
|
||||
```
|
||||
|
||||
On the device:
|
||||
- Launch Thermograph — it must open **full-screen with no browser address bar**
|
||||
(that confirms assetlinks verified).
|
||||
- Subscribe to a city (or trigger a test alert) and confirm the "unusual weather"
|
||||
push arrives as a **native Android notification**.
|
||||
- Grant the notifications permission when prompted (Android 13+).
|
||||
|
||||
## Keep the keystore safe
|
||||
`android-keystore.jks` is required for **every future update** and for the
|
||||
eventual Play Store listing. Losing it means a new package identity and a fresh
|
||||
install for every user. Store it out of the repo, backed up.
|
||||
|
||||
## Deferred
|
||||
- **Play Store listing** ($25 one-time Google Play developer account). Upload the
|
||||
`.aab`; the store review for a well-formed TWA is light. Not part of this pass.
|
||||
|
||||
## Why no backend changes
|
||||
Android delegates the TWA's web notifications to the OS, so `backend/push.py`
|
||||
(VAPID) and `backend/notify.py` deliver push to the app unchanged. Nothing in
|
||||
`backend/` is touched by this track.
|
||||
33
infra/deploy/twa/twa-manifest.json
Normal file
33
infra/deploy/twa/twa-manifest.json
Normal file
|
|
@ -0,0 +1,33 @@
|
|||
{
|
||||
"_comment": "Bubblewrap config for the Thermograph Trusted Web Activity. `bubblewrap init --manifest https://thermograph.org/manifest.webmanifest` scaffolds a file like this; the values below are the ones that matter for us (packageId must match /.well-known/assetlinks.json, enableNotifications wires web-push delegation, colors come from manifest.webmanifest). See README.md in this folder.",
|
||||
"packageId": "org.thermograph.twa",
|
||||
"host": "thermograph.org",
|
||||
"name": "Thermograph",
|
||||
"launcherName": "Thermograph",
|
||||
"display": "standalone",
|
||||
"orientation": "default",
|
||||
"themeColor": "#f0803c",
|
||||
"themeColorDark": "#0f1216",
|
||||
"navigationColor": "#171b21",
|
||||
"navigationColorDark": "#0f1216",
|
||||
"navigationColorDivider": "#2a323c",
|
||||
"navigationColorDividerDark": "#2a323c",
|
||||
"backgroundColor": "#171b21",
|
||||
"enableNotifications": true,
|
||||
"startUrl": "/",
|
||||
"webManifestUrl": "https://thermograph.org/manifest.webmanifest",
|
||||
"iconUrl": "https://thermograph.org/logo.png?v=3",
|
||||
"maskableIconUrl": "https://thermograph.org/logo-maskable-512.png?v=3",
|
||||
"monochromeIconUrl": "https://thermograph.org/favicon.svg?v=3",
|
||||
"fallbackType": "customtabs",
|
||||
"features": {},
|
||||
"signingKey": {
|
||||
"path": "./android-keystore.jks",
|
||||
"alias": "thermograph"
|
||||
},
|
||||
"appVersionName": "1.0.0",
|
||||
"appVersionCode": 1,
|
||||
"minSdkVersion": 23,
|
||||
"shortcuts": [],
|
||||
"generatorApp": "bubblewrap-cli"
|
||||
}
|
||||
47
infra/docker-compose.dev.yml
Normal file
47
infra/docker-compose.dev.yml
Normal file
|
|
@ -0,0 +1,47 @@
|
|||
# Dev overlay for the LAN dev server (deploy/deploy-dev.sh):
|
||||
#
|
||||
# docker compose -f docker-compose.yml -f docker-compose.dev.yml up -d
|
||||
#
|
||||
# Split-repo adaptation: the monorepo's docker-compose.dev.yml (see
|
||||
# thermograph/docker-compose.dev.yml) overlaid a base file where backend and
|
||||
# frontend had `build: .` -- dev's whole point there was `--build` in place from
|
||||
# the working checkout. This infra repo holds no Dockerfile at all (each service's
|
||||
# lives in its own app repo, thermograph-backend / thermograph-frontend), so the
|
||||
# base docker-compose.yml already has NO `build:` for either service, only
|
||||
# `image: .../${BACKEND_IMAGE_PATH}:${BACKEND_IMAGE_TAG}` (and the frontend
|
||||
# equivalent) -- same registry-pull model as prod/beta, just pointed at a dev tag
|
||||
# by deploy-dev.sh. This overlay must NOT reintroduce `build:`; it only relaxes
|
||||
# resource caps and LAN-exposes a port, same as the monorepo overlay did.
|
||||
#
|
||||
# Differences from the prod stack (unchanged intent from the monorepo overlay):
|
||||
# 1. backend is published on ALL interfaces (0.0.0.0:8137), not loopback, so
|
||||
# phones and other devices on the Wi-Fi can reach the dev server directly --
|
||||
# dev has no Caddy in front (prod does, which is why the base file binds
|
||||
# 127.0.0.1 only).
|
||||
# 2. frontend's port publish is dropped entirely -- dev has no Caddy to reach it
|
||||
# directly, so it stays compose-internal-only, reached solely through
|
||||
# backend's own reverse-proxy fallback (THERMOGRAPH_FRONTEND_BASE_INTERNAL,
|
||||
# see backend/web/app.py's _proxy_to_frontend in the backend repo). This
|
||||
# keeps the dev stack serving the one URL it always has, at :8137.
|
||||
# 3. The CPU caps are removed -- dev runs UNCAPPED (no thread/CPU limits), unlike
|
||||
# prod's backend=4 / frontend=2 / db=2 allocation. `!reset` drops the base
|
||||
# value (both the top-level `cpus:` and the Swarm-style `deploy.resources`
|
||||
# block the base file carries for parity).
|
||||
services:
|
||||
backend:
|
||||
ports: !override
|
||||
- "8137:8137"
|
||||
cpus: !reset null
|
||||
deploy: !reset null
|
||||
frontend:
|
||||
ports: !reset null
|
||||
cpus: !reset null
|
||||
deploy: !reset null
|
||||
db:
|
||||
cpus: !reset null
|
||||
deploy: !reset null
|
||||
# Uncapped memory on dev too (prod ceilings it at 8g). The Postgres/DuckDB
|
||||
# memory *budget* is still the ~8 GB derived by deploy/db/init/20-tuning.sh from
|
||||
# the default DB_MEMORY (raise DB_MEMORY to give dev more); shm_size stays (it's
|
||||
# required shared memory for parallel query, not a limit).
|
||||
mem_limit: !reset null
|
||||
72
infra/docker-compose.openmeteo.yml
Normal file
72
infra/docker-compose.openmeteo.yml
Normal file
|
|
@ -0,0 +1,72 @@
|
|||
# Self-hosted Open-Meteo overlay — serves the ERA5 archive locally so the app is
|
||||
# off the rate-limited public archive API. Enable it only on the self-hosting host:
|
||||
#
|
||||
# docker compose -f docker-compose.yml -f docker-compose.openmeteo.yml up -d
|
||||
#
|
||||
# The hourly .om data (copernicus_era5_land 0.1° + copernicus_era5 0.25°) lives in
|
||||
# OBJECT STORAGE, surfaced on the host as a local directory by an rclone FUSE mount
|
||||
# (a host systemd unit — see deploy/openmeteo/README.md). OM_DATA_DIR points the
|
||||
# containers at that mount; it defaults to ./data/om-archive so a local smoke test
|
||||
# works against a plain directory. Object storage is never bind-mounted into the app
|
||||
# — only Open-Meteo reads it; the app just talks HTTP to open-meteo-api and keeps its
|
||||
# own daily-per-cell climate record in the TimescaleDB `db` service.
|
||||
#
|
||||
# One-time backfill (writes ~1–1.5 TB of .om to object storage — run once before the
|
||||
# app is flipped over): `make om-backfill`.
|
||||
|
||||
services:
|
||||
# Serves /v1/archive on the compose network. No host port: only `backend`
|
||||
# reaches it, as http://open-meteo-api:8080. Reads .om from the object-storage
|
||||
# mount, on demand for any point, returning JSON byte-identical to the public
|
||||
# archive API.
|
||||
open-meteo-api:
|
||||
image: ${OM_IMAGE:-ghcr.io/open-meteo/open-meteo}
|
||||
command: ["serve", "--env", "production", "--hostname", "0.0.0.0", "--port", "8080"]
|
||||
volumes:
|
||||
- ${OM_DATA_DIR:-./data/om-archive}:/app/data
|
||||
restart: unless-stopped
|
||||
|
||||
# Rolling keep-current worker for ERA5-Land (0.1°): the surface variables it
|
||||
# carries. --past-days 14 re-syncs the recent tail every day, appending new days and
|
||||
# absorbing ERA5T→final corrections. dew_point_2m → relative_humidity and
|
||||
# shortwave_radiation → apparent_temperature are derived server-side at query time.
|
||||
open-meteo-sync-land:
|
||||
image: ${OM_IMAGE:-ghcr.io/open-meteo/open-meteo}
|
||||
command:
|
||||
- "sync"
|
||||
- "copernicus_era5_land"
|
||||
- "temperature_2m,dew_point_2m,precipitation,shortwave_radiation,wind_u_component_10m,wind_v_component_10m"
|
||||
- "--past-days"
|
||||
- "${OM_SYNC_PAST_DAYS:-14}"
|
||||
- "--repeat-interval"
|
||||
- "1440"
|
||||
volumes:
|
||||
- ${OM_DATA_DIR:-./data/om-archive}:/app/data
|
||||
restart: unless-stopped
|
||||
|
||||
# Rolling keep-current worker for ERA5 (0.25°): wind gusts — absent from ERA5-Land —
|
||||
# plus the seamless-blend fallback over water/coastline where ERA5-Land has no data.
|
||||
open-meteo-sync-era5:
|
||||
image: ${OM_IMAGE:-ghcr.io/open-meteo/open-meteo}
|
||||
command:
|
||||
- "sync"
|
||||
- "copernicus_era5"
|
||||
- "wind_gusts_10m,temperature_2m,dew_point_2m,precipitation,shortwave_radiation,wind_u_component_10m,wind_v_component_10m"
|
||||
- "--past-days"
|
||||
- "${OM_SYNC_PAST_DAYS:-14}"
|
||||
- "--repeat-interval"
|
||||
- "1440"
|
||||
volumes:
|
||||
- ${OM_DATA_DIR:-./data/om-archive}:/app/data
|
||||
restart: unless-stopped
|
||||
|
||||
# Point backend's historical fetches at the local instance (frontend never
|
||||
# fetches climate data itself). Set here (not in the base file) so it
|
||||
# applies only when this overlay is active — and so an unset var never
|
||||
# reaches the app as an empty string, which would defeat climate.py's default.
|
||||
backend:
|
||||
depends_on:
|
||||
open-meteo-api:
|
||||
condition: service_started
|
||||
environment:
|
||||
THERMOGRAPH_ARCHIVE_URL: http://open-meteo-api:8080/v1/archive
|
||||
196
infra/docker-compose.yml
Normal file
196
infra/docker-compose.yml
Normal file
|
|
@ -0,0 +1,196 @@
|
|||
# Thermograph production stack: backend (FastAPI/API + TestClient(app.py)),
|
||||
# frontend (the SSR content service), and TimescaleDB (PostgreSQL 18).
|
||||
#
|
||||
# FE/BE CI-CD split (finishes repo-split Stage 6/7): backend and frontend are
|
||||
# two containers, each running its OWN image published by its OWN repo's
|
||||
# build-push.yml -- backend = ${BACKEND_IMAGE_PATH}, frontend =
|
||||
# ${FRONTEND_IMAGE_PATH}, tagged independently by BACKEND_IMAGE_TAG /
|
||||
# FRONTEND_IMAGE_TAG. This replaces the earlier Stage-4 model where both
|
||||
# containers shared the single emi/thermograph/app image and
|
||||
# THERMOGRAPH_SERVICE_ROLE picked the process; the split Dockerfiles now start
|
||||
# the right process directly (frontend `uvicorn app:app`, backend
|
||||
# entrypoint.sh), so a backend deploy and a frontend deploy are fully
|
||||
# independent -- deploy.sh rolls one service without touching the other's tag.
|
||||
# Both get their own loopback-published port since prod/beta's host Caddy
|
||||
# path-splits directly to each; backend also gets a reverse-proxy fallback to
|
||||
# frontend (THERMOGRAPH_FRONTEND_BASE_INTERNAL) for any environment with no
|
||||
# Caddy in front (LAN dev, bare-metal run.sh) -- see backend/web/app.py's
|
||||
# _proxy_to_frontend.
|
||||
#
|
||||
# docker compose up -d --build # or: make up
|
||||
#
|
||||
# POSTGRES_PASSWORD must be set at `docker compose` time — compose reads it from
|
||||
# the repo-root .env for local runs (copy .env.example -> .env), and in prod the
|
||||
# systemd unit's EnvironmentFile=/etc/thermograph.env puts it in the environment
|
||||
# so `docker compose up` can interpolate it. It is used BOTH to initialize the db
|
||||
# container and to build backend's THERMOGRAPH_DATABASE_URL below.
|
||||
|
||||
services:
|
||||
db:
|
||||
# TimescaleDB on PostgreSQL 18 (the stock image already sets
|
||||
# shared_preload_libraries=timescaledb). The app's climate record — the full
|
||||
# daily archive and the hourly recent+forecast bundle — lives in hypertables
|
||||
# here (see backend/data/climate_store.py), so the DB, not the filesystem, is the
|
||||
# source of truth. The init script CREATE EXTENSIONs timescaledb on a fresh volume
|
||||
# (Alembic also does, idempotently, at app boot).
|
||||
#
|
||||
# TIMESCALEDB_TAG defaults to the floating latest-pg18 tag (today's behavior,
|
||||
# unchanged) so a plain `docker compose up` keeps working with no setup. Pin it
|
||||
# to an exact minor (e.g. 2.17.2-pg18) before any host of this stack could ever
|
||||
# replicate with another — a floating tag risks two hosts landing on different
|
||||
# extension minors, which blocks a physical replica and risks compressed-chunk
|
||||
# corruption on restore. docker-stack.yml (the Swarm interim stack) REQUIRES an
|
||||
# exact pin for exactly this reason; use the SAME tag on both once you set one.
|
||||
image: timescale/timescaledb:${TIMESCALEDB_TAG:-latest-pg18}
|
||||
environment:
|
||||
POSTGRES_USER: thermograph
|
||||
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD}
|
||||
POSTGRES_DB: thermograph
|
||||
# The init tuning script (deploy/db/init/20-tuning.sh) scales the Postgres +
|
||||
# DuckDB memory GUCs from this budget, matching the mem_limit below. One knob
|
||||
# per host: Terraform sets DB_MEMORY (prod 16g), local/beta default 8g.
|
||||
DB_MEMORY: ${DB_MEMORY:-8g}
|
||||
volumes:
|
||||
# Mount the volume at the PARENT of the data dir and let the image pick its own
|
||||
# PGDATA subdir under it (the timescaledb image defaults to
|
||||
# /var/lib/postgresql/data). Pinning PGDATA directly at the mountpoint trips an
|
||||
# initdb chmod on some Docker setups; the whole tree still persists this way.
|
||||
- pgdata:/var/lib/postgresql
|
||||
- ./deploy/db/init:/docker-entrypoint-initdb.d
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "pg_isready -U thermograph -d thermograph"]
|
||||
interval: 5s
|
||||
timeout: 5s
|
||||
retries: 10
|
||||
# Give Postgres room to cache + process. mem_limit is the hard ceiling; the actual
|
||||
# budget is tuned in deploy/db/init/20-tuning.sh, which scales shared_buffers (25%),
|
||||
# effective_cache_size (75%), work_mem and maintenance_work_mem from DB_MEMORY — so
|
||||
# raising DB_MEMORY raises both the cap and the tuning together. shm_size backs
|
||||
# parallel-query shared memory (the 64MB docker default is
|
||||
# too small once shared_buffers/parallelism grow).
|
||||
# Sized via env (Terraform sets DB_CPUS/DB_MEMORY per host); defaults match the
|
||||
# historical 2 CPU / 8 GB budget so a plain `docker compose up` is unchanged.
|
||||
mem_limit: ${DB_MEMORY:-8g}
|
||||
shm_size: 1gb
|
||||
# Cap the DB at DB_CPUS CPUs. Compose v2 honors the top-level `cpus:`; the
|
||||
# deploy.resources block is the Swarm-style equivalent, kept for parity.
|
||||
cpus: ${DB_CPUS:-2}
|
||||
deploy:
|
||||
resources:
|
||||
limits:
|
||||
cpus: "${DB_CPUS:-2}"
|
||||
# Must match mem_limit above (compose rejects distinct values).
|
||||
memory: ${DB_MEMORY:-8g}
|
||||
restart: unless-stopped
|
||||
# No host port on purpose: the app reaches Postgres as db:5432 on the
|
||||
# compose network. Nothing outside the stack should touch the database.
|
||||
|
||||
backend:
|
||||
# Backend's OWN image, published by thermograph-backend's build-push.yml.
|
||||
# No `build:` here -- infra holds no Dockerfile; each service's Dockerfile
|
||||
# lives in its own app repo. deploy.sh sets BACKEND_IMAGE_TAG to the SHA
|
||||
# build-push.yml pushed for the backend commit being deployed; local dev
|
||||
# builds `emi/thermograph-backend/app:local` from the backend repo (see
|
||||
# infra Makefile) and this pulls/uses it. BACKEND_IMAGE_PATH/TAG are
|
||||
# independent of the frontend's, so the two services deploy separately.
|
||||
image: ${REGISTRY_HOST:-git.thermograph.org}/${BACKEND_IMAGE_PATH:-emi/thermograph-backend/app}:${BACKEND_IMAGE_TAG:-local}
|
||||
depends_on:
|
||||
db:
|
||||
condition: service_healthy
|
||||
environment:
|
||||
# Built from POSTGRES_PASSWORD; this `environment` value wins over anything
|
||||
# in env_file, so the URL always matches the db container's password.
|
||||
THERMOGRAPH_DATABASE_URL: postgresql+asyncpg://thermograph:${POSTGRES_PASSWORD}@db:5432/thermograph
|
||||
THERMOGRAPH_BASE: /
|
||||
PORT: 8137
|
||||
THERMOGRAPH_SERVICE_ROLE: backend
|
||||
# Docker's built-in service-name DNS -- reachable only inside the compose
|
||||
# network. See backend/web/app.py's _proxy_to_frontend.
|
||||
THERMOGRAPH_FRONTEND_BASE_INTERNAL: http://frontend:8080
|
||||
# Worker count is env-driven so Terraform can raise it on a bigger host;
|
||||
# defaults to 4 to keep a plain `docker compose up` identical to before.
|
||||
WORKERS: ${WORKERS:-4}
|
||||
# One worker wins this lock and runs the subscription notifier / homepage
|
||||
# sweep; it lives on the appdata volume so it's shared across workers.
|
||||
THERMOGRAPH_DATA_DIR: /state
|
||||
# notifier.lock lives in the state volume, which is now mounted at /state
|
||||
# (NOT /app/data) so it never shadows the Python `data/` package -- see the
|
||||
# volumes: note below and thermograph-backend paths.py.
|
||||
THERMOGRAPH_SINGLETON_LOCK: /state/notifier.lock
|
||||
# Prod secrets live in /etc/thermograph.env: POSTGRES_PASSWORD,
|
||||
# THERMOGRAPH_AUTH_SECRET, THERMOGRAPH_VAPID_PRIVATE_KEY/_PUBLIC_KEY,
|
||||
# THERMOGRAPH_COOKIE_SECURE=1, mail/Discord keys, ... (see
|
||||
# deploy/thermograph.env.example). `required: false` so local `docker compose
|
||||
# up` works without that file — it reads POSTGRES_PASSWORD from repo-root .env.
|
||||
env_file:
|
||||
- path: /etc/thermograph.env
|
||||
required: false
|
||||
volumes:
|
||||
# Parquet cache, notifier.lock, homepage.json, vapid.json persist here.
|
||||
# /state, NOT /app/data: after the repo split the backend's Python package
|
||||
# `data/` sits at /app/data, so mounting the runtime volume there erased
|
||||
# data/*.py and broke `import data.climate` at boot. THERMOGRAPH_DATA_DIR=/state
|
||||
# (above) points runtime state here instead, clear of the code.
|
||||
- appdata:/state
|
||||
- applogs:/app/logs
|
||||
# No compose-level healthcheck override -- the image's own Dockerfile
|
||||
# HEALTHCHECK (port-aware via ${PORT}) already covers this, and frontend's
|
||||
# depends_on below reads it the same way.
|
||||
# Sized via env (Terraform sets APP_CPUS per host); defaults to 4 CPUs. The
|
||||
# deploy.resources block mirrors the top-level `cpus:` for Swarm parity; the
|
||||
# dev overlay drops both so dev runs uncapped.
|
||||
cpus: ${APP_CPUS:-4}
|
||||
deploy:
|
||||
resources:
|
||||
limits:
|
||||
cpus: "${APP_CPUS:-4}"
|
||||
ports:
|
||||
# Loopback only — host Caddy terminates TLS and reverse-proxies to this
|
||||
# (and, in prod/beta, directly to frontend's own port below too).
|
||||
- "127.0.0.1:8137:8137"
|
||||
restart: unless-stopped
|
||||
|
||||
frontend:
|
||||
# Frontend's OWN image, published by thermograph-frontend's build-push.yml
|
||||
# from its own Dockerfile (which starts `uvicorn app:app` directly -- no
|
||||
# THERMOGRAPH_SERVICE_ROLE process-picking anymore). Independent
|
||||
# FRONTEND_IMAGE_TAG so a frontend deploy never disturbs the backend's tag.
|
||||
image: ${REGISTRY_HOST:-git.thermograph.org}/${FRONTEND_IMAGE_PATH:-emi/thermograph-frontend/app}:${FRONTEND_IMAGE_TAG:-local}
|
||||
# Its own register() fetches the IndexNow key from backend at boot -- must
|
||||
# wait for a real, healthy backend, not just a started container.
|
||||
depends_on:
|
||||
backend:
|
||||
condition: service_healthy
|
||||
environment:
|
||||
THERMOGRAPH_BASE: /
|
||||
PORT: 8080
|
||||
THERMOGRAPH_SERVICE_ROLE: frontend
|
||||
THERMOGRAPH_API_BASE_INTERNAL: http://backend:8137
|
||||
# No volumes -- stateless, holds no data of its own.
|
||||
cpus: ${FRONTEND_CPUS:-2}
|
||||
deploy:
|
||||
resources:
|
||||
limits:
|
||||
cpus: "${FRONTEND_CPUS:-2}"
|
||||
ports:
|
||||
- "127.0.0.1:8080:8080"
|
||||
restart: unless-stopped
|
||||
|
||||
volumes:
|
||||
pgdata: {}
|
||||
appdata: {}
|
||||
applogs: {}
|
||||
|
||||
networks:
|
||||
# Pin the default network's subnet + gateway so the host Postfix can rely on a
|
||||
# stable address. The app sends verification email by speaking SMTP to the
|
||||
# gateway (172.19.0.1), where the host Postfix listens and relays out (see
|
||||
# deploy/provision-mail.sh + THERMOGRAPH_SMTP_HOST in thermograph.env). Without
|
||||
# the pin, Docker picks a subnet from its pool and the gateway could move.
|
||||
# 172.19.0.0/16 matches what prod and beta are live on today, so applying this
|
||||
# is a no-op there.
|
||||
default:
|
||||
ipam:
|
||||
config:
|
||||
- subnet: 172.19.0.0/16
|
||||
gateway: 172.19.0.1
|
||||
173
infra/docker-stack.yml
Normal file
173
infra/docker-stack.yml
Normal file
|
|
@ -0,0 +1,173 @@
|
|||
# Docker Swarm stack for the hop-1 interim cutover — see
|
||||
# thermograph-docs/runbooks/hop1-forgejo-registry-cutover.md (which this file implements) and
|
||||
# thermograph-docs/architecture/repo-topology-and-infrastructure.md §6/§9.
|
||||
#
|
||||
# Distinct from docker-compose.yml (today's plain-compose deploy, unaffected by
|
||||
# this file): Swarm pulls a pre-built image from the registry (IMAGE_TAG) rather
|
||||
# than building in place, and needs Swarm-specific keys (deploy:, networks:,
|
||||
# secrets:, no host-published ports on the overlay-fronted services).
|
||||
#
|
||||
# docker stack deploy -c docker-stack.yml thermograph
|
||||
#
|
||||
# `docker stack deploy` only interpolates from the invoking shell's environment,
|
||||
# not a .env file — export the required vars first (or `set -a; . ./stack.env;
|
||||
# set +a; docker stack deploy ...`). Required shell vars: IMAGE_TAG (the tag CI
|
||||
# pushed to the registry), TIMESCALEDB_TAG (an EXACT pinned minor — see the db
|
||||
# service below, hazard #7). POSTGRES_PASSWORD is NOT a shell var here: unlike
|
||||
# docker-compose.yml (which interpolates it into THERMOGRAPH_DATABASE_URL at
|
||||
# compose time), this file reads it as the `postgres_password` Swarm secret —
|
||||
# POSTGRES_PASSWORD_FILE for db (native support in the postgres/timescaledb
|
||||
# image), and the chunk-3 entrypoint shim builds THERMOGRAPH_DATABASE_URL for
|
||||
# app/worker from the same secret file at container start. All `postgres_password`
|
||||
# / `thermograph_*` secrets below must already exist in the Swarm (`docker secret
|
||||
# create <name> -`) before the first deploy — Track B step 7 in
|
||||
# thermograph-docs/runbooks/implementation-handoff.md.
|
||||
#
|
||||
# Migrations are NOT run inline by app/worker here (RUN_MIGRATIONS=0) — the
|
||||
# runbook's Stage F brings the schema to head via a one-shot task BEFORE scaling
|
||||
# app up, so multiple replicas never race Alembic and nothing runs DDL against a
|
||||
# still-read-only standby mid-cutover (hazard #9):
|
||||
#
|
||||
# docker run --rm --network thermograph_internal \
|
||||
# -e THERMOGRAPH_DATABASE_URL=... <image> migrate
|
||||
|
||||
services:
|
||||
db:
|
||||
# Pin the EXACT TimescaleDB minor — never latest-pg18 here. A floating tag can
|
||||
# give two hosts different extension minors, which blocks a physical replica
|
||||
# (a newer .so over an older catalog won't start) and risks compressed-chunk
|
||||
# corruption on timescaledb_post_restore() (hazard #7). Get the exact X.Y.Z
|
||||
# from the source database: SELECT extversion FROM pg_extension WHERE
|
||||
# extname='timescaledb' — use the SAME tag docker-compose.yml's TIMESCALEDB_TAG
|
||||
# is pinned to, on every host that could ever replicate with this one.
|
||||
image: timescale/timescaledb:${TIMESCALEDB_TAG:?set TIMESCALEDB_TAG to an exact pinned minor, e.g. 2.17.2-pg18 -- not latest-pg18}
|
||||
environment:
|
||||
POSTGRES_USER: thermograph
|
||||
POSTGRES_PASSWORD_FILE: /run/secrets/postgres_password
|
||||
POSTGRES_DB: thermograph
|
||||
DB_MEMORY: ${DB_MEMORY:-8g}
|
||||
volumes:
|
||||
- pgdata:/var/lib/postgresql
|
||||
- ./deploy/db/init:/docker-entrypoint-initdb.d
|
||||
networks:
|
||||
- internal
|
||||
secrets:
|
||||
- postgres_password
|
||||
deploy:
|
||||
# Pin to the labelled DB node so replication/IO never crosses the slow WG
|
||||
# uplink (hazard #14): `docker node update --label-add db=true <node>` once,
|
||||
# on whichever node holds pgdata (Track B).
|
||||
placement:
|
||||
constraints: ["node.labels.db == true"]
|
||||
resources:
|
||||
limits:
|
||||
cpus: "${DB_CPUS:-2}"
|
||||
memory: ${DB_MEMORY:-8g}
|
||||
restart_policy:
|
||||
condition: on-failure
|
||||
# No published port: reachable only as db:5432 on the `internal` overlay —
|
||||
# Postgres must never be public (design doc §6).
|
||||
|
||||
# Stateless, freely-replicable web tier. ROLE=web means _should_run_notifier()
|
||||
# is False unconditionally (web/app.py) — it never starts the notifier or the
|
||||
# worker scheduler even if it would otherwise win leader election.
|
||||
app:
|
||||
image: ${IMAGE_TAG:?set IMAGE_TAG to the image CI pushed, e.g. forge.example/thermograph/app:sha-abc123}
|
||||
environment:
|
||||
THERMOGRAPH_BASE: /
|
||||
PORT: 8137
|
||||
WORKERS: ${WORKERS:-4}
|
||||
THERMOGRAPH_ROLE: web
|
||||
RUN_MIGRATIONS: "0"
|
||||
volumes:
|
||||
- appdata:/app/data
|
||||
- applogs:/app/logs
|
||||
networks:
|
||||
- internal
|
||||
secrets:
|
||||
- postgres_password
|
||||
- thermograph_auth_secret
|
||||
- thermograph_vapid_private_key
|
||||
- thermograph_vapid_public_key
|
||||
deploy:
|
||||
replicas: 1 # single replica this hop -- homepage.json now lives in
|
||||
# Postgres (chunk 4), but the notifier/scheduler split (chunk
|
||||
# 2/5) is what actually lets this go multi-replica in Phase 2
|
||||
resources:
|
||||
limits:
|
||||
cpus: "${APP_CPUS:-4}"
|
||||
restart_policy:
|
||||
condition: on-failure
|
||||
update_config:
|
||||
order: start-first # new task must pass the image's HEALTHCHECK (now
|
||||
# GET /healthz) before the old one is stopped
|
||||
# No `ports:` — 127.0.0.1:8137:8137 (docker-compose.yml) has no Swarm
|
||||
# equivalent: Swarm's routing mesh publishes on 0.0.0.0, which would expose
|
||||
# the plaintext app un-fronted (hazard #6). Only the host's Caddy reaches
|
||||
# `app`, over the `internal` overlay network — see the Caddyfile templates'
|
||||
# health-gated reverse_proxy.
|
||||
|
||||
# Owns the notifier + worker scheduler (chunks 1/2/5). Exactly one replica —
|
||||
# the leader-election guard is belt-and-suspenders, not a substitute for it.
|
||||
worker:
|
||||
image: ${IMAGE_TAG:?set IMAGE_TAG to the image CI pushed, e.g. forge.example/thermograph/app:sha-abc123}
|
||||
environment:
|
||||
THERMOGRAPH_BASE: /
|
||||
WORKERS: "1"
|
||||
THERMOGRAPH_ROLE: worker
|
||||
# Cluster-wide advisory lock, not the flock: the flock only arbitrates
|
||||
# workers on ONE host, so under Swarm it guards nothing across replicas.
|
||||
THERMOGRAPH_SINGLETON_PG: "1"
|
||||
RUN_MIGRATIONS: "0"
|
||||
volumes:
|
||||
- appdata:/app/data
|
||||
- applogs:/app/logs
|
||||
networks:
|
||||
- internal
|
||||
secrets:
|
||||
- postgres_password
|
||||
- thermograph_auth_secret
|
||||
- thermograph_vapid_private_key
|
||||
- thermograph_vapid_public_key
|
||||
- thermograph_discord_webhook
|
||||
deploy:
|
||||
replicas: 1
|
||||
resources:
|
||||
limits:
|
||||
cpus: "${WORKER_CPUS:-1}"
|
||||
restart_policy:
|
||||
condition: on-failure
|
||||
# No `ports:` — the worker serves no public traffic; GET /healthz on its own
|
||||
# container port is for the image's own HEALTHCHECK (Swarm task health) only.
|
||||
|
||||
networks:
|
||||
internal:
|
||||
driver: overlay
|
||||
driver_opts:
|
||||
# VXLAN-over-WireGuard double encapsulation needs a lower MTU than the
|
||||
# ~1450 overlay default, or large payloads (a /cell bundle, the homepage
|
||||
# feed) silently stall while small packets (health checks) keep passing
|
||||
# (hazard #13). Validate with a real large-payload transfer, not just a
|
||||
# ping, once the mesh is up (Track B).
|
||||
com.docker.network.driver.mtu: "1370"
|
||||
|
||||
volumes:
|
||||
pgdata: {}
|
||||
appdata: {}
|
||||
applogs: {}
|
||||
|
||||
# Declared here, provisioned externally (docker secret create <name> - < file) —
|
||||
# never by this stack, and never committed. See Stage 0 of the cutover runbook
|
||||
# for which values are continuity-critical (AUTH_SECRET, VAPID, POSTGRES_PASSWORD
|
||||
# must be the EXISTING live values, not freshly generated ones).
|
||||
secrets:
|
||||
postgres_password:
|
||||
external: true
|
||||
thermograph_auth_secret:
|
||||
external: true
|
||||
thermograph_vapid_private_key:
|
||||
external: true
|
||||
thermograph_vapid_public_key:
|
||||
external: true
|
||||
thermograph_discord_webhook:
|
||||
external: true
|
||||
16
infra/terraform/.gitignore
vendored
Normal file
16
infra/terraform/.gitignore
vendored
Normal file
|
|
@ -0,0 +1,16 @@
|
|||
# Terraform working dir + plugins
|
||||
.terraform/
|
||||
|
||||
# Local state (holds secrets — never commit)
|
||||
*.tfstate
|
||||
*.tfstate.*
|
||||
crash.log
|
||||
crash.*.log
|
||||
|
||||
# tfvars carry real secrets — ignore all except the checked-in example
|
||||
*.tfvars
|
||||
!*.tfvars.example
|
||||
|
||||
# The dependency lock IS committed (pins provider versions across machines) — keep it.
|
||||
!.terraform.lock.hcl
|
||||
*.plan
|
||||
85
infra/terraform/.terraform.lock.hcl
Normal file
85
infra/terraform/.terraform.lock.hcl
Normal file
|
|
@ -0,0 +1,85 @@
|
|||
# This file is maintained automatically by "terraform init".
|
||||
# Manual edits may be lost in future updates.
|
||||
|
||||
provider "registry.terraform.io/hashicorp/google" {
|
||||
version = "6.50.0"
|
||||
constraints = "~> 6.0"
|
||||
hashes = [
|
||||
"h1:faTJQOetP9/RYuHwA3r2SWnuYoyzQNm4tUWZrZggcgY=",
|
||||
"zh:1f3513fcfcbf7ca53d667a168c5067a4dd91a4d4cccd19743e248ff31065503c",
|
||||
"zh:3da7db8fc2c51a77dd958ea8baaa05c29cd7f829bd8941c26e2ea9cb3aadc1e5",
|
||||
"zh:3e09ac3f6ca8111cbb659d38c251771829f4347ab159a12db195e211c76068bb",
|
||||
"zh:7bb9e41c568df15ccf1a8946037355eefb4dfb4e35e3b190808bb7c4abae547d",
|
||||
"zh:81e5d78bdec7778e6d67b5c3544777505db40a826b6eb5abe9b86d4ba396866b",
|
||||
"zh:8d309d020fb321525883f5c4ea864df3d5942b6087f6656d6d8b3a1377f340fc",
|
||||
"zh:93e112559655ab95a523193158f4a4ac0f2bfed7eeaa712010b85ebb551d5071",
|
||||
"zh:d3efe589ffd625b300cef5917c4629513f77e3a7b111c9df65075f76a46a63c7",
|
||||
"zh:d4a4d672bbef756a870d8f32b35925f8ce2ef4f6bbd5b71a3cb764f1b6c85421",
|
||||
"zh:e13a86bca299ba8a118e80d5f84fbdd708fe600ecdceea1a13d4919c068379fe",
|
||||
"zh:f569b65999264a9416862bca5cd2a6177d94ccb0424f3a4ef424428912b9cb3c",
|
||||
"zh:fec30c095647b583a246c39d557704947195a1b7d41f81e369ba377d997faef6",
|
||||
]
|
||||
}
|
||||
|
||||
provider "registry.terraform.io/hashicorp/local" {
|
||||
version = "2.9.0"
|
||||
constraints = "~> 2.5"
|
||||
hashes = [
|
||||
"h1:9rBZCMNpxKwMlRbWH2QpwD3kqUCAejdOZQ/aiiDObXQ=",
|
||||
"zh:0baa4566cf77f1ff52f4293d1c8536202dd23edc197c3196413a28343c3ac3a0",
|
||||
"zh:16b5559c3c07088ddad11a9bb9e9c0799999363c2958e9a5be2bcbbf2cd9ca64",
|
||||
"zh:197c79015a10d1cce904a8ea722cbc750c42aeae2da53f44a6a0751d9fd1aa90",
|
||||
"zh:29d0b03e5343a80677ebfeb2e2c31cbe4b1f65e736e53417454a4277fec2544c",
|
||||
"zh:4896bfa6cf1d2fd562b47ef2e87f47862ae92a04f8ad5d764380f0c6653473b8",
|
||||
"zh:531f8529cbca49f681883e57761a05a8398afaef6d1ab0d205d26bf12f4428e8",
|
||||
"zh:6aaf5011d83161c86d2bfb80c0923ec934e578288758da2f37acb7aec129004b",
|
||||
"zh:7430275253d3d3c40aa6179e0ec0d63212874dbbc06c5a51b9d07ec590f9756c",
|
||||
"zh:78d5eefdd9e494defcb3c68d282b8f96630502cac21d1ea161f53cfe9bb483b3",
|
||||
"zh:be17dc611e95e26cdf6cad79dfccf1064f0e32032a2efeb939a9bbe7fb1cbfe9",
|
||||
"zh:f0e3b0aa644202e1d79d2000dca91f6019425da71e9800fa23f27e51c034f195",
|
||||
"zh:f62bae4519e4ead49182ddc8afe8cf61e2a4c3ba3973b0fbba967736a2696aa3",
|
||||
"zh:fcafa360a5b0b96244f26f4e3a6d642b716a376557142c2442ff2fb12d11da18",
|
||||
]
|
||||
}
|
||||
|
||||
provider "registry.terraform.io/hashicorp/null" {
|
||||
version = "3.3.0"
|
||||
constraints = "~> 3.2"
|
||||
hashes = [
|
||||
"h1:l+dm3lhmu4ys7GbvIldfn544olSPH0DOiYruuFSfQkY=",
|
||||
"zh:021748b5ea3b5f6956f2e75c42c5cdc113b391fb98ac71364a4965d23b37000f",
|
||||
"zh:3b27956f8541d46704fda234e0d535c2ae2a4b33411848b1ee262a1ec03568b0",
|
||||
"zh:3de4ed47d6d0f4d8edba4a5092c7c9799950eda63989d8d0d2586e6afcb0aa20",
|
||||
"zh:57ed8935c7d56dbc91cf2673534582cacfaab7a2f105f51d9f797e99df0c0c47",
|
||||
"zh:58e176ba1d142827089e30e0711e007309a9f2726e8881986da5026e9778fdf4",
|
||||
"zh:5949c4a3d4a93f841f155cdb7e991c087e637145c1630572e21948224f8f4923",
|
||||
"zh:76d60f366b743003c1b085afa769b45b2198ee919927e45807d7d44fb42c067d",
|
||||
"zh:78d5eefdd9e494defcb3c68d282b8f96630502cac21d1ea161f53cfe9bb483b3",
|
||||
"zh:79cd1bab1261a07f84e917191d7ddc4340ac5f5524283767256f7ffd7f87caf0",
|
||||
"zh:8ec9083038cf710b30e319eaa467c9df7fa52bbd9969b61053a35bc2cdd2e0a6",
|
||||
"zh:a6e502cb579685ab7aeb886c2bb11ddd9cfed74b41008592d57cbc3351a9218b",
|
||||
"zh:acb74d6b4f66ff6acfcda315df802a7432170ef3955c9b432cb4580767004006",
|
||||
"zh:f0ce55d8d9ffdb33dab612b1246f9bab060a9d54fc32ce2b4a038646155660af",
|
||||
]
|
||||
}
|
||||
|
||||
provider "registry.terraform.io/hashicorp/time" {
|
||||
version = "0.14.0"
|
||||
constraints = "~> 0.12"
|
||||
hashes = [
|
||||
"h1:4EThC3ocCFiFPMZQSUvSGSxoJqBcGWxMcFYmL67uS7Y=",
|
||||
"zh:12abfd6b800e4d7fa6db7310dec8ffd440b31993861ef188c7ed5260b3073937",
|
||||
"zh:23005521e800bb19e1597bf755c5f70d675d30b685d4255001ed5fa47d9df3f1",
|
||||
"zh:2fea249b582ae97cd1cc10385187ea50993bb47c28cc5df0305e57ceaabf0a10",
|
||||
"zh:322018d3b987b7aad08697178029a2bb667bed699e88328f0c89c52a2fd41341",
|
||||
"zh:32a08e98fce2d273cb9b2c89d6c54727cc9f0a32e15bfd896be4e02cc6b48f95",
|
||||
"zh:3db89aabd0e619616bd4b0f8b373a7586dfe60feffcea12a84a0bdbc445714b3",
|
||||
"zh:7488f56c81d742dc020f29063626c8f07ca188aa97be61e7307e8d62397020a2",
|
||||
"zh:78d5eefdd9e494defcb3c68d282b8f96630502cac21d1ea161f53cfe9bb483b3",
|
||||
"zh:7cb4067f2e7559b13f7562ef722f948950901eb37834873e98360ab28f66e9d7",
|
||||
"zh:9d552c8345f61e1b7db8e725144981345f18ac1014d58d6f5ddf0928a195fffb",
|
||||
"zh:a8e69fb6b97fc9d86fb19a9f4d42abe33c4a68e700b15387ce2e17d2b9934bed",
|
||||
"zh:aeeb900eb8dd0f790c60ea5c0e0c8d42bd6e4a54f391681d4decca15b544394b",
|
||||
"zh:c239c619101a8c95e1f14061eb973c57a8d15fa0e68878ced5bbd76858ee5b79",
|
||||
]
|
||||
}
|
||||
164
infra/terraform/README.md
Normal file
164
infra/terraform/README.md
Normal file
|
|
@ -0,0 +1,164 @@
|
|||
# Thermograph — Terraform (host provisioning)
|
||||
|
||||
Terraform that **provisions and configures VPS hosts** and hands the app off to
|
||||
`docker compose`, which pulls the published app image (built + pushed by the app
|
||||
repo's CI) and runs it — no app source is ever built or checked out on a host.
|
||||
By default it does *not* create servers: `modules/thermograph-host` is
|
||||
SSH-provisioner-driven against a host that already exists. An optional
|
||||
`modules/gcp-host` can additionally *create* the VM on GCP first (see "GCP
|
||||
scaffold" below) — scaffold-only today, no live resources until you opt in.
|
||||
|
||||
## What it manages
|
||||
|
||||
One reusable module (`modules/thermograph-host`) is instantiated per host via
|
||||
`for_each` (`main.tf`'s `local.all_hosts`, merging `var.hosts` — SSH-managed,
|
||||
already-existing boxes — with any `var.gcp_hosts` Terraform created itself).
|
||||
This config manages **two VPS hosts** today:
|
||||
|
||||
| key | role | VPS | branch | domain | notes |
|
||||
|--------|--------|-------------------------|-----------|-------------------|------------------------------------|
|
||||
| `prod` | prod | NEW 48 GB / 12-core box (`169.58.46.181`) | `release` (of the APP repos — see `backend_image_tag` / `frontend_image_tag`) | `thermograph.org` | Caddy TLS; sized up (8/8/4/16g) |
|
||||
| `beta` | beta | old box `75.119.132.91` | `main` (of the APP repo) | `beta.thermograph.org` | Caddy TLS; also hosts Forgejo |
|
||||
|
||||
Each host's own checkout on disk (`app_dir`, `git_branch`) is **this infra
|
||||
repo**, not the app repo — the "branch" column above is which app-repo tag a
|
||||
host is meant to track conceptually; the actual pinned versions are
|
||||
`var.hosts[*].backend_image_tag` / `frontend_image_tag` (e.g. `"sha-<12 hex>"` each —
|
||||
the app is two separately-published images, `emi/thermograph-backend/app` and
|
||||
`emi/thermograph-frontend/app`), since the host has no app checkout to derive a tag
|
||||
from. The LAN dev server is **out of scope here** — it
|
||||
builds from source via the app repo's `deploy/deploy-dev.sh` (a self-hosted
|
||||
Forgejo Actions runner), not Terraform.
|
||||
|
||||
Per host, over SSH provisioners, Terraform:
|
||||
|
||||
- installs Docker + the compose plugin if missing;
|
||||
- configures a `ufw` firewall (22/80/443 always; on a host with **no** domain it also
|
||||
opens the app port `8137`);
|
||||
- ensures the git checkout at `app_dir` exists (clones **this repo** on a fresh box)
|
||||
and resets it to the host's `git_branch`;
|
||||
- renders `/etc/thermograph-topology.env` from Terraform variables (non-secret sizing
|
||||
only — `WORKERS`/`APP_CPUS`/`DB_CPUS`/`DB_MEMORY`/etc.), then runs
|
||||
`deploy/render-secrets.sh` (from that same freshly-synced checkout) to render
|
||||
`/etc/thermograph.env` from the SOPS+age vault — **Terraform itself never sees or
|
||||
carries an app secret** (see "Secrets" below);
|
||||
- for a host **with** a domain, installs a rendered Caddyfile and reloads Caddy;
|
||||
- brings the stack up on the explicit per-service tags (exports `BACKEND_IMAGE_TAG` /
|
||||
`FRONTEND_IMAGE_TAG` for the compose file): `docker login` to the registry,
|
||||
`docker compose <-f each compose file> pull backend frontend`, then
|
||||
`... up -d --remove-orphans`, running docker as root with both env files sourced;
|
||||
- health-checks `http://127.0.0.1:8137/`.
|
||||
|
||||
A change to the rendered topology env, the compose files, the branch, the app image
|
||||
tag, or the sizing flips the `null_resource` trigger and re-runs the provisioners on
|
||||
next apply.
|
||||
|
||||
## GCP scaffold (no live resources)
|
||||
|
||||
`modules/gcp-host` can create a GCE VM (+ a dedicated VPC, subnet, and a firewall
|
||||
opening 22/80/443) instead of assuming the host already exists — see `main.tf`'s
|
||||
`module.gcp_vm`. It contributes *nothing* to provisioning: its only output
|
||||
(`external_ip`) feeds into the exact same `thermograph-host` module every
|
||||
SSH-managed host uses, so there's one provisioning path regardless of how a
|
||||
host came to exist. `var.gcp_hosts` defaults to `{}`, so by default no
|
||||
`google_*` resource is ever planned and the `google` provider is never
|
||||
invoked — `terraform plan`/`validate` succeed with no GCP credentials
|
||||
configured at all. To actually use it: populate an entry in
|
||||
`terraform.tfvars` (see the commented example in
|
||||
`terraform.tfvars.example`) and authenticate via `gcloud auth
|
||||
application-default login` or `GOOGLE_APPLICATION_CREDENTIALS`. This is the
|
||||
same create-then-provision composition a future Proxmox module (the
|
||||
architecture doc's longer-term target — see §6/§9 there) would use.
|
||||
|
||||
### Container sizing is env-driven
|
||||
|
||||
`docker-compose.yml` reads `WORKERS`, `APP_CPUS`, `DB_CPUS`, and `DB_MEMORY` from the
|
||||
environment (defaults `4 / 4 / 2 / 8g`, identical to before). Terraform sets them per
|
||||
host through `/etc/thermograph-topology.env`, so the big prod box can run larger caps
|
||||
without a compose edit. The Postgres *internal* memory budget (`shared_buffers`,
|
||||
`effective_cache_size`, `work_mem`, `maintenance_work_mem`) is derived from the same
|
||||
`DB_MEMORY` by `deploy/db/init/20-tuning.sh` — so raising `db_memory` scales the
|
||||
container cap and the tuning together (prod 16g → shared_buffers 4 GB). The tuning
|
||||
applies on a fresh DB volume; on an existing volume re-run it by hand (see the script
|
||||
header).
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Terraform >= 1.6 (v1.15 is installed).
|
||||
- SSH key access to **both** hosts as a **sudo-capable** user (the live boxes use the
|
||||
dedicated `agent` account, `~/.ssh/thermograph_agent_ed25519` — see
|
||||
`deploy/provision-agent-access.sh`). Point `ssh_private_key_path` at that key (`~`
|
||||
is expanded).
|
||||
- The hosts are Debian/Ubuntu with `apt` and outbound internet (Docker/Caddy installs
|
||||
pull from the network). Docker may already be present — installs are conditional.
|
||||
- For the `prod` host: DNS for `thermograph.org` must point at the new box before apply,
|
||||
or Caddy's first-request cert issuance will fail.
|
||||
|
||||
## Use
|
||||
|
||||
```sh
|
||||
cd terraform
|
||||
cp terraform.tfvars.example terraform.tfvars # then edit: real IPs + secrets
|
||||
terraform init
|
||||
terraform plan
|
||||
terraform apply
|
||||
```
|
||||
|
||||
Target one host with `-target='module.host["beta"]'` if you want to apply to just one.
|
||||
|
||||
## Secrets
|
||||
|
||||
App secrets (`POSTGRES_PASSWORD`, `THERMOGRAPH_AUTH_SECRET`, VAPID keys,
|
||||
`REGISTRY_TOKEN`, Discord/SMTP credentials, …) are **not** Terraform variables
|
||||
— they live in the git-native SOPS+age vault at `../deploy/secrets/*.yaml`
|
||||
(committed, encrypted) and are rendered into `/etc/thermograph.env` at deploy
|
||||
time by `../deploy/render-secrets.sh`, which the provisioner's deploy step
|
||||
runs from the freshly-synced checkout. Terraform only renders the non-secret
|
||||
`/etc/thermograph-topology.env` (sizing/routing). See
|
||||
`../deploy/secrets/README.md` to rotate a secret or add a new one — it's a
|
||||
`sops edit` + commit + deploy, no Terraform apply involved.
|
||||
|
||||
`om_rclone_conf` (object-storage bucket credentials for the self-hosted ERA5
|
||||
archive) is the one exception still supplied via Terraform, in
|
||||
`terraform.tfvars` — folding it into the vault too is a reasonable future
|
||||
step, not done here.
|
||||
|
||||
## Local state
|
||||
|
||||
The backend is **local**: `terraform.tfstate` is written next to the config.
|
||||
It's gitignored (`terraform/.gitignore` and the root `.gitignore`) — keep it
|
||||
off shared disks and back it up somewhere private. `terraform.tfvars` is
|
||||
likewise gitignored (it may carry a `repo_url` credential and always carries
|
||||
`om_rclone_conf`); only `terraform.tfvars.example` (dummy values) is
|
||||
committed. `.terraform.lock.hcl` **is** committed so provider versions are
|
||||
pinned across machines.
|
||||
|
||||
## WARNING — applying against live prod
|
||||
|
||||
`terraform apply` runs `remote-exec` **on the server**: it resets the checkout to the
|
||||
branch, renders topology config + secrets, and runs `docker compose pull && up -d`
|
||||
(pulling the pinned `backend_image_tag` / `frontend_image_tag` and recreating
|
||||
containers — a brief app restart).
|
||||
Against the live production host this is a real deploy. Review the plan, apply in a
|
||||
maintenance window, and prefer `-target` to touch one host at a time.
|
||||
|
||||
This is **separate from the Postgres data cutover** in `deploy/POSTGRES-MIGRATION.md`.
|
||||
Terraform provisions the host and starts the stack; it does **not** migrate the
|
||||
SQLite→Postgres accounts data. Sequence them deliberately: for a first cutover on a
|
||||
host, follow the migration doc's freeze/backup/copy steps around the point where
|
||||
Terraform brings the stack up — don't let Terraform recreate containers mid-migration.
|
||||
|
||||
## Assumptions / notes
|
||||
|
||||
- **beta has no public domain by default.** With `compose_files = ["docker-compose.yml"]`
|
||||
the app binds `127.0.0.1:8137` (loopback), so opening the port in `ufw` alone does not
|
||||
expose it. Reach beta via an SSH tunnel, or set `domain = "beta.thermograph.org"` (adds
|
||||
Caddy TLS) — or add the `0.0.0.0`-publishing dev overlay to `compose_files` — to make
|
||||
it reachable. `COOKIE_SECURE` is auto-set to `0` when there's no domain (a Secure
|
||||
cookie is never sent over plain HTTP) and `1` behind Caddy TLS.
|
||||
- The rendered Caddyfile only reverse-proxies the app. The repo's `deploy/Caddyfile`
|
||||
additionally serves the `emigriffith.dev` portfolio and legacy redirects; those are
|
||||
host-specific and not templated here.
|
||||
- Provisioner-based by design: the hosts already exist, so this is not a
|
||||
create-from-scratch cloud config. Re-applying is idempotent (installs are guarded,
|
||||
`git reset --hard`, compose `up` reconciles).
|
||||
117
infra/terraform/main.tf
Normal file
117
infra/terraform/main.tf
Normal file
|
|
@ -0,0 +1,117 @@
|
|||
# No project/region default here on purpose: each var.gcp_hosts entry carries its
|
||||
# own project + zone (a fleet could span projects), and with var.gcp_hosts empty
|
||||
# this provider is never invoked, so there's nothing to default. When you do
|
||||
# populate var.gcp_hosts, authenticate via Application Default Credentials
|
||||
# (`gcloud auth application-default login`) or GOOGLE_APPLICATION_CREDENTIALS --
|
||||
# no credentials are configured in this repo.
|
||||
provider "google" {}
|
||||
|
||||
locals {
|
||||
# Repo root (one level above this terraform/ dir). The module hashes the compose
|
||||
# files here so a compose change re-triggers the remote deploy, and this is the
|
||||
# tree the host's checkout mirrors over git.
|
||||
repo_root = abspath("${path.root}/..")
|
||||
|
||||
# Reusable size presets — a host can reference one by name (hosts.<name>.size)
|
||||
# instead of hand-picking workers/app_cpus/db_cpus/db_memory separately. Mirrors
|
||||
# the target Proxmox sizing-tier model (architecture doc §6: nano/small/medium/
|
||||
# large -> {cores, mem}) ahead of actually provisioning VMs — Proxmox itself is
|
||||
# deferred; today these tiers just size the container caps on the existing
|
||||
# SSH-managed VPS hosts. "large" matches the current 48 GB prod box's numbers.
|
||||
sizes = {
|
||||
nano = { workers = 1, app_cpus = 1, db_cpus = 1, db_memory = "1g" }
|
||||
small = { workers = 4, app_cpus = 4, db_cpus = 2, db_memory = "8g" }
|
||||
medium = { workers = 6, app_cpus = 6, db_cpus = 3, db_memory = "12g" }
|
||||
large = { workers = 8, app_cpus = 8, db_cpus = 4, db_memory = "16g" }
|
||||
}
|
||||
}
|
||||
|
||||
# GCP-created hosts (empty by default — see variables.tf and modules/gcp-host).
|
||||
# Creates only the VM + minimal networking; every provisioning behavior comes from
|
||||
# the SAME thermograph-host module every SSH-managed host below uses, via
|
||||
# local.all_hosts merging its output IP in below. No live resources exist until
|
||||
# var.gcp_hosts is populated.
|
||||
module "gcp_vm" {
|
||||
source = "./modules/gcp-host"
|
||||
for_each = var.gcp_hosts
|
||||
|
||||
name = each.key
|
||||
project = each.value.project
|
||||
zone = each.value.zone
|
||||
machine_type = each.value.machine_type
|
||||
ssh_user = each.value.ssh_user
|
||||
ssh_public_key_path = each.value.ssh_public_key_path
|
||||
}
|
||||
|
||||
locals {
|
||||
# Every host this config manages, SSH-only (var.hosts) plus GCP-created (whose
|
||||
# `host` is filled in from the VM Terraform just created) — one unified map so a
|
||||
# single `module.host` for_each below handles both without duplicating any
|
||||
# provisioning logic. GCP hosts always use a named size tier (simpler than
|
||||
# exposing the four raw sizing fields on that variable too).
|
||||
all_hosts = merge(
|
||||
var.hosts,
|
||||
{
|
||||
for name, h in var.gcp_hosts : name => merge(h, {
|
||||
host = module.gcp_vm[name].external_ip
|
||||
workers = null
|
||||
app_cpus = null
|
||||
db_cpus = null
|
||||
db_memory = null
|
||||
openmeteo = false
|
||||
om_data_dir = "/mnt/om-archive"
|
||||
})
|
||||
}
|
||||
)
|
||||
|
||||
# Per-host resolved sizing: a named tier (host.size) wins when set; otherwise the
|
||||
# host's own workers/app_cpus/db_cpus/db_memory fields (each individually
|
||||
# defaulted in variables.tf) — so existing tfvars with explicit numbers are
|
||||
# unaffected, and a tier is purely an opt-in shortcut.
|
||||
host_sizing = {
|
||||
for name, h in local.all_hosts : name => h.size != null ? local.sizes[h.size] : {
|
||||
workers = h.workers
|
||||
app_cpus = h.app_cpus
|
||||
db_cpus = h.db_cpus
|
||||
db_memory = h.db_memory
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
# One module instance per host (SSH-managed or GCP-created — see local.all_hosts).
|
||||
# The module is entirely SSH-provisioner driven: it configures an already-running
|
||||
# VM (however it came to exist) and hands the app off to docker compose.
|
||||
module "host" {
|
||||
source = "./modules/thermograph-host"
|
||||
for_each = local.all_hosts
|
||||
|
||||
# Per-host config
|
||||
name = each.key
|
||||
host = each.value.host
|
||||
ssh_user = each.value.ssh_user
|
||||
ssh_private_key_path = each.value.ssh_private_key_path
|
||||
role = each.value.role
|
||||
git_branch = each.value.git_branch
|
||||
backend_image_tag = each.value.backend_image_tag
|
||||
frontend_image_tag = each.value.frontend_image_tag
|
||||
domain = each.value.domain
|
||||
compose_files = each.value.compose_files
|
||||
app_dir = each.value.app_dir
|
||||
workers = local.host_sizing[each.key].workers
|
||||
app_cpus = local.host_sizing[each.key].app_cpus
|
||||
db_cpus = local.host_sizing[each.key].db_cpus
|
||||
db_memory = local.host_sizing[each.key].db_memory
|
||||
timescaledb_tag = each.value.timescaledb_tag
|
||||
openmeteo = each.value.openmeteo
|
||||
om_data_dir = each.value.om_data_dir
|
||||
|
||||
# Shared infra config
|
||||
repo_root = local.repo_root
|
||||
repo_url = var.repo_url
|
||||
app_port = var.app_port
|
||||
|
||||
# Shared object-storage config (only used where openmeteo = true)
|
||||
om_bucket_remote = var.om_bucket_remote
|
||||
om_rclone_conf = var.om_rclone_conf
|
||||
om_vfs_cache_max = var.om_vfs_cache_max
|
||||
}
|
||||
77
infra/terraform/modules/gcp-host/main.tf
Normal file
77
infra/terraform/modules/gcp-host/main.tf
Normal file
|
|
@ -0,0 +1,77 @@
|
|||
# Creates a GCE VM (+ the minimal networking it needs) for the thermograph-host
|
||||
# module to then provision over SSH — this module contributes NOTHING to
|
||||
# provisioning logic (no Docker install, no compose, no secrets rendering); it
|
||||
# only stands the box up and hands its IP back. See ../../main.tf's module.gcp_vm
|
||||
# / module.host composition. Scaffold only: nothing here is applied until a
|
||||
# caller populates var.gcp_hosts (root variables.tf) with a real entry.
|
||||
|
||||
# A dedicated VPC rather than the project's "default" network, so this doesn't
|
||||
# depend on (or clutter) whatever else may already exist in the project.
|
||||
resource "google_compute_network" "this" {
|
||||
project = var.project
|
||||
name = "thermograph-${var.name}"
|
||||
auto_create_subnetworks = false
|
||||
}
|
||||
|
||||
resource "google_compute_subnetwork" "this" {
|
||||
project = var.project
|
||||
name = "thermograph-${var.name}"
|
||||
network = google_compute_network.this.id
|
||||
region = join("-", slice(split("-", var.zone), 0, 2))
|
||||
ip_cidr_range = "10.20.0.0/24"
|
||||
}
|
||||
|
||||
# 22/80/443 only -- matches the SSH-managed hosts' own ufw rules (thermograph-host
|
||||
# module's setup step). The app port itself is opened by that same module's ufw
|
||||
# step when the host has no domain, same as any other host it provisions.
|
||||
resource "google_compute_firewall" "allow_ingress" {
|
||||
project = var.project
|
||||
name = "thermograph-${var.name}-allow-ingress"
|
||||
network = google_compute_network.this.id
|
||||
source_ranges = ["0.0.0.0/0"]
|
||||
|
||||
allow {
|
||||
protocol = "tcp"
|
||||
ports = ["22", "80", "443"]
|
||||
}
|
||||
}
|
||||
|
||||
resource "google_compute_instance" "this" {
|
||||
project = var.project
|
||||
name = "thermograph-${var.name}"
|
||||
zone = var.zone
|
||||
machine_type = var.machine_type
|
||||
|
||||
boot_disk {
|
||||
initialize_params {
|
||||
# Debian: matches the apt-based setup script thermograph-host's
|
||||
# provisioner runs (Docker's get.docker.com installer, ufw, Caddy's apt repo).
|
||||
image = "debian-cloud/debian-12"
|
||||
size = 30
|
||||
}
|
||||
}
|
||||
|
||||
network_interface {
|
||||
subnetwork = google_compute_subnetwork.this.id
|
||||
access_config {} # ephemeral public IP
|
||||
}
|
||||
|
||||
metadata = {
|
||||
ssh-keys = "${var.ssh_user}:${file(pathexpand(var.ssh_public_key_path))}"
|
||||
}
|
||||
}
|
||||
|
||||
# A brief pause before the caller's thermograph-host module opens an SSH
|
||||
# connection: a fresh GCE instance's sshd is not always immediately reachable
|
||||
# the instant the API reports the instance RUNNING (cloud-init/sshd startup is a
|
||||
# known race here, distinct from the null_resource provisioner's own timeout).
|
||||
resource "time_sleep" "wait_for_ssh" {
|
||||
create_duration = "30s"
|
||||
depends_on = [google_compute_instance.this]
|
||||
}
|
||||
|
||||
output "external_ip" {
|
||||
description = "Public IP of the created instance, once it (and a short settle delay) exist."
|
||||
value = google_compute_instance.this.network_interface[0].access_config[0].nat_ip
|
||||
depends_on = [time_sleep.wait_for_ssh]
|
||||
}
|
||||
31
infra/terraform/modules/gcp-host/variables.tf
Normal file
31
infra/terraform/modules/gcp-host/variables.tf
Normal file
|
|
@ -0,0 +1,31 @@
|
|||
variable "name" {
|
||||
description = "Short host key (e.g. \"gcp-uat\"), used to name the created resources."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "project" {
|
||||
description = "GCP project ID to create the instance in."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "zone" {
|
||||
description = "GCP zone, e.g. \"us-west1-a\"."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "machine_type" {
|
||||
description = "GCE machine type."
|
||||
type = string
|
||||
default = "e2-medium"
|
||||
}
|
||||
|
||||
variable "ssh_user" {
|
||||
description = "Login user provisioned via instance metadata (must be able to sudo)."
|
||||
type = string
|
||||
default = "deploy"
|
||||
}
|
||||
|
||||
variable "ssh_public_key_path" {
|
||||
description = "Path to the PUBLIC key installed into the instance's ssh-keys metadata for ssh_user."
|
||||
type = string
|
||||
}
|
||||
12
infra/terraform/modules/gcp-host/versions.tf
Normal file
12
infra/terraform/modules/gcp-host/versions.tf
Normal file
|
|
@ -0,0 +1,12 @@
|
|||
terraform {
|
||||
required_providers {
|
||||
google = {
|
||||
source = "hashicorp/google"
|
||||
version = "~> 6.0"
|
||||
}
|
||||
time = {
|
||||
source = "hashicorp/time"
|
||||
version = "~> 0.12"
|
||||
}
|
||||
}
|
||||
}
|
||||
307
infra/terraform/modules/thermograph-host/main.tf
Normal file
307
infra/terraform/modules/thermograph-host/main.tf
Normal file
|
|
@ -0,0 +1,307 @@
|
|||
locals {
|
||||
# A Secure cookie is only sent over HTTPS, so enable it only where Caddy terminates
|
||||
# TLS (domain set). A plain-HTTP dev box (domain "") would otherwise drop its login
|
||||
# cookie and no one could stay signed in.
|
||||
cookie_secure = var.domain != "" ? "1" : "0"
|
||||
|
||||
# Public base URL: the domain over HTTPS, else the raw host:port for a Caddy-less box.
|
||||
base_url = var.domain != "" ? "https://${var.domain}" : "http://${var.host}:${var.app_port}"
|
||||
|
||||
# Layer the self-hosted Open-Meteo overlay on hosts that self-host the archive.
|
||||
effective_compose_files = var.openmeteo ? concat(var.compose_files, ["docker-compose.openmeteo.yml"]) : var.compose_files
|
||||
|
||||
# `-f a -f b` for the compose invocations (dev layers the dev overlay).
|
||||
compose_flags = join(" ", [for f in local.effective_compose_files : "-f ${f}"])
|
||||
|
||||
# Hash the local compose files so a compose edit re-triggers the remote deploy.
|
||||
compose_files_sha = join(",", [for f in local.effective_compose_files : filesha256("${var.repo_root}/${f}")])
|
||||
|
||||
# Rendered /etc/thermograph-topology.env — non-secret sizing/routing config only.
|
||||
# Every app secret is rendered separately, at deploy time, from the SOPS+age
|
||||
# vault (see remote-exec step 3 below) — Terraform never sees or carries them.
|
||||
topology_env_content = templatefile("${path.module}/templates/thermograph-topology.env.tftpl", {
|
||||
app_port = var.app_port
|
||||
workers = var.workers
|
||||
app_cpus = var.app_cpus
|
||||
db_cpus = var.db_cpus
|
||||
db_memory = var.db_memory
|
||||
timescaledb_tag = var.timescaledb_tag
|
||||
base = "/"
|
||||
base_url = local.base_url
|
||||
cookie_secure = local.cookie_secure
|
||||
openmeteo = var.openmeteo
|
||||
om_data_dir = var.om_data_dir
|
||||
})
|
||||
|
||||
# Caddyfile is only meaningful on a host with a public domain.
|
||||
caddy_content = var.domain != "" ? templatefile("${path.module}/templates/Caddyfile.tftpl", {
|
||||
domain = var.domain
|
||||
port = var.app_port
|
||||
frontend_port = var.frontend_port
|
||||
}) : "# No public domain on this host; Caddy is not managed here.\n"
|
||||
|
||||
# systemd unit that keeps the object-storage bucket rclone-mounted at om_data_dir,
|
||||
# so the Open-Meteo containers read the ERA5 .om archive from it. Only installed on
|
||||
# openmeteo hosts; --allow-other lets the container (root) read the FUSE mount.
|
||||
rclone_unit = <<-UNIT
|
||||
[Unit]
|
||||
Description=rclone mount ERA5 archive (Thermograph)
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=notify
|
||||
ExecStartPre=/bin/mkdir -p ${var.om_data_dir}
|
||||
ExecStart=/usr/bin/rclone mount ${var.om_bucket_remote} ${var.om_data_dir} --config /etc/rclone/rclone.conf --vfs-cache-mode full --vfs-cache-max-size ${var.om_vfs_cache_max} --dir-cache-time 12h --allow-other --umask 000
|
||||
ExecStop=/bin/fusermount -u ${var.om_data_dir}
|
||||
Restart=on-failure
|
||||
RestartSec=5
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
UNIT
|
||||
|
||||
rclone_unit_content = var.openmeteo ? local.rclone_unit : "# openmeteo disabled on this host\n"
|
||||
}
|
||||
|
||||
resource "null_resource" "host" {
|
||||
# Re-provision when the rendered topology env, the compose files, the branch,
|
||||
# sizing, the app image tag, or the Caddyfile change.
|
||||
triggers = {
|
||||
topology_env_sha = sha256(local.topology_env_content)
|
||||
compose_sha = local.compose_files_sha
|
||||
compose_flags = local.compose_flags
|
||||
caddy_sha = sha256(local.caddy_content)
|
||||
branch = var.git_branch
|
||||
backend_image_tag = var.backend_image_tag
|
||||
frontend_image_tag = var.frontend_image_tag
|
||||
app_dir = var.app_dir
|
||||
sizing = "${var.workers}/${var.app_cpus}/${var.db_cpus}/${var.db_memory}/${var.timescaledb_tag}"
|
||||
# Re-provision when the archive self-hosting config changes. The rclone.conf is
|
||||
# hashed (nonsensitive on a one-way digest) so a credential rotation redeploys.
|
||||
openmeteo = "${var.openmeteo}/${var.om_data_dir}/${var.om_bucket_remote}/${var.om_vfs_cache_max}"
|
||||
om_conf = var.openmeteo ? nonsensitive(sha256(var.om_rclone_conf)) : "off"
|
||||
}
|
||||
|
||||
connection {
|
||||
type = "ssh"
|
||||
host = var.host
|
||||
user = var.ssh_user
|
||||
private_key = file(pathexpand(var.ssh_private_key_path))
|
||||
timeout = "5m"
|
||||
}
|
||||
|
||||
# Push the rendered topology env via `content` (consistent with the other
|
||||
# `file` provisioners here, though nothing in it is secret).
|
||||
provisioner "file" {
|
||||
content = local.topology_env_content
|
||||
destination = "/tmp/thermograph-topology.env"
|
||||
}
|
||||
|
||||
provisioner "file" {
|
||||
content = local.caddy_content
|
||||
destination = "/tmp/thermograph.Caddyfile"
|
||||
}
|
||||
|
||||
# rclone config (bucket credentials) + the mount unit. Pushed via `content` so the
|
||||
# secret never touches local disk; harmless placeholders on non-openmeteo hosts.
|
||||
provisioner "file" {
|
||||
content = var.openmeteo ? var.om_rclone_conf : "# openmeteo disabled on this host\n"
|
||||
destination = "/tmp/thermograph.rclone.conf"
|
||||
}
|
||||
|
||||
provisioner "file" {
|
||||
content = local.rclone_unit_content
|
||||
destination = "/tmp/rclone-om.service"
|
||||
}
|
||||
|
||||
# 1. Host setup: Docker + compose plugin, ufw firewall, and (domain hosts) Caddy.
|
||||
provisioner "remote-exec" {
|
||||
inline = [
|
||||
<<-EOT
|
||||
set -eu
|
||||
echo "[${var.name}] setup: docker, compose plugin, firewall"
|
||||
if ! command -v docker >/dev/null 2>&1; then
|
||||
curl -fsSL https://get.docker.com | sudo sh
|
||||
fi
|
||||
sudo usermod -aG docker "$(id -un)" || true
|
||||
if ! sudo docker compose version >/dev/null 2>&1; then
|
||||
sudo apt-get update -y
|
||||
sudo apt-get install -y docker-compose-plugin
|
||||
fi
|
||||
if ! command -v ufw >/dev/null 2>&1; then
|
||||
sudo apt-get update -y
|
||||
sudo apt-get install -y ufw
|
||||
fi
|
||||
sudo ufw allow 22/tcp
|
||||
sudo ufw allow 80/tcp
|
||||
sudo ufw allow 443/tcp
|
||||
DOMAIN='${var.domain}'
|
||||
if [ -z "$DOMAIN" ]; then
|
||||
sudo ufw allow ${var.app_port}/tcp
|
||||
fi
|
||||
sudo ufw --force enable
|
||||
if [ -n "$DOMAIN" ]; then
|
||||
if ! command -v caddy >/dev/null 2>&1; then
|
||||
sudo apt-get install -y debian-keyring debian-archive-keyring apt-transport-https curl
|
||||
curl -1sLf 'https://dl.cloudsmith.io/public/caddy/stable/gpg.key' | sudo gpg --batch --yes --dearmor -o /usr/share/keyrings/caddy-stable-archive-keyring.gpg
|
||||
curl -1sLf 'https://dl.cloudsmith.io/public/caddy/stable/debian.deb.txt' | sudo tee /etc/apt/sources.list.d/caddy-stable.list >/dev/null
|
||||
sudo apt-get update -y
|
||||
sudo apt-get install -y caddy
|
||||
fi
|
||||
sudo install -m 0644 /tmp/thermograph.Caddyfile /etc/caddy/Caddyfile
|
||||
sudo systemctl reload caddy || sudo systemctl restart caddy
|
||||
fi
|
||||
rm -f /tmp/thermograph.Caddyfile
|
||||
EOT
|
||||
]
|
||||
}
|
||||
|
||||
# 1b. Self-hosted archive: rclone-mount the ERA5 bucket before compose up so
|
||||
# open-meteo-api can read .om from object storage. Skipped on non-openmeteo hosts.
|
||||
provisioner "remote-exec" {
|
||||
inline = [
|
||||
<<-EOT
|
||||
set -eu
|
||||
if [ "${var.openmeteo}" != "true" ]; then
|
||||
rm -f /tmp/thermograph.rclone.conf /tmp/rclone-om.service
|
||||
# Undo any prior openmeteo setup so a toggled-off host doesn't wait on a mount.
|
||||
if [ -e /etc/systemd/system/docker.service.d/10-wait-rclone.conf ]; then
|
||||
sudo rm -f /etc/systemd/system/docker.service.d/10-wait-rclone.conf
|
||||
sudo systemctl disable --now rclone-om >/dev/null 2>&1 || true
|
||||
sudo systemctl daemon-reload
|
||||
fi
|
||||
echo "[${var.name}] openmeteo: disabled"
|
||||
exit 0
|
||||
fi
|
||||
echo "[${var.name}] openmeteo: rclone mount ${var.om_bucket_remote} -> ${var.om_data_dir}"
|
||||
if ! command -v rclone >/dev/null 2>&1; then
|
||||
curl -fsSL https://rclone.org/install.sh | sudo bash
|
||||
fi
|
||||
# FUSE allow_other so the container (root) can read a mount owned by this user.
|
||||
if ! grep -q '^user_allow_other' /etc/fuse.conf 2>/dev/null; then
|
||||
echo user_allow_other | sudo tee -a /etc/fuse.conf >/dev/null
|
||||
fi
|
||||
sudo install -d -m 0755 /etc/rclone
|
||||
sudo install -m 0600 /tmp/thermograph.rclone.conf /etc/rclone/rclone.conf
|
||||
sudo install -m 0644 /tmp/rclone-om.service /etc/systemd/system/rclone-om.service
|
||||
rm -f /tmp/thermograph.rclone.conf /tmp/rclone-om.service
|
||||
sudo install -d -m 0755 ${var.om_data_dir}
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl enable rclone-om
|
||||
sudo systemctl restart rclone-om
|
||||
# Wait for the mount before compose bind-mounts it.
|
||||
i=0
|
||||
while [ "$i" -lt 30 ]; do
|
||||
if mountpoint -q ${var.om_data_dir}; then break; fi
|
||||
i=$((i + 1)); sleep 2
|
||||
done
|
||||
if ! mountpoint -q ${var.om_data_dir}; then
|
||||
echo "[${var.name}] RCLONE MOUNT FAILED" >&2
|
||||
sudo systemctl status rclone-om --no-pager || true
|
||||
exit 1
|
||||
fi
|
||||
# Order Docker after the mount on every boot, so the restart-policy containers
|
||||
# never bind an empty mount point. rclone-om is Type=notify, so `After` waits
|
||||
# until the mount is actually ready — not merely that the unit was launched.
|
||||
sudo install -d -m 0755 /etc/systemd/system/docker.service.d
|
||||
printf '[Unit]\nWants=rclone-om.service\nAfter=rclone-om.service\n' \
|
||||
| sudo tee /etc/systemd/system/docker.service.d/10-wait-rclone.conf >/dev/null
|
||||
sudo systemctl daemon-reload
|
||||
echo "[${var.name}] openmeteo: mounted at ${var.om_data_dir}"
|
||||
EOT
|
||||
]
|
||||
}
|
||||
|
||||
# 2. Sync the checkout to the host's branch (cloning first on a fresh box). This
|
||||
# is now the INFRA repo's checkout (compose files, db init, Caddy templates,
|
||||
# the secrets vault, deploy.sh itself) — the app's own source is never checked
|
||||
# out on the host; only its published registry image is pulled (step 3).
|
||||
provisioner "remote-exec" {
|
||||
inline = [
|
||||
<<-EOT
|
||||
set -eu
|
||||
echo "[${var.name}] code: ${var.app_dir} -> origin/${var.git_branch}"
|
||||
sudo mkdir -p ${var.app_dir}
|
||||
sudo chown -R "$(id -un)":"$(id -gn)" ${var.app_dir}
|
||||
if [ ! -d ${var.app_dir}/.git ]; then
|
||||
git clone ${var.repo_url} ${var.app_dir}
|
||||
fi
|
||||
cd ${var.app_dir}
|
||||
git fetch --prune origin ${var.git_branch}
|
||||
git reset --hard origin/${var.git_branch}
|
||||
EOT
|
||||
]
|
||||
}
|
||||
|
||||
# 3. Install /etc/thermograph-topology.env, render /etc/thermograph.env from the
|
||||
# SOPS+age vault (deploy/render-secrets.sh, now part of this same checkout),
|
||||
# then bring the stack up on the EXPLICIT per-service app image tags — there's no
|
||||
# app checkout on this host to derive a tag from (see var.backend_image_tag /
|
||||
# frontend_image_tag). The compose file reads BACKEND_IMAGE_TAG / FRONTEND_IMAGE_TAG
|
||||
# (for emi/thermograph-backend/app and emi/thermograph-frontend/app), so we export
|
||||
# those, not the old single IMAGE_PATH/IMAGE_TAG. docker runs as root (sources both
|
||||
# env files in the same shell) so it never depends on the docker group membership
|
||||
# taking effect in this session.
|
||||
provisioner "remote-exec" {
|
||||
inline = [
|
||||
<<-EOT
|
||||
set -eu
|
||||
echo "[${var.name}] deploy: install topology env + render secrets + docker compose pull + up"
|
||||
sudo install -m 0640 -o root -g root /tmp/thermograph-topology.env /etc/thermograph-topology.env
|
||||
rm -f /tmp/thermograph-topology.env
|
||||
sudo bash -c '
|
||||
set -eu
|
||||
cd ${var.app_dir}
|
||||
. deploy/render-secrets.sh
|
||||
render_thermograph_secrets ${var.app_dir}
|
||||
set -a; . /etc/thermograph-topology.env; . /etc/thermograph.env; set +a
|
||||
export REGISTRY_HOST="git.thermograph.org" BACKEND_IMAGE_TAG="${var.backend_image_tag}" FRONTEND_IMAGE_TAG="${var.frontend_image_tag}"
|
||||
echo "$REGISTRY_TOKEN" | docker login "$REGISTRY_HOST" --username emi --password-stdin
|
||||
docker compose ${local.compose_flags} pull backend frontend
|
||||
docker compose ${local.compose_flags} up -d --remove-orphans
|
||||
'
|
||||
EOT
|
||||
]
|
||||
}
|
||||
|
||||
# 4. Health check backend on loopback -- a plain "/" here already exercises
|
||||
# the whole chain end to end even without Caddy in front (backend's own
|
||||
# reverse-proxy fallback forwards to frontend), so one curl covers both
|
||||
# services in every topology this module supports (repo-split Stage 4).
|
||||
provisioner "remote-exec" {
|
||||
inline = [
|
||||
<<-EOT
|
||||
set -eu
|
||||
echo "[${var.name}] health: http://127.0.0.1:${var.app_port}/"
|
||||
ok=0
|
||||
i=0
|
||||
while [ "$i" -lt 30 ]; do
|
||||
if curl -fsS -o /dev/null "http://127.0.0.1:${var.app_port}/"; then ok=1; break; fi
|
||||
i=$((i + 1))
|
||||
sleep 2
|
||||
done
|
||||
if [ "$ok" != 1 ]; then
|
||||
echo "[${var.name}] HEALTH CHECK FAILED" >&2
|
||||
sudo bash -c 'cd ${var.app_dir} && docker compose ${local.compose_flags} ps; docker compose ${local.compose_flags} logs --tail=50 backend; docker compose ${local.compose_flags} logs --tail=50 frontend' || true
|
||||
exit 1
|
||||
fi
|
||||
echo "[${var.name}] OK: serving on 127.0.0.1:${var.app_port}"
|
||||
if ! curl -fsS -o /dev/null "http://127.0.0.1:${var.frontend_port}/healthz"; then
|
||||
echo "[${var.name}] frontend's own /healthz failed directly (backend's proxy to it may still work) -- check separately" >&2
|
||||
sudo bash -c 'cd ${var.app_dir} && docker compose ${local.compose_flags} logs --tail=50 frontend' || true
|
||||
fi
|
||||
EOT
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
output "host" {
|
||||
description = "Address this module manages."
|
||||
value = var.host
|
||||
}
|
||||
|
||||
output "role" {
|
||||
description = "Role of this host."
|
||||
value = var.role
|
||||
}
|
||||
|
|
@ -0,0 +1,45 @@
|
|||
# /etc/caddy/Caddyfile — RENDERED BY TERRAFORM for ${domain}.
|
||||
# Caddy provisions a Let's Encrypt cert on first request and auto-renews; the domain's
|
||||
# A/AAAA record must already point at this host and ports 80/443 must be open.
|
||||
#
|
||||
# The app owns the domain root and runs with THERMOGRAPH_BASE=/. Repo-split
|
||||
# Stage 4: backend (127.0.0.1:${port}) and frontend (127.0.0.1:${frontend_port})
|
||||
# are two containers now -- path-split directly to whichever owns a given
|
||||
# path. Repo-split Stage 7a flipped which side owns the enumerated list:
|
||||
# backend owns the short, stable set below (mirrors backend/web/app.py's own
|
||||
# routing exactly -- a single catch-all proxy to frontend for everything
|
||||
# else), so the two never disagree about who owns what.
|
||||
#
|
||||
# NOTE: the repo's deploy/Caddyfile additionally serves the emigriffith.dev portfolio
|
||||
# and legacy redirects; those are host-specific and intentionally not templated here.
|
||||
${domain} {
|
||||
encode zstd gzip
|
||||
|
||||
@backend_paths path /api/* /digest /discord/interactions
|
||||
|
||||
# Active health check on the same cheap /healthz route each container's own
|
||||
# HEALTHCHECK uses (Dockerfile) — so a `docker compose up -d --build` deploy
|
||||
# that's still restarting/booting never gets proxied into (a reload alone has
|
||||
# no gate, hop-1 runbook hazard #10). health_uri is relative to the upstream.
|
||||
handle @backend_paths {
|
||||
reverse_proxy 127.0.0.1:${port} {
|
||||
health_uri /healthz
|
||||
health_interval 5s
|
||||
health_timeout 3s
|
||||
health_status 2xx
|
||||
}
|
||||
}
|
||||
|
||||
handle {
|
||||
reverse_proxy 127.0.0.1:${frontend_port} {
|
||||
health_uri /healthz
|
||||
health_interval 5s
|
||||
health_timeout 3s
|
||||
health_status 2xx
|
||||
}
|
||||
}
|
||||
|
||||
log {
|
||||
output file /var/log/caddy/thermograph.log
|
||||
}
|
||||
}
|
||||
|
|
@ -0,0 +1,33 @@
|
|||
# /etc/thermograph-topology.env — RENDERED BY TERRAFORM. Do not edit on the host;
|
||||
# change the tfvars and re-apply. Non-secret sizing/routing config ONLY; every app
|
||||
# secret (POSTGRES_PASSWORD, THERMOGRAPH_AUTH_SECRET, VAPID keys, REGISTRY_TOKEN,
|
||||
# Discord/SMTP creds, ...) lives in /etc/thermograph.env instead, rendered at deploy
|
||||
# time from the SOPS+age vault by deploy/render-secrets.sh (see deploy.sh). Both
|
||||
# files are sourced before `docker compose` runs — this one first, so a same-named
|
||||
# vault value (there shouldn't be one) would still win.
|
||||
|
||||
# Port uvicorn binds inside the container (compose publishes it on the host).
|
||||
PORT=${app_port}
|
||||
|
||||
# --- Container sizing (consumed by docker-compose.yml interpolation) -------------
|
||||
# Terraform sizes the containers per host. Defaults in compose are 4 / 4 / 2 / 8g.
|
||||
WORKERS=${workers}
|
||||
APP_CPUS=${app_cpus}
|
||||
DB_CPUS=${db_cpus}
|
||||
DB_MEMORY=${db_memory}
|
||||
# "latest-pg18" (the default) matches today's behavior. Pin an exact minor before
|
||||
# any host of this stack could ever replicate with another (docker-compose.yml).
|
||||
TIMESCALEDB_TAG=${timescaledb_tag}
|
||||
%{ if openmeteo ~}
|
||||
|
||||
# --- Self-hosted Open-Meteo archive (docker-compose.openmeteo.yml) ---------------
|
||||
# Host rclone mount of the ERA5 object-storage bucket; the overlay bind-mounts it
|
||||
# into the Open-Meteo containers. THERMOGRAPH_ARCHIVE_URL is set by the overlay.
|
||||
OM_DATA_DIR=${om_data_dir}
|
||||
%{ endif ~}
|
||||
|
||||
# --- Serving --------------------------------------------------------------------
|
||||
THERMOGRAPH_BASE=${base}
|
||||
THERMOGRAPH_BASE_URL=${base_url}
|
||||
# Secure cookie is only sent over HTTPS: 1 behind Caddy TLS, 0 on a plain-HTTP host.
|
||||
THERMOGRAPH_COOKIE_SECURE=${cookie_secure}
|
||||
140
infra/terraform/modules/thermograph-host/variables.tf
Normal file
140
infra/terraform/modules/thermograph-host/variables.tf
Normal file
|
|
@ -0,0 +1,140 @@
|
|||
# Per-host inputs (all supplied by the root module's for_each).
|
||||
|
||||
variable "name" {
|
||||
description = "Short host key (e.g. \"prod\", \"dev\"), used in log lines."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "host" {
|
||||
description = "IP or hostname to SSH to."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "ssh_user" {
|
||||
description = "SSH login user (must be able to sudo)."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "ssh_private_key_path" {
|
||||
description = "Path to the private key file for ssh_user."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "role" {
|
||||
description = "\"prod\" | \"dev\" — informational."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "git_branch" {
|
||||
description = "This INFRA repo's branch the host checkout is reset to (independent of which app images are deployed — see backend_image_tag / frontend_image_tag)."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "backend_image_tag" {
|
||||
description = "Backend image tag to pull (emi/thermograph-backend/app), e.g. \"sha-<12 hex>\" (build-push.yml's tag for the backend-repo commit) or a semver tag. The host has no app-repo checkout to derive this from, so it's always explicit."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "frontend_image_tag" {
|
||||
description = "Frontend image tag to pull (emi/thermograph-frontend/app), e.g. \"sha-<12 hex>\" (build-push.yml's tag for the frontend-repo commit) or a semver tag. Always explicit, same as backend_image_tag."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "domain" {
|
||||
description = "Public domain. \"\" => no Caddy/TLS (open the app port instead)."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "compose_files" {
|
||||
description = "Compose files to layer, in order (dev appends docker-compose.dev.yml)."
|
||||
type = list(string)
|
||||
}
|
||||
|
||||
variable "openmeteo" {
|
||||
description = "Self-host the ERA5 archive: layer docker-compose.openmeteo.yml + provision the host rclone mount."
|
||||
type = bool
|
||||
default = false
|
||||
}
|
||||
|
||||
variable "om_data_dir" {
|
||||
description = "Host rclone mount point for the archive bucket (OM_DATA_DIR the overlay bind-mounts)."
|
||||
type = string
|
||||
default = "/mnt/om-archive"
|
||||
}
|
||||
|
||||
variable "om_bucket_remote" {
|
||||
description = "rclone remote:path for the archive bucket (mounted at om_data_dir)."
|
||||
type = string
|
||||
default = ""
|
||||
}
|
||||
|
||||
variable "om_rclone_conf" {
|
||||
description = "rclone.conf contents installed to /etc/rclone/rclone.conf. Sensitive."
|
||||
type = string
|
||||
default = ""
|
||||
sensitive = true
|
||||
}
|
||||
|
||||
variable "om_vfs_cache_max" {
|
||||
description = "rclone --vfs-cache-max-size for the mount's on-disk hot cache."
|
||||
type = string
|
||||
default = "80G"
|
||||
}
|
||||
|
||||
variable "app_dir" {
|
||||
description = "Checkout path on the host."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "repo_root" {
|
||||
description = "Local repo root, used to hash the compose files for the re-apply trigger."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "repo_url" {
|
||||
description = "Git remote to clone from if the host has no checkout yet."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "app_port" {
|
||||
description = "Port backend binds / is health-checked on."
|
||||
type = number
|
||||
}
|
||||
|
||||
variable "frontend_port" {
|
||||
description = "Port the frontend SSR service binds / is health-checked on (repo-split Stage 4). Loopback-only, never opened in ufw -- reached via Caddy's path-split or backend's own reverse-proxy fallback, never directly."
|
||||
type = number
|
||||
default = 8080
|
||||
}
|
||||
|
||||
# ---- Sizing -------------------------------------------------------------------
|
||||
variable "workers" {
|
||||
description = "uvicorn worker count (WORKERS)."
|
||||
type = number
|
||||
}
|
||||
|
||||
variable "app_cpus" {
|
||||
description = "App container CPU cap (APP_CPUS)."
|
||||
type = number
|
||||
}
|
||||
|
||||
variable "db_cpus" {
|
||||
description = "DB container CPU cap (DB_CPUS)."
|
||||
type = number
|
||||
}
|
||||
|
||||
variable "db_memory" {
|
||||
description = "DB container memory cap (DB_MEMORY), e.g. \"8g\"."
|
||||
type = string
|
||||
}
|
||||
|
||||
variable "timescaledb_tag" {
|
||||
description = "TimescaleDB image tag (TIMESCALEDB_TAG), e.g. \"2.17.2-pg18\". \"latest-pg18\" (the default) matches today's behavior; pin an exact minor before any host could ever replicate with another."
|
||||
type = string
|
||||
default = "latest-pg18"
|
||||
}
|
||||
|
||||
# Secrets (POSTGRES_PASSWORD, THERMOGRAPH_AUTH_SECRET, VAPID keys, REGISTRY_TOKEN,
|
||||
# Discord/SMTP credentials, ...) are no longer Terraform variables -- they're
|
||||
# rendered at deploy time from the SOPS+age vault (deploy/secrets/*.yaml) by
|
||||
# deploy/render-secrets.sh, which deploy.sh calls. See main.tf's remote-exec step 3.
|
||||
11
infra/terraform/modules/thermograph-host/versions.tf
Normal file
11
infra/terraform/modules/thermograph-host/versions.tf
Normal file
|
|
@ -0,0 +1,11 @@
|
|||
terraform {
|
||||
required_version = ">= 1.6"
|
||||
|
||||
required_providers {
|
||||
# The module drives everything through a null_resource + SSH provisioners.
|
||||
null = {
|
||||
source = "hashicorp/null"
|
||||
version = "~> 3.2"
|
||||
}
|
||||
}
|
||||
}
|
||||
14
infra/terraform/outputs.tf
Normal file
14
infra/terraform/outputs.tf
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
output "prod_url" {
|
||||
description = "Public URL of the production site."
|
||||
value = "https://thermograph.org"
|
||||
}
|
||||
|
||||
output "hosts" {
|
||||
description = "Per-host summary: the address managed and its role."
|
||||
value = {
|
||||
for name, mod in module.host : name => {
|
||||
host = mod.host
|
||||
role = mod.role
|
||||
}
|
||||
}
|
||||
}
|
||||
125
infra/terraform/terraform.tfvars.example
Normal file
125
infra/terraform/terraform.tfvars.example
Normal file
|
|
@ -0,0 +1,125 @@
|
|||
# Copy to terraform.tfvars and fill in real IPs + credentials.
|
||||
# cp terraform.tfvars.example terraform.tfvars
|
||||
# terraform.tfvars is gitignored (repo_url may carry a credential, and om_rclone_conf
|
||||
# always does — both land in local state). NEVER commit real values.
|
||||
#
|
||||
# App secrets (POSTGRES_PASSWORD, THERMOGRAPH_AUTH_SECRET, VAPID keys,
|
||||
# REGISTRY_TOKEN, Discord/SMTP creds, ...) are NOT here anymore -- they live in
|
||||
# the SOPS+age vault (../deploy/secrets/*.yaml), rendered at deploy time. See
|
||||
# deploy/secrets/README.md to rotate or add one.
|
||||
|
||||
# ---------------------------------------------------------------------------------
|
||||
# Hosts
|
||||
# ---------------------------------------------------------------------------------
|
||||
# Two VPS hosts. (The `dev` branch deploys to the LAN dev server via the app repo's
|
||||
# deploy/deploy-dev.sh — that box is NOT managed by Terraform.)
|
||||
hosts = {
|
||||
# Production: the NEW 48 GB / 12-core VPS serving thermograph.org.
|
||||
prod = {
|
||||
host = "REPLACE_WITH_NEW_VPS_IP" # <-- the new prod VPS IP/hostname
|
||||
ssh_user = "agent"
|
||||
ssh_private_key_path = "~/.ssh/thermograph_agent_ed25519"
|
||||
role = "prod"
|
||||
git_branch = "main" # this INFRA repo's branch -- see *_image_tag for the app versions
|
||||
backend_image_tag = "REPLACE_WITH_BACKEND_IMAGE_TAG" # e.g. "sha-abcdef012345" -- from thermograph-backend build-push.yml
|
||||
frontend_image_tag = "REPLACE_WITH_FRONTEND_IMAGE_TAG" # e.g. "sha-012345abcdef" -- from thermograph-frontend build-push.yml
|
||||
domain = "thermograph.org" # Caddy TLS in front, app on loopback
|
||||
compose_files = ["docker-compose.yml"]
|
||||
app_dir = "/opt/thermograph"
|
||||
# "large" is the named size tier for this box (locals.sizes in main.tf) — same
|
||||
# numbers as hand-picking workers=8/app_cpus=8/db_cpus=4/db_memory="16g" below,
|
||||
# via the shortcut. The Postgres internal budget scales from db_memory
|
||||
# automatically (deploy/db/init/20-tuning.sh); no separate tuning edit.
|
||||
size = "large"
|
||||
# Self-host the ERA5 archive here: layers docker-compose.openmeteo.yml and
|
||||
# provisions the rclone mount of the object-storage bucket (om_* vars below).
|
||||
openmeteo = true
|
||||
om_data_dir = "/mnt/om-archive"
|
||||
}
|
||||
|
||||
# Beta / testing: the OLD VPS, repurposed.
|
||||
beta = {
|
||||
host = "75.119.132.91"
|
||||
ssh_user = "agent"
|
||||
ssh_private_key_path = "~/.ssh/thermograph_agent_ed25519"
|
||||
role = "beta"
|
||||
git_branch = "main"
|
||||
backend_image_tag = "REPLACE_WITH_BACKEND_IMAGE_TAG"
|
||||
frontend_image_tag = "REPLACE_WITH_FRONTEND_IMAGE_TAG"
|
||||
# No public domain by default: no Caddy/TLS, firewall opens the app port. NOTE:
|
||||
# with compose_files = ["docker-compose.yml"] the app binds 127.0.0.1 only, so
|
||||
# until you either set a domain (e.g. "beta.thermograph.org", which fronts it with
|
||||
# Caddy) or add the 0.0.0.0-publishing dev overlay, reach it via an SSH tunnel.
|
||||
domain = ""
|
||||
compose_files = ["docker-compose.yml"]
|
||||
app_dir = "/opt/thermograph"
|
||||
# Explicit numbers, not a size tier — both styles work on any host; a tier is
|
||||
# purely an opt-in shortcut (see prod's `size = "large"` above).
|
||||
workers = 4
|
||||
app_cpus = 4
|
||||
db_cpus = 2
|
||||
db_memory = "8g"
|
||||
}
|
||||
|
||||
# UAT: an ephemeral, single-node environment — same images/topology shape as
|
||||
# prod, not prod's scale (design doc §6/§9). Uncomment once a UAT box exists;
|
||||
# not managed until then. "nano" keeps it cheap since it's destroyed when idle.
|
||||
# uat = {
|
||||
# host = "REPLACE_WITH_UAT_VM_IP"
|
||||
# ssh_user = "agent"
|
||||
# ssh_private_key_path = "~/.ssh/thermograph_agent_ed25519"
|
||||
# role = "uat"
|
||||
# git_branch = "main"
|
||||
# backend_image_tag = "REPLACE_WITH_BACKEND_IMAGE_TAG"
|
||||
# frontend_image_tag = "REPLACE_WITH_FRONTEND_IMAGE_TAG"
|
||||
# domain = ""
|
||||
# compose_files = ["docker-compose.yml"]
|
||||
# app_dir = "/opt/thermograph"
|
||||
# size = "nano"
|
||||
# }
|
||||
}
|
||||
|
||||
# GCP-created hosts — SCAFFOLD ONLY, empty by default. Populate an entry to have
|
||||
# Terraform actually create a GCP Compute Engine VM (see modules/gcp-host); until
|
||||
# then no google_* resource is planned and no GCP credentials are needed. Example:
|
||||
# gcp_hosts = {
|
||||
# gcp-uat = {
|
||||
# project = "REPLACE_WITH_GCP_PROJECT_ID"
|
||||
# zone = "us-west1-a"
|
||||
# machine_type = "e2-medium"
|
||||
# ssh_user = "agent"
|
||||
# ssh_public_key_path = "~/.ssh/thermograph_agent_ed25519.pub"
|
||||
# ssh_private_key_path = "~/.ssh/thermograph_agent_ed25519"
|
||||
# role = "uat"
|
||||
# git_branch = "main"
|
||||
# backend_image_tag = "REPLACE_WITH_BACKEND_IMAGE_TAG"
|
||||
# frontend_image_tag = "REPLACE_WITH_FRONTEND_IMAGE_TAG"
|
||||
# size = "nano"
|
||||
# }
|
||||
# }
|
||||
|
||||
# Optional overrides (shown with their defaults):
|
||||
# repo_url = "https://git.thermograph.org/emi/thermograph-infra.git"
|
||||
# app_port = 8137
|
||||
# frontend_port = 8080
|
||||
#
|
||||
# thermograph-infra is a PRIVATE repo, so a host cloning it for the first time
|
||||
# needs read credentials embedded in repo_url, e.g. a Forgejo deploy token:
|
||||
# repo_url = "https://deploy:REPLACE_WITH_TOKEN@git.thermograph.org/emi/thermograph-infra.git"
|
||||
|
||||
# ---------------------------------------------------------------------------------
|
||||
# Self-hosted Open-Meteo archive (only used by hosts with openmeteo = true) --------
|
||||
# ---------------------------------------------------------------------------------
|
||||
# The ERA5 .om archive lives in an object-storage bucket, rclone-mounted on the host.
|
||||
# om_rclone_conf holds bucket credentials (sensitive; lands in state — keep out of git).
|
||||
# See deploy/openmeteo/README.md for the bucket + mount setup.
|
||||
om_bucket_remote = "om-archive:REPLACE_WITH_BUCKET_NAME"
|
||||
om_vfs_cache_max = "80G"
|
||||
om_rclone_conf = <<-RCLONE
|
||||
[om-archive]
|
||||
type = s3
|
||||
provider = Cloudflare
|
||||
endpoint = https://REPLACE.r2.cloudflarestorage.com
|
||||
access_key_id = REPLACE_WITH_ACCESS_KEY
|
||||
secret_access_key = REPLACE_WITH_SECRET_KEY
|
||||
RCLONE
|
||||
145
infra/terraform/variables.tf
Normal file
145
infra/terraform/variables.tf
Normal file
|
|
@ -0,0 +1,145 @@
|
|||
# ---------------------------------------------------------------------------------
|
||||
# Hosts
|
||||
# ---------------------------------------------------------------------------------
|
||||
# One entry per VPS. The same module is instantiated for each (see main.tf's
|
||||
# for_each). This config manages TWO hosts: prod (new 48 GB / 12-core VPS, branch
|
||||
# `release`, thermograph.org) and beta (the old VPS 75.119.132.91, branch `main`,
|
||||
# no public domain by default). The `dev` branch deploys to the LAN dev server via
|
||||
# deploy/deploy-dev.sh (a self-hosted runner, in the app repo) and is NOT managed here.
|
||||
#
|
||||
# A host with `domain = ""` gets no Caddy/TLS: the app port is opened on the firewall
|
||||
# and the app is reached directly (beta does this by default; set a domain to front
|
||||
# it with Caddy TLS).
|
||||
#
|
||||
# Resource sizing (workers / app_cpus / db_cpus / db_memory) defaults to the historical
|
||||
# 4 / 4 / 2 / 8g budget (right for beta). The prod box is 48 GB / 12 cores — raise these
|
||||
# there (the example uses app_cpus 8, db_cpus 4, db_memory "16g"). The Postgres internal
|
||||
# budget scales from db_memory automatically (deploy/db/init/20-tuning.sh), so no
|
||||
# separate tuning edit is needed to exploit the RAM.
|
||||
variable "hosts" {
|
||||
description = "Map of hosts to manage, keyed by a short name (e.g. \"prod\", \"beta\")."
|
||||
type = map(object({
|
||||
host = string # IP or hostname to SSH to
|
||||
ssh_user = optional(string, "deploy") # SSH login user
|
||||
ssh_private_key_path = string # path to the private key for that user
|
||||
role = string # "prod" | "beta" (informational + outputs)
|
||||
git_branch = string # this INFRA repo's branch the checkout is reset to
|
||||
# Which app images to run, e.g. "sha-<12 hex>" (each matching build-push.yml's tag
|
||||
# for the commit that app repo built) or a semver tag on a release push. The app is
|
||||
# TWO separately-published images now — emi/thermograph-backend/app and
|
||||
# emi/thermograph-frontend/app — pinned independently. Required, no default: the
|
||||
# host's checkout is this infra repo, not the app repos, so there's no "current
|
||||
# commit" to derive a tag from; both must be explicit. Bump these (via a normal
|
||||
# tfvars edit + apply) whenever an app repo ships a commit you want this host
|
||||
# running; the infra repo's own git_branch is independent and rarely needs to change.
|
||||
backend_image_tag = string
|
||||
frontend_image_tag = string
|
||||
domain = optional(string, "") # public domain; "" => no Caddy/TLS
|
||||
compose_files = optional(list(string), ["docker-compose.yml"])
|
||||
app_dir = optional(string, "/opt/thermograph") # checkout path on the host
|
||||
workers = optional(number, 4) # uvicorn workers (WORKERS)
|
||||
app_cpus = optional(number, 4) # app container CPU cap (APP_CPUS)
|
||||
db_cpus = optional(number, 2) # db container CPU cap (DB_CPUS)
|
||||
db_memory = optional(string, "8g") # db container memory cap (DB_MEMORY)
|
||||
# A named size tier (see locals.sizes in main.tf: nano/small/medium/large) —
|
||||
# when set, overrides the four fields above with the tier's preset. Leave
|
||||
# null (default) to keep hand-picking workers/app_cpus/db_cpus/db_memory,
|
||||
# as prod/beta already do below.
|
||||
size = optional(string, null)
|
||||
# The floating tag matches today's behavior everywhere until you pin it. Pin to
|
||||
# an exact minor (SELECT extversion FROM pg_extension WHERE extname='timescaledb'
|
||||
# on the live DB) before any host of this stack could ever replicate with
|
||||
# another — a floating tag risks mismatched extension minors, which blocks a
|
||||
# physical replica (hop-1 cutover runbook hazard #7). Use the SAME tag everywhere.
|
||||
timescaledb_tag = optional(string, "latest-pg18")
|
||||
# Self-host the ERA5 archive (docker-compose.openmeteo.yml + a host rclone mount
|
||||
# of the object-storage bucket). Only the self-hosting host (prod) sets true.
|
||||
openmeteo = optional(bool, false)
|
||||
om_data_dir = optional(string, "/mnt/om-archive") # host rclone mount point (OM_DATA_DIR)
|
||||
}))
|
||||
}
|
||||
|
||||
# GCP-created hosts, keyed the same way as `hosts`. Default {} => zero GCP resources
|
||||
# planned and the google provider is never actually invoked (see versions.tf and
|
||||
# modules/gcp-host). Populate an entry to have Terraform create the VM itself; its
|
||||
# output IP then feeds into the SAME thermograph-host module every SSH-managed host
|
||||
# uses (main.tf), so provisioning logic is never duplicated between providers.
|
||||
variable "gcp_hosts" {
|
||||
description = "Map of hosts for Terraform to CREATE on GCP (Compute Engine), keyed the same way as `hosts`. Empty by default -- no live GCP resources exist yet; this is a scaffold for future use. See modules/gcp-host."
|
||||
type = map(object({
|
||||
project = string # GCP project ID
|
||||
zone = string # e.g. "us-west1-a"
|
||||
machine_type = optional(string, "e2-medium")
|
||||
ssh_user = optional(string, "deploy")
|
||||
ssh_public_key_path = string # path to the PUBLIC key installed on the instance
|
||||
ssh_private_key_path = string # path to the matching PRIVATE key (for the module's provisioner)
|
||||
role = string
|
||||
git_branch = string
|
||||
backend_image_tag = string
|
||||
frontend_image_tag = string
|
||||
domain = optional(string, "")
|
||||
compose_files = optional(list(string), ["docker-compose.yml"])
|
||||
app_dir = optional(string, "/opt/thermograph")
|
||||
size = optional(string, "small")
|
||||
timescaledb_tag = optional(string, "latest-pg18")
|
||||
}))
|
||||
default = {}
|
||||
}
|
||||
|
||||
variable "repo_url" {
|
||||
description = "Git remote to clone from when a host has no checkout yet. Points at THIS repo (thermograph-infra) now, not the app repo -- the app's own source is never checked out on a host; only its published registry images are pulled (see var.hosts[*].backend_image_tag / frontend_image_tag). thermograph-infra is a private repo, so this typically needs embedded read credentials, e.g. a Forgejo deploy token: \"https://<token-name>:<token>@git.thermograph.org/emi/thermograph-infra.git\"."
|
||||
type = string
|
||||
default = "https://git.thermograph.org/emi/thermograph-infra.git"
|
||||
sensitive = true
|
||||
}
|
||||
|
||||
variable "app_port" {
|
||||
description = "Port the app binds inside the container / is health-checked on."
|
||||
type = number
|
||||
default = 8137
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------------------------
|
||||
# Self-hosted Open-Meteo (object storage) — consumed only by hosts with openmeteo=true
|
||||
# ---------------------------------------------------------------------------------
|
||||
# The ERA5 .om archive lives in an object-storage bucket, surfaced on the host by an
|
||||
# rclone FUSE mount at each host's om_data_dir. These describe that bucket + mount.
|
||||
# om_rclone_conf holds credentials, so it's sensitive and lands in state — keep the
|
||||
# real value in terraform.tfvars (gitignored), never committed.
|
||||
variable "om_bucket_remote" {
|
||||
description = "rclone remote:path for the archive bucket, e.g. \"om-archive:thermograph-era5\" (matches a [remote] in om_rclone_conf)."
|
||||
type = string
|
||||
default = ""
|
||||
}
|
||||
|
||||
variable "om_rclone_conf" {
|
||||
description = "Full rclone.conf contents defining the archive remote (installed to /etc/rclone/rclone.conf, 0600). Sensitive."
|
||||
type = string
|
||||
default = ""
|
||||
sensitive = true
|
||||
}
|
||||
|
||||
variable "om_vfs_cache_max" {
|
||||
description = "rclone --vfs-cache-max-size: bounds the on-disk hot cache for the mount (keep within the disk budget)."
|
||||
type = string
|
||||
default = "80G"
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------------------------
|
||||
# Secrets: owned by the SOPS+age vault now, NOT Terraform
|
||||
# ---------------------------------------------------------------------------------
|
||||
# POSTGRES_PASSWORD, THERMOGRAPH_AUTH_SECRET, THERMOGRAPH_METRICS_TOKEN,
|
||||
# THERMOGRAPH_INDEXNOW_KEY, THERMOGRAPH_VAPID_*, REGISTRY_TOKEN, Discord/SMTP
|
||||
# credentials, and every other app secret are no longer Terraform variables --
|
||||
# they're rendered at deploy time from deploy/secrets/*.yaml (SOPS-encrypted,
|
||||
# committed) by deploy/render-secrets.sh, which deploy.sh calls before `docker
|
||||
# compose up`. Terraform's job here shrank to topology/sizing only (see
|
||||
# modules/thermograph-host/templates/thermograph-topology.env.tftpl) plus
|
||||
# triggering the deploy. See deploy/secrets/README.md to rotate or add a secret.
|
||||
#
|
||||
# The four secrets Terraform used to generate itself (postgres_password,
|
||||
# auth_secret, metrics_token, indexnow_key) still exist as values -- they just
|
||||
# live in the vault now, seeded once from Terraform's own generated values via
|
||||
# deploy/secrets/seed-from-live.sh so the handoff changed nothing in use. Rotate
|
||||
# them the same way as any other vault secret from here on (sops edit + commit +
|
||||
# deploy), not via a Terraform keeper change.
|
||||
29
infra/terraform/versions.tf
Normal file
29
infra/terraform/versions.tf
Normal file
|
|
@ -0,0 +1,29 @@
|
|||
terraform {
|
||||
# Terraform v1.15 is installed; the optional() object defaults used here need >= 1.3.
|
||||
required_version = ">= 1.6"
|
||||
|
||||
# Hosts under var.hosts already exist and are driven over SSH by the null
|
||||
# provider's provisioners (local for rendering). google is scaffold-only: with
|
||||
# var.gcp_hosts left at its default {}, no google_* resource is ever planned and
|
||||
# the provider is never actually invoked -- `terraform plan`/`validate` work with
|
||||
# no GCP credentials configured at all. Populate var.gcp_hosts (and the usual
|
||||
# GOOGLE_APPLICATION_CREDENTIALS / gcloud auth) to actually create a GCP host.
|
||||
required_providers {
|
||||
null = {
|
||||
source = "hashicorp/null"
|
||||
version = "~> 3.2"
|
||||
}
|
||||
local = {
|
||||
source = "hashicorp/local"
|
||||
version = "~> 2.5"
|
||||
}
|
||||
google = {
|
||||
source = "hashicorp/google"
|
||||
version = "~> 6.0"
|
||||
}
|
||||
time = {
|
||||
source = "hashicorp/time"
|
||||
version = "~> 0.12"
|
||||
}
|
||||
}
|
||||
}
|
||||
Loading…
Reference in a new issue