2026-07-11 00:29:47 +00:00
|
|
|
# /etc/caddy/Caddyfile on the VPS.
|
2026-07-15 19:58:32 +00:00
|
|
|
# Each domain's A/AAAA record must already point at this VPS — Caddy provisions a
|
2026-07-11 03:01:16 +00:00
|
|
|
# Let's Encrypt cert on first request and auto-renews. Nothing else to do for TLS
|
|
|
|
|
# (just make sure ports 80 and 443 are open).
|
|
|
|
|
#
|
2026-07-15 19:58:32 +00:00
|
|
|
# Layout:
|
2026-07-21 20:01:30 +00:00
|
|
|
# thermograph.org/* -> the Thermograph app (path-split across
|
|
|
|
|
# backend/frontend -- see below)
|
2026-07-15 19:58:32 +00:00
|
|
|
# emigriffith.dev/ -> static portfolio site (served straight from disk)
|
|
|
|
|
# emigriffith.dev/thermograph* -> permanent redirect to thermograph.org (the app moved)
|
2026-07-11 03:01:16 +00:00
|
|
|
#
|
2026-07-21 20:01:30 +00:00
|
|
|
# Thermograph now owns thermograph.org's root, so both services run with
|
|
|
|
|
# THERMOGRAPH_BASE=/ (see /etc/thermograph.env) — pages, assets and API all sit
|
|
|
|
|
# at "/" with no sub-path prefix. Repo-split Stage 4: backend and frontend are
|
|
|
|
|
# two containers now (docker-compose.yml), each on its own loopback port --
|
|
|
|
|
# Caddy path-splits directly to whichever owns a given path, so the browser
|
2026-07-21 22:48:59 +00:00
|
|
|
# still sees one apparent origin. Repo-split Stage 7a flipped which side owns
|
|
|
|
|
# the enumerated list: frontend now owns everything (content pages, the
|
|
|
|
|
# interactive tool's SPA shells, every static asset, the dynamic IndexNow key
|
|
|
|
|
# file) except the short, stable set below, which mirrors backend/web/app.py's
|
|
|
|
|
# own routing exactly (a single catch-all proxy to frontend for everything
|
|
|
|
|
# else) -- unlike frontend's paths, backend's don't grow every time a new
|
|
|
|
|
# static asset filename is added. A gap in this list still just degrades to
|
|
|
|
|
# "one extra hop" through backend's own proxy fallback, never a 404.
|
2026-07-15 19:58:32 +00:00
|
|
|
|
|
|
|
|
thermograph.org {
|
|
|
|
|
encode zstd gzip
|
|
|
|
|
|
2026-07-21 22:48:59 +00:00
|
|
|
@backend_paths path /api/* /digest /discord/interactions
|
2026-07-21 20:01:30 +00:00
|
|
|
|
|
|
|
|
# Active health check on the same cheap /healthz route each container's own
|
Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235)
* Split web/worker duties with THERMOGRAPH_ROLE
Background work (the subscription notifier) is welded to the same process
that serves requests, so scaling the web tier to N replicas would also scale
notifier instances unless something restricts it further than leader
election alone.
Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process
behavior). Every replica runs the same image; ROLE only gates whether a
process is allowed to own the notifier at all, layered on top of the
existing leader election: web replicas never start it even if they'd win
leader election, worker replicas start it if they win. The decision is
pulled into _should_run_notifier() so it's unit-testable without booting the
full app (DB init, places index, neighbor warmer).
Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE)
so a worker replica - which serves no real traffic - still has something
Swarm can health-check.
* Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate
Three changes toward the hop-1 interim cutover, all inert until Track B
stands up the platform:
docker-stack.yml: the Swarm stack file for the interim cutover, distinct
from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a
pre-built image (IMAGE_TAG) instead of building in place; app/worker publish
no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing
mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted),
reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard
needs a smaller MTU or large payloads silently stall); db is placement-
pinned to a labelled node; app/worker skip inline migrations
(RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing
that ever runs Alembic; secrets are real Swarm secrets mounted at
/run/secrets, read by the entrypoint shim rather than plain env vars.
TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default
latest-pg18, today's behavior unchanged), wired through Terraform
(timescaledb_tag, default "latest-pg18") so it can actually be pinned to an
exact minor without hand-editing the host - required before any host of the
stack could replicate with another (a floating tag risks mismatched
extension minors, which blocks a physical replica and risks compressed-
chunk corruption on restore).
Caddy active health-gate: both the Terraform-rendered Caddyfile and the live
deploy/Caddyfile now health-check the app on the same cheap /healthz route
its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage,
so it's cheap enough for a tight interval and works identically for a
worker replica, which serves no public traffic at all) - Caddy won't
forward into a container that's still booting or unhealthy.
Verified live: built and booted the real image via docker compose - both
containers report healthy via the new /healthz-based HEALTHCHECK, and GET /
still renders the full SSR homepage unchanged. Both Caddyfiles validated
with the real caddy binary. docker-stack.yml validated with docker compose
config (required-var guards fire with clear messages; secrets correctly
mount at /run/secrets/<name>, matching the entrypoint shim's mapping).
docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside
the existing openmeteo overlay. terraform validate + fmt clean.
2026-07-21 00:39:48 +00:00
|
|
|
# HEALTHCHECK uses (Dockerfile) — so a deploy that's still restarting/booting
|
|
|
|
|
# never gets proxied into (a reload alone has no gate, hop-1 runbook hazard #10).
|
2026-07-24 04:37:41 +00:00
|
|
|
# 15s (was 5s): plenty responsive for a process that only restarts on a deploy,
|
|
|
|
|
# and a quarter of the polling load.
|
2026-07-21 22:48:59 +00:00
|
|
|
handle @backend_paths {
|
|
|
|
|
reverse_proxy 127.0.0.1:8137 {
|
2026-07-21 20:01:30 +00:00
|
|
|
health_uri /healthz
|
2026-07-24 04:37:41 +00:00
|
|
|
health_interval 15s
|
2026-07-21 20:01:30 +00:00
|
|
|
health_timeout 3s
|
|
|
|
|
health_status 2xx
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
handle {
|
2026-07-21 22:48:59 +00:00
|
|
|
reverse_proxy 127.0.0.1:8080 {
|
2026-07-21 20:01:30 +00:00
|
|
|
health_uri /healthz
|
2026-07-24 04:37:41 +00:00
|
|
|
health_interval 15s
|
2026-07-21 20:01:30 +00:00
|
|
|
health_timeout 3s
|
|
|
|
|
health_status 2xx
|
|
|
|
|
}
|
Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate (#235)
* Split web/worker duties with THERMOGRAPH_ROLE
Background work (the subscription notifier) is welded to the same process
that serves requests, so scaling the web tier to N replicas would also scale
notifier instances unless something restricts it further than leader
election alone.
Add THERMOGRAPH_ROLE (web|worker|all, default all - unchanged single-process
behavior). Every replica runs the same image; ROLE only gates whether a
process is allowed to own the notifier at all, layered on top of the
existing leader election: web replicas never start it even if they'd win
leader election, worker replicas start it if they win. The decision is
pulled into _should_run_notifier() so it's unit-testable without booting the
full app (DB init, places index, neighbor warmer).
Add a minimal /healthz liveness route (no DB/upstream I/O, not under BASE)
so a worker replica - which serves no real traffic - still has something
Swarm can health-check.
* Add the Swarm interim stack file, a pinnable TimescaleDB tag, and a Caddy health-gate
Three changes toward the hop-1 interim cutover, all inert until Track B
stands up the platform:
docker-stack.yml: the Swarm stack file for the interim cutover, distinct
from docker-compose.yml (today's plain-compose deploy, unaffected). Pulls a
pre-built image (IMAGE_TAG) instead of building in place; app/worker publish
no host port (127.0.0.1:8137:8137 has no Swarm equivalent - Swarm's routing
mesh publishes on 0.0.0.0, which would expose the plaintext app un-fronted),
reaching Caddy only over an MTU-lowered overlay network (VXLAN-over-WireGuard
needs a smaller MTU or large payloads silently stall); db is placement-
pinned to a labelled node; app/worker skip inline migrations
(RUN_MIGRATIONS=0) so the runbook's one-shot migrate task is the only thing
that ever runs Alembic; secrets are real Swarm secrets mounted at
/run/secrets, read by the entrypoint shim rather than plain env vars.
TIMESCALEDB_TAG: docker-compose.yml's db image now reads this (default
latest-pg18, today's behavior unchanged), wired through Terraform
(timescaledb_tag, default "latest-pg18") so it can actually be pinned to an
exact minor without hand-editing the host - required before any host of the
stack could replicate with another (a floating tag risks mismatched
extension minors, which blocks a physical replica and risks compressed-
chunk corruption on restore).
Caddy active health-gate: both the Terraform-rendered Caddyfile and the live
deploy/Caddyfile now health-check the app on the same cheap /healthz route
its own Docker HEALTHCHECK uses (now /healthz instead of the SSR homepage,
so it's cheap enough for a tight interval and works identically for a
worker replica, which serves no public traffic at all) - Caddy won't
forward into a container that's still booting or unhealthy.
Verified live: built and booted the real image via docker compose - both
containers report healthy via the new /healthz-based HEALTHCHECK, and GET /
still renders the full SSR homepage unchanged. Both Caddyfiles validated
with the real caddy binary. docker-stack.yml validated with docker compose
config (required-var guards fire with clear messages; secrets correctly
mount at /run/secrets/<name>, matching the entrypoint shim's mapping).
docker-compose.yml validated with and without TIMESCALEDB_TAG set, alongside
the existing openmeteo overlay. terraform validate + fmt clean.
2026-07-21 00:39:48 +00:00
|
|
|
}
|
2026-07-15 19:58:32 +00:00
|
|
|
|
2026-07-24 04:37:41 +00:00
|
|
|
# Access-log hygiene: the default JSON encoder serializes full request headers,
|
|
|
|
|
# the TLS block, and response headers on every line (measured ~1,133B/line) --
|
|
|
|
|
# strip those with the `filter` format encoder. Also strip the query string from
|
|
|
|
|
# the logged URI: Caddy's default logger records request.uri *including* the
|
|
|
|
|
# query string, so every `?q=<search text>` a visitor typed sat in Loki next to
|
|
|
|
|
# their client IP for the full 30-day retention -- a real privacy leak, not just
|
|
|
|
|
# noise. Bot/crawler skipping stays out of here: `log_skip` needs Caddy >= 2.7
|
|
|
|
|
# and an upgrade is out of scope, so that's handled downstream in Alloy's
|
|
|
|
|
# loki.process "caddy" stage instead (see observability/alloy/config.alloy).
|
2026-07-15 19:58:32 +00:00
|
|
|
log {
|
2026-07-24 04:37:41 +00:00
|
|
|
output file /var/log/caddy/thermograph.log {
|
|
|
|
|
roll_size 20MiB
|
|
|
|
|
roll_keep 5
|
|
|
|
|
}
|
|
|
|
|
format filter {
|
|
|
|
|
wrap json
|
|
|
|
|
fields {
|
|
|
|
|
request>headers delete
|
|
|
|
|
request>tls delete
|
|
|
|
|
resp_headers delete
|
|
|
|
|
request>uri regexp \?.* ""
|
|
|
|
|
}
|
|
|
|
|
}
|
2026-07-15 19:58:32 +00:00
|
|
|
}
|
|
|
|
|
}
|
2026-07-11 00:29:47 +00:00
|
|
|
|
2026-07-11 03:01:16 +00:00
|
|
|
emigriffith.dev {
|
2026-07-11 00:29:47 +00:00
|
|
|
encode zstd gzip
|
|
|
|
|
|
2026-07-15 19:58:32 +00:00
|
|
|
# Thermograph moved to its own domain. Send the old sub-path there with a
|
|
|
|
|
# permanent redirect, stripping the /thermograph prefix so deep links map
|
|
|
|
|
# straight across (…/thermograph/calendar -> thermograph.org/calendar). The
|
|
|
|
|
# bare /thermograph (no trailing slash) goes to the new root.
|
|
|
|
|
handle_path /thermograph/* {
|
|
|
|
|
redir https://thermograph.org{uri} permanent
|
|
|
|
|
}
|
|
|
|
|
handle /thermograph {
|
|
|
|
|
redir https://thermograph.org/ permanent
|
2026-07-11 03:01:16 +00:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
# Portfolio at the root. Point `root` at the built static site (for the Astro
|
|
|
|
|
# portfolio that's its `dist/` output). file_server serves index.html for
|
|
|
|
|
# directories and returns a real 404 for missing paths.
|
|
|
|
|
handle {
|
|
|
|
|
root * /var/www/emigriffith
|
|
|
|
|
file_server
|
|
|
|
|
}
|
2026-07-11 00:29:47 +00:00
|
|
|
|
|
|
|
|
log {
|
2026-07-11 03:01:16 +00:00
|
|
|
output file /var/log/caddy/emigriffith.log
|
2026-07-11 00:29:47 +00:00
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
2026-07-11 03:40:08 +00:00
|
|
|
# Old bookmarks to the raw IP (the pre-domain URL) would otherwise get bounced to
|
2026-07-15 19:58:32 +00:00
|
|
|
# HTTPS-on-the-IP, which has no cert and fails. Redirect them to the portfolio domain.
|
2026-07-11 03:40:08 +00:00
|
|
|
http://75.119.132.91 {
|
|
|
|
|
redir https://emigriffith.dev{uri} permanent
|
|
|
|
|
}
|
|
|
|
|
|
2026-07-15 19:58:32 +00:00
|
|
|
# Optional: redirect www -> apex for either domain. Add the www CNAME/A record
|
|
|
|
|
# first, then uncomment the matching block.
|
2026-07-11 03:01:16 +00:00
|
|
|
# www.emigriffith.dev {
|
|
|
|
|
# redir https://emigriffith.dev{uri} permanent
|
2026-07-11 00:29:47 +00:00
|
|
|
# }
|
2026-07-15 19:58:32 +00:00
|
|
|
# www.thermograph.org {
|
|
|
|
|
# redir https://thermograph.org{uri} permanent
|
|
|
|
|
# }
|