Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
name: Ops cron (backup + IndexNow)
|
|
|
|
|
|
|
|
|
|
# Scheduled operational jobs that don't belong in the app's own worker-tier
|
|
|
|
|
# scheduler (notifications/scheduler.py, Track A chunk 5) because they're
|
|
|
|
|
# infra/ops concerns rather than app-domain background work -- see the job
|
|
|
|
|
# classification table in
|
2026-07-21 21:29:14 +00:00
|
|
|
# thermograph-docs/architecture/repo-topology-and-infrastructure.md §7.
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
#
|
2026-07-26 06:56:38 +00:00
|
|
|
# The jobs SSH into a host and run inside the already-running stack (docker
|
|
|
|
|
# exec), the same way deploy.sh runs its own post-deploy IndexNow ping -- no new
|
ci: add an always-on Actions runner on vps2
The estate had exactly one registered runner, on the desktop. It went offline
at 2026-07-31 16:31Z; for the next 21 hours no PR could satisfy a required
check, no deploy could run, and the 03:00Z ops-cron -- the only backup for both
application databases and for Forgejo -- did not fire. Forgejo queued that
scheduled run rather than dropping it, so it completed on reconnect and nothing
was lost. A longer outage would have meant real gaps.
Three files claimed an "always-on Swarm-hosted runner" existed and that the
estate therefore no longer depended on the desktop. It did not exist: an early
revision of docker-stack.yml ran one as a Docker-in-Docker sidecar and it was
removed. That claim is why a single point of failure sat unnoticed. Corrected
in docker-stack.yml, forgejo/README.md and register-lan-runner.sh.
The new runner is a plain restart:always container, not a Swarm service: a
Swarm-scheduled runner cannot redeploy the Swarm that schedules it, so CI would
be gone exactly when the cluster is what is broken. It runs from
/opt/forgejo-runner rather than in place, because the checkout is reset on
every prod deploy and one `git clean -fdx` there would destroy the
registration.
vps2 runs prod, so the socket mount is bounded rather than assumed benign:
capacity 1, --cpus=2/--memory=4g on job containers, valid_volumes empty so no
job can bind-mount /etc/thermograph.env, and no thermograph network joined.
This defends against accident, not against a hostile workflow author -- stated
plainly in the compose header rather than implied.
The desktop runner stays registered as extra capacity. Nothing may assume it
is up.
2026-08-01 14:54:28 +00:00
|
|
|
# network exposure, no separate dependency install. Runs on the `docker` label,
|
|
|
|
|
# served by the vps2 runner (infra/deploy/forgejo/runner-vps2/) and, when it is
|
|
|
|
|
# up, the desktop.
|
|
|
|
|
#
|
|
|
|
|
# THIS WORKFLOW IS THE ESTATE'S ONLY BACKUP, so its dependency on a runner is
|
|
|
|
|
# load-bearing. On 2026-07-31 the desktop was the only registered runner and
|
|
|
|
|
# went offline for 21 hours; this schedule did not fire at 03:00Z. Forgejo held
|
|
|
|
|
# the run in `waiting` and it completed on reconnect, so nothing was lost that
|
|
|
|
|
# time. The vps2 runner exists so the next outage is not a coin flip.
|
2026-07-22 23:26:35 +00:00
|
|
|
#
|
2026-07-26 06:56:38 +00:00
|
|
|
# WHICH BOX EACH JOB TALKS TO, and why the secret names say so:
|
|
|
|
|
#
|
|
|
|
|
# backup, indexnow -> VPS2_SSH_* (vps2 = 169.58.46.181: prod AND beta, the
|
|
|
|
|
# shared TimescaleDB, Postfix, the backups)
|
|
|
|
|
# forgejo-backup -> VPS1_SSH_* (vps1 = 75.119.132.91: Forgejo, Grafana,
|
|
|
|
|
# Loki, and the dev environment)
|
|
|
|
|
#
|
|
|
|
|
# The secrets are keyed by HOST rather than by environment because an
|
|
|
|
|
# environment no longer implies a machine: vps2 runs two of them. The previous
|
|
|
|
|
# names encoded the opposite assumption and had already caused one incident --
|
|
|
|
|
# an early revision used SSH_* here, which meant "beta", so the job labelled
|
|
|
|
|
# "prod backup" was silently dumping beta while prod had no backup at all. The
|
|
|
|
|
# same conflation is why Forgejo's backup used to be described as running "on
|
|
|
|
|
# beta": beta and Forgejo happened to share a box, so one secret served both
|
|
|
|
|
# meanings. It does not any more.
|
|
|
|
|
#
|
|
|
|
|
# BOTH APPLICATION DATABASES ARE BACKED UP. prod and beta are separate databases
|
|
|
|
|
# (`thermograph` and `thermograph_beta`) on one shared instance. Dumping only
|
|
|
|
|
# the prod database would recreate the original failure in a new shape: a job
|
|
|
|
|
# that looks like "the backup" while one environment's data is silently
|
|
|
|
|
# uncovered. The loop below dumps each, to its own off-box prefix.
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
|
2026-07-23 05:11:33 +00:00
|
|
|
# Monorepo port: the compose file now lives under infra/ of the host's
|
|
|
|
|
# /opt/thermograph monorepo checkout, so every docker-compose invocation cd's
|
|
|
|
|
# into /opt/thermograph/infra (compose also pins `name: thermograph` so the
|
|
|
|
|
# project name no longer depends on the directory). Everything else unchanged.
|
|
|
|
|
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
on:
|
|
|
|
|
schedule:
|
|
|
|
|
# 03:00 UTC daily -- a low-traffic window for both jobs.
|
|
|
|
|
- cron: '0 3 * * *'
|
|
|
|
|
workflow_dispatch: {}
|
|
|
|
|
|
|
|
|
|
jobs:
|
|
|
|
|
backup:
|
|
|
|
|
name: pg_dump backup
|
Reconcile Forgejo workflows with the real infrastructure now in place (#238)
The Swarm/Forgejo standup (a parallel infra track, see INFRA.md) landed real
.forgejo/workflows/{build,pr-build,deploy,deploy-dev}.yml before this PR
merged, making ci.yml a redundant duplicate of their build.yml (same build
gate, same job). Drop it.
Fix the remaining two files to match the real, already-deployed conventions
those files established rather than the guesses this PR shipped with:
runs-on: [self-hosted, thermograph] -> docker (the actual Docker-in-Docker
Swarm-hosted runner label), and appleboy/ssh-action referenced by full GitHub
URL (confirmed not mirrored on this Forgejo instance's default action
registry, per deploy.yml's own header comment) rather than the short form.
actions/checkout needed no change - already confirmed to resolve unchanged
from Forgejo's default mirror.
build-push.yml (image registry push) and ops-cron.yml (backup + IndexNow)
stay: neither duplicates anything in the real mirror, which faithfully
replicates the OLD git-checkout-and-build-in-place deploy model rather than
the registry-based one, and has no scheduled ops jobs at all.
2026-07-21 01:08:09 +00:00
|
|
|
runs-on: docker
|
2026-07-21 20:09:32 +00:00
|
|
|
# Guards against the schedule and a manual workflow_dispatch landing
|
|
|
|
|
# close together (or two manual triggers): queue rather than cancel, so
|
|
|
|
|
# a workflow_dispatch never aborts a pg_dump mid-write and leaves a
|
|
|
|
|
# truncated .dump file -- same cancel-in-progress:false rationale as
|
|
|
|
|
# deploy.yml/deploy-dev.yml's own SSH-script jobs.
|
|
|
|
|
concurrency:
|
|
|
|
|
group: ops-backup
|
|
|
|
|
cancel-in-progress: false
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
steps:
|
2026-07-26 06:56:38 +00:00
|
|
|
- name: Dump both application databases over SSH
|
Reconcile Forgejo workflows with the real infrastructure now in place (#238)
The Swarm/Forgejo standup (a parallel infra track, see INFRA.md) landed real
.forgejo/workflows/{build,pr-build,deploy,deploy-dev}.yml before this PR
merged, making ci.yml a redundant duplicate of their build.yml (same build
gate, same job). Drop it.
Fix the remaining two files to match the real, already-deployed conventions
those files established rather than the guesses this PR shipped with:
runs-on: [self-hosted, thermograph] -> docker (the actual Docker-in-Docker
Swarm-hosted runner label), and appleboy/ssh-action referenced by full GitHub
URL (confirmed not mirrored on this Forgejo instance's default action
registry, per deploy.yml's own header comment) rather than the short form.
actions/checkout needed no change - already confirmed to resolve unchanged
from Forgejo's default mirror.
build-push.yml (image registry push) and ops-cron.yml (backup + IndexNow)
stay: neither duplicates anything in the real mirror, which faithfully
replicates the OLD git-checkout-and-build-in-place deploy model rather than
the registry-based one, and has no scheduled ops jobs at all.
2026-07-21 01:08:09 +00:00
|
|
|
uses: https://github.com/appleboy/ssh-action@v1.2.0
|
2026-07-23 20:42:09 +00:00
|
|
|
# S3 creds for the encrypted off-box copy to Contabo Object Storage, passed
|
2026-07-26 06:56:38 +00:00
|
|
|
# into the remote script via `envs:` (a host env file isn't a reliable
|
|
|
|
|
# source across boxes -- on vps1 it isn't readable by the deploy user --
|
|
|
|
|
# so the CI secret is the uniform home, like the VPS*_SSH_* pairs).
|
2026-07-23 20:42:09 +00:00
|
|
|
env:
|
|
|
|
|
S3_ENDPOINT: ${{ secrets.S3_ENDPOINT }}
|
|
|
|
|
S3_BUCKET: ${{ secrets.S3_BUCKET }}
|
|
|
|
|
S3_ACCESS_KEY: ${{ secrets.S3_ACCESS_KEY }}
|
|
|
|
|
S3_SECRET_KEY: ${{ secrets.S3_SECRET_KEY }}
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
with:
|
2026-07-26 06:56:38 +00:00
|
|
|
host: ${{ secrets.VPS2_SSH_HOST }}
|
|
|
|
|
username: ${{ secrets.VPS2_SSH_USER }}
|
|
|
|
|
key: ${{ secrets.VPS2_SSH_KEY }}
|
|
|
|
|
port: ${{ secrets.VPS2_SSH_PORT }}
|
2026-07-23 20:42:09 +00:00
|
|
|
envs: S3_ENDPOINT,S3_BUCKET,S3_ACCESS_KEY,S3_SECRET_KEY
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
script: |
|
|
|
|
|
set -euo pipefail
|
2026-07-23 05:11:33 +00:00
|
|
|
cd /opt/thermograph/infra
|
2026-07-22 23:26:35 +00:00
|
|
|
# Source the env so `docker compose` can interpolate POSTGRES_PASSWORD;
|
|
|
|
|
# without it compose fails to parse docker-compose.yml, the redirect
|
|
|
|
|
# still creates the target, and the job leaves a 0-byte .dump and exits
|
|
|
|
|
# non-zero (exactly how this backup was silently failing). Mirrors the
|
|
|
|
|
# IndexNow job below, which already sources it.
|
|
|
|
|
set -a; . /etc/thermograph.env 2>/dev/null || true; set +a
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
backup_dir="$HOME/thermograph-backups"
|
|
|
|
|
mkdir -p "$backup_dir"
|
|
|
|
|
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
|
2026-07-26 06:56:38 +00:00
|
|
|
# The db may run under plain compose OR as a Swarm stack task;
|
|
|
|
|
# resolve the container either way so the backup survives a
|
|
|
|
|
# deploy-mode switch. One instance now serves both environments.
|
2026-07-23 04:13:12 +00:00
|
|
|
dbc=$(docker ps -q --filter "label=com.docker.swarm.service.name=thermograph_db" | head -1)
|
2026-07-23 05:11:33 +00:00
|
|
|
[ -z "$dbc" ] && dbc=$(cd /opt/thermograph/infra && docker compose ps -q db 2>/dev/null | head -1)
|
2026-07-23 04:13:12 +00:00
|
|
|
[ -n "$dbc" ] || { echo "!! no db container found (compose or stack)"; exit 1; }
|
2026-07-26 06:56:38 +00:00
|
|
|
# env:database pairs. Both are dumped every night: they are separate
|
|
|
|
|
# databases on one instance, and backing up only one would leave the
|
|
|
|
|
# other silently uncovered.
|
|
|
|
|
for pair in prod:thermograph beta:thermograph_beta; do
|
|
|
|
|
envname="${pair%%:*}"; dbname="${pair##*:}"
|
|
|
|
|
# A missing database is a HARD failure, never a quiet skip -- a
|
|
|
|
|
# backup job that shrugs off an absent database is exactly how an
|
|
|
|
|
# environment ends up with no backups and nobody noticing.
|
|
|
|
|
if ! docker exec "$dbc" psql -U thermograph -d postgres -tAc \
|
|
|
|
|
"select 1 from pg_database where datname='$dbname'" | grep -q 1; then
|
|
|
|
|
echo "!! database '$dbname' ($envname) does not exist on this instance."
|
|
|
|
|
if [ "$envname" = prod ]; then
|
|
|
|
|
echo "!! That is prod's own database on prod's own instance — this is not a"
|
|
|
|
|
echo "!! provisioning gap, something is badly wrong. Do not 'fix' it by creating"
|
|
|
|
|
echo "!! an empty database; find out where the real one went."
|
|
|
|
|
else
|
|
|
|
|
echo "!! provision it first: sudo bash /opt/thermograph-beta/infra/deploy/db/provision-env-db.sh $envname"
|
|
|
|
|
fi
|
|
|
|
|
exit 1
|
|
|
|
|
fi
|
|
|
|
|
out="$backup_dir/$dbname-$stamp.dump"
|
|
|
|
|
# Write to a .partial and rename on success so a mid-dump failure can
|
|
|
|
|
# never leave a truncated file that looks like a good backup.
|
|
|
|
|
docker exec "$dbc" pg_dump -U thermograph -d "$dbname" \
|
|
|
|
|
--format=custom > "$out.partial"
|
|
|
|
|
mv "$out.partial" "$out"
|
|
|
|
|
echo "wrote $out ($(du -h "$out" | cut -f1))"
|
|
|
|
|
dumps="${dumps:-} $envname:$out"
|
|
|
|
|
done
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
# The dumps are the disaster-recovery copy, not a versioned
|
|
|
|
|
# archive -- keep the last 14 days and let the rest age out.
|
2026-07-26 06:56:38 +00:00
|
|
|
find "$backup_dir" -name '*.dump' -mtime +14 -delete
|
|
|
|
|
find "$backup_dir" -name '*.dump.partial' -mtime +1 -delete
|
2026-07-23 20:42:09 +00:00
|
|
|
# --- off-box copy to Contabo Object Storage (age-encrypted) ---
|
|
|
|
|
# Streams the just-written dump through age (to the vault's age
|
|
|
|
|
# recipient -- the same key each host renders secrets with, so
|
|
|
|
|
# /etc/thermograph/age.key decrypts it for restore) straight to S3. No
|
|
|
|
|
# local intermediate, no plaintext leaves the box. Skips cleanly if the
|
|
|
|
|
# S3 secrets aren't set. Keeps 30 days of encrypted dumps off-box.
|
|
|
|
|
if [ -n "${S3_ACCESS_KEY:-}" ]; then
|
|
|
|
|
command -v rclone >/dev/null 2>&1 || sudo apt-get install -y -qq rclone
|
|
|
|
|
command -v age >/dev/null 2>&1 || sudo apt-get install -y -qq age
|
|
|
|
|
export RCLONE_CONFIG_ARCHIVE_TYPE=s3 RCLONE_CONFIG_ARCHIVE_PROVIDER=Other \
|
|
|
|
|
RCLONE_CONFIG_ARCHIVE_ENDPOINT="$S3_ENDPOINT" \
|
|
|
|
|
RCLONE_CONFIG_ARCHIVE_ACCESS_KEY_ID="$S3_ACCESS_KEY" \
|
|
|
|
|
RCLONE_CONFIG_ARCHIVE_SECRET_ACCESS_KEY="$S3_SECRET_KEY" \
|
|
|
|
|
RCLONE_CONFIG_ARCHIVE_REGION=default RCLONE_CONFIG_ARCHIVE_FORCE_PATH_STYLE=true
|
|
|
|
|
recip=age1xx4dzs0dxlwvkv9sjuqzsphl7lfrxannkfken374yu2qvvcte9sqzktqt2
|
2026-07-26 06:56:38 +00:00
|
|
|
# One prefix per environment, so a restore never has to guess which
|
|
|
|
|
# database a dump came from: backups/db/prod/ and backups/db/beta/.
|
|
|
|
|
for entry in $dumps; do
|
|
|
|
|
envname="${entry%%:*}"; f="${entry##*:}"
|
|
|
|
|
s3base="$S3_BUCKET/backups/db/$envname"
|
|
|
|
|
age -r "$recip" < "$f" | rclone rcat "archive:$s3base/$(basename "$f").age"
|
|
|
|
|
echo "off-box: uploaded $(basename "$f").age to $s3base"
|
|
|
|
|
rclone delete --min-age 30d "archive:$s3base/" 2>/dev/null || true
|
|
|
|
|
done
|
2026-07-23 20:42:09 +00:00
|
|
|
else
|
|
|
|
|
echo "!! S3_* secrets unset -- skipped off-box push (add them as repo secrets)"
|
|
|
|
|
fi
|
|
|
|
|
|
|
|
|
|
forgejo-backup:
|
2026-07-26 06:56:38 +00:00
|
|
|
name: Forgejo backup (vps1) -> S3
|
2026-07-23 20:42:09 +00:00
|
|
|
runs-on: docker
|
2026-07-26 06:56:38 +00:00
|
|
|
# Forgejo (the git host + all CI history) lives on vps1 and had NO backup at
|
2026-07-23 20:42:09 +00:00
|
|
|
# all. Dumps its Postgres db (consistent) + tars its data volume (repos, LFS,
|
|
|
|
|
# config, avatars), age-encrypts both, pushes off-box, keeps 30 days. Runs as
|
2026-07-26 06:56:38 +00:00
|
|
|
# vps1's SSH user (VPS1_SSH_*), which is in the docker group -- so docker
|
|
|
|
|
# exec/run need no host sudo, and the data tar runs inside a throwaway
|
|
|
|
|
# container (the volume dir is root-owned on the host). rclone + age must be
|
|
|
|
|
# present on vps1.
|
|
|
|
|
#
|
|
|
|
|
# This job follows FORGEJO, not beta. It used to use the same secret as the
|
|
|
|
|
# beta deploy purely because Forgejo and beta shared a box; beta has since
|
|
|
|
|
# moved to vps2 and Forgejo has not moved at all.
|
2026-07-23 20:42:09 +00:00
|
|
|
concurrency:
|
|
|
|
|
group: ops-forgejo-backup
|
|
|
|
|
cancel-in-progress: false
|
|
|
|
|
steps:
|
|
|
|
|
- name: Back up Forgejo (db + data volume) over SSH
|
|
|
|
|
uses: https://github.com/appleboy/ssh-action@v1.2.0
|
|
|
|
|
env:
|
|
|
|
|
S3_ENDPOINT: ${{ secrets.S3_ENDPOINT }}
|
|
|
|
|
S3_BUCKET: ${{ secrets.S3_BUCKET }}
|
|
|
|
|
S3_ACCESS_KEY: ${{ secrets.S3_ACCESS_KEY }}
|
|
|
|
|
S3_SECRET_KEY: ${{ secrets.S3_SECRET_KEY }}
|
|
|
|
|
with:
|
2026-07-26 06:56:38 +00:00
|
|
|
host: ${{ secrets.VPS1_SSH_HOST }}
|
|
|
|
|
username: ${{ secrets.VPS1_SSH_USER }}
|
|
|
|
|
key: ${{ secrets.VPS1_SSH_KEY }}
|
|
|
|
|
port: ${{ secrets.VPS1_SSH_PORT }}
|
2026-07-23 20:42:09 +00:00
|
|
|
envs: S3_ENDPOINT,S3_BUCKET,S3_ACCESS_KEY,S3_SECRET_KEY
|
|
|
|
|
script: |
|
|
|
|
|
set -euo pipefail
|
|
|
|
|
[ -n "${S3_ACCESS_KEY:-}" ] || { echo "!! S3_* secrets unset"; exit 1; }
|
2026-07-26 06:56:38 +00:00
|
|
|
command -v rclone >/dev/null 2>&1 || { echo "!! rclone missing on vps1"; exit 1; }
|
|
|
|
|
command -v age >/dev/null 2>&1 || { echo "!! age missing on vps1"; exit 1; }
|
2026-07-23 20:42:09 +00:00
|
|
|
export RCLONE_CONFIG_ARCHIVE_TYPE=s3 RCLONE_CONFIG_ARCHIVE_PROVIDER=Other \
|
|
|
|
|
RCLONE_CONFIG_ARCHIVE_ENDPOINT="$S3_ENDPOINT" \
|
|
|
|
|
RCLONE_CONFIG_ARCHIVE_ACCESS_KEY_ID="$S3_ACCESS_KEY" \
|
|
|
|
|
RCLONE_CONFIG_ARCHIVE_SECRET_ACCESS_KEY="$S3_SECRET_KEY" \
|
|
|
|
|
RCLONE_CONFIG_ARCHIVE_REGION=default RCLONE_CONFIG_ARCHIVE_FORCE_PATH_STYLE=true
|
|
|
|
|
recip=age1xx4dzs0dxlwvkv9sjuqzsphl7lfrxannkfken374yu2qvvcte9sqzktqt2
|
|
|
|
|
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
|
|
|
|
|
base="$S3_BUCKET/backups/forgejo"
|
|
|
|
|
fdb=$(docker ps -qf name=forgejo_db | head -1)
|
|
|
|
|
[ -n "$fdb" ] || { echo "!! no forgejo_db container"; exit 1; }
|
|
|
|
|
# 1) Forgejo Postgres (consistent dump); creds read inside the container
|
|
|
|
|
docker exec "$fdb" sh -c 'PGPASSWORD=$POSTGRES_PASSWORD pg_dump -U $POSTGRES_USER -d $POSTGRES_DB --format=custom' \
|
|
|
|
|
| age -r "$recip" | rclone rcat "archive:$base/db/forgejo-$stamp.dump.age"
|
|
|
|
|
# 2) Forgejo data volume via a throwaway container (no host sudo needed)
|
|
|
|
|
docker run --rm -v forgejo_forgejo_data:/data:ro busybox tar -C /data -cf - . \
|
|
|
|
|
| age -r "$recip" | rclone rcat "archive:$base/data/forgejo-data-$stamp.tar.age"
|
|
|
|
|
echo "forgejo backup uploaded ($stamp)"
|
|
|
|
|
rclone delete --min-age 30d "archive:$base/db/" 2>/dev/null || true
|
|
|
|
|
rclone delete --min-age 30d "archive:$base/data/" 2>/dev/null || true
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
|
openbao: make the parity gate actually work, and run it nightly
Three fixes, all on the path between "the migration looks ready" and "the
migration is ready".
verify-parity.sh never sourced env-topology.sh, so TG_BAO_APPROLE was unset and
render-secrets-openbao.sh fell through to the bare default -- prod's credential
-- for every environment. On vps2 that made `--env beta` authenticate as
tg-prod and take a 403 from tg-host-prod.hcl on thermograph/data/env/beta.
`--all` failed the same way, taking dev with it. deploy.sh and
deploy-stack.sh always sourced it; only the verifier did not, which is the
worst place for the omission: the tool whose job is to notice divergence was
itself diverging from the path it verifies. It now calls thermograph_topology
per environment and takes TG_SKIP_COMMON from there rather than re-deriving it.
infra-sync.yml renders all three env files on every push touching infra/**, and
none of its three jobs sourced env-topology.sh either. TG_SECRETS_BACKEND was
therefore unset and render-secrets.sh:44 defaulted to sops. Flipping the
selector would have changed deploy.sh and deploy-stack.sh but not this
workflow, which would have kept re-rendering from SOPS -- silent while parity
holds, and a hard failure the moment the SOPS files are retired. The one-line
cutover the README describes was never sufficient on its own.
ops-cron.yml gains a secrets-parity job: prod and beta from vps2, dev from
vps1, nightly, failing loudly on a mismatch. README.md called this the gate for
a cutover; nothing implemented it. A gate that exists only in prose gets
satisfied by assertion. It grants CI no vault access -- the host renders, CI
only asks it to.
First measured run, immediately after the verifier fix:
dev 12 keys PASS, byte-identical
beta 24 keys FAIL, 3 values differ
prod 32 keys FAIL, 3 values differ
The same three keys on both: THERMOGRAPH_S3_SECRET_KEY,
THERMOGRAPH_LAKE_S3_SECRET_KEY, THERMOGRAPH_VAPID_CONTACT. All three live in
common.yaml, seeded 2026-07-31 02:30Z -- before the Contabo rotation landed.
dev passes because dev never layers common. One `seed-from-sops.sh --env
common` fixes beta and prod together, and the 7-day clock starts after that.
2026-08-01 15:36:35 +00:00
|
|
|
secrets-parity:
|
|
|
|
|
name: OpenBao/SOPS parity gate
|
|
|
|
|
runs-on: docker
|
|
|
|
|
# The gate infra/openbao/README.md specifies for a cutover -- "7 consecutive
|
|
|
|
|
# green parity runs" -- and which, until now, nothing implemented. A gate
|
|
|
|
|
# that exists only in prose gets satisfied by assertion.
|
|
|
|
|
#
|
|
|
|
|
# verify-parity.sh renders each environment through BOTH backends and diffs
|
|
|
|
|
# the KEY=value sets after a last-wins collapse. It prints key NAMES and
|
|
|
|
|
# counts only, never values, so these logs are safe to read and to paste.
|
|
|
|
|
#
|
|
|
|
|
# A mismatch FAILS the job, deliberately and loudly. The failure modes on
|
|
|
|
|
# the other side of a bad cutover are all silent: a missing
|
|
|
|
|
# THERMOGRAPH_AUTH_SECRET makes every worker mint its own, so logins start
|
|
|
|
|
# working intermittently; a half-rendered VAPID pair makes push.py generate a
|
|
|
|
|
# new keypair and delete every subscription row as "gone". Red here is the
|
|
|
|
|
# only warning that arrives before those.
|
|
|
|
|
#
|
|
|
|
|
# Runs under sudo because the age key is 0400 root and the AppRole files are
|
|
|
|
|
# 0400 agent/root -- the SSH user cannot read them directly. This grants CI
|
|
|
|
|
# no vault access of its own: the HOST renders, CI only asks it to.
|
|
|
|
|
#
|
|
|
|
|
# The run history in Forgejo IS the record of consecutive green days. Do not
|
|
|
|
|
# count a day this job was skipped or the runner was down as green.
|
|
|
|
|
concurrency:
|
|
|
|
|
group: ops-secrets-parity
|
|
|
|
|
cancel-in-progress: false
|
|
|
|
|
steps:
|
|
|
|
|
- name: prod + beta (vps2)
|
|
|
|
|
uses: https://github.com/appleboy/ssh-action@v1.2.0
|
|
|
|
|
with:
|
|
|
|
|
host: ${{ secrets.VPS2_SSH_HOST }}
|
|
|
|
|
username: ${{ secrets.VPS2_SSH_USER }}
|
|
|
|
|
key: ${{ secrets.VPS2_SSH_KEY }}
|
|
|
|
|
port: ${{ secrets.VPS2_SSH_PORT }}
|
|
|
|
|
script: |
|
|
|
|
|
set -euo pipefail
|
|
|
|
|
# Each environment is checked from its OWN checkout. The trees are
|
|
|
|
|
# identical, but running beta's from beta's directory also proves
|
|
|
|
|
# that checkout is current rather than assuming it.
|
|
|
|
|
sudo /opt/thermograph/infra/openbao/verify-parity.sh --env prod
|
|
|
|
|
sudo /opt/thermograph-beta/infra/openbao/verify-parity.sh --env beta
|
|
|
|
|
|
|
|
|
|
- name: dev (vps1)
|
|
|
|
|
uses: https://github.com/appleboy/ssh-action@v1.2.0
|
|
|
|
|
with:
|
|
|
|
|
host: ${{ secrets.VPS1_SSH_HOST }}
|
|
|
|
|
username: ${{ secrets.VPS1_SSH_USER }}
|
|
|
|
|
key: ${{ secrets.VPS1_SSH_KEY }}
|
|
|
|
|
port: ${{ secrets.VPS1_SSH_PORT }}
|
|
|
|
|
script: |
|
|
|
|
|
set -euo pipefail
|
|
|
|
|
# dev renders dev.yaml ALONE -- verify-parity.sh takes that from
|
|
|
|
|
# env-topology.sh's TG_SKIP_COMMON rather than re-deriving it, so the
|
|
|
|
|
# comparison matches what deploy-dev.sh actually does. vps1 must
|
|
|
|
|
# never see common.yaml: it runs Forgejo, its CI, and dev's
|
|
|
|
|
# unreviewed branch.
|
|
|
|
|
sudo /opt/thermograph-dev/infra/openbao/verify-parity.sh --env dev
|
|
|
|
|
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
indexnow:
|
|
|
|
|
name: IndexNow ping
|
Reconcile Forgejo workflows with the real infrastructure now in place (#238)
The Swarm/Forgejo standup (a parallel infra track, see INFRA.md) landed real
.forgejo/workflows/{build,pr-build,deploy,deploy-dev}.yml before this PR
merged, making ci.yml a redundant duplicate of their build.yml (same build
gate, same job). Drop it.
Fix the remaining two files to match the real, already-deployed conventions
those files established rather than the guesses this PR shipped with:
runs-on: [self-hosted, thermograph] -> docker (the actual Docker-in-Docker
Swarm-hosted runner label), and appleboy/ssh-action referenced by full GitHub
URL (confirmed not mirrored on this Forgejo instance's default action
registry, per deploy.yml's own header comment) rather than the short form.
actions/checkout needed no change - already confirmed to resolve unchanged
from Forgejo's default mirror.
build-push.yml (image registry push) and ops-cron.yml (backup + IndexNow)
stay: neither duplicates anything in the real mirror, which faithfully
replicates the OLD git-checkout-and-build-in-place deploy model rather than
the registry-based one, and has no scheduled ops jobs at all.
2026-07-21 01:08:09 +00:00
|
|
|
runs-on: docker
|
2026-07-21 20:09:32 +00:00
|
|
|
# Same overlap guard as the backup job above (schedule vs. manual
|
|
|
|
|
# dispatch); --if-changed already makes a second ping a cheap no-op, but
|
|
|
|
|
# queueing avoids two SSH sessions racing on the host regardless.
|
|
|
|
|
concurrency:
|
|
|
|
|
group: ops-indexnow
|
|
|
|
|
cancel-in-progress: false
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
steps:
|
|
|
|
|
- name: Ping IndexNow if the URL set changed
|
Reconcile Forgejo workflows with the real infrastructure now in place (#238)
The Swarm/Forgejo standup (a parallel infra track, see INFRA.md) landed real
.forgejo/workflows/{build,pr-build,deploy,deploy-dev}.yml before this PR
merged, making ci.yml a redundant duplicate of their build.yml (same build
gate, same job). Drop it.
Fix the remaining two files to match the real, already-deployed conventions
those files established rather than the guesses this PR shipped with:
runs-on: [self-hosted, thermograph] -> docker (the actual Docker-in-Docker
Swarm-hosted runner label), and appleboy/ssh-action referenced by full GitHub
URL (confirmed not mirrored on this Forgejo instance's default action
registry, per deploy.yml's own header comment) rather than the short form.
actions/checkout needed no change - already confirmed to resolve unchanged
from Forgejo's default mirror.
build-push.yml (image registry push) and ops-cron.yml (backup + IndexNow)
stay: neither duplicates anything in the real mirror, which faithfully
replicates the OLD git-checkout-and-build-in-place deploy model rather than
the registry-based one, and has no scheduled ops jobs at all.
2026-07-21 01:08:09 +00:00
|
|
|
uses: https://github.com/appleboy/ssh-action@v1.2.0
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
with:
|
2026-07-26 06:56:38 +00:00
|
|
|
host: ${{ secrets.VPS2_SSH_HOST }}
|
|
|
|
|
username: ${{ secrets.VPS2_SSH_USER }}
|
|
|
|
|
key: ${{ secrets.VPS2_SSH_KEY }}
|
|
|
|
|
port: ${{ secrets.VPS2_SSH_PORT }}
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
script: |
|
|
|
|
|
set -euo pipefail
|
2026-07-23 05:11:33 +00:00
|
|
|
cd /opt/thermograph/infra
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
set -a; . /etc/thermograph.env 2>/dev/null || true; set +a
|
2026-07-23 04:13:12 +00:00
|
|
|
bec=$(docker ps -q --filter "label=com.docker.swarm.service.name=thermograph_web" | head -1)
|
2026-07-23 05:11:33 +00:00
|
|
|
[ -z "$bec" ] && bec=$(cd /opt/thermograph/infra && docker compose ps -q backend 2>/dev/null | head -1)
|
2026-07-23 04:13:12 +00:00
|
|
|
[ -n "$bec" ] || { echo "!! no backend/web container found"; exit 1; }
|
|
|
|
|
docker exec "$bec" python indexnow.py --if-changed \
|
Add Forgejo Actions workflows: CI, image build+push, ops cron (#236)
Three workflows for the self-hosted Forgejo instance (Track A chunk 7 of the
infra-design implementation handoff), mirroring/extending the existing
.github/workflows without touching them - GitHub stays the primary repo and
its own build.yml/ci-cd.yml keep gating auto-merge.
ci.yml mirrors build.yml's build gate (deps, backend tests, frontend JS
syntax check, boot + page/API health check) on a Forgejo runner.
build-push.yml builds the app image and pushes it to Forgejo's built-in
registry, tagged by git SHA (every push) and semver (version tags) - the
registry half of the hop-1 cutover's build-once/deploy-everywhere model.
Runs alongside the existing deploy.yml/deploy.sh (git-checkout-and-build-in-
place) without touching it, per the cutover runbook's Stage C.
ops-cron.yml runs a daily pg_dump backup and an IndexNow --if-changed ping,
both via SSH into the prod host (the same SSH secrets deploy.yml already
uses) using the app's already-running compose stack - no new network
exposure, no separate dependency install. These are ops/infra concerns
(scheduled via Forgejo Actions cron), distinct from the app-domain jobs the
worker's own APScheduler runs.
All three need a runner registered with the `thermograph` label and the
registry's public/mesh exposure resolved (Track B).
Verified: raw YAML syntax valid; actionlint clean except the pre-existing,
already-present-on-.github/workflows "unknown custom runner label" warning
(not something these files introduce - the existing thermograph-lan label
triggers the identical warning). One real finding fixed pre-commit: an
unused shellcheck-flagged loop variable in the boot-check retry loop.
2026-07-21 00:45:44 +00:00
|
|
|
"${THERMOGRAPH_BASE_URL:-https://thermograph.org}"
|