All checks were successful
secrets-guard / encrypted (pull_request) Successful in 5s
shell-lint / shellcheck (pull_request) Successful in 8s
PR build (required check) / changes (pull_request) Successful in 13s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 7s
secrets-guard / encrypted (push) Successful in 7s
shell-lint / shellcheck (push) Successful in 7s
The estate had exactly one registered runner, on the desktop. It went offline at 2026-07-31 16:31Z; for the next 21 hours no PR could satisfy a required check, no deploy could run, and the 03:00Z ops-cron -- the only backup for both application databases and for Forgejo -- did not fire. Forgejo queued that scheduled run rather than dropping it, so it completed on reconnect and nothing was lost. A longer outage would have meant real gaps. Three files claimed an "always-on Swarm-hosted runner" existed and that the estate therefore no longer depended on the desktop. It did not exist: an early revision of docker-stack.yml ran one as a Docker-in-Docker sidecar and it was removed. That claim is why a single point of failure sat unnoticed. Corrected in docker-stack.yml, forgejo/README.md and register-lan-runner.sh. The new runner is a plain restart:always container, not a Swarm service: a Swarm-scheduled runner cannot redeploy the Swarm that schedules it, so CI would be gone exactly when the cluster is what is broken. It runs from /opt/forgejo-runner rather than in place, because the checkout is reset on every prod deploy and one `git clean -fdx` there would destroy the registration. vps2 runs prod, so the socket mount is bounded rather than assumed benign: capacity 1, --cpus=2/--memory=4g on job containers, valid_volumes empty so no job can bind-mount /etc/thermograph.env, and no thermograph network joined. This defends against accident, not against a hostile workflow author -- stated plainly in the compose header rather than implied. The desktop runner stays registered as extra capacity. Nothing may assume it is up.
127 lines
5.7 KiB
Bash
Executable file
127 lines
5.7 KiB
Bash
Executable file
#!/usr/bin/env bash
|
|
# Re-points the existing LAN dev self-hosted runner from GitHub Actions to
|
|
# Forgejo Actions. Run on the SAME machine that already runs the GitHub
|
|
# runner (see DEPLOY-DEV.md) — this replaces that runner, it doesn't add a
|
|
# second one. Sudo-free, systemd --user, same pattern as the app service.
|
|
#
|
|
# It registers with BOTH labels the workflows historically needed: general
|
|
# CI/build/deploy jobs (`docker`, containerized via this machine's own
|
|
# already-installed Docker — no Docker-in-Docker sidecar needed, since a real
|
|
# host with a real Docker install needs no such indirection) and the LAN-deploy
|
|
# job (`thermograph-lan`, bare/host-native — it wrote to ~/thermograph-dev and
|
|
# restarted a systemd --user service, which only worked running directly on the
|
|
# host, not inside a container).
|
|
#
|
|
# `thermograph-lan` IS NOW OBSOLETE. Dev moved off this machine to vps1 and is
|
|
# deployed over SSH by the same `Deploy` workflow that ships beta and prod;
|
|
# there is no host-native LAN deploy job left for that label to serve. Nothing
|
|
# breaks by keeping it registered — no workflow requests it — but do not build
|
|
# anything new on it.
|
|
#
|
|
# This machine keeps serving the `docker` label as EXTRA capacity, and that is
|
|
# now a settled call rather than an open one. The claim it used to make here —
|
|
# that a Swarm-hosted runner was always-on, so the estate no longer depended on
|
|
# this box — was false: no such runner existed. This was the only registered
|
|
# runner in the estate until 2026-08-01, and when it went offline on 2026-07-31
|
|
# every merge, deploy and backup stopped for 21 hours.
|
|
#
|
|
# The runner the estate depends on is deploy/forgejo/runner-vps2/. Keeping this
|
|
# one is worthwhile (a second runner means a queue drains instead of stalling),
|
|
# but nothing may be built on the assumption that this box is up.
|
|
#
|
|
# bash deploy/forgejo/register-lan-runner.sh <forgejo_url> <registration_token>
|
|
#
|
|
# Get <registration_token> from the Forgejo web UI:
|
|
# repo -> Settings -> Actions -> Runners -> Create new Runner
|
|
# (or an org/instance-level runner page, if you want it to serve more than
|
|
# this one repo — same as the GitHub runner did).
|
|
set -euo pipefail
|
|
|
|
FORGEJO_URL="${1:?usage: $0 <forgejo_url> <registration_token>}"
|
|
TOKEN="${2:?}"
|
|
RUNNER_DIR="${RUNNER_DIR:-$HOME/forgejo-runner}"
|
|
LABELS="${LABELS:-docker:docker://git.thermograph.org/emi/thermograph/ci-runner:v2,thermograph-lan}"
|
|
|
|
echo "==> Stopping and disabling the old GitHub Actions runner service, if present"
|
|
systemctl --user stop github-actions-runner 2>/dev/null || true
|
|
systemctl --user disable github-actions-runner 2>/dev/null || true
|
|
|
|
if ! id -nG "$USER" 2>/dev/null | grep -qw docker; then
|
|
echo "WARNING: $USER is not in the 'docker' group — the 'docker:docker://...'"
|
|
echo "labeled jobs (general CI) will fail to start a container until this is"
|
|
echo "fixed: sudo usermod -aG docker $USER && (log out and back in)."
|
|
fi
|
|
|
|
echo "==> Installing forgejo-runner into $RUNNER_DIR"
|
|
mkdir -p "$RUNNER_DIR"
|
|
cd "$RUNNER_DIR"
|
|
if [ ! -x ./forgejo-runner ]; then
|
|
ARCH="$(uname -m)"
|
|
case "$ARCH" in
|
|
x86_64) BIN_ARCH=amd64 ;;
|
|
aarch64) BIN_ARCH=arm64 ;;
|
|
*) echo "Unsupported arch: $ARCH — download the right binary by hand from" \
|
|
"https://code.forgejo.org/forgejo/runner/releases" >&2; exit 1 ;;
|
|
esac
|
|
VER="${FORGEJO_RUNNER_VERSION:-6.3.1}"
|
|
curl -fsSL -o forgejo-runner \
|
|
"https://code.forgejo.org/forgejo/runner/releases/download/v${VER}/forgejo-runner-${VER}-linux-${BIN_ARCH}"
|
|
chmod +x forgejo-runner
|
|
fi
|
|
|
|
echo "==> Registering with $FORGEJO_URL (label: $LABELS)"
|
|
./forgejo-runner register --no-interactive \
|
|
--instance "$FORGEJO_URL" \
|
|
--token "$TOKEN" \
|
|
--name "thermograph-lan-$(hostname -s)" \
|
|
--labels "$LABELS"
|
|
|
|
echo "==> Runner config (docker.sock automount so docker:-labeled jobs like"
|
|
echo " build-push.yml can actually run 'docker build/push'; capacity raised"
|
|
echo " from the default of 1 -- a single PR push fires 3+ independent"
|
|
echo " workflows (pr-build, secrets-guard, shell-lint) simultaneously, so"
|
|
echo " anything less than that serializes jobs that could run in parallel)"
|
|
CAPACITY="${CAPACITY:-8}"
|
|
./forgejo-runner generate-config \
|
|
| sed -e 's/docker_host: "-"/docker_host: "automount"/' \
|
|
-e "s/capacity: 1/capacity: ${CAPACITY}/" \
|
|
> "${RUNNER_DIR}/config.yaml"
|
|
|
|
echo "==> systemd --user unit"
|
|
mkdir -p "$HOME/.config/systemd/user"
|
|
cat > "$HOME/.config/systemd/user/forgejo-runner.service" <<EOF
|
|
[Unit]
|
|
Description=Forgejo Actions runner (docker + thermograph-lan)
|
|
After=network-online.target
|
|
|
|
[Service]
|
|
WorkingDirectory=${RUNNER_DIR}
|
|
# CI job steps shell out to user-installed tools (sops, age, ...) that live in
|
|
# ~/.local/bin, not on systemd --user's default PATH. Without this, any
|
|
# workflow that needs one (e.g. render-secrets.sh once a host is
|
|
# SOPS-configured) fails with "not installed" even though it plainly is.
|
|
Environment=PATH=%h/.local/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/usr/games:/usr/local/games:/snap/bin
|
|
ExecStart=${RUNNER_DIR}/forgejo-runner daemon -c config.yaml
|
|
Restart=on-failure
|
|
RestartSec=5
|
|
|
|
[Install]
|
|
WantedBy=default.target
|
|
EOF
|
|
|
|
systemctl --user daemon-reload
|
|
systemctl --user enable --now forgejo-runner
|
|
loginctl enable-linger "$USER" 2>/dev/null || true
|
|
|
|
cat <<EOF
|
|
|
|
Done. The runner now serves Forgejo, not GitHub — one runner, both labels
|
|
($LABELS), replacing what used to be a separate Swarm-hosted runner for
|
|
general CI plus this machine's own GitHub runner for LAN deploys.
|
|
status: systemctl --user status forgejo-runner
|
|
logs: journalctl --user -u forgejo-runner -f
|
|
restart: systemctl --user restart forgejo-runner
|
|
|
|
The old github-actions-runner unit was stopped and disabled but not deleted —
|
|
remove ~/actions-runner by hand once you've confirmed Forgejo deploys work.
|
|
EOF
|