thermograph/infra/deploy/forgejo/register-lan-runner.sh
Emi Griffith f25466ca76
All checks were successful
PR build (required check) / changes (pull_request) Successful in 9s
shell-lint / shellcheck (pull_request) Successful in 9s
secrets-guard / encrypted (pull_request) Successful in 11s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
forgejo: make the desktop runner unit restart-always, fix the vps2 config path
Both runners exist and are up; these are the two defects found while verifying
that.

The desktop's systemd --user unit had Restart=on-failure. That is the same
distinction that took Forgejo down for 27 hours on 2026-07-29 (docker-stack.yml
records it for the db service): a daemon that exits 0 is not "finished
successfully", and on-failure cannot tell that from a clean shutdown. The
failure mode is silent — unit inactive (dead), runner offline in the UI, and
every protected-branch merge blocked on a check nothing will produce.

The StartLimit directives that bound that retry loop were in [Service], where
systemd accepts them without complaint and ignores them; the live unit was
running the 10s default rather than the intended 300s. Moved to [Unit], which
is where they are read.

runner-vps2/README told you to copy config.yaml next to docker-compose.yml, but
the compose file mounts ./data:/data and loads --config /data/config.yaml, so a
config there is invisible to the container — the daemon starts on defaults with
no --add-host, and every registry push then fails as if the credential were
wrong. vps2 had a stray copy at the documented path proving the instruction had
been followed.
2026-08-01 11:50:42 -07:00

143 lines
6.8 KiB
Bash
Executable file

#!/usr/bin/env bash
# Re-points the existing LAN dev self-hosted runner from GitHub Actions to
# Forgejo Actions. Run on the SAME machine that already runs the GitHub
# runner (see DEPLOY-DEV.md) — this replaces that runner, it doesn't add a
# second one. Sudo-free, systemd --user, same pattern as the app service.
#
# It registers with BOTH labels the workflows historically needed: general
# CI/build/deploy jobs (`docker`, containerized via this machine's own
# already-installed Docker — no Docker-in-Docker sidecar needed, since a real
# host with a real Docker install needs no such indirection) and the LAN-deploy
# job (`thermograph-lan`, bare/host-native — it wrote to ~/thermograph-dev and
# restarted a systemd --user service, which only worked running directly on the
# host, not inside a container).
#
# `thermograph-lan` IS NOW OBSOLETE. Dev moved off this machine to vps1 and is
# deployed over SSH by the same `Deploy` workflow that ships beta and prod;
# there is no host-native LAN deploy job left for that label to serve. Nothing
# breaks by keeping it registered — no workflow requests it — but do not build
# anything new on it.
#
# This machine keeps serving the `docker` label as EXTRA capacity, and that is
# now a settled call rather than an open one. The claim it used to make here —
# that a Swarm-hosted runner was always-on, so the estate no longer depended on
# this box — was false: no such runner existed. This was the only registered
# runner in the estate until 2026-08-01, and when it went offline on 2026-07-31
# every merge, deploy and backup stopped for 21 hours.
#
# The runner the estate depends on is deploy/forgejo/runner-vps2/. Keeping this
# one is worthwhile (a second runner means a queue drains instead of stalling),
# but nothing may be built on the assumption that this box is up.
#
# bash deploy/forgejo/register-lan-runner.sh <forgejo_url> <registration_token>
#
# Get <registration_token> from the Forgejo web UI:
# repo -> Settings -> Actions -> Runners -> Create new Runner
# (or an org/instance-level runner page, if you want it to serve more than
# this one repo — same as the GitHub runner did).
set -euo pipefail
FORGEJO_URL="${1:?usage: $0 <forgejo_url> <registration_token>}"
TOKEN="${2:?}"
RUNNER_DIR="${RUNNER_DIR:-$HOME/forgejo-runner}"
LABELS="${LABELS:-docker:docker://git.thermograph.org/jinemi/thermograph/ci-runner:v2,thermograph-lan}"
echo "==> Stopping and disabling the old GitHub Actions runner service, if present"
systemctl --user stop github-actions-runner 2>/dev/null || true
systemctl --user disable github-actions-runner 2>/dev/null || true
if ! id -nG "$USER" 2>/dev/null | grep -qw docker; then
echo "WARNING: $USER is not in the 'docker' group — the 'docker:docker://...'"
echo "labeled jobs (general CI) will fail to start a container until this is"
echo "fixed: sudo usermod -aG docker $USER && (log out and back in)."
fi
echo "==> Installing forgejo-runner into $RUNNER_DIR"
mkdir -p "$RUNNER_DIR"
cd "$RUNNER_DIR"
if [ ! -x ./forgejo-runner ]; then
ARCH="$(uname -m)"
case "$ARCH" in
x86_64) BIN_ARCH=amd64 ;;
aarch64) BIN_ARCH=arm64 ;;
*) echo "Unsupported arch: $ARCH — download the right binary by hand from" \
"https://code.forgejo.org/forgejo/runner/releases" >&2; exit 1 ;;
esac
VER="${FORGEJO_RUNNER_VERSION:-6.3.1}"
curl -fsSL -o forgejo-runner \
"https://code.forgejo.org/forgejo/runner/releases/download/v${VER}/forgejo-runner-${VER}-linux-${BIN_ARCH}"
chmod +x forgejo-runner
fi
echo "==> Registering with $FORGEJO_URL (label: $LABELS)"
./forgejo-runner register --no-interactive \
--instance "$FORGEJO_URL" \
--token "$TOKEN" \
--name "thermograph-lan-$(hostname -s)" \
--labels "$LABELS"
echo "==> Runner config (docker.sock automount so docker:-labeled jobs like"
echo " build-push.yml can actually run 'docker build/push'; capacity raised"
echo " from the default of 1 -- a single PR push fires 3+ independent"
echo " workflows (pr-build, secrets-guard, shell-lint) simultaneously, so"
echo " anything less than that serializes jobs that could run in parallel)"
CAPACITY="${CAPACITY:-8}"
./forgejo-runner generate-config \
| sed -e 's/docker_host: "-"/docker_host: "automount"/' \
-e "s/capacity: 1/capacity: ${CAPACITY}/" \
> "${RUNNER_DIR}/config.yaml"
echo "==> systemd --user unit"
mkdir -p "$HOME/.config/systemd/user"
cat > "$HOME/.config/systemd/user/forgejo-runner.service" <<EOF
[Unit]
Description=Forgejo Actions runner (docker + thermograph-lan)
After=network-online.target
# These bound the Restart=always loop below, so a genuinely broken config (bad
# token, docker unreachable) ends as a stopped unit rather than spinning
# forever. They belong in [Unit]: systemd accepts them in [Service] without
# complaint but IGNORES them, leaving the 10s default — \`systemctl --user show
# forgejo-runner -p StartLimitIntervalUSec\` is how you tell which one you got.
StartLimitIntervalSec=300
StartLimitBurst=5
[Service]
WorkingDirectory=${RUNNER_DIR}
# CI job steps shell out to user-installed tools (sops, age, ...) that live in
# ~/.local/bin, not on systemd --user's default PATH. Without this, any
# workflow that needs one (e.g. render-secrets.sh once a host is
# SOPS-configured) fails with "not installed" even though it plainly is.
Environment=PATH=%h/.local/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/usr/games:/usr/local/games:/snap/bin
ExecStart=${RUNNER_DIR}/forgejo-runner daemon -c config.yaml
# \`always\`, NOT \`on-failure\` — the same distinction that took Forgejo down for
# 27 hours on 2026-07-29 (see docker-stack.yml's restart_policy comment): an
# always-on daemon that exits **0** is not "finished successfully", it is a
# daemon that stopped and must come back. \`on-failure\` cannot tell those apart,
# and forgejo-runner exits 0 on several paths — a lost instance connection it
# gives up on, or a SIGTERM from a Docker restart it interprets as a clean
# shutdown. The failure mode is silent: the unit sits \`inactive (dead)\`, the UI
# shows the runner offline, and every protected-branch merge blocks on a check
# that will never be produced.
Restart=always
RestartSec=5
[Install]
WantedBy=default.target
EOF
systemctl --user daemon-reload
systemctl --user enable --now forgejo-runner
loginctl enable-linger "$USER" 2>/dev/null || true
cat <<EOF
Done. The runner now serves Forgejo, not GitHub — one runner, both labels
($LABELS), replacing what used to be a separate Swarm-hosted runner for
general CI plus this machine's own GitHub runner for LAN deploys.
status: systemctl --user status forgejo-runner
logs: journalctl --user -u forgejo-runner -f
restart: systemctl --user restart forgejo-runner
The old github-actions-runner unit was stopped and disabled but not deleted —
remove ~/actions-runner by hand once you've confirmed Forgejo deploys work.
EOF