All checks were successful
PR build (required check) / changes (pull_request) Successful in 9s
shell-lint / shellcheck (pull_request) Successful in 9s
secrets-guard / encrypted (pull_request) Successful in 11s
PR build (required check) / build-backend (pull_request) Has been skipped
PR build (required check) / build-frontend (pull_request) Has been skipped
PR build (required check) / validate-observability (pull_request) Has been skipped
PR build (required check) / gate (pull_request) Successful in 2s
Both runners exist and are up; these are the two defects found while verifying that. The desktop's systemd --user unit had Restart=on-failure. That is the same distinction that took Forgejo down for 27 hours on 2026-07-29 (docker-stack.yml records it for the db service): a daemon that exits 0 is not "finished successfully", and on-failure cannot tell that from a clean shutdown. The failure mode is silent — unit inactive (dead), runner offline in the UI, and every protected-branch merge blocked on a check nothing will produce. The StartLimit directives that bound that retry loop were in [Service], where systemd accepts them without complaint and ignores them; the live unit was running the 10s default rather than the intended 300s. Moved to [Unit], which is where they are read. runner-vps2/README told you to copy config.yaml next to docker-compose.yml, but the compose file mounts ./data:/data and loads --config /data/config.yaml, so a config there is invisible to the container — the daemon starts on defaults with no --add-host, and every registry push then fails as if the credential were wrong. vps2 had a stray copy at the documented path proving the instruction had been followed.
143 lines
6.8 KiB
Bash
Executable file
143 lines
6.8 KiB
Bash
Executable file
#!/usr/bin/env bash
|
|
# Re-points the existing LAN dev self-hosted runner from GitHub Actions to
|
|
# Forgejo Actions. Run on the SAME machine that already runs the GitHub
|
|
# runner (see DEPLOY-DEV.md) — this replaces that runner, it doesn't add a
|
|
# second one. Sudo-free, systemd --user, same pattern as the app service.
|
|
#
|
|
# It registers with BOTH labels the workflows historically needed: general
|
|
# CI/build/deploy jobs (`docker`, containerized via this machine's own
|
|
# already-installed Docker — no Docker-in-Docker sidecar needed, since a real
|
|
# host with a real Docker install needs no such indirection) and the LAN-deploy
|
|
# job (`thermograph-lan`, bare/host-native — it wrote to ~/thermograph-dev and
|
|
# restarted a systemd --user service, which only worked running directly on the
|
|
# host, not inside a container).
|
|
#
|
|
# `thermograph-lan` IS NOW OBSOLETE. Dev moved off this machine to vps1 and is
|
|
# deployed over SSH by the same `Deploy` workflow that ships beta and prod;
|
|
# there is no host-native LAN deploy job left for that label to serve. Nothing
|
|
# breaks by keeping it registered — no workflow requests it — but do not build
|
|
# anything new on it.
|
|
#
|
|
# This machine keeps serving the `docker` label as EXTRA capacity, and that is
|
|
# now a settled call rather than an open one. The claim it used to make here —
|
|
# that a Swarm-hosted runner was always-on, so the estate no longer depended on
|
|
# this box — was false: no such runner existed. This was the only registered
|
|
# runner in the estate until 2026-08-01, and when it went offline on 2026-07-31
|
|
# every merge, deploy and backup stopped for 21 hours.
|
|
#
|
|
# The runner the estate depends on is deploy/forgejo/runner-vps2/. Keeping this
|
|
# one is worthwhile (a second runner means a queue drains instead of stalling),
|
|
# but nothing may be built on the assumption that this box is up.
|
|
#
|
|
# bash deploy/forgejo/register-lan-runner.sh <forgejo_url> <registration_token>
|
|
#
|
|
# Get <registration_token> from the Forgejo web UI:
|
|
# repo -> Settings -> Actions -> Runners -> Create new Runner
|
|
# (or an org/instance-level runner page, if you want it to serve more than
|
|
# this one repo — same as the GitHub runner did).
|
|
set -euo pipefail
|
|
|
|
FORGEJO_URL="${1:?usage: $0 <forgejo_url> <registration_token>}"
|
|
TOKEN="${2:?}"
|
|
RUNNER_DIR="${RUNNER_DIR:-$HOME/forgejo-runner}"
|
|
LABELS="${LABELS:-docker:docker://git.thermograph.org/jinemi/thermograph/ci-runner:v2,thermograph-lan}"
|
|
|
|
echo "==> Stopping and disabling the old GitHub Actions runner service, if present"
|
|
systemctl --user stop github-actions-runner 2>/dev/null || true
|
|
systemctl --user disable github-actions-runner 2>/dev/null || true
|
|
|
|
if ! id -nG "$USER" 2>/dev/null | grep -qw docker; then
|
|
echo "WARNING: $USER is not in the 'docker' group — the 'docker:docker://...'"
|
|
echo "labeled jobs (general CI) will fail to start a container until this is"
|
|
echo "fixed: sudo usermod -aG docker $USER && (log out and back in)."
|
|
fi
|
|
|
|
echo "==> Installing forgejo-runner into $RUNNER_DIR"
|
|
mkdir -p "$RUNNER_DIR"
|
|
cd "$RUNNER_DIR"
|
|
if [ ! -x ./forgejo-runner ]; then
|
|
ARCH="$(uname -m)"
|
|
case "$ARCH" in
|
|
x86_64) BIN_ARCH=amd64 ;;
|
|
aarch64) BIN_ARCH=arm64 ;;
|
|
*) echo "Unsupported arch: $ARCH — download the right binary by hand from" \
|
|
"https://code.forgejo.org/forgejo/runner/releases" >&2; exit 1 ;;
|
|
esac
|
|
VER="${FORGEJO_RUNNER_VERSION:-6.3.1}"
|
|
curl -fsSL -o forgejo-runner \
|
|
"https://code.forgejo.org/forgejo/runner/releases/download/v${VER}/forgejo-runner-${VER}-linux-${BIN_ARCH}"
|
|
chmod +x forgejo-runner
|
|
fi
|
|
|
|
echo "==> Registering with $FORGEJO_URL (label: $LABELS)"
|
|
./forgejo-runner register --no-interactive \
|
|
--instance "$FORGEJO_URL" \
|
|
--token "$TOKEN" \
|
|
--name "thermograph-lan-$(hostname -s)" \
|
|
--labels "$LABELS"
|
|
|
|
echo "==> Runner config (docker.sock automount so docker:-labeled jobs like"
|
|
echo " build-push.yml can actually run 'docker build/push'; capacity raised"
|
|
echo " from the default of 1 -- a single PR push fires 3+ independent"
|
|
echo " workflows (pr-build, secrets-guard, shell-lint) simultaneously, so"
|
|
echo " anything less than that serializes jobs that could run in parallel)"
|
|
CAPACITY="${CAPACITY:-8}"
|
|
./forgejo-runner generate-config \
|
|
| sed -e 's/docker_host: "-"/docker_host: "automount"/' \
|
|
-e "s/capacity: 1/capacity: ${CAPACITY}/" \
|
|
> "${RUNNER_DIR}/config.yaml"
|
|
|
|
echo "==> systemd --user unit"
|
|
mkdir -p "$HOME/.config/systemd/user"
|
|
cat > "$HOME/.config/systemd/user/forgejo-runner.service" <<EOF
|
|
[Unit]
|
|
Description=Forgejo Actions runner (docker + thermograph-lan)
|
|
After=network-online.target
|
|
# These bound the Restart=always loop below, so a genuinely broken config (bad
|
|
# token, docker unreachable) ends as a stopped unit rather than spinning
|
|
# forever. They belong in [Unit]: systemd accepts them in [Service] without
|
|
# complaint but IGNORES them, leaving the 10s default — \`systemctl --user show
|
|
# forgejo-runner -p StartLimitIntervalUSec\` is how you tell which one you got.
|
|
StartLimitIntervalSec=300
|
|
StartLimitBurst=5
|
|
|
|
[Service]
|
|
WorkingDirectory=${RUNNER_DIR}
|
|
# CI job steps shell out to user-installed tools (sops, age, ...) that live in
|
|
# ~/.local/bin, not on systemd --user's default PATH. Without this, any
|
|
# workflow that needs one (e.g. render-secrets.sh once a host is
|
|
# SOPS-configured) fails with "not installed" even though it plainly is.
|
|
Environment=PATH=%h/.local/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/usr/games:/usr/local/games:/snap/bin
|
|
ExecStart=${RUNNER_DIR}/forgejo-runner daemon -c config.yaml
|
|
# \`always\`, NOT \`on-failure\` — the same distinction that took Forgejo down for
|
|
# 27 hours on 2026-07-29 (see docker-stack.yml's restart_policy comment): an
|
|
# always-on daemon that exits **0** is not "finished successfully", it is a
|
|
# daemon that stopped and must come back. \`on-failure\` cannot tell those apart,
|
|
# and forgejo-runner exits 0 on several paths — a lost instance connection it
|
|
# gives up on, or a SIGTERM from a Docker restart it interprets as a clean
|
|
# shutdown. The failure mode is silent: the unit sits \`inactive (dead)\`, the UI
|
|
# shows the runner offline, and every protected-branch merge blocks on a check
|
|
# that will never be produced.
|
|
Restart=always
|
|
RestartSec=5
|
|
|
|
[Install]
|
|
WantedBy=default.target
|
|
EOF
|
|
|
|
systemctl --user daemon-reload
|
|
systemctl --user enable --now forgejo-runner
|
|
loginctl enable-linger "$USER" 2>/dev/null || true
|
|
|
|
cat <<EOF
|
|
|
|
Done. The runner now serves Forgejo, not GitHub — one runner, both labels
|
|
($LABELS), replacing what used to be a separate Swarm-hosted runner for
|
|
general CI plus this machine's own GitHub runner for LAN deploys.
|
|
status: systemctl --user status forgejo-runner
|
|
logs: journalctl --user -u forgejo-runner -f
|
|
restart: systemctl --user restart forgejo-runner
|
|
|
|
The old github-actions-runner unit was stopped and disabled but not deleted —
|
|
remove ~/actions-runner by hand once you've confirmed Forgejo deploys work.
|
|
EOF
|