thermograph/infra/deploy/forgejo/register-lan-runner.sh
Emi Griffith 53e2eb7c84
All checks were successful
Sync infra to hosts / sync-beta (push) Has been skipped
Sync infra to hosts / sync-prod (push) Has been skipped
Sync infra to hosts / sync-centralis (push) Has been skipped
Sync infra to hosts / sync-dev (push) Successful in 9s
secrets-guard / encrypted (push) Successful in 5s
Validate observability stack / validate (push) Successful in 12s
shell-lint / shellcheck (push) Successful in 8s
Build + push images (Forgejo registry) / build-push (frontend) (push) Successful in 26s
Build + push images (Forgejo registry) / build-push (backend) (push) Successful in 1m10s
Deploy / deploy (backend) (push) Successful in 1m41s
Deploy / deploy (frontend) (push) Successful in 1m48s
secrets-guard / encrypted (pull_request) Successful in 8s
shell-lint / shellcheck (pull_request) Successful in 9s
PR build (required check) / changes (pull_request) Successful in 20s
PR build (required check) / validate-observability (pull_request) Successful in 15s
PR build (required check) / build-frontend (pull_request) Successful in 17s
PR build (required check) / build-backend (pull_request) Successful in 2m26s
PR build (required check) / gate (pull_request) Successful in 1s
registry: move the registry host to dev.jinemi.com, MCP to mcp.jinemi.com
Completes the Forgejo domain migration. ROOT_URL moved to dev.jinemi.com
earlier; the registry half was deliberately deferred. Every image name,
REGISTRY_HOST default, runner label and --add-host pin now names
dev.jinemi.com, so the registry host and the bearer-token realm agree again.

Both names address the same Forgejo, so no image needs re-pushing and a
rollback to a tag pushed under the old prefix still resolves.
git.thermograph.org therefore stays served off the same Caddy site block --
one block, so the /v2/* mesh-only matcher keeps covering both names -- for
pre-migration tags and for runners holding it as their registered instance
URL.

Mesh clients now pin both names in /etc/hosts: the new one as registry host
and token realm, the old one for pre-migration tags. runner-vps2/config.yaml
carries both --add-host entries for the same reason.

Also renames Centralis' endpoint to mcp.jinemi.com in the two places this
repo names it; Centralis itself is provisioned outside this repo.

Host-side steps this cannot do (documented in deploy/forgejo/README.md,
"Host-side steps"): the Forgejo Actions variable REGISTRY_HOST, docker login
against the new host, and the /etc/hosts pins.
2026-08-01 16:06:22 -07:00

143 lines
6.8 KiB
Bash
Executable file

#!/usr/bin/env bash
# Re-points the existing LAN dev self-hosted runner from GitHub Actions to
# Forgejo Actions. Run on the SAME machine that already runs the GitHub
# runner (see DEPLOY-DEV.md) — this replaces that runner, it doesn't add a
# second one. Sudo-free, systemd --user, same pattern as the app service.
#
# It registers with BOTH labels the workflows historically needed: general
# CI/build/deploy jobs (`docker`, containerized via this machine's own
# already-installed Docker — no Docker-in-Docker sidecar needed, since a real
# host with a real Docker install needs no such indirection) and the LAN-deploy
# job (`thermograph-lan`, bare/host-native — it wrote to ~/thermograph-dev and
# restarted a systemd --user service, which only worked running directly on the
# host, not inside a container).
#
# `thermograph-lan` IS NOW OBSOLETE. Dev moved off this machine to vps1 and is
# deployed over SSH by the same `Deploy` workflow that ships beta and prod;
# there is no host-native LAN deploy job left for that label to serve. Nothing
# breaks by keeping it registered — no workflow requests it — but do not build
# anything new on it.
#
# This machine keeps serving the `docker` label as EXTRA capacity, and that is
# now a settled call rather than an open one. The claim it used to make here —
# that a Swarm-hosted runner was always-on, so the estate no longer depended on
# this box — was false: no such runner existed. This was the only registered
# runner in the estate until 2026-08-01, and when it went offline on 2026-07-31
# every merge, deploy and backup stopped for 21 hours.
#
# The runner the estate depends on is deploy/forgejo/runner-vps2/. Keeping this
# one is worthwhile (a second runner means a queue drains instead of stalling),
# but nothing may be built on the assumption that this box is up.
#
# bash deploy/forgejo/register-lan-runner.sh <forgejo_url> <registration_token>
#
# Get <registration_token> from the Forgejo web UI:
# repo -> Settings -> Actions -> Runners -> Create new Runner
# (or an org/instance-level runner page, if you want it to serve more than
# this one repo — same as the GitHub runner did).
set -euo pipefail
FORGEJO_URL="${1:?usage: $0 <forgejo_url> <registration_token>}"
TOKEN="${2:?}"
RUNNER_DIR="${RUNNER_DIR:-$HOME/forgejo-runner}"
LABELS="${LABELS:-docker:docker://dev.jinemi.com/jinemi/thermograph/ci-runner:v2,thermograph-lan}"
echo "==> Stopping and disabling the old GitHub Actions runner service, if present"
systemctl --user stop github-actions-runner 2>/dev/null || true
systemctl --user disable github-actions-runner 2>/dev/null || true
if ! id -nG "$USER" 2>/dev/null | grep -qw docker; then
echo "WARNING: $USER is not in the 'docker' group — the 'docker:docker://...'"
echo "labeled jobs (general CI) will fail to start a container until this is"
echo "fixed: sudo usermod -aG docker $USER && (log out and back in)."
fi
echo "==> Installing forgejo-runner into $RUNNER_DIR"
mkdir -p "$RUNNER_DIR"
cd "$RUNNER_DIR"
if [ ! -x ./forgejo-runner ]; then
ARCH="$(uname -m)"
case "$ARCH" in
x86_64) BIN_ARCH=amd64 ;;
aarch64) BIN_ARCH=arm64 ;;
*) echo "Unsupported arch: $ARCH — download the right binary by hand from" \
"https://code.forgejo.org/forgejo/runner/releases" >&2; exit 1 ;;
esac
VER="${FORGEJO_RUNNER_VERSION:-6.3.1}"
curl -fsSL -o forgejo-runner \
"https://code.forgejo.org/forgejo/runner/releases/download/v${VER}/forgejo-runner-${VER}-linux-${BIN_ARCH}"
chmod +x forgejo-runner
fi
echo "==> Registering with $FORGEJO_URL (label: $LABELS)"
./forgejo-runner register --no-interactive \
--instance "$FORGEJO_URL" \
--token "$TOKEN" \
--name "thermograph-lan-$(hostname -s)" \
--labels "$LABELS"
echo "==> Runner config (docker.sock automount so docker:-labeled jobs like"
echo " build-push.yml can actually run 'docker build/push'; capacity raised"
echo " from the default of 1 -- a single PR push fires 3+ independent"
echo " workflows (pr-build, secrets-guard, shell-lint) simultaneously, so"
echo " anything less than that serializes jobs that could run in parallel)"
CAPACITY="${CAPACITY:-8}"
./forgejo-runner generate-config \
| sed -e 's/docker_host: "-"/docker_host: "automount"/' \
-e "s/capacity: 1/capacity: ${CAPACITY}/" \
> "${RUNNER_DIR}/config.yaml"
echo "==> systemd --user unit"
mkdir -p "$HOME/.config/systemd/user"
cat > "$HOME/.config/systemd/user/forgejo-runner.service" <<EOF
[Unit]
Description=Forgejo Actions runner (docker + thermograph-lan)
After=network-online.target
# These bound the Restart=always loop below, so a genuinely broken config (bad
# token, docker unreachable) ends as a stopped unit rather than spinning
# forever. They belong in [Unit]: systemd accepts them in [Service] without
# complaint but IGNORES them, leaving the 10s default — \`systemctl --user show
# forgejo-runner -p StartLimitIntervalUSec\` is how you tell which one you got.
StartLimitIntervalSec=300
StartLimitBurst=5
[Service]
WorkingDirectory=${RUNNER_DIR}
# CI job steps shell out to user-installed tools (sops, age, ...) that live in
# ~/.local/bin, not on systemd --user's default PATH. Without this, any
# workflow that needs one (e.g. render-secrets.sh once a host is
# SOPS-configured) fails with "not installed" even though it plainly is.
Environment=PATH=%h/.local/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/usr/games:/usr/local/games:/snap/bin
ExecStart=${RUNNER_DIR}/forgejo-runner daemon -c config.yaml
# \`always\`, NOT \`on-failure\` — the same distinction that took Forgejo down for
# 27 hours on 2026-07-29 (see docker-stack.yml's restart_policy comment): an
# always-on daemon that exits **0** is not "finished successfully", it is a
# daemon that stopped and must come back. \`on-failure\` cannot tell those apart,
# and forgejo-runner exits 0 on several paths — a lost instance connection it
# gives up on, or a SIGTERM from a Docker restart it interprets as a clean
# shutdown. The failure mode is silent: the unit sits \`inactive (dead)\`, the UI
# shows the runner offline, and every protected-branch merge blocks on a check
# that will never be produced.
Restart=always
RestartSec=5
[Install]
WantedBy=default.target
EOF
systemctl --user daemon-reload
systemctl --user enable --now forgejo-runner
loginctl enable-linger "$USER" 2>/dev/null || true
cat <<EOF
Done. The runner now serves Forgejo, not GitHub — one runner, both labels
($LABELS), replacing what used to be a separate Swarm-hosted runner for
general CI plus this machine's own GitHub runner for LAN deploys.
status: systemctl --user status forgejo-runner
logs: journalctl --user -u forgejo-runner -f
restart: systemctl --user restart forgejo-runner
The old github-actions-runner unit was stopped and disabled but not deleted —
remove ~/actions-runner by hand once you've confirmed Forgejo deploys work.
EOF