runner-controller counts zombie VMs (ACTIVE in OpenStack, dead in Forgejo) as capacity → CI job starves up to 2h #170
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#170
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Severity: HIGH (CI can wedge for up to 2h). Found live 2026-07-20 while troubleshooting a stuck job on PR #169 (run 105 "Waiting to run" ~1h).
Symptom
A queued
cijob never starts. Controller logs, every loop:So it thinks capacity exists and never boots a working runner.
Root cause
The controller's
activecount comes only from OpenStack server status (reconcileStep 2: a VM counts as active unlessSHUTOFF/ERRORorage > reap_max_age). It does not check whether the VM's Forgejo runner is actually alive.Confirmed both sides at the time:
ephemeral-runner-68dbb936(7d874baf-…), ACTIVE, age ~3642s,metadata.forgejo_runner_id=3422.GET /api/v1/admin/actions/runners: 0 runners.So the VM was a zombie — VM up, act_runner gone. Trigger: run 104 was canceled, which exited the one-shot (
ephemeral: true) runner (it deregistered → 0 runners) but the VM never powered off (stayed ACTIVE). The zombie is counted as active →desired == active→create=0, and it isn't reaped until the 2hreap_max_agebackstop. The queued job starves for up to 2h.Fix (this issue)
In the reap step, also reap a VM that has a
forgejo_runner_idin its metadata which is no longer present in the liveadmin/actions/runnerslist, once it is past a short boot+register grace (register_grace_seconds, default 300s — longer than the ~22s boot + registration, so a mid-registering VM is never caught). Best-effort: if the runners-list read fails, skip only the zombie rule that cycle (never mass-reap on a transient Forgejo error — existing safety contract). This makesactivereflect real runner liveness, so a fresh runner is created immediately.Immediate remediation (ops)
Delete the zombie VM by id (
openstack server delete <id>, or via the controller's reaper) → next loop creates a working runner; or wait for the 2h backstop.Epic: gitborg/gitborg-docs#1 (CI). Relates to ADR 0021.