reap zombie VMs with no live Forgejo runner (#170) #171
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!171
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "feat/170-reap-zombie-runners"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Closes #170. Found live while troubleshooting run 105 on PR #169 (stuck "Waiting to run" ~1h).
Bug: the controller's
activecount comes only from OpenStack server status (not SHUTOFF/ERROR, age < reap_max_age). A VM that is ACTIVE in OpenStack but whose Forgejo runner registration is gone — e.g. after a canceled job exits the one-shot (ephemeral: true) runner but the VM never powers off — is a zombie. It's counted as active →desired == active→create=0, and isn't reaped until the 2hreap_max_agebackstop, so no working runner is booted and the queued job starves for up to 2h.Fix: in the reap step, also reap a VM whose
metadata.forgejo_runner_idis no longer present in the liveGET /api/v1/admin/actions/runnerslist, once it's older thanregister_grace_seconds(default 300s — comfortably longer than the ~22s boot + registration, so a VM still mid-registration is never falsely reaped). Best-effort: if the runners-list read fails, the zombie rule is skipped that cycle (never mass-reap on a transient Forgejo error — the existing safety contract).activenow reflects real runner liveness, so a fresh runner is created immediately.New knob:
runner_controller_register_grace_seconds(default 300).Verified:
py_compile+ an 8-case logic test (zombie / healthy-registered / young-within-grace / forgejo-read-fail / no-id / SHUTOFF / ERROR / age-backstop) + ansible-lint clean.Note: deploying this rebuilds the runner-controller image (content-hash tag) and restarts the container — it also self-heals the current live zombie on its first loop. Refs ADR 0021.