reap zombie VMs with no live Forgejo runner (#170) #171

Sammanfogat
supernaut sammanfogade 1 incheckning från feat/170-reap-zombie-runners in i main 2026-07-20 14:31:51 +00:00
Ägare

Closes #170. Found live while troubleshooting run 105 on PR #169 (stuck "Waiting to run" ~1h).

Bug: the controller's active count comes only from OpenStack server status (not SHUTOFF/ERROR, age < reap_max_age). A VM that is ACTIVE in OpenStack but whose Forgejo runner registration is gone — e.g. after a canceled job exits the one-shot (ephemeral: true) runner but the VM never powers off — is a zombie. It's counted as active → desired == active → create=0, and isn't reaped until the 2h reap_max_age backstop, so no working runner is booted and the queued job starves for up to 2h.

Fix: in the reap step, also reap a VM whose metadata.forgejo_runner_id is no longer present in the live GET /api/v1/admin/actions/runners list, once it's older than register_grace_seconds (default 300s — comfortably longer than the ~22s boot + registration, so a VM still mid-registration is never falsely reaped). Best-effort: if the runners-list read fails, the zombie rule is skipped that cycle (never mass-reap on a transient Forgejo error — the existing safety contract). active now reflects real runner liveness, so a fresh runner is created immediately.

New knob: runner_controller_register_grace_seconds (default 300).

Verified: py_compile + an 8-case logic test (zombie / healthy-registered / young-within-grace / forgejo-read-fail / no-id / SHUTOFF / ERROR / age-backstop) + ansible-lint clean.

Note: deploying this rebuilds the runner-controller image (content-hash tag) and restarts the container — it also self-heals the current live zombie on its first loop. Refs ADR 0021.

Closes #170. Found live while troubleshooting run 105 on PR #169 (stuck "Waiting to run" ~1h). **Bug:** the controller's `active` count comes only from OpenStack server status (not SHUTOFF/ERROR, age < reap_max_age). A VM that is ACTIVE in OpenStack but whose Forgejo runner registration is gone — e.g. after a **canceled job** exits the one-shot (`ephemeral: true`) runner but the VM never powers off — is a **zombie**. It's counted as active → `desired == active` → `create=0`, and isn't reaped until the 2h `reap_max_age` backstop, so no working runner is booted and the queued job starves for up to 2h. **Fix:** in the reap step, also reap a VM whose `metadata.forgejo_runner_id` is no longer present in the live `GET /api/v1/admin/actions/runners` list, once it's older than `register_grace_seconds` (default 300s — comfortably longer than the ~22s boot + registration, so a VM still mid-registration is never falsely reaped). Best-effort: if the runners-list read fails, the zombie rule is **skipped that cycle** (never mass-reap on a transient Forgejo error — the existing safety contract). `active` now reflects real runner liveness, so a fresh runner is created immediately. New knob: `runner_controller_register_grace_seconds` (default 300). **Verified:** `py_compile` + an 8-case logic test (zombie / healthy-registered / young-within-grace / forgejo-read-fail / no-id / SHUTOFF / ERROR / age-backstop) + ansible-lint clean. **Note:** deploying this rebuilds the runner-controller image (content-hash tag) and restarts the container — it also self-heals the current live zombie on its first loop. Refs ADR 0021.
supernaut lade till 1 incheckning 2026-07-20 14:11:46 +00:00
fix(runner-controller): reap zombie VMs with no live Forgejo runner (#170)
Alla kontroller lyckades
ci / ci (pull_request) Successful in 2m39s
a6bcff39d2
A queued CI job could starve for up to 2h. The controller's capacity math
counts a VM as active purely from OpenStack status (not SHUTOFF/ERROR, age <
reap_max_age), ignoring whether its Forgejo runner is alive. When a job is
canceled, the one-shot (ephemeral) runner exits and deregisters but the VM can
stay ACTIVE — a zombie. It's counted active → desired==active → create=0, and
isn't reaped until the 2h reap_max_age backstop, so no working runner is booted.

Fix: in the reap step, also reap a VM whose metadata.forgejo_runner_id is no
longer in the live admin/actions/runners list, once past register_grace_seconds
(default 300s — boot+register headroom). Best-effort: if the runners-list read
fails, the zombie rule is skipped that cycle (never mass-reap on a transient
Forgejo error — existing safety contract). This makes active reflect real runner
liveness so a fresh runner is created immediately.

py_compile + 8-case logic test (zombie/healthy/young/read-fail/no-id/shutoff/
error/backstop) + ansible-lint clean. Closes #170.
supernaut sammanfogade incheckning c0adb49110 till main 2026-07-20 14:31:51 +00:00
supernaut tog bort grenen feat/170-reap-zombie-runners 2026-07-20 14:31:51 +00:00
supernaut refererade denna ändringsförfrågan från en incheckning 2026-07-20 14:31:53 +00:00
supernaut refererade denna ändringsförfrågan från en incheckning 2026-08-03 09:41:34 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!171
Ingen beskrivning angiven.