runner-controller counts zombie VMs (ACTIVE in OpenStack, dead in Forgejo) as capacity → CI job starves up to 2h #170

Stängd
öppnade 2026-07-20 14:08:42 +00:00 av supernaut · 0 kommentarer
Ägare

Severity: HIGH (CI can wedge for up to 2h). Found live 2026-07-20 while troubleshooting a stuck job on PR #169 (run 105 "Waiting to run" ~1h).

Symptom

A queued ci job never starts. Controller logs, every loop:

reconcile: 1 active ephemeral VM(s) after reap
reconcile: 1 queued job(s) for label=ci
reconcile: active=1 desired=1 (queued=1 + min_idle=0, cap=4) → create=0

So it thinks capacity exists and never boots a working runner.

Root cause

The controller's active count comes only from OpenStack server status (reconcile Step 2: a VM counts as active unless SHUTOFF/ERROR or age > reap_max_age). It does not check whether the VM's Forgejo runner is actually alive.

Confirmed both sides at the time:

  • OpenStack: 1 ephemeral VM ephemeral-runner-68dbb936 (7d874baf-…), ACTIVE, age ~3642s, metadata.forgejo_runner_id=3422.
  • Forgejo GET /api/v1/admin/actions/runners: 0 runners.

So the VM was a zombie — VM up, act_runner gone. Trigger: run 104 was canceled, which exited the one-shot (ephemeral: true) runner (it deregistered → 0 runners) but the VM never powered off (stayed ACTIVE). The zombie is counted as active → desired == active → create=0, and it isn't reaped until the 2h reap_max_age backstop. The queued job starves for up to 2h.

Fix (this issue)

In the reap step, also reap a VM that has a forgejo_runner_id in its metadata which is no longer present in the live admin/actions/runners list, once it is past a short boot+register grace (register_grace_seconds, default 300s — longer than the ~22s boot + registration, so a mid-registering VM is never caught). Best-effort: if the runners-list read fails, skip only the zombie rule that cycle (never mass-reap on a transient Forgejo error — existing safety contract). This makes active reflect real runner liveness, so a fresh runner is created immediately.

Immediate remediation (ops)

Delete the zombie VM by id (openstack server delete <id>, or via the controller's reaper) → next loop creates a working runner; or wait for the 2h backstop.

Epic: gitborg/gitborg-docs#1 (CI). Relates to ADR 0021.

**Severity: HIGH** (CI can wedge for up to 2h). Found live 2026-07-20 while troubleshooting a stuck job on PR #169 (run 105 "Waiting to run" ~1h). ## Symptom A queued `ci` job never starts. Controller logs, every loop: ``` reconcile: 1 active ephemeral VM(s) after reap reconcile: 1 queued job(s) for label=ci reconcile: active=1 desired=1 (queued=1 + min_idle=0, cap=4) → create=0 ``` So it thinks capacity exists and never boots a working runner. ## Root cause The controller's `active` count comes **only from OpenStack server status** (`reconcile` Step 2: a VM counts as active unless `SHUTOFF`/`ERROR` or `age > reap_max_age`). It does **not** check whether the VM's Forgejo runner is actually alive. Confirmed both sides at the time: - OpenStack: 1 ephemeral VM `ephemeral-runner-68dbb936` (`7d874baf-…`), **ACTIVE**, age ~3642s, `metadata.forgejo_runner_id=3422`. - Forgejo `GET /api/v1/admin/actions/runners`: **0 runners**. So the VM was a **zombie** — VM up, act_runner gone. Trigger: **run 104 was canceled**, which exited the one-shot (`ephemeral: true`) runner (it deregistered → 0 runners) but the VM never powered off (stayed ACTIVE). The zombie is counted as active → `desired == active` → `create=0`, and it isn't reaped until the 2h `reap_max_age` backstop. The queued job starves for up to 2h. ## Fix (this issue) In the reap step, also reap a VM that has a `forgejo_runner_id` in its metadata which is **no longer present** in the live `admin/actions/runners` list, once it is past a short boot+register grace (`register_grace_seconds`, default 300s — longer than the ~22s boot + registration, so a mid-registering VM is never caught). Best-effort: if the runners-list read fails, skip only the zombie rule that cycle (never mass-reap on a transient Forgejo error — existing safety contract). This makes `active` reflect real runner liveness, so a fresh runner is created immediately. ## Immediate remediation (ops) Delete the zombie VM by id (`openstack server delete <id>`, or via the controller's reaper) → next loop creates a working runner; or wait for the 2h backstop. Epic: gitborg/gitborg-docs#1 (CI). Relates to ADR 0021.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#170
Ingen beskrivning angiven.