fix(runner-controller): busy VMs are counted as spare capacity, starving queued jobs #327

Stängd
öppnade 2026-08-02 12:30:45 +00:00 av supernaut · 0 kommentarer
Ägare

Problem

reconcile() in ansible/roles/runner-controller/files/controller.py computes:

desired   = min(max_total, queued_jobs + min_idle)
to_create = max(0, desired - active_count)

active_count includes every ephemeral VM that is not being reaped — including VMs that have
already claimed a job and are busy running it
. With min_idle: 0, that means a newly queued job
creates no capacity at all while any other VM is busy: it waits for that VM to finish, power off,
be seen SHUTOFF and be reaped before a VM is booted for it.

Evidence

From the controller journal (one CI job, cap 4, two of four slots free the whole time):

11:12:42Z  reconcile: active=0 desired=2 (queued=2 + min_idle=0, cap=4) → create=2
11:13:52Z  reconcile: active=2 desired=1 (queued=1 + min_idle=0, cap=4) → create=0   ← job waiting
…  six identical cycles, 61 s …
11:14:52Z  reap: … status=SHUTOFF (x2)
11:14:53Z  reconcile: active=0 desired=1 (queued=1 + min_idle=0, cap=4) → create=1
11:14:55Z  create: booted …

The job's total wait was 93 s, of which 61 s was the controller declining to boot into free
capacity
. The cold start itself was a normal 24 s.

Impact (measured over 7 days, 345 workflow runs)

  • Pickup latency is bimodal: median 33 s, but p90 137 s and 20 % of runs (70/345) take
    over 60 s with a mean of 187 s.
  • 301 controller cycles carried the unambiguous starvation signature (active > desired ≥ 1,
    create=0) — a lower bound, since the active == desired case is indistinguishable in the log.
  • Characteristic fingerprint: two runs triggered in the same second, one starting at ~33 s and the
    other at ~155 s (≈ one job duration + reap + cold start).

RunnerQueueStalled does not catch this — its for: window is minutes and these stalls are
60–200 s.

Proposed fix

Count only idle capacity. Jobs in flight = waiting + running:

desired   = min(max_total, queued_jobs + running_jobs + min_idle)
to_create = max(0, desired - active_count)

_list_running_jobs(session, forgejo_url, label) already exists (it feeds the orphaned-run
cancellation path in #78) and returns exactly this. Preserve the safety contract: if that read
returns None, fall back to the current conservative formula rather than guessing.

Preferred alternative, if we want to retire the whole class of bug: use the per-attempt
handle that _list_queued_jobs already parses — boot one VM per queued handle and start the
runner with one-job --handle <handle> instead of --wait. That makes the job↔VM mapping exact,
removes the surplus-VM waste, and makes the capacity question trivial. Noted as a TODO in
_list_queued_jobs and in runner-userdata.yaml.tmpl already.

Acceptance

  • Unit test in files/test_controller_logic.py: 1 queued job + 2 busy VMs + max_total=4
    → create=1.
  • Unit test: the running-jobs read failing falls back to the current formula (no over-boot).
  • Unit test: max_total is still a hard ceiling.
  • After apply, p90 pickup latency drops from ~137 s to ~35 s and no active=N desired=M
    cycle with M < N and M ≥ 1 persists for more than one cycle.

Part of gitborg/gitborg-docs#68.

## Problem `reconcile()` in `ansible/roles/runner-controller/files/controller.py` computes: ```python desired = min(max_total, queued_jobs + min_idle) to_create = max(0, desired - active_count) ``` `active_count` includes every ephemeral VM that is not being reaped — **including VMs that have already claimed a job and are busy running it**. With `min_idle: 0`, that means a newly queued job creates no capacity at all while any other VM is busy: it waits for that VM to finish, power off, be seen SHUTOFF and be reaped before a VM is booted for it. ## Evidence From the controller journal (one CI job, cap 4, two of four slots free the whole time): ``` 11:12:42Z reconcile: active=0 desired=2 (queued=2 + min_idle=0, cap=4) → create=2 11:13:52Z reconcile: active=2 desired=1 (queued=1 + min_idle=0, cap=4) → create=0 ← job waiting … six identical cycles, 61 s … 11:14:52Z reap: … status=SHUTOFF (x2) 11:14:53Z reconcile: active=0 desired=1 (queued=1 + min_idle=0, cap=4) → create=1 11:14:55Z create: booted … ``` The job's total wait was 93 s, of which **61 s was the controller declining to boot into free capacity**. The cold start itself was a normal 24 s. ## Impact (measured over 7 days, 345 workflow runs) - Pickup latency is bimodal: median **33 s**, but **p90 137 s** and 20 % of runs (70/345) take over 60 s with a mean of **187 s**. - 301 controller cycles carried the unambiguous starvation signature (`active > desired ≥ 1`, `create=0`) — a lower bound, since the `active == desired` case is indistinguishable in the log. - Characteristic fingerprint: two runs triggered in the same second, one starting at ~33 s and the other at ~155 s (≈ one job duration + reap + cold start). `RunnerQueueStalled` does not catch this — its `for:` window is minutes and these stalls are 60–200 s. ## Proposed fix Count only *idle* capacity. Jobs in flight = waiting + running: ```python desired = min(max_total, queued_jobs + running_jobs + min_idle) to_create = max(0, desired - active_count) ``` `_list_running_jobs(session, forgejo_url, label)` already exists (it feeds the orphaned-run cancellation path in #78) and returns exactly this. Preserve the safety contract: if that read returns `None`, fall back to the current conservative formula rather than guessing. **Preferred alternative, if we want to retire the whole class of bug:** use the per-attempt `handle` that `_list_queued_jobs` already parses — boot one VM per queued handle and start the runner with `one-job --handle <handle>` instead of `--wait`. That makes the job↔VM mapping exact, removes the surplus-VM waste, and makes the capacity question trivial. Noted as a TODO in `_list_queued_jobs` and in `runner-userdata.yaml.tmpl` already. ## Acceptance - [ ] Unit test in `files/test_controller_logic.py`: 1 queued job + 2 busy VMs + `max_total=4` → `create=1`. - [ ] Unit test: the running-jobs read failing falls back to the current formula (no over-boot). - [ ] Unit test: `max_total` is still a hard ceiling. - [ ] After apply, p90 pickup latency drops from ~137 s to ~35 s and no `active=N desired=M` cycle with `M < N` and `M ≥ 1` persists for more than one cycle. Part of gitborg/gitborg-docs#68.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#327
Ingen beskrivning angiven.