fix(runner-controller): busy VMs are counted as spare capacity, starving queued jobs #327
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#327
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Problem
reconcile()inansible/roles/runner-controller/files/controller.pycomputes:active_countincludes every ephemeral VM that is not being reaped — including VMs that havealready claimed a job and are busy running it. With
min_idle: 0, that means a newly queued jobcreates no capacity at all while any other VM is busy: it waits for that VM to finish, power off,
be seen SHUTOFF and be reaped before a VM is booted for it.
Evidence
From the controller journal (one CI job, cap 4, two of four slots free the whole time):
The job's total wait was 93 s, of which 61 s was the controller declining to boot into free
capacity. The cold start itself was a normal 24 s.
Impact (measured over 7 days, 345 workflow runs)
over 60 s with a mean of 187 s.
active > desired ≥ 1,create=0) — a lower bound, since theactive == desiredcase is indistinguishable in the log.other at ~155 s (≈ one job duration + reap + cold start).
RunnerQueueStalleddoes not catch this — itsfor:window is minutes and these stalls are60–200 s.
Proposed fix
Count only idle capacity. Jobs in flight = waiting + running:
_list_running_jobs(session, forgejo_url, label)already exists (it feeds the orphaned-runcancellation path in #78) and returns exactly this. Preserve the safety contract: if that read
returns
None, fall back to the current conservative formula rather than guessing.Preferred alternative, if we want to retire the whole class of bug: use the per-attempt
handlethat_list_queued_jobsalready parses — boot one VM per queued handle and start therunner with
one-job --handle <handle>instead of--wait. That makes the job↔VM mapping exact,removes the surplus-VM waste, and makes the capacity question trivial. Noted as a TODO in
_list_queued_jobsand inrunner-userdata.yaml.tmplalready.Acceptance
files/test_controller_logic.py: 1 queued job + 2 busy VMs +max_total=4→
create=1.max_totalis still a hard ceiling.active=N desired=Mcycle with
M < NandM ≥ 1persists for more than one cycle.Part of gitborg/gitborg-docs#68.
supernaut refererade till detta ärende från bitborg/bitborg-docs2026-08-02 12:32:47 +00:00