fix(runner-controller): stop queued jobs waiting behind busy capacity, and measure pickup latency #345
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!345
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/runner-capacity-and-latency"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Closes #327, closes #328.
#327 — queued jobs were waiting behind busy capacity.
reconcile()subtractedactive_countfrom the desired pool size, andactive_countcounts VMsthat have already claimed a job and are busy running it. So a queued job created no capacity while
any other VM was busy: it waited for that VM to finish, power off, be seen SHUTOFF and be reaped
before one was booted for it.
Measured over 7 days and 345 runs: median pickup 33 s, but p90 137 s, with 20 % of runs
averaging 187 s. 301 cycles carried the starvation signature. One run waited 93 s of which 61 s
was the controller declining to boot into free capacity with half the pool idle.
One deliberate deviation from the formula in the issue. The issue proposes
min(cap, queued + running + min_idle) - active; this implements the stated intent — count onlyidle capacity:
Identical in every ordinary state. They diverge when
running > active— a job leftrunningafterits VM was reaped. The additive form boots a VM for a job no runner can ever claim, which then idles
until the two-hour backstop. There is a test locking that in. A failed running-jobs read reproduces
the old conservative formula exactly, and
max_totalremains a hard ceiling in every branch.The per-attempt
handlealternative was assessed and declined for now: it changes the cloud-initcontract, depends on semantics unverified against this instance, and is incompatible with the warm
buffer, since nothing can pre-boot against a handle that does not exist yet. Recorded as a TODO; it
plausibly belongs after the controller extraction.
The running-jobs snapshot is now read once and shared with orphan-cancel, so no extra API call.
#328 — nothing measured latency.
Every number above had to be reconstructed by hand, which is why a stall affecting a fifth of all
runs went unnoticed for weeks despite a dedicated dashboard and two alert rules. Adds boot, ready and
pickup durations, all derived from reads the loop already makes — no new API calls per cycle —
exported through the existing atomic textfile writer, plus a dashboard row and a companion alert for
sub-minute systematic stalls, which a duration-based queue alert is structurally blind to.
"Ready" keys on the runner's status leaving
offlinerather than on registration membership: theregistration is minted before the VM boots, so membership would time nothing.
Verification
47 checks pass, up from 17, written before the implementation. The three named acceptance criteria
pass explicitly.
reconcile()was also driven end-to-end against fakes and logs the acceptance caseverbatim.
ruffclean,ansible-lint0 failures across 191 files, syntax-check clean, textfilesmoke test 12 ok.
Caveats, documented in code and the runbook
Sampling is once per cycle, so every figure over-estimates by up to one interval; and a job cancelled
while queued records as a short pickup. Correcting the latter costs a second read per cycle for no
operational gain. #327's acceptance criterion that p90 drops to ~35 s is only verifiable after an
apply.
b155bce53ba8292cda9da8292cda9d256b381db4