CI runner pickup latency: queued jobs wait behind busy capacity #68
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-docs#68
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Time from triggering a workflow to the job actually starting is bimodal. Measured over 7 days
and 345 workflow runs, correlating the Forgejo API against the controller journal:
The tail is a defect, not physics.
reconcile()computesdesired = min(max_total, queued + min_idle)and subtractsactive_count, which counts VMs that have already claimed a job and arebusy. So a queued job gets no VM while any other VM is running one, even with most of the pool
free. One run waited 93 s, of which 61 s was six consecutive cycles declining to boot into free
capacity. 301 cycles carried that signature over the week — a lower bound. The existing
RunnerQueueStalledrule cannot see it: its window is minutes and these stalls are 60–200 s.The 33 s median decomposes as ~8 s waiting for the next queue poll plus a remarkably consistent
~24 s cold start (p25 22 s, median 24 s, p90 26 s, max 27 s). None of the cold start is
instrumented — the create call does not wait, no ACTIVE transition is ever logged, and runner VM
journals are not drained. Every number above had to be reconstructed by hand.
Scope
hand-correlation
boot-volume size
Not doing
9 s, the fastest pickups in the whole week — but the capacity fix alone is expected to take p90
from ~137 s to ~35 s, which is most of the available win. Revisit only if the residual cold start
proves to matter.
substantially larger change; the git host has no job-queued webhook, only completion, so a poll
would remain as backstop regardless.
Done when
p90 pickup latency is close to the median, no cycle declines to boot into free capacity, and the
per-stage timings are on a dashboard.
Status — 2026-08-02
Both fixes are merged, applied and — unusually for a latency claim — measured in production
rather than projected.
A burst of real CI (four merges, each firing
pull_requestpluspush) produced exactly thecontention the fix targets, and the new instrumentation caught it:
The capacity formula is visible doing its job under load —
active=2 desired=4 (queued=2 + running=2, cap=4) → create=2— counting busy VMs as not spare, thenrespecting the cap. #327's acceptance criterion is met with evidence.
Outstanding: #329, the cheap head-of-distribution win (poll interval, concurrency ceiling,
boot-volume size). Worth revisiting now that the tail is gone and the remaining ~30 s is almost all
cold start.
supernaut refererade till detta ärende från bitborg/bitborg-infra2026-08-02 12:30:46 +00:00
Shipping. The done-when is met and measured in production (status above): pickup p90 39.9 s against p50 30.1 s, down from 137 s, with no starvation cycles, and per-stage timings instrumented.
The remaining task, bitborg/bitborg-infra#329 (poll interval and concurrency ceiling), is an optional head-of-distribution win. It stays in the infra backlog on its own.