CI runner pickup latency: queued jobs wait behind busy capacity #68

Stängd
öppnade 2026-08-02 12:24:12 +00:00 av supernaut · 1 kommentar
Ägare

Time from triggering a workflow to the job actually starting is bimodal. Measured over 7 days
and 345 workflow runs, correlating the Forgejo API against the controller journal:

  • End-to-end: median 33 s, p75 42 s, p90 137 s, mean 64 s.
  • About 80 % of runs sit in a tight 24–45 s band; 20 % (70/345) take over 60 s, averaging 187 s.

The tail is a defect, not physics. reconcile() computes desired = min(max_total, queued + min_idle) and subtracts active_count, which counts VMs that have already claimed a job and are
busy
. So a queued job gets no VM while any other VM is running one, even with most of the pool
free. One run waited 93 s, of which 61 s was six consecutive cycles declining to boot into free
capacity. 301 cycles carried that signature over the week — a lower bound. The existing
RunnerQueueStalled rule cannot see it: its window is minutes and these stalls are 60–200 s.

The 33 s median decomposes as ~8 s waiting for the next queue poll plus a remarkably consistent
~24 s cold start (p25 22 s, median 24 s, p90 26 s, max 27 s). None of the cold start is
instrumented
— the create call does not wait, no ACTIVE transition is ever logged, and runner VM
journals are not drained. Every number above had to be reconstructed by hand.

Scope

Not doing

  • A warm standby pool. It works — the two pre-booted VMs in the sample were claimed in 3 s and
    9 s, the fastest pickups in the whole week — but the capacity fix alone is expected to take p90
    from ~137 s to ~35 s, which is most of the available win. Revisit only if the residual cold start
    proves to matter.
  • An event-driven trigger. The poll is already 10 s, so a webhook would save roughly 5 s for a
    substantially larger change; the git host has no job-queued webhook, only completion, so a poll
    would remain as backstop regardless.

Done when

p90 pickup latency is close to the median, no cycle declines to boot into free capacity, and the
per-stage timings are on a dashboard.

Status — 2026-08-02

Both fixes are merged, applied and — unusually for a latency claim — measured in production
rather than projected
.

A burst of real CI (four merges, each firing pull_request plus push) produced exactly the
contention the fix targets, and the new instrumentation caught it:

pickup  p90 39.9 s   p50 30.1 s   samples 25      (was p90 137 s)
boot    p90 20.4 s   ready p90 30.5 s
starvation-signature cycles: 0 of 267 examined    (was 301 in one week)

The capacity formula is visible doing its job under load —
active=2 desired=4 (queued=2 + running=2, cap=4) → create=2 — counting busy VMs as not spare, then
respecting the cap. #327's acceptance criterion is met with evidence.

Outstanding: #329, the cheap head-of-distribution win (poll interval, concurrency ceiling,
boot-volume size). Worth revisiting now that the tail is gone and the remaining ~30 s is almost all
cold start.

Time from triggering a workflow to the job actually starting is **bimodal**. Measured over 7 days and 345 workflow runs, correlating the Forgejo API against the controller journal: - End-to-end: median **33 s**, p75 42 s, **p90 137 s**, mean 64 s. - About 80 % of runs sit in a tight 24–45 s band; **20 % (70/345) take over 60 s, averaging 187 s**. The tail is a defect, not physics. `reconcile()` computes `desired = min(max_total, queued + min_idle)` and subtracts `active_count`, which counts VMs that have **already claimed a job and are busy**. So a queued job gets no VM while any other VM is running one, even with most of the pool free. One run waited 93 s, of which 61 s was six consecutive cycles declining to boot into free capacity. 301 cycles carried that signature over the week — a lower bound. The existing `RunnerQueueStalled` rule cannot see it: its window is minutes and these stalls are 60–200 s. The 33 s median decomposes as ~8 s waiting for the next queue poll plus a remarkably consistent ~24 s cold start (p25 22 s, median 24 s, p90 26 s, max 27 s). **None of the cold start is instrumented** — the create call does not wait, no ACTIVE transition is ever logged, and runner VM journals are not drained. Every number above had to be reconstructed by hand. ## Scope - [x] gitborg/gitborg-infra#327 — count only idle capacity, so a queued job boots into a free slot - [x] gitborg/gitborg-infra#328 — instrument the stages, so a latency regression is visible without hand-correlation - [ ] gitborg/gitborg-infra#329 (optional, continues in the infra backlog) — shorten the queue poll and re-examine the concurrency ceiling and boot-volume size ## Not doing - **A warm standby pool.** It works — the two pre-booted VMs in the sample were claimed in 3 s and 9 s, the fastest pickups in the whole week — but the capacity fix alone is expected to take p90 from ~137 s to ~35 s, which is most of the available win. Revisit only if the residual cold start proves to matter. - **An event-driven trigger.** The poll is already 10 s, so a webhook would save roughly 5 s for a substantially larger change; the git host has no job-queued webhook, only completion, so a poll would remain as backstop regardless. ## Done when p90 pickup latency is close to the median, no cycle declines to boot into free capacity, and the per-stage timings are on a dashboard. ## Status — 2026-08-02 Both fixes are merged, applied and — unusually for a latency claim — **measured in production rather than projected**. A burst of real CI (four merges, each firing `pull_request` plus `push`) produced exactly the contention the fix targets, and the new instrumentation caught it: ```text pickup p90 39.9 s p50 30.1 s samples 25 (was p90 137 s) boot p90 20.4 s ready p90 30.5 s starvation-signature cycles: 0 of 267 examined (was 301 in one week) ``` The capacity formula is visible doing its job under load — `active=2 desired=4 (queued=2 + running=2, cap=4) → create=2` — counting busy VMs as not spare, then respecting the cap. #327's acceptance criterion is met with evidence. Outstanding: #329, the cheap head-of-distribution win (poll interval, concurrency ceiling, boot-volume size). Worth revisiting now that the tail is gone and the remaining ~30 s is almost all cold start.
supernaut lade till detta till projektet Bitborg Roadmap 2026-08-02 12:27:12 +00:00
Upphovsperson
Ägare

Shipping. The done-when is met and measured in production (status above): pickup p90 39.9 s against p50 30.1 s, down from 137 s, with no starvation cycles, and per-stage timings instrumented.

The remaining task, bitborg/bitborg-infra#329 (poll interval and concurrency ceiling), is an optional head-of-distribution win. It stays in the infra backlog on its own.

Shipping. The done-when is met and measured in production (status above): pickup p90 39.9 s against p50 30.1 s, down from 137 s, with no starvation cycles, and per-stage timings instrumented. The remaining task, bitborg/bitborg-infra#329 (poll interval and concurrency ceiling), is an optional head-of-distribution win. It stays in the infra backlog on its own.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-docs#68
Ingen beskrivning angiven.