perf(runner-controller): shorten the queue poll interval and raise the concurrency ceiling #329
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#329
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Problem
Two config values add avoidable latency:
runner_controller_poll_interval_seconds: 10— a uniformly-arriving job waits a mean of 5 s justfor the next poll. Measured contribution of this stage (run created → VM boot call): median 8 s,
p90 12 s, max 16 s.
runner_controller_max_total: 4— the cap was saturated for 690 s over the last 7 days (0.1 % ofthe week). Small, but it is a hard stop on burst pickup.
Proposal
runner_controller_poll_interval_seconds: 10 → 3. The controller reaches Forgejo over theinternal podman network, so it never crosses the Caddy edge and is not subject to the
forgejo_apirate-limit zone (which exempts private ranges anyway — see #302). The reconcileloop already completes well inside 1 s, so cycles cannot overlap.
servers(),volumes(),snapshots(),get_volume_limits()run every cycle). Consider putting the volume sweep and quota samplingon a slower sub-cadence (~60 s) and letting only the queue read run at 3 s.
runner_controller_max_total: 4 → 6, only after confirming Cinder volume-quota headroom(6 concurrent VMs × the boot-volume size — this is the quota that caused the outage in #133).
runner_controller_os_boot_volume_size: 40 → 20— verify the bakedgitborg-runnerimage fits in 20 GiB first. Smaller volumes create faster, halve the quota exposure from #133,
and cost less.
Notes
Do not expect this to fix the long tail — that is a separate defect (see the capacity-formula
issue). This is the cheap head-of-distribution win: median ~33 s → ~28 s.
Acceptance
--check→ confirm → apply → verify gates.gitborg_runner_controller_last_loop_okstays 1 and loop duration stays well under the newinterval for 24 h after the change.
Part of gitborg/gitborg-docs#68.
Epic: bitborg-docs#1
supernaut refererade till detta ärende från bitborg/bitborg-docs2026-08-02 12:32:47 +00:00
Fresh evidence for item 2 (
runner_controller_max_total4 → 6), fromRunnerPickupSlowfiring 2026-10-01 20:10 UTC:active=4 desired=4 (cap=4) → create=0. Replacements booted in the same cycle as each reap.active<4,queued≥1,create=0hadqueued ≤ active − running.So the alert is real queueing at the cap, not a bug. It self-resolves when the samples leave the 6 h window. Raising the cap trades VM-hours for burst latency; not a launch item.
More evidence for item 2, 2026-10-02: 36 jobs queued at 17:20 UTC against
max_total: 4(a manual Renovate run rebasing its PRs, plus CI onmainin seven repos at once). The queue drained by 18:30 UTC; the worst pickup was 2020 s. Pickup was back to about 30 s from 19:00 UTC. No controller fault: every cycle loggedok=1.