perf(runner-controller): shorten the queue poll interval and raise the concurrency ceiling #329

Öppen
öppnade 2026-08-02 12:30:46 +00:00 av supernaut · 2 kommentarer
Ägare

Problem

Two config values add avoidable latency:

  • runner_controller_poll_interval_seconds: 10 — a uniformly-arriving job waits a mean of 5 s just
    for the next poll. Measured contribution of this stage (run created → VM boot call): median 8 s,
    p90 12 s, max 16 s.
  • runner_controller_max_total: 4 — the cap was saturated for 690 s over the last 7 days (0.1 % of
    the week). Small, but it is a hard stop on burst pickup.

Proposal

  1. runner_controller_poll_interval_seconds: 10 → 3. The controller reaches Forgejo over the
    internal podman network, so it never crosses the Caddy edge and is not subject to the
    forgejo_api rate-limit zone (which exempts private ranges anyway — see #302). The reconcile
    loop already completes well inside 1 s, so cycles cannot overlap.
    • Watch item: OpenStack API volume triples (servers(), volumes(), snapshots(),
      get_volume_limits() run every cycle). Consider putting the volume sweep and quota sampling
      on a slower sub-cadence (~60 s) and letting only the queue read run at 3 s.
  2. runner_controller_max_total: 4 → 6, only after confirming Cinder volume-quota headroom
    (6 concurrent VMs × the boot-volume size — this is the quota that caused the outage in #133).
  3. Evaluate runner_controller_os_boot_volume_size: 40 → 20 — verify the baked gitborg-runner
    image fits in 20 GiB first. Smaller volumes create faster, halve the quota exposure from #133,
    and cost less.

Notes

Do not expect this to fix the long tail — that is a separate defect (see the capacity-formula
issue). This is the cheap head-of-distribution win: median ~33 s → ~28 s.

Acceptance

  • Applied via the standard validate → --check → confirm → apply → verify gates.
  • gitborg_runner_controller_last_loop_ok stays 1 and loop duration stays well under the new
    interval for 24 h after the change.
  • No increase in OpenStack API errors in the controller journal.

Part of gitborg/gitborg-docs#68.

Epic: bitborg-docs#1

## Problem Two config values add avoidable latency: - `runner_controller_poll_interval_seconds: 10` — a uniformly-arriving job waits a mean of 5 s just for the next poll. Measured contribution of this stage (run created → VM boot call): median 8 s, p90 12 s, max 16 s. - `runner_controller_max_total: 4` — the cap was saturated for 690 s over the last 7 days (0.1 % of the week). Small, but it is a hard stop on burst pickup. ## Proposal 1. `runner_controller_poll_interval_seconds: 10 → 3`. The controller reaches Forgejo over the **internal** podman network, so it never crosses the Caddy edge and is not subject to the `forgejo_api` rate-limit zone (which exempts private ranges anyway — see #302). The reconcile loop already completes well inside 1 s, so cycles cannot overlap. - Watch item: OpenStack API volume triples (`servers()`, `volumes()`, `snapshots()`, `get_volume_limits()` run every cycle). Consider putting the volume sweep and quota sampling on a slower sub-cadence (~60 s) and letting only the queue read run at 3 s. 2. `runner_controller_max_total: 4 → 6`, **only after** confirming Cinder volume-quota headroom (6 concurrent VMs × the boot-volume size — this is the quota that caused the outage in #133). 3. Evaluate `runner_controller_os_boot_volume_size: 40 → 20` — verify the baked `gitborg-runner` image fits in 20 GiB first. Smaller volumes create faster, halve the quota exposure from #133, and cost less. ## Notes Do **not** expect this to fix the long tail — that is a separate defect (see the capacity-formula issue). This is the cheap head-of-distribution win: median ~33 s → ~28 s. ## Acceptance - [ ] Applied via the standard validate → `--check` → confirm → apply → verify gates. - [ ] `gitborg_runner_controller_last_loop_ok` stays 1 and loop duration stays well under the new interval for 24 h after the change. - [ ] No increase in OpenStack API errors in the controller journal. Part of gitborg/gitborg-docs#68. Epic: bitborg-docs#1
Upphovsperson
Ägare

Fresh evidence for item 2 (runner_controller_max_total 4 → 6), from RunnerPickupSlow firing 2026-10-01 20:10 UTC:

  • About 11 jobs arrived 19:48 to 19:54 UTC; demand peaked at 7 (3 queued + 4 running).
  • Every cycle with waiting jobs read active=4 desired=4 (cap=4) → create=0. Replacements booted in the same cycle as each reap.
  • Four pickups of 110 to 140 s, all in that burst. p90 went 30 s → 130 s. p50 stayed 30 s.
  • Boot p90 20 s, ready p90 30 s, no OpenStack errors, no controller restarts. Cinder 520 of 5000 GB.
  • No #327-class starvation: every cycle with active<4, queued≥1, create=0 had queued ≤ active − running.

So the alert is real queueing at the cap, not a bug. It self-resolves when the samples leave the 6 h window. Raising the cap trades VM-hours for burst latency; not a launch item.

Fresh evidence for item 2 (`runner_controller_max_total` 4 → 6), from `RunnerPickupSlow` firing 2026-10-01 20:10 UTC: - About 11 jobs arrived 19:48 to 19:54 UTC; demand peaked at 7 (3 queued + 4 running). - Every cycle with waiting jobs read `active=4 desired=4 (cap=4) → create=0`. Replacements booted in the same cycle as each reap. - Four pickups of 110 to 140 s, all in that burst. p90 went 30 s → 130 s. p50 stayed 30 s. - Boot p90 20 s, ready p90 30 s, no OpenStack errors, no controller restarts. Cinder 520 of 5000 GB. - No #327-class starvation: every cycle with `active<4`, `queued≥1`, `create=0` had `queued ≤ active − running`. So the alert is real queueing at the cap, not a bug. It self-resolves when the samples leave the 6 h window. Raising the cap trades VM-hours for burst latency; not a launch item.
Upphovsperson
Ägare

More evidence for item 2, 2026-10-02: 36 jobs queued at 17:20 UTC against max_total: 4 (a manual Renovate run rebasing its PRs, plus CI on main in seven repos at once). The queue drained by 18:30 UTC; the worst pickup was 2020 s. Pickup was back to about 30 s from 19:00 UTC. No controller fault: every cycle logged ok=1.

More evidence for item 2, 2026-10-02: 36 jobs queued at 17:20 UTC against `max_total: 4` (a manual Renovate run rebasing its PRs, plus CI on `main` in seven repos at once). The queue drained by 18:30 UTC; the worst pickup was 2020 s. Pickup was back to about 30 s from 19:00 UTC. No controller fault: every cycle logged `ok=1`.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#329
Ingen beskrivning angiven.