feat(monitoring): critical paging tier for sustained CI queue stall (#196 follow-up) #206

Sammanfogat
supernaut sammanfogade 1 incheckning från feat/196-followup-runnerqueue-critical in i main 2026-07-22 11:08:35 +00:00
Ägare

Follow-up to #196 (noted in its close): review RunnerQueueStalled severity/lag.

Finding (data-driven)

Checked gitborg_runner_controller_queued_jobs across the 2026-07-21 incident. The metric was clean, not flapping — mostly 0 with isolated single-sample blips (healthy jobs draining in one ~1-2 min cycle), then a steady 1 from 21:20→22:04 (run 133). RunnerQueueStalled (for: 15m) fired ~15 min into that sustained stall (~21:35) — working as designed. My earlier '~3h lag' was wrong; the boot-loop churned idle VMs but no job was sustained-queued until 21:20.

The real gap: severity

It fired warning, which doesn't page — so the stall was visible but the operator found the stuck run manually. No flap-robustness needed.

Change

Add RunnerQueueStalledCritical — queued_jobs > 0 for: 30m, severity critical → pages. The 15-min warning tier is unchanged (early, non-paging heads-up); a queue still stalled at 30 min is an unattended CI outage (blocks deploys) and pages. Applied to the monitoring host + verified in the deployed rules; vmalert reloaded cleanly.

Follow-up to #196 (noted in its close): review RunnerQueueStalled severity/lag. ## Finding (data-driven) Checked `gitborg_runner_controller_queued_jobs` across the 2026-07-21 incident. The metric was **clean, not flapping** — mostly 0 with isolated single-sample blips (healthy jobs draining in one ~1-2 min cycle), then a **steady 1 from 21:20→22:04** (run 133). `RunnerQueueStalled` (`for: 15m`) fired ~15 min into that sustained stall (~21:35) — **working as designed**. My earlier '~3h lag' was wrong; the boot-loop churned idle VMs but no job was *sustained*-queued until 21:20. ## The real gap: severity It fired **warning**, which doesn't page — so the stall was visible but the operator found the stuck run manually. No flap-robustness needed. ## Change Add **RunnerQueueStalledCritical** — `queued_jobs > 0` `for: 30m`, **severity critical** → pages. The 15-min warning tier is unchanged (early, non-paging heads-up); a queue still stalled at 30 min is an unattended CI outage (blocks deploys) and pages. Applied to the monitoring host + verified in the deployed rules; vmalert reloaded cleanly.
supernaut lade till 1 incheckning 2026-07-22 07:37:51 +00:00
feat(monitoring): critical paging tier for a sustained CI queue stall (#196 follow-up)
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m20s
5cf67e915a
RunnerQueueStalled fired correctly ~15 min into the 2026-07-21 sustained stall
(metric was clean, not flapping — verified against history), but it is warning-only
so it never paged; the operator caught the stuck run manually. Add
RunnerQueueStalledCritical (queued_jobs > 0 for 30m, severity critical) so an
unattended stall that blocks deploys pages. Warning tier unchanged (early, visible).
supernaut sammanfogade incheckning 41defc16c0 till main 2026-07-22 11:08:35 +00:00
supernaut tog bort grenen feat/196-followup-runnerqueue-critical 2026-07-22 11:08:35 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!206
Ingen beskrivning angiven.