feat(monitoring): critical paging tier for sustained CI queue stall (#196 follow-up) #206
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!206
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "feat/196-followup-runnerqueue-critical"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Follow-up to #196 (noted in its close): review RunnerQueueStalled severity/lag.
Finding (data-driven)
Checked
gitborg_runner_controller_queued_jobsacross the 2026-07-21 incident. The metric was clean, not flapping — mostly 0 with isolated single-sample blips (healthy jobs draining in one ~1-2 min cycle), then a steady 1 from 21:20→22:04 (run 133).RunnerQueueStalled(for: 15m) fired ~15 min into that sustained stall (~21:35) — working as designed. My earlier '~3h lag' was wrong; the boot-loop churned idle VMs but no job was sustained-queued until 21:20.The real gap: severity
It fired warning, which doesn't page — so the stall was visible but the operator found the stuck run manually. No flap-robustness needed.
Change
Add RunnerQueueStalledCritical —
queued_jobs > 0for: 30m, severity critical → pages. The 15-min warning tier is unchanged (early, non-paging heads-up); a queue still stalled at 30 min is an unattended CI outage (blocks deploys) and pages. Applied to the monitoring host + verified in the deployed rules; vmalert reloaded cleanly.