fix(runner-controller): alert on volumes past the sweep grace, not on ones still inside it #371
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!371
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/runner-orphan-volume-alert-grace"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
RunnerOrphanVolumesUnsweptfired on healthy behaviour. Fixes the false-positive half of #320, andrecords what the investigation turned up about the leak itself.
Applied and verified on both hosts:
changed=4each,changed=0on the follow-up--check, and thealert has cleared — only
Watchdog(severitynone) is firing now.The sweep was never lagging
It waits out
orphan_volume_max_age_seconds(1h) before deleting an orphan boot volume, so it cannever catch one mid-attach to a booting runner. Every deletion on 2026-08-04 landed at age≈3600s — it
was working perfectly, on every volume:
But the alert read
class="sweepable" > 0for 2h, andsweepablecounted volumes still serving outthat grace. On a busy CI day the grace windows overlap continuously, so the expression never went
false long enough to reset the
for— an alert named Unswept was really reading not yet due forsweeping. Three PRs plus pushes and Renovate kept it firing all afternoon.
This is the mirror image of the error #320 fixed. That change made the detector able to see a slow
leak the quota gauge could not; this one stops it seeing normal traffic as a leak.
Fix
Split
sweepableby age, and point the alert at the new class:sweepable— unnamed, no snapshot, inside the grace. Normal CI traffic; not alerted.sweepable_overdue— past the grace and still present after the sweep ran. The realbacklog: the sweep is failing, refusing, or something is producing volumes it cannot reclaim.
Threshold stays at 0. Raising it instead would hide precisely the ~1.8 volumes/week accumulation
#320 was filed about, and lengthening
for:does nothing because continuous CI keeps the oldexpression true indefinitely.
deleted_idsexcludes volumes the same cycle just deleted — the survey is computed from the volumelist read at the TOP of the sweep, so without it every successful sweep would report its own
just-deleted volume as overdue for one cycle, blipping the metric on exactly the event that proves the
sweep works.
Age never overrides the other classes: snapshot-pinned is unsweepable by design however old, and a
named volume is deliberate retention however old.
What the investigation also established
stranded volumes and the sweep collected them.
snapshot_pinnedis down to the 2 retainedgenerations,
namedis 0, Cinder usage 440 GB against the ~640–720 GB floor #320 described.reasoned the orphans probably were not runner volumes because
controller.pyboots 40 GB whileevery orphan was 20 GB, and pointed at the weekly backup-drill VM. The sweep logs show one 20 GiB
orphan created and reclaimed per CI run, continuously, by the runner path. So
runner_controller_os_boot_volume_size: 40is not what those VMs get — worth a separate look, sincea declared-vs-actual mismatch is what sent this investigation down the wrong path in the first place.
Verification
69/69 tests pass, six new: the grace split in both directions, age-vs-snapshot and age-vs-named
precedence, the same-cycle exclusion, and the overdue GiB. On the host after applying:
The role hashes
Containerfile + controller.pyinto the image tag, so the apply rebuilt the imageon-host and restarted the container — no CI dependency. Note the restart resets
_sweep_stuck(inmemory by design), so the two snapshot-pinned volumes are re-checked and re-skipped once afterwards;
that is expected, not a regression.
The dashboard panel gains the overdue series, and its description no longer tells the reader that
sweepableshould be 0.