fix(runner-controller): alert on volumes past the sweep grace, not on ones still inside it #371

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/runner-orphan-volume-alert-grace in i main 2026-08-04 14:36:00 +00:00
Ägare

RunnerOrphanVolumesUnswept fired on healthy behaviour. Fixes the false-positive half of #320, and
records what the investigation turned up about the leak itself.

Applied and verified on both hosts: changed=4 each, changed=0 on the follow-up --check, and the
alert has cleared — only Watchdog (severity none) is firing now.

The sweep was never lagging

It waits out orphan_volume_max_age_seconds (1h) before deleting an orphan boot volume, so it can
never catch one mid-attach to a booting runner. Every deletion on 2026-08-04 landed at age≈3600s — it
was working perfectly, on every volume:

11:02:23 sweep: deleted orphan boot volume 694bb730… (20 GiB, age=3603s)
11:05:13 sweep: deleted orphan boot volume e1f885ef… (20 GiB, age=3603s)
11:14:23 sweep: deleted 2 orphan boot volume(s)      (age=3603s, 3601s)
13:16:11 sweep: deleted orphan boot volume 6bac70d7… (20 GiB, age=3605s)

But the alert read class="sweepable" > 0 for 2h, and sweepable counted volumes still serving out
that grace
. On a busy CI day the grace windows overlap continuously, so the expression never went
false long enough to reset the for — an alert named Unswept was really reading not yet due for
sweeping
. Three PRs plus pushes and Renovate kept it firing all afternoon.

This is the mirror image of the error #320 fixed. That change made the detector able to see a slow
leak the quota gauge could not; this one stops it seeing normal traffic as a leak.

Fix

Split sweepable by age, and point the alert at the new class:

  • sweepable — unnamed, no snapshot, inside the grace. Normal CI traffic; not alerted.
  • sweepable_overdue — past the grace and still present after the sweep ran. The real
    backlog: the sweep is failing, refusing, or something is producing volumes it cannot reclaim.

Threshold stays at 0. Raising it instead would hide precisely the ~1.8 volumes/week accumulation
#320 was filed about, and lengthening for: does nothing because continuous CI keeps the old
expression true indefinitely.

deleted_ids excludes volumes the same cycle just deleted — the survey is computed from the volume
list read at the TOP of the sweep, so without it every successful sweep would report its own
just-deleted volume as overdue for one cycle, blipping the metric on exactly the event that proves the
sweep works.

Age never overrides the other classes: snapshot-pinned is unsweepable by design however old, and a
named volume is deliberate retention however old.

What the investigation also established

  • The 160 GiB is reclaimed. Pruning the bake snapshots earlier that day unpinned the eight
    stranded volumes and the sweep collected them. snapshot_pinned is down to the 2 retained
    generations, named is 0, Cinder usage 440 GB against the ~640–720 GB floor #320 described.
  • #320 step 1, "identify the producer", is answered — and its hypothesis was wrong. The issue
    reasoned the orphans probably were not runner volumes because controller.py boots 40 GB while
    every orphan was 20 GB, and pointed at the weekly backup-drill VM. The sweep logs show one 20 GiB
    orphan created and reclaimed per CI run, continuously, by the runner path. So
    runner_controller_os_boot_volume_size: 40 is not what those VMs get — worth a separate look, since
    a declared-vs-actual mismatch is what sent this investigation down the wrong path in the first place.

Verification

69/69 tests pass, six new: the grace split in both directions, age-vs-snapshot and age-vs-named
precedence, the same-cycle exclusion, and the overdue GiB. On the host after applying:

gitborg_runner_controller_os_volumes_available{class="sweepable"} 0
gitborg_runner_controller_os_volumes_available{class="sweepable_overdue"} 0
gitborg_runner_controller_os_volumes_available{class="snapshot_pinned"} 2
gitborg_runner_controller_os_volumes_available{class="named"} 0

The role hashes Containerfile + controller.py into the image tag, so the apply rebuilt the image
on-host and restarted the container — no CI dependency. Note the restart resets _sweep_stuck (in
memory by design), so the two snapshot-pinned volumes are re-checked and re-skipped once afterwards;
that is expected, not a regression.

The dashboard panel gains the overdue series, and its description no longer tells the reader that
sweepable should be 0.

`RunnerOrphanVolumesUnswept` fired on healthy behaviour. Fixes the false-positive half of #320, and records what the investigation turned up about the leak itself. Applied and verified on both hosts: `changed=4` each, `changed=0` on the follow-up `--check`, and the alert has cleared — only `Watchdog` (severity `none`) is firing now. ## The sweep was never lagging It waits out `orphan_volume_max_age_seconds` (1h) before deleting an orphan boot volume, so it can never catch one mid-attach to a booting runner. Every deletion on 2026-08-04 landed at age≈3600s — it was working perfectly, on every volume: ``` 11:02:23 sweep: deleted orphan boot volume 694bb730… (20 GiB, age=3603s) 11:05:13 sweep: deleted orphan boot volume e1f885ef… (20 GiB, age=3603s) 11:14:23 sweep: deleted 2 orphan boot volume(s) (age=3603s, 3601s) 13:16:11 sweep: deleted orphan boot volume 6bac70d7… (20 GiB, age=3605s) ``` But the alert read `class="sweepable" > 0` for 2h, and `sweepable` counted volumes **still serving out that grace**. On a busy CI day the grace windows overlap continuously, so the expression never went false long enough to reset the `for` — an alert named *Unswept* was really reading *not yet due for sweeping*. Three PRs plus pushes and Renovate kept it firing all afternoon. This is the mirror image of the error #320 fixed. That change made the detector able to see a slow leak the quota gauge could not; this one stops it seeing normal traffic as a leak. ## Fix Split `sweepable` by age, and point the alert at the new class: - **`sweepable`** — unnamed, no snapshot, inside the grace. Normal CI traffic; not alerted. - **`sweepable_overdue`** — past the grace **and still present after the sweep ran**. The real backlog: the sweep is failing, refusing, or something is producing volumes it cannot reclaim. Threshold stays at **0**. Raising it instead would hide precisely the ~1.8 volumes/week accumulation #320 was filed about, and lengthening `for:` does nothing because continuous CI keeps the old expression true indefinitely. `deleted_ids` excludes volumes the same cycle just deleted — the survey is computed from the volume list read at the TOP of the sweep, so without it every successful sweep would report its own just-deleted volume as overdue for one cycle, blipping the metric on exactly the event that proves the sweep works. Age never overrides the other classes: snapshot-pinned is unsweepable by design however old, and a named volume is deliberate retention however old. ## What the investigation also established - **The 160 GiB is reclaimed.** Pruning the bake snapshots earlier that day *unpinned* the eight stranded volumes and the sweep collected them. `snapshot_pinned` is down to the 2 retained generations, `named` is 0, Cinder usage 440 GB against the ~640–720 GB floor #320 described. - **#320 step 1, "identify the producer", is answered — and its hypothesis was wrong.** The issue reasoned the orphans probably were not runner volumes because `controller.py` boots 40 GB while every orphan was 20 GB, and pointed at the weekly backup-drill VM. The sweep logs show one **20 GiB** orphan created and reclaimed per CI run, continuously, by the runner path. So `runner_controller_os_boot_volume_size: 40` is not what those VMs get — worth a separate look, since a declared-vs-actual mismatch is what sent this investigation down the wrong path in the first place. ## Verification 69/69 tests pass, six new: the grace split in both directions, age-vs-snapshot and age-vs-named precedence, the same-cycle exclusion, and the overdue GiB. On the host after applying: ``` gitborg_runner_controller_os_volumes_available{class="sweepable"} 0 gitborg_runner_controller_os_volumes_available{class="sweepable_overdue"} 0 gitborg_runner_controller_os_volumes_available{class="snapshot_pinned"} 2 gitborg_runner_controller_os_volumes_available{class="named"} 0 ``` The role hashes `Containerfile + controller.py` into the image tag, so the apply rebuilt the image on-host and restarted the container — no CI dependency. Note the restart resets `_sweep_stuck` (in memory by design), so the two snapshot-pinned volumes are re-checked and re-skipped once afterwards; that is expected, not a regression. The dashboard panel gains the overdue series, and its description no longer tells the reader that `sweepable` should be 0.
supernaut lade till 1 incheckning 2026-08-04 14:30:08 +00:00
`RunnerOrphanVolumesUnswept` fired on healthy behaviour. Closes the false-positive half of #320.

### What was actually happening

The sweep waits out `orphan_volume_max_age_seconds` (1h) before deleting an orphan boot volume, so it
can never catch one mid-attach to a booting runner. Measured on 2026-08-04, every deletion landed at
age≈3600s — the sweep was working perfectly, on every volume:

    11:02:23 sweep: deleted orphan boot volume … (20 GiB, age=3603s)
    11:14:23 sweep: deleted 2 orphan boot volume(s)   (age=3603s, 3601s)
    13:16:11 sweep: deleted orphan boot volume … (20 GiB, age=3605s)

But the alert read `gitborg_runner_controller_os_volumes_available{class="sweepable"} > 0` for 2h, and
`sweepable` counted volumes still legitimately serving out that grace. On a busy CI day the grace
windows overlap continuously, so the expression never goes false long enough to reset the `for` — and
an alert named "Unswept" was really reading "not yet due for sweeping". Three PRs plus pushes and
Renovate was enough to keep it firing all afternoon.

This is the opposite error to the one #320 fixed. That change made the detector able to see a slow
leak the quota gauge could not; this one stops it seeing normal traffic as a leak.

### The fix

Split `sweepable` by age into `sweepable` (inside the grace — normal) and `sweepable_overdue` (past
the grace AND still present after the sweep ran — the real backlog). The alert now reads the latter,
with the threshold left at **0**, so a single genuinely-stuck volume is still caught: raising the
threshold instead would hide exactly the 1–2 volumes/week accumulation #320 was filed about, and
lengthening `for:` does nothing because continuous CI keeps the old expression true indefinitely.

`deleted_ids` excludes volumes the same cycle just deleted. The survey is computed from the volume
list read at the TOP of the sweep, so without it every successful sweep would report its own
just-deleted volume as overdue for one cycle — the metric would blip on precisely the event that
proves the sweep works.

Age never overrides the other classes: a snapshot-pinned volume is unsweepable by design however old,
and a named volume is deliberate retention however old. Both are covered by new tests.

### While confirming this, #320's open question resolved itself

#320 step 1 was "identify the producer", and it reasoned the orphans probably were NOT ephemeral
runner volumes because `controller.py` boots 40 GB while every orphan was 20 GB. The sweep logs now
show the controller deleting **20 GiB** orphan boot volumes continuously, one per CI run — so the
producer IS the runner path, and the 40 GB default is not what these VMs get. Recorded on the issue.

The leak itself is also gone: `snapshot_pinned` is down to the 2 retained bake generations after the
2026-08-04 snapshot prune (which unpinned the eight stranded volumes so the sweep could take them),
`named`=0, and Cinder usage is 440 GB against the ~640–720 GB floor #320 described.

Tests: 69/69, six of them new (grace split both ways, age-vs-snapshot precedence, age-vs-named
precedence, same-cycle exclusion, overdue GiB). The dashboard panel gains the overdue series and its
description no longer tells the reader that `sweepable` should be 0.
supernaut sammanfogade incheckning 6ab4ff4406 till main 2026-08-04 14:36:00 +00:00
supernaut tog bort grenen fix/runner-orphan-volume-alert-grace 2026-08-04 14:36:00 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!371
Ingen beskrivning angiven.