8 orphaned 20 GB boot volumes (160 GB) accumulating since June — the leak metric cannot see it #320

Stängd
öppnade 2026-08-01 19:16:13 +00:00 av supernaut · 6 kommentarer
Ägare

8 unattached bootable volumes, 160 GB, accumulating since 2026-06-28. Found while tearing down a
throwaway VM on 2026-08-01. This answers the open question #305 recorded and left unchased, and it
contradicts the "volumes are not leaking" conclusion in #311 — for a reason that is worth
understanding rather than just correcting.

What is there

Every one is identical: 20 GB, bootable, sourced from Debian13, available (attached to
nothing), and unnamed.

Created (UTC) Size Image
2026-08-01 18:21:43 20 GB Debian13
2026-08-01 18:17:32 20 GB Debian13
2026-07-29 21:09:44 20 GB Debian13
2026-07-22 14:12:04 20 GB Debian13
2026-07-20 17:58:02 20 GB Debian13
2026-07-20 17:40:25 20 GB Debian13
2026-06-28 22:04:27 20 GB Debian13
2026-06-28 14:41:14 20 GB Debian13

Everything else in the project is legitimately attached: gitborg-prod-root (20), -data (60),
-backup (150), -lfs (60), gitborg-prod-monitoring-root (30) — 320 GB in use.

Why the existing metric cannot see it

#311 concluded volumes were not leaking because gitborg_runner_controller_os_volume_gb_used
oscillates between 640 and 1020 GB over 7 days and returns to a 640–720 baseline without climbing.
That reasoning is sound for a fast leak and blind to this one: 160 GB over five weeks is roughly
1.8 volumes/week, ~36 GB/week, which disappears inside a band that swings by ~380 GB during
normal CI churn. The metric is not wrong; it is not sensitive enough to separate slow accumulation
from churn. A leak detector whose noise band is ten times the weekly leak rate will always read
clean.

This is also most of the unexplained baseline #305 flagged: 320 GB attached + 160 GB orphaned =
480 GB against an observed ~640–720 GB floor. Not the whole gap, but the largest identified piece.

These are probably NOT ephemeral runner volumes

The obvious suspect is the runner controller, given the prior outage with 227 orphaned boot volumes.
The sizes argue against it: controller.py boots runners with
"volume_size": int(os_cfg.get("boot_volume_size", 40)) — 40 GB — and every orphan here is
20 GB. Unless that default is overridden to 20 somewhere, these come from a different producer.

Worth checking first, roughly in order of suspicion:

  • the weekly backup-drill VM (backup-drill boots a throwaway server reusing the controller's
    create_server shape), whose cadence is closest to the observed ~1–2/week;
  • any manual/one-off server create — note that a zero-disk flavour makes a plain create fail with
    Only volume-backed servers are allowed for flavors with zero disk, and it is worth confirming
    whether a rejected create can still leave a volume behind;
  • tofu replacements of the services or monitoring host.

The creation timestamps are the evidence to correlate against — several cluster in pairs hours
apart, which does not obviously match a weekly timer.

Disclosure

A ninth such volume existed and was created during the investigation itself, inside the window of a
throwaway VM booted for the #315 rollback rehearsal, matching that VM's image and size. It was
deleted immediately. That VM's declared boot volume cascaded correctly on
delete_on_termination=true, so how a second one appeared is unexplained — possibly a create that
Nova rejected after Cinder had already provisioned. Flagged because it may be the same mechanism
producing the eight, and because two of the eight (18:17, 18:21) predate any command run that day.

Suggested

  1. Identify the producer before deleting anything — the eight are currently the only evidence.
    Cinder retains no attachment history for a deleted server, so correlation has to come from the
    timestamps against CI runs, drill runs and applies.
  2. Then reclaim the 160 GB.
  3. Make the detector able to see this class of leak. An alert on
    count of available volumes (or their total size) is a direct signal and does not depend on
    separating a slow trend from CI churn — unattached-and-unnamed is an unambiguous state in a
    project where every legitimate volume is named and attached.

Refs #305, #311.

**8 unattached bootable volumes, 160 GB, accumulating since 2026-06-28.** Found while tearing down a throwaway VM on 2026-08-01. This answers the open question #305 recorded and left unchased, and it contradicts the "volumes are not leaking" conclusion in #311 — for a reason that is worth understanding rather than just correcting. ## What is there Every one is identical: **20 GB, bootable, sourced from `Debian13`, `available` (attached to nothing), and unnamed.** | Created (UTC) | Size | Image | | ------------------- | ----- | -------- | | 2026-08-01 18:21:43 | 20 GB | Debian13 | | 2026-08-01 18:17:32 | 20 GB | Debian13 | | 2026-07-29 21:09:44 | 20 GB | Debian13 | | 2026-07-22 14:12:04 | 20 GB | Debian13 | | 2026-07-20 17:58:02 | 20 GB | Debian13 | | 2026-07-20 17:40:25 | 20 GB | Debian13 | | 2026-06-28 22:04:27 | 20 GB | Debian13 | | 2026-06-28 14:41:14 | 20 GB | Debian13 | Everything else in the project is legitimately attached: `gitborg-prod-root` (20), `-data` (60), `-backup` (150), `-lfs` (60), `gitborg-prod-monitoring-root` (30) — 320 GB in use. ## Why the existing metric cannot see it #311 concluded volumes were not leaking because `gitborg_runner_controller_os_volume_gb_used` oscillates between 640 and 1020 GB over 7 days and returns to a 640–720 baseline without climbing. That reasoning is sound for a *fast* leak and blind to this one: 160 GB over five weeks is roughly **1.8 volumes/week, ~36 GB/week**, which disappears inside a band that swings by ~380 GB during normal CI churn. The metric is not wrong; it is not sensitive enough to separate slow accumulation from churn. A leak detector whose noise band is ten times the weekly leak rate will always read clean. This is also most of the unexplained baseline #305 flagged: 320 GB attached + 160 GB orphaned = 480 GB against an observed ~640–720 GB floor. Not the whole gap, but the largest identified piece. ## These are probably NOT ephemeral runner volumes The obvious suspect is the runner controller, given the prior outage with 227 orphaned boot volumes. The sizes argue against it: `controller.py` boots runners with `"volume_size": int(os_cfg.get("boot_volume_size", 40))` — **40 GB** — and every orphan here is **20 GB**. Unless that default is overridden to 20 somewhere, these come from a different producer. Worth checking first, roughly in order of suspicion: - the **weekly backup-drill VM** (`backup-drill` boots a throwaway server reusing the controller's `create_server` shape), whose cadence is closest to the observed ~1–2/week; - any manual/one-off `server create` — note that a zero-disk flavour makes a *plain* create fail with `Only volume-backed servers are allowed for flavors with zero disk`, and it is worth confirming whether a rejected create can still leave a volume behind; - `tofu` replacements of the services or monitoring host. The creation timestamps are the evidence to correlate against — several cluster in pairs hours apart, which does not obviously match a weekly timer. ## Disclosure A ninth such volume existed and was created during the investigation itself, inside the window of a throwaway VM booted for the #315 rollback rehearsal, matching that VM's image and size. It was deleted immediately. That VM's *declared* boot volume cascaded correctly on `delete_on_termination=true`, so how a second one appeared is unexplained — possibly a create that Nova rejected after Cinder had already provisioned. Flagged because it may be the same mechanism producing the eight, and because two of the eight (18:17, 18:21) predate any command run that day. ## Suggested 1. **Identify the producer** before deleting anything — the eight are currently the only evidence. Cinder retains no attachment history for a deleted server, so correlation has to come from the timestamps against CI runs, drill runs and applies. 2. **Then reclaim the 160 GB.** 3. **Make the detector able to see this class of leak.** An alert on *count of `available` volumes* (or their total size) is a direct signal and does not depend on separating a slow trend from CI churn — unattached-and-unnamed is an unambiguous state in a project where every legitimate volume is named and attached. Refs #305, #311.
Upphovsperson
Ägare

Producer identified. This answers the "identify the producer before deleting anything" step.

They are image-bake leftovers, not runner-job leaks. Baking the gitborg-runner Glance image
snapshots the staging VM's boot volume; the snapshot keeps a reference, so Cinder refuses to delete
the volume, and the controller's sweep then correctly skips it — _sweep_orphan_volumes excludes
any volume that has snapshots, precisely so it cannot destroy image or backup infrastructure.

The controller logs the refusal on every restart:

sweep: skipping volume e610f170-… — it has snapshot(s), so it is image/backup infra,
not an ephemeral leak (will not retry this session)

Each stranded volume pairs one-to-one with a bake snapshot, and the timestamps line up:

Volume Created Paired snapshot
f1ce224d 2026-07-29 21:09 snapshot for gitborg-runner-staging-20260729-230936
7b4588d3 2026-07-22 14:12 snapshot for gitborg-runner-staging-20260722-161157
e8dd5800 2026-07-20 17:58 snapshot for gitborg-runner-staging-20260720-195756
4d4a7cd2 2026-07-20 17:40 snapshot for gitborg-runner-staging-20260720-194019
40781d11 2026-06-28 22:04 snapshot for gitborg-runner
e610f170 2026-06-28 14:41 snapshot for gitborg-runner

This also resolves the 40 GB vs 20 GB puzzle raised above, which was good reasoning from a wrong
premise. The controller does request 40, but the gitborg-runner image is bdm_v2 and its
block_device_mapping declares "volume_size": 20 — the image's mapping wins. Every live
runner boot volume is 20 GB, so runner_controller_os_boot_volume_size: 40 is inert and
misleading. The Debian13 source image is consistent with this too: the staging VM boots Debian13,
is provisioned, and is then snapshotted to produce gitborg-runner.

The two 2026-08-01 volumes in the table above are gone — the sweep reclaimed them normally, which
confirms the sweep works and that only snapshot-referenced volumes persist.

One is load-bearing. The live Glance image boots from snapshot_id 4daabf1b-…, which is the
2026-07-29 row. That snapshot and volume f1ce224d must not be deleted — doing so leaves the
image unbootable and breaks all CI. Verify against the image's block_device_mapping immediately
before deleting rather than trusting this table. The other five pairs are safe.

Current consumption for scale: 500 GiB across 14 volumes, 280 GiB across 8 snapshots. Reclaiming
the five stale pairs returns roughly 200 GiB.

Suggestion 3 above still stands and is the right detector, with one refinement: a count of
available volumes would have caught this, but should be paired with the snapshot-refused count,
since these volumes are unsweepable by design rather than merely unswept.

Separately: predrill-backup-20260709 (100 GiB) and predrill-data-20260709 (60 GiB) have been
retained since 2026-07-09 — same shape of problem, another 160 GiB, out of scope here.

**Producer identified.** This answers the "identify the producer before deleting anything" step. They are **image-bake leftovers**, not runner-job leaks. Baking the `gitborg-runner` Glance image snapshots the staging VM's boot volume; the snapshot keeps a reference, so Cinder refuses to delete the volume, and the controller's sweep then correctly skips it — `_sweep_orphan_volumes` excludes any volume that has snapshots, precisely so it cannot destroy image or backup infrastructure. The controller logs the refusal on every restart: ```text sweep: skipping volume e610f170-… — it has snapshot(s), so it is image/backup infra, not an ephemeral leak (will not retry this session) ``` Each stranded volume pairs one-to-one with a bake snapshot, and the timestamps line up: | Volume | Created | Paired snapshot | | --- | --- | --- | | `f1ce224d` | 2026-07-29 21:09 | `snapshot for gitborg-runner-staging-20260729-230936` | | `7b4588d3` | 2026-07-22 14:12 | `snapshot for gitborg-runner-staging-20260722-161157` | | `e8dd5800` | 2026-07-20 17:58 | `snapshot for gitborg-runner-staging-20260720-195756` | | `4d4a7cd2` | 2026-07-20 17:40 | `snapshot for gitborg-runner-staging-20260720-194019` | | `40781d11` | 2026-06-28 22:04 | `snapshot for gitborg-runner` | | `e610f170` | 2026-06-28 14:41 | `snapshot for gitborg-runner` | This also resolves the 40 GB vs 20 GB puzzle raised above, which was good reasoning from a wrong premise. The controller does request 40, but the `gitborg-runner` image is `bdm_v2` and its `block_device_mapping` declares `"volume_size": 20` — **the image's mapping wins**. Every live runner boot volume is 20 GB, so `runner_controller_os_boot_volume_size: 40` is inert and misleading. The `Debian13` source image is consistent with this too: the staging VM boots Debian13, is provisioned, and is then snapshotted to produce `gitborg-runner`. The two 2026-08-01 volumes in the table above are gone — the sweep reclaimed them normally, which confirms the sweep works and that only snapshot-referenced volumes persist. **One is load-bearing.** The live Glance image boots from `snapshot_id 4daabf1b-…`, which is the **2026-07-29** row. That snapshot and volume `f1ce224d` must not be deleted — doing so leaves the image unbootable and breaks all CI. Verify against the image's `block_device_mapping` immediately before deleting rather than trusting this table. The other five pairs are safe. Current consumption for scale: 500 GiB across 14 volumes, 280 GiB across 8 snapshots. Reclaiming the five stale pairs returns roughly 200 GiB. Suggestion 3 above still stands and is the right detector, with one refinement: a count of `available` volumes would have caught this, but should be paired with the snapshot-refused count, since these volumes are unsweepable by design rather than merely unswept. Separately: `predrill-backup-20260709` (100 GiB) and `predrill-data-20260709` (60 GiB) have been retained since 2026-07-09 — same shape of problem, another 160 GiB, out of scope here.
supernaut refererade till detta ärende från en incheckning 2026-08-02 17:05:45 +00:00
Upphovsperson
Ägare

Status: detection is live in production; the reclaim is deliberately deferred.

#346 and #353 are applied. The controller now classifies unattached volumes and publishes the
counts, so this is measured rather than inferred:

gitborg_runner_controller_os_volumes_available{class="snapshot_pinned"}  6    (120 GiB)
gitborg_runner_controller_os_volumes_available{class="sweepable"}        2     (40 GiB)
gitborg_runner_controller_os_volumes_available{class="named"}            0

RunnerSnapshotPinnedVolumesHigh (> 2, for: 6h) will therefore fire, by decision — the standing
waste is left visible rather than the threshold being widened to accept it.

The reclaim, pre-verified 2026-08-02 against the live API

scripts/bake-runner-image.sh --prune-only would retire four generations. Every safety condition was
checked independently of the script's own reporting:

Volume Status Name Attachments Snapshot
e8dd5800-0003-4651-b0aa-3d2dce6c1030 available (unnamed) 0 d301abd1-…
4d4a7cd2-c735-406f-9166-184c68a0859f available (unnamed) 0 9eb201af-…
40781d11-63e5-48e5-83e0-b692691ae87a available (unnamed) 0 9a3daded-…
e610f170-221a-4c60-a81b-2aa594b31b9b available (unnamed) 0 30aad4e5-…

Retained, and must stay:

  • 4daabf1b-… / volume f1ce224d-… — the live gitborg-runner image's block_device_mapping
    points at this snapshot. Confirmed directly from image show. Deleting it leaves the image
    unbootable and breaks all CI.
  • a500b034-… / volume 7b4588d3-… — kept by RETAIN_GENERATIONS=2 as the rollback generation.

Those two protections are independent: the live image's snapshot is excluded by the
block-device-mapping check even if the retention count were lowered to 1.

Reclaim total: 160 GiB (4 snapshots + 4 volumes at 20 GiB each). Afterwards the pinned count is
2, which satisfies the alert threshold exactly — note that leaves no slack, so the next bake trips
it again until pruning becomes routine.

Worth knowing before running it

#353 fixed a defect in the prune that would have made this look successful while doing half the
job: the volume column was read under the wrong key and a tab-delimited read then disguised the empty
field as a timestamp, so every volume reported as "already gone". Re-run the dry run and confirm each
retired generation prints a real volume UUID before running for real — that output is the check.

**Status: detection is live in production; the reclaim is deliberately deferred.** #346 and #353 are applied. The controller now classifies unattached volumes and publishes the counts, so this is measured rather than inferred: ```text gitborg_runner_controller_os_volumes_available{class="snapshot_pinned"} 6 (120 GiB) gitborg_runner_controller_os_volumes_available{class="sweepable"} 2 (40 GiB) gitborg_runner_controller_os_volumes_available{class="named"} 0 ``` `RunnerSnapshotPinnedVolumesHigh` (`> 2`, `for: 6h`) will therefore fire, by decision — the standing waste is left visible rather than the threshold being widened to accept it. ## The reclaim, pre-verified 2026-08-02 against the live API `scripts/bake-runner-image.sh --prune-only` would retire four generations. Every safety condition was checked independently of the script's own reporting: | Volume | Status | Name | Attachments | Snapshot | | --- | --- | --- | --- | --- | | `e8dd5800-0003-4651-b0aa-3d2dce6c1030` | available | *(unnamed)* | 0 | `d301abd1-…` | | `4d4a7cd2-c735-406f-9166-184c68a0859f` | available | *(unnamed)* | 0 | `9eb201af-…` | | `40781d11-63e5-48e5-83e0-b692691ae87a` | available | *(unnamed)* | 0 | `9a3daded-…` | | `e610f170-221a-4c60-a81b-2aa594b31b9b` | available | *(unnamed)* | 0 | `30aad4e5-…` | **Retained, and must stay:** - `4daabf1b-…` / volume `f1ce224d-…` — the live `gitborg-runner` image's `block_device_mapping` points at this snapshot. Confirmed directly from `image show`. Deleting it leaves the image unbootable and breaks all CI. - `a500b034-…` / volume `7b4588d3-…` — kept by `RETAIN_GENERATIONS=2` as the rollback generation. Those two protections are independent: the live image's snapshot is excluded by the block-device-mapping check even if the retention count were lowered to 1. Reclaim total: **160 GiB** (4 snapshots + 4 volumes at 20 GiB each). Afterwards the pinned count is 2, which satisfies the alert threshold exactly — note that leaves no slack, so the next bake trips it again until pruning becomes routine. ## Worth knowing before running it `#353` fixed a defect in the prune that would have made this look successful while doing half the job: the volume column was read under the wrong key and a tab-delimited read then disguised the empty field as a timestamp, so every volume reported as "already gone". Re-run the dry run and confirm each retired generation prints a real volume UUID before running for real — that output is the check.
Upphovsperson
Ägare

Status 2026-08-04: the leak is gone, and the alert now firing is a false positive

The 160 GiB is reclaimed

Current unattached-volume survey on gitborg-prod:

class count
sweepable 1 (inside its grace window)
snapshot_pinned 2 — the retained bake generations, correct steady state
named 0

Cinder usage is 440 GB, against the ~640–720 GB floor this issue described. The eight stranded
volumes went when the bake snapshots were pruned earlier today: pruning unpinned them, and the
sweep then collected them on its next pass. That is exactly the mechanism the
RunnerSnapshotPinnedVolumesHigh comment predicted — "eight stranded 20 GiB volumes, each pinned by
a bake snapshot nobody pruned"
.

Step 1 — "identify the producer" — is answered, and the hypothesis in this issue was wrong

The issue reasoned the orphans were probably not ephemeral runner volumes, because
controller.py boots with boot_volume_size 40 GB while every orphan was 20 GB, and pointed at the
weekly backup-drill VM as the likeliest producer given its ~1–2/week cadence.

The sweep's own logs settle it. Today, one 20 GiB orphan boot volume was created and reclaimed per
CI run, continuously, by the runner path:

11:02:23 sweep: deleted orphan boot volume 694bb730… (20 GiB, age=3603s)
11:05:13 sweep: deleted orphan boot volume e1f885ef… (20 GiB, age=3603s)
11:05:23 sweep: deleted orphan boot volume 5d236561… (20 GiB, age=3603s)
11:14:23 sweep: deleted 2 orphan boot volume(s)      (age=3603s, 3601s)
13:03:41 sweep: deleted orphan boot volume 8bed814a… (20 GiB, age=3605s)
13:16:11 sweep: deleted orphan boot volume 6bac70d7… (20 GiB, age=3605s)

So the producer is the ephemeral runner path, not the drill, and the runners are not getting the
40 GB the default implies. The size mismatch that argued against the runner hypothesis is real but
means something different: runner_controller_os_boot_volume_size: 40 is not what these VMs end up
with. Worth a separate look — it does not affect the leak, which the sweep now handles, but the
declared and actual boot volume sizes disagree.

Why the alert re-fired, and why that is not a leak

Every deletion above lands at age ≈ 3600s, because orphan_volume_max_age_seconds is 3600: a
deliberate grace so the sweep can never catch a volume mid-attach to a booting runner. The sweep is
not lagging — it is waiting, exactly as designed, on every volume.

But RunnerOrphanVolumesUnswept read class="sweepable" > 0 for 2h, and sweepable counted
volumes still serving out that grace
. On a busy CI day the grace windows overlap continuously, so
the expression never went false long enough to reset the for, and the alert fired on healthy
behaviour. An alert named Unswept was reading not yet due for sweeping. Today's load — three PRs
plus pushes and Renovate — was enough to keep it firing all afternoon.

This is the mirror image of the error this issue fixed. Suggestion 3 here ("make the detector able to
see this class of leak") was implemented and overshot: it can now also see normal traffic.

Fix

Split sweepable by age and point the alert at the new class:

  • sweepable — unnamed, no snapshot, inside the grace. Normal CI traffic; not alerted.
  • sweepable_overdue — past the grace and still present after the sweep ran. The real backlog:
    the sweep is failing, refusing, or something is producing volumes it cannot reclaim.

Threshold stays at 0, so a single genuinely-stuck volume is still caught — raising it would hide
precisely the ~1.8 volumes/week accumulation this issue was filed about, and lengthening for: does
nothing because continuous CI keeps the old expression true indefinitely.

One subtlety worth recording: the survey is computed from the volume list read at the top of the
sweep, so volumes deleted during that same cycle are excluded explicitly. Without that, every
successful sweep would report its own just-deleted volume as overdue for one cycle — the metric would
blip on exactly the event proving the sweep works.

Age never overrides the other classes: snapshot-pinned is unsweepable by design however old, and a
named volume is deliberate retention however old. Both covered by tests (69/69 passing, six new).

The dashboard panel gains the overdue series, and its description no longer tells the reader that
sweepable should be 0.

## Status 2026-08-04: the leak is gone, and the alert now firing is a false positive ### The 160 GiB is reclaimed Current unattached-volume survey on `gitborg-prod`: | class | count | | ----------------- | ----- | | `sweepable` | 1 (inside its grace window) | | `snapshot_pinned` | 2 — the retained bake generations, correct steady state | | `named` | 0 | Cinder usage is **440 GB**, against the ~640–720 GB floor this issue described. The eight stranded volumes went when the bake snapshots were pruned earlier today: pruning **unpinned** them, and the sweep then collected them on its next pass. That is exactly the mechanism the `RunnerSnapshotPinnedVolumesHigh` comment predicted — *"eight stranded 20 GiB volumes, each pinned by a bake snapshot nobody pruned"*. ### Step 1 — "identify the producer" — is answered, and the hypothesis in this issue was wrong The issue reasoned the orphans were probably **not** ephemeral runner volumes, because `controller.py` boots with `boot_volume_size` 40 GB while every orphan was 20 GB, and pointed at the weekly backup-drill VM as the likeliest producer given its ~1–2/week cadence. The sweep's own logs settle it. Today, one **20 GiB** orphan boot volume was created and reclaimed per CI run, continuously, by the runner path: ``` 11:02:23 sweep: deleted orphan boot volume 694bb730… (20 GiB, age=3603s) 11:05:13 sweep: deleted orphan boot volume e1f885ef… (20 GiB, age=3603s) 11:05:23 sweep: deleted orphan boot volume 5d236561… (20 GiB, age=3603s) 11:14:23 sweep: deleted 2 orphan boot volume(s) (age=3603s, 3601s) 13:03:41 sweep: deleted orphan boot volume 8bed814a… (20 GiB, age=3605s) 13:16:11 sweep: deleted orphan boot volume 6bac70d7… (20 GiB, age=3605s) ``` So the producer is the ephemeral runner path, not the drill, and the runners are not getting the 40 GB the default implies. The size mismatch that argued against the runner hypothesis is real but means something different: `runner_controller_os_boot_volume_size: 40` is not what these VMs end up with. Worth a separate look — it does not affect the leak, which the sweep now handles, but the declared and actual boot volume sizes disagree. ### Why the alert re-fired, and why that is not a leak Every deletion above lands at **age ≈ 3600s**, because `orphan_volume_max_age_seconds` is 3600: a deliberate grace so the sweep can never catch a volume mid-attach to a booting runner. The sweep is not lagging — it is waiting, exactly as designed, on every volume. But `RunnerOrphanVolumesUnswept` read `class="sweepable" > 0` for 2h, and `sweepable` **counted volumes still serving out that grace**. On a busy CI day the grace windows overlap continuously, so the expression never went false long enough to reset the `for`, and the alert fired on healthy behaviour. An alert named *Unswept* was reading *not yet due for sweeping*. Today's load — three PRs plus pushes and Renovate — was enough to keep it firing all afternoon. This is the mirror image of the error this issue fixed. Suggestion 3 here ("make the detector able to see this class of leak") was implemented and overshot: it can now also see normal traffic. ### Fix Split `sweepable` by age and point the alert at the new class: - `sweepable` — unnamed, no snapshot, **inside** the grace. Normal CI traffic; not alerted. - `sweepable_overdue` — **past** the grace and still present after the sweep ran. The real backlog: the sweep is failing, refusing, or something is producing volumes it cannot reclaim. Threshold stays at **0**, so a single genuinely-stuck volume is still caught — raising it would hide precisely the ~1.8 volumes/week accumulation this issue was filed about, and lengthening `for:` does nothing because continuous CI keeps the old expression true indefinitely. One subtlety worth recording: the survey is computed from the volume list read at the *top* of the sweep, so volumes deleted during that same cycle are excluded explicitly. Without that, every successful sweep would report its own just-deleted volume as overdue for one cycle — the metric would blip on exactly the event proving the sweep works. Age never overrides the other classes: snapshot-pinned is unsweepable by design however old, and a named volume is deliberate retention however old. Both covered by tests (69/69 passing, six new). The dashboard panel gains the overdue series, and its description no longer tells the reader that `sweepable` should be 0.
Upphovsperson
Ägare

Closing as done. The orphan sweep and its alerts shipped between 2026-08-01 and 2026-08-04 (the runbook's orphan-volume section records the 80 GiB reclaim) and have run cleanly since. The sweep deletes each orphan at its one-hour grace.

What is left is why a 20 GiB volume is created per CI run at all. That is a separate question and stays open as #372.

Closing as done. The orphan sweep and its alerts shipped between 2026-08-01 and 2026-08-04 (the runbook's orphan-volume section records the 80 GiB reclaim) and have run cleanly since. The sweep deletes each orphan at its one-hour grace. What is left is why a 20 GiB volume is created per CI run at all. That is a separate question and stays open as #372.
Upphovsperson
Ägare

Correction to the closing comment, from a live openstack volume list on 2026-09-24 13:06 UTC.

  • Six unnamed 20 GiB volumes created 12:51 to 13:04 that day are CI runs inside the sweep's one-hour grace. Expected, and the subject of #372.
  • f1ce224d (2026-07-29) is not an orphan. Its snapshot 4daabf1b backs the live gitborg-runner image. Deleting it would break the runner image.
  • 7b4588d3 (2026-07-22) and its snapshot a500b034 are a leftover from an older image bake. No image references them. They are the last real orphan from the table above (40 GiB, about 43 kr/month). The sweep cannot delete a volume that still has a snapshot, which is why it survived.

Leaving this closed. Deleting that last pair is a one-off manual step, pending the operator's go.

Correction to the closing comment, from a live `openstack volume list` on 2026-09-24 13:06 UTC. - Six unnamed 20 GiB volumes created 12:51 to 13:04 that day are CI runs inside the sweep's one-hour grace. Expected, and the subject of #372. - `f1ce224d` (2026-07-29) is **not** an orphan. Its snapshot `4daabf1b` backs the live `gitborg-runner` image. Deleting it would break the runner image. - `7b4588d3` (2026-07-22) and its snapshot `a500b034` are a leftover from an older image bake. No image references them. They are the last real orphan from the table above (40 GiB, about 43 kr/month). The sweep cannot delete a volume that still has a snapshot, which is why it survived. Leaving this closed. Deleting that last pair is a one-off manual step, pending the operator's go.
Upphovsperson
Ägare

Deleted the last leftover pair on 2026-09-28: snapshot a500b034 and volume 7b4588d3, the 2026-07-22 rollback generation.

I used the bake script's prune, not an ad-hoc delete:

RETAIN_GENERATIONS=1 scripts/bake-runner-image.sh --prune-only --dry-run   # then without --dry-run

Checked before and after:

  • The gitborg-runner image boots from 4daabf1b-0721-4968-86fc-64e0c20015d4, and no other image references a snapshot.
  • The dry run retired only a500b034 / 7b4588d3 and kept 4daabf1b.
  • Afterwards openstack volume snapshot list shows only 4daabf1b, volume show 7b4588d3 returns No Volume found, and the image still maps 4daabf1b.

Side effect: there is no rollback runner image until the next bake. The next bake with the default RETAIN_GENERATIONS=2 restores one.

The runbook's image-reference check printed nothing even for the live image. That is fixed in #502.

Deleted the last leftover pair on 2026-09-28: snapshot `a500b034` and volume `7b4588d3`, the 2026-07-22 rollback generation. I used the bake script's prune, not an ad-hoc delete: ```bash RETAIN_GENERATIONS=1 scripts/bake-runner-image.sh --prune-only --dry-run # then without --dry-run ``` Checked before and after: - The `gitborg-runner` image boots from `4daabf1b-0721-4968-86fc-64e0c20015d4`, and no other image references a snapshot. - The dry run retired only `a500b034` / `7b4588d3` and kept `4daabf1b`. - Afterwards `openstack volume snapshot list` shows only `4daabf1b`, `volume show 7b4588d3` returns `No Volume found`, and the image still maps `4daabf1b`. Side effect: there is no rollback runner image until the next bake. The next bake with the default `RETAIN_GENERATIONS=2` restores one. The runbook's image-reference check printed nothing even for the live image. That is fixed in #502.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#320
Ingen beskrivning angiven.