8 orphaned 20 GB boot volumes (160 GB) accumulating since June — the leak metric cannot see it #320
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#320
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
8 unattached bootable volumes, 160 GB, accumulating since 2026-06-28. Found while tearing down a
throwaway VM on 2026-08-01. This answers the open question #305 recorded and left unchased, and it
contradicts the "volumes are not leaking" conclusion in #311 — for a reason that is worth
understanding rather than just correcting.
What is there
Every one is identical: 20 GB, bootable, sourced from
Debian13,available(attached tonothing), and unnamed.
Everything else in the project is legitimately attached:
gitborg-prod-root(20),-data(60),-backup(150),-lfs(60),gitborg-prod-monitoring-root(30) — 320 GB in use.Why the existing metric cannot see it
#311 concluded volumes were not leaking because
gitborg_runner_controller_os_volume_gb_usedoscillates between 640 and 1020 GB over 7 days and returns to a 640–720 baseline without climbing.
That reasoning is sound for a fast leak and blind to this one: 160 GB over five weeks is roughly
1.8 volumes/week, ~36 GB/week, which disappears inside a band that swings by ~380 GB during
normal CI churn. The metric is not wrong; it is not sensitive enough to separate slow accumulation
from churn. A leak detector whose noise band is ten times the weekly leak rate will always read
clean.
This is also most of the unexplained baseline #305 flagged: 320 GB attached + 160 GB orphaned =
480 GB against an observed ~640–720 GB floor. Not the whole gap, but the largest identified piece.
These are probably NOT ephemeral runner volumes
The obvious suspect is the runner controller, given the prior outage with 227 orphaned boot volumes.
The sizes argue against it:
controller.pyboots runners with"volume_size": int(os_cfg.get("boot_volume_size", 40))— 40 GB — and every orphan here is20 GB. Unless that default is overridden to 20 somewhere, these come from a different producer.
Worth checking first, roughly in order of suspicion:
backup-drillboots a throwaway server reusing the controller'screate_servershape), whose cadence is closest to the observed ~1–2/week;server create— note that a zero-disk flavour makes a plain create fail withOnly volume-backed servers are allowed for flavors with zero disk, and it is worth confirmingwhether a rejected create can still leave a volume behind;
tofureplacements of the services or monitoring host.The creation timestamps are the evidence to correlate against — several cluster in pairs hours
apart, which does not obviously match a weekly timer.
Disclosure
A ninth such volume existed and was created during the investigation itself, inside the window of a
throwaway VM booted for the #315 rollback rehearsal, matching that VM's image and size. It was
deleted immediately. That VM's declared boot volume cascaded correctly on
delete_on_termination=true, so how a second one appeared is unexplained — possibly a create thatNova rejected after Cinder had already provisioned. Flagged because it may be the same mechanism
producing the eight, and because two of the eight (18:17, 18:21) predate any command run that day.
Suggested
Cinder retains no attachment history for a deleted server, so correlation has to come from the
timestamps against CI runs, drill runs and applies.
count of
availablevolumes (or their total size) is a direct signal and does not depend onseparating a slow trend from CI churn — unattached-and-unnamed is an unambiguous state in a
project where every legitimate volume is named and attached.
Refs #305, #311.
Producer identified. This answers the "identify the producer before deleting anything" step.
They are image-bake leftovers, not runner-job leaks. Baking the
gitborg-runnerGlance imagesnapshots the staging VM's boot volume; the snapshot keeps a reference, so Cinder refuses to delete
the volume, and the controller's sweep then correctly skips it —
_sweep_orphan_volumesexcludesany volume that has snapshots, precisely so it cannot destroy image or backup infrastructure.
The controller logs the refusal on every restart:
Each stranded volume pairs one-to-one with a bake snapshot, and the timestamps line up:
f1ce224dsnapshot for gitborg-runner-staging-20260729-2309367b4588d3snapshot for gitborg-runner-staging-20260722-161157e8dd5800snapshot for gitborg-runner-staging-20260720-1957564d4a7cd2snapshot for gitborg-runner-staging-20260720-19401940781d11snapshot for gitborg-runnere610f170snapshot for gitborg-runnerThis also resolves the 40 GB vs 20 GB puzzle raised above, which was good reasoning from a wrong
premise. The controller does request 40, but the
gitborg-runnerimage isbdm_v2and itsblock_device_mappingdeclares"volume_size": 20— the image's mapping wins. Every liverunner boot volume is 20 GB, so
runner_controller_os_boot_volume_size: 40is inert andmisleading. The
Debian13source image is consistent with this too: the staging VM boots Debian13,is provisioned, and is then snapshotted to produce
gitborg-runner.The two 2026-08-01 volumes in the table above are gone — the sweep reclaimed them normally, which
confirms the sweep works and that only snapshot-referenced volumes persist.
One is load-bearing. The live Glance image boots from
snapshot_id 4daabf1b-…, which is the2026-07-29 row. That snapshot and volume
f1ce224dmust not be deleted — doing so leaves theimage unbootable and breaks all CI. Verify against the image's
block_device_mappingimmediatelybefore deleting rather than trusting this table. The other five pairs are safe.
Current consumption for scale: 500 GiB across 14 volumes, 280 GiB across 8 snapshots. Reclaiming
the five stale pairs returns roughly 200 GiB.
Suggestion 3 above still stands and is the right detector, with one refinement: a count of
availablevolumes would have caught this, but should be paired with the snapshot-refused count,since these volumes are unsweepable by design rather than merely unswept.
Separately:
predrill-backup-20260709(100 GiB) andpredrill-data-20260709(60 GiB) have beenretained since 2026-07-09 — same shape of problem, another 160 GiB, out of scope here.
Status: detection is live in production; the reclaim is deliberately deferred.
#346 and #353 are applied. The controller now classifies unattached volumes and publishes the
counts, so this is measured rather than inferred:
RunnerSnapshotPinnedVolumesHigh(> 2,for: 6h) will therefore fire, by decision — the standingwaste is left visible rather than the threshold being widened to accept it.
The reclaim, pre-verified 2026-08-02 against the live API
scripts/bake-runner-image.sh --prune-onlywould retire four generations. Every safety condition waschecked independently of the script's own reporting:
e8dd5800-0003-4651-b0aa-3d2dce6c1030d301abd1-…4d4a7cd2-c735-406f-9166-184c68a0859f9eb201af-…40781d11-63e5-48e5-83e0-b692691ae87a9a3daded-…e610f170-221a-4c60-a81b-2aa594b31b9b30aad4e5-…Retained, and must stay:
4daabf1b-…/ volumef1ce224d-…— the livegitborg-runnerimage'sblock_device_mappingpoints at this snapshot. Confirmed directly from
image show. Deleting it leaves the imageunbootable and breaks all CI.
a500b034-…/ volume7b4588d3-…— kept byRETAIN_GENERATIONS=2as the rollback generation.Those two protections are independent: the live image's snapshot is excluded by the
block-device-mapping check even if the retention count were lowered to 1.
Reclaim total: 160 GiB (4 snapshots + 4 volumes at 20 GiB each). Afterwards the pinned count is
2, which satisfies the alert threshold exactly — note that leaves no slack, so the next bake trips
it again until pruning becomes routine.
Worth knowing before running it
#353fixed a defect in the prune that would have made this look successful while doing half thejob: the volume column was read under the wrong key and a tab-delimited read then disguised the empty
field as a timestamp, so every volume reported as "already gone". Re-run the dry run and confirm each
retired generation prints a real volume UUID before running for real — that output is the check.
Status 2026-08-04: the leak is gone, and the alert now firing is a false positive
The 160 GiB is reclaimed
Current unattached-volume survey on
gitborg-prod:sweepablesnapshot_pinnednamedCinder usage is 440 GB, against the ~640–720 GB floor this issue described. The eight stranded
volumes went when the bake snapshots were pruned earlier today: pruning unpinned them, and the
sweep then collected them on its next pass. That is exactly the mechanism the
RunnerSnapshotPinnedVolumesHighcomment predicted — "eight stranded 20 GiB volumes, each pinned bya bake snapshot nobody pruned".
Step 1 — "identify the producer" — is answered, and the hypothesis in this issue was wrong
The issue reasoned the orphans were probably not ephemeral runner volumes, because
controller.pyboots withboot_volume_size40 GB while every orphan was 20 GB, and pointed at theweekly backup-drill VM as the likeliest producer given its ~1–2/week cadence.
The sweep's own logs settle it. Today, one 20 GiB orphan boot volume was created and reclaimed per
CI run, continuously, by the runner path:
So the producer is the ephemeral runner path, not the drill, and the runners are not getting the
40 GB the default implies. The size mismatch that argued against the runner hypothesis is real but
means something different:
runner_controller_os_boot_volume_size: 40is not what these VMs end upwith. Worth a separate look — it does not affect the leak, which the sweep now handles, but the
declared and actual boot volume sizes disagree.
Why the alert re-fired, and why that is not a leak
Every deletion above lands at age ≈ 3600s, because
orphan_volume_max_age_secondsis 3600: adeliberate grace so the sweep can never catch a volume mid-attach to a booting runner. The sweep is
not lagging — it is waiting, exactly as designed, on every volume.
But
RunnerOrphanVolumesUnsweptreadclass="sweepable" > 0for 2h, andsweepablecountedvolumes still serving out that grace. On a busy CI day the grace windows overlap continuously, so
the expression never went false long enough to reset the
for, and the alert fired on healthybehaviour. An alert named Unswept was reading not yet due for sweeping. Today's load — three PRs
plus pushes and Renovate — was enough to keep it firing all afternoon.
This is the mirror image of the error this issue fixed. Suggestion 3 here ("make the detector able to
see this class of leak") was implemented and overshot: it can now also see normal traffic.
Fix
Split
sweepableby age and point the alert at the new class:sweepable— unnamed, no snapshot, inside the grace. Normal CI traffic; not alerted.sweepable_overdue— past the grace and still present after the sweep ran. The real backlog:the sweep is failing, refusing, or something is producing volumes it cannot reclaim.
Threshold stays at 0, so a single genuinely-stuck volume is still caught — raising it would hide
precisely the ~1.8 volumes/week accumulation this issue was filed about, and lengthening
for:doesnothing because continuous CI keeps the old expression true indefinitely.
One subtlety worth recording: the survey is computed from the volume list read at the top of the
sweep, so volumes deleted during that same cycle are excluded explicitly. Without that, every
successful sweep would report its own just-deleted volume as overdue for one cycle — the metric would
blip on exactly the event proving the sweep works.
Age never overrides the other classes: snapshot-pinned is unsweepable by design however old, and a
named volume is deliberate retention however old. Both covered by tests (69/69 passing, six new).
The dashboard panel gains the overdue series, and its description no longer tells the reader that
sweepableshould be 0.Closing as done. The orphan sweep and its alerts shipped between 2026-08-01 and 2026-08-04 (the runbook's orphan-volume section records the 80 GiB reclaim) and have run cleanly since. The sweep deletes each orphan at its one-hour grace.
What is left is why a 20 GiB volume is created per CI run at all. That is a separate question and stays open as #372.
Correction to the closing comment, from a live
openstack volume liston 2026-09-24 13:06 UTC.f1ce224d(2026-07-29) is not an orphan. Its snapshot4daabf1bbacks the livegitborg-runnerimage. Deleting it would break the runner image.7b4588d3(2026-07-22) and its snapshota500b034are a leftover from an older image bake. No image references them. They are the last real orphan from the table above (40 GiB, about 43 kr/month). The sweep cannot delete a volume that still has a snapshot, which is why it survived.Leaving this closed. Deleting that last pair is a one-off manual step, pending the operator's go.
Deleted the last leftover pair on 2026-09-28: snapshot
a500b034and volume7b4588d3, the 2026-07-22 rollback generation.I used the bake script's prune, not an ad-hoc delete:
Checked before and after:
gitborg-runnerimage boots from4daabf1b-0721-4968-86fc-64e0c20015d4, and no other image references a snapshot.a500b034/7b4588d3and kept4daabf1b.openstack volume snapshot listshows only4daabf1b,volume show 7b4588d3returnsNo Volume found, and the image still maps4daabf1b.Side effect: there is no rollback runner image until the next bake. The next bake with the default
RETAIN_GENERATIONS=2restores one.The runbook's image-reference check printed nothing even for the live image. That is fixed in #502.