Each runner image rebake strands a 20 GiB boot volume and its snapshot #340
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#340
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Problem
Every rebake of the
gitborg-runnerGlance image permanently strands a 20 GiB Cinder volume plusits 20 GiB snapshot. Neither is ever reclaimed, and the pair is invisible to the orphan sweep by
design.
Baking the image snapshots the staging VM's boot volume. The snapshot keeps a reference to that
volume, so Cinder refuses to delete it — and the sweep in
ansible/roles/runner-controller/files/controller.pycorrectly skips any volume that has snapshots,because that heuristic is what stops it destroying image and backup infrastructure (added in #133).
The sweep is behaving exactly as designed. Nothing here is a sweep defect. The gap is that
nothing owns the other end: the bake process never cleans up after itself.
Evidence
Nine volumes are currently
available(detached), all 20 GiB and bootable. Three were createdtoday and are simply inside the sweep's age window — they will be reclaimed normally. The other six
are permanent, and each pairs with an image-bake snapshot:
f1ce224dsnapshot for gitborg-runner-staging-20260729-2309367b4588d3snapshot for gitborg-runner-staging-20260722-161157e8dd5800snapshot for gitborg-runner-staging-20260720-1957564d4a7cd2snapshot for gitborg-runner-staging-20260720-19401940781d11snapshot for gitborg-runnere610f170snapshot for gitborg-runnerThe controller logs the refusal on every restart, and has done since at least 2026-07-29:
Current block-storage consumption is 500 GiB across 14 volumes and 280 GiB across 8 snapshots.
The six stranded pairs account for roughly 200 GiB of that — about a quarter of all allocated block
storage, holding nothing.
Why it matters
This is the same quota that caused the CI outage in #133: once volume-gigabytes is exhausted, every
runner boot fails for lack of a boot volume, self-reinforcingly and silently. The growth rate here
is slow — one pair per rebake — but it is unbounded and it consumes the exact resource that has
already taken CI down once.
Careful — one pair is live
The current Glance image boots from
snapshot_id 4daabf1b-…, which issnapshot for gitborg-runner-staging-20260729-230936— the 2026-07-29 row above. That snapshotand its volume
f1ce224dare load-bearing and must not be deleted. Deleting them would leave theimage unbootable and break all CI.
The five older pairs are safe to reclaim.
Proposed work
block_device_mappingimmediately beforehand rather than trusting the table above. Roughly200 GiB returned.
the staging VM's volume and the intermediate snapshot, keeping only what the image references.
Whichever step bakes the image should own this — a leak that needs a human to notice is not
fixed.
generation back is a reasonable rollback story; keeping five is an accident. Make the retention
count explicit.
gitborg_runner_controller_os_volume_gb_*is already exported and dashboarded.Add a gauge for detached-but-unsweepable volumes, so the next thing that accumulates behind the
snapshot heuristic is visible rather than discovered by hand.
Also worth a look while in here
Two unrelated snapshots,
predrill-backup-20260709(100 GiB) andpredrill-data-20260709(60 GiB), have been retained since 2026-07-09. If the restore drill they belong to is long finished,
that is another 160 GiB. Out of scope for this issue, but it is the same failure shape: a safety
artefact that nothing is responsible for removing.
Acceptance
generation.
Duplicate of #320, which predates this and already carries the accumulation evidence and the metric-sensitivity analysis. The root-cause analysis from this issue has been moved to #320, where it answers that issue's open "identify the producer" step. Closing in favour of #320.