An undeclared 20 GiB boot volume is provisioned once per CI run — matching neither the runner's 40 GB nor the drill's 80 GB #372
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#372
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
What is observed
The orphan sweep reclaims one unnamed, unattached 20 GiB volume per CI run, continuously. From
gitborg-prodon 2026-08-04 (each deleted at the sweep's 1-hour grace, so the sweep is healthy):20 GiB matches nothing we declare
gitborg-runnerrunner_controller_os_boot_volume_size)gitborg-runnerbackup_drill_os_boot_volume_size)Confirmed against the live config on the host, not just the role defaults:
So the volumes being reclaimed are not the declared boot volume of either producer. #320 recorded the
eight historical orphans the same way — 20 GB, bootable, sourced from
Debian13— which is notthe image either producer boots from.
Something is provisioning an undeclared 20 GiB volume roughly once per CI run.
Why file it when the sweep already reclaims them
Not cost: the sweep takes each one an hour later and Cinder usage is stable at 440 GB. Two other
reasons:
the same shape as the class that exhausted the volume quota and stopped CI in #133. If the sweep
ever regresses, this refills the quota on its own.
be runner volumes because the controller declares 40 GB and the orphans were 20 GB, and named the
weekly backup-drill VM as the likeliest producer on cadence grounds. The per-CI-run rhythm above
rules the drill out. A declared-vs-actual gap that survives is a trap for the next person.
Suggested diagnostic
The volumes exist for a full hour before the sweep takes them, so there is a comfortable window to
catch one live. While
gitborg_runner_controller_os_volumes_available{class="sweepable"}is 1:Then correlate against the runner VM active at that timestamp — specifically whether its own
declared boot volume (40 GB,
delete_on_termination=true) cascaded correctly, i.e. whether this is asecond volume rather than the VM's own.
#320's disclosure hypothesised exactly that and is worth testing first: "That VM's declared boot
volume cascaded correctly on
delete_on_termination=true, so how a second one appeared isunexplained — possibly a create that Nova rejected after Cinder had already provisioned." If Nova
rejects a create after Cinder has provisioned, the abandoned volume would be image-sized rather than
request-sized, which would explain 20 GiB against a 40 GB request.
Refs #320, #133.