fix(runner-image): the prune reads the wrong snapshot column, so no boot volume is ever deleted #352
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#352
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Found by running the reclaim dry run that #346's description recommends:
The volume field holds a timestamp, the created field is empty, and every generation reports its
volume as "already gone". Run for real, the prune would delete the snapshots and leave every boot
volume behind — half the reclaim, and the half that holds 20 GiB apiece.
Two bugs, compounding
1. The jq reads a column the CLI does not return.
openstack volume snapshot list --long -f jsonnames it
Volume;bake_generations()reads."Volume ID" // .volume_id // "", so the value isalways empty. Verified:
2. A tab-delimited read collapses the empty field. Tab is IFS whitespace, so bash treats runs
of tabs as one delimiter and drops empties — the timestamp slides into
vol_id:Bug 1 empties the field; bug 2 disguises it as a plausible-looking value, which is why it reads as
"already gone" rather than as an obvious error.
Fix
Read
.Volumefirst (keeping the other spellings as fallbacks for other CLI versions), and switchthe row delimiter from
@tsvto a pipe, which is not IFS whitespace so empty fields survive. Neithera UUID nor an ISO-8601 timestamp can contain one.
After the fix the dry run resolves real volumes and still protects what it must:
Those four are exactly the stranded volumes identified in #320.
Note on blast radius
Not silently destructive: the two safety conditions in
prune_generation_volume(available andunnamed) are unaffected, and a snapshot referenced by a live image is still never touched. The
failure mode was under-deletion, not over-deletion. The volumes would in fact have been reclaimed
eventually — once their snapshot was gone they become unpinned, and the controller's sweep takes
unnamed available volumes — but by accident, via a different mechanism, and the script's own output
would have said the job was done when it was not.
controller.pyis unaffected: it readsvolume_idoff the openstacksdk object, which is the correctattribute there, and its published counts match the live state.
Done when
column
shellcheckandbash -nclean