fix(runner-image): read the snapshot's real volume column when pruning #353
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!353
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/bake-prune-volume-field"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Closes #352. Found by running the reclaim dry run that #346's own description recommends.
Before — the volume column holds a timestamp, created is empty, and every generation claims its
volume is already gone:
Run for real, that deletes the snapshots and leaves every boot volume behind — the half of the
reclaim that actually frees quota.
Two bugs, compounding
The jq reads a column the CLI does not return.
openstack volume snapshot list --long -f jsonnames it
Volume; the code read."Volume ID" // .volume_id // "", so the value was always empty.Verified against the live API:
has "Volume ID": false,has "Volume": true.A tab-delimited read then hid it. Tab is IFS whitespace, so bash collapses runs of tabs and
drops empty fields — the timestamp slid one position left into
vol_id:The first bug empties the field; the second disguises it as a plausible value. That is why it
surfaced as a calm "already gone" rather than an error.
Fix
Read
.Volumefirst, keeping the other spellings as fallbacks for other CLI versions, and delimitrows with a pipe instead of
@tsv. A pipe is not IFS whitespace, so empty fields survive, andneither a UUID nor an ISO-8601 timestamp can contain one.
After
Those four are exactly the stranded volumes identified in #320, and the live image's pair is still
protected.
Blast radius
Not silently destructive — the failure mode was under-deletion.
prune_generation_volume's twosafety conditions (available and unnamed) are untouched, and a snapshot referenced by a live
image is still never a candidate. The volumes would in fact have been reclaimed eventually, since
deleting the snapshot unpins them and the controller's sweep takes unnamed available volumes — but
by accident, and the script would have reported success while leaving 80 GiB behind.
controller.pyis unaffected: it readsvolume_idoff the openstacksdk object, the correctattribute there, and its published counts match live state (
snapshot_pinned 6, 120 GiB).Checks
shellcheckexit 0 (the first attempt at this fix regressed it — a comment containing shell tabsyntax inside the jq string tripped SC1012; reworded),
bash -nclean, Prettier clean, and the dryrun re-run against production to confirm the output above. Nothing was deleted.