fix(runner-controller): sweep skips snapshot-backed volumes, no retry-spam (#133) #136
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!136
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/runner-controller-sweep-skip-snapshots"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Third and final refinement to the #133 sweep (follows #134 sweep+alerts, #135 tz fix).
What post-#135 verification found: the sweep now ages orphans correctly (the tz fix works) and deletes plain runner leaks — but it hammered two unnamed detached volumes every 10s with:
Those two aren't ephemeral leaks: a runner boot volume never has a snapshot. They carry snapshots — the
gitborg-runnerimage's source snapshot and thepredrill-*2026-07-09 safety snapshots — i.e. image/backup infrastructure. OpenStack correctly refused; the blank-name heuristic was just too broad.Fix:
snapshots()list call per cycle, not per-volume; falls back gracefully if that call fails).Validated:
py_compile+ansible-lint(production) clean, and a stub-connection unit test of the actual function asserts it deletes a plain old orphan, skips+skip-lists a snapshot-backed volume (no re-attempt on the next cycle), and never touches named/in-use/too-young volumes.Needs image rebuild +
site.ymlapply. After it, the 10 genuine runner-leak orphans (currently aging past the 1h threshold) sweep cleanly and the two infra volumes are left alone silently.