fix(runner-controller): naive OpenStack timestamps disable the volume sweep (#133) #135
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!135
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/runner-controller-age-naive-tz"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Critical follow-up to #134 — without this the orphan-volume sweep is inert.
Verifying #134 in prod after apply: the new metrics are live and the controller is healthy, but
orphan_volumes_sweptstayed 0 with 11 detached orphans present. Root cause: OpenStack returnscreated_atwithout a timezone (2026-07-19T08:07:59.000000).datetime.fromisoformatparses that as a naive datetime; subtracting it from an awarenow()raisesTypeError, which_iso_age_seconds'sexceptswallowed into0.0. So every volume looked 0s old →age <= max_agewas always true → the sweep never deleted anything. The same bug sat in_server_age_seconds, silently disabling thereap_max_ageserver backstop.Fix: stamp a missing
tzinfoas UTC (OpenStack timestamps are UTC), and route_server_age_secondsthrough the same helper.Verified by exec'ing the actual function from source against real formats:
py_compileclean. Needs the runner-controller image rebuilt +site.ymlapply, same as #134. Once live, the sweep will clear the 11 leftover orphans as they cross the 1h age threshold.