runner-controller: the reaper logs a permanent WARNING for a detach that can never succeed #305
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#305
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
The runner-controller's reaper logs a
WARNINGevery few minutes that reads like a cleanup failure butis not one. Found while checking production before the public announcement.
Volumes are NOT leaking — stated up front, because the log implies otherwise
Checked before filing, so nobody re-runs the investigation:
gitborg_runner_controller_os_volume_gb_usedover 7 days oscillates between 640 and 1020 GB andreturns to a 640–720 baseline — it does not climb. Quota is 5000 GB, so ~14–20% used.
gitborg_runner_controller_last_loop_ok= 1,active_vms= 0,boot_error_vms= 0,queued_jobs= 0.IDs each time — so it is not one stuck volume retried forever.
The volumes evidently do get released, presumably because deleting the server cascades to its root
device. The detach attempt that precedes it can never succeed: OpenStack refuses to detach a root
device volume by design.
Why it is still worth fixing
A prior CI outage was caused by orphaned runner boot volumes (227 of them), and the orphan sweep is the
control that stops that recurring. This message is emitted at
WARNINGin that exact code path, severaltimes an hour, permanently. A genuine reap failure would arrive in the middle of a steady stream of
identical-looking warnings that everyone has learned to ignore. That is the condition under which the
next volume leak goes unnoticed, and the fix is cheap.
Suggested
Do not attempt to detach a volume that is the server's root device — delete the server and let the root
volume cascade, which is what already happens. If an explicit detach is wanted for non-root volumes,
skip root devices rather than calling and catching the 400. Either way the reaper should stop logging a
WARNINGfor an outcome that is normal and expected.Open question, not investigated
The baseline of ~640 GB used with zero active runner VMs is larger than the known persistent
volumes obviously account for. It may be entirely legitimate (host root volumes, the monitoring host,
images or snapshots counted against the same project quota) — but it was not chased down, and if it is
not fully explained it deserves its own look. Recorded so the number is not mistaken for verified.