runner-controller: the reaper logs a permanent WARNING for a detach that can never succeed #305

Stängd
öppnade 2026-08-01 14:12:41 +00:00 av supernaut · 0 kommentarer
Ägare

The runner-controller's reaper logs a WARNING every few minutes that reads like a cleanup failure but
is not one. Found while checking production before the public announcement.

WARNING [controller] reap: could not delete volume ade7c40a-…: BadRequestException: 400:
Client Error for url: https://bcp.bahnhof.cloud:8774/v2.1/servers/926e0760-…/os-volume_attachments/ade7c40a-…,
Cannot detach a root device volume

Volumes are NOT leaking — stated up front, because the log implies otherwise

Checked before filing, so nobody re-runs the investigation:

  • gitborg_runner_controller_os_volume_gb_used over 7 days oscillates between 640 and 1020 GB and
    returns to a 640–720 baseline
    — it does not climb. Quota is 5000 GB, so ~14–20% used.
  • gitborg_runner_controller_last_loop_ok = 1, active_vms = 0, boot_error_vms = 0,
    queued_jobs = 0.
  • Rate is steady at roughly 2–7 occurrences per hour over 24 h, with different volume and server
    IDs each time — so it is not one stuck volume retried forever.

The volumes evidently do get released, presumably because deleting the server cascades to its root
device. The detach attempt that precedes it can never succeed: OpenStack refuses to detach a root
device volume by design.

Why it is still worth fixing

A prior CI outage was caused by orphaned runner boot volumes (227 of them), and the orphan sweep is the
control that stops that recurring. This message is emitted at WARNING in that exact code path, several
times an hour, permanently. A genuine reap failure would arrive in the middle of a steady stream of
identical-looking warnings that everyone has learned to ignore. That is the condition under which the
next volume leak goes unnoticed, and the fix is cheap.

Suggested

Do not attempt to detach a volume that is the server's root device — delete the server and let the root
volume cascade, which is what already happens. If an explicit detach is wanted for non-root volumes,
skip root devices rather than calling and catching the 400. Either way the reaper should stop logging a
WARNING for an outcome that is normal and expected.

Open question, not investigated

The baseline of ~640 GB used with zero active runner VMs is larger than the known persistent
volumes obviously account for. It may be entirely legitimate (host root volumes, the monitoring host,
images or snapshots counted against the same project quota) — but it was not chased down, and if it is
not fully explained it deserves its own look. Recorded so the number is not mistaken for verified.

The runner-controller's reaper logs a `WARNING` every few minutes that reads like a cleanup failure but is not one. Found while checking production before the public announcement. ```text WARNING [controller] reap: could not delete volume ade7c40a-…: BadRequestException: 400: Client Error for url: https://bcp.bahnhof.cloud:8774/v2.1/servers/926e0760-…/os-volume_attachments/ade7c40a-…, Cannot detach a root device volume ``` ## Volumes are NOT leaking — stated up front, because the log implies otherwise Checked before filing, so nobody re-runs the investigation: - `gitborg_runner_controller_os_volume_gb_used` over 7 days **oscillates between 640 and 1020 GB and returns to a 640–720 baseline** — it does not climb. Quota is 5000 GB, so ~14–20% used. - `gitborg_runner_controller_last_loop_ok` = 1, `active_vms` = 0, `boot_error_vms` = 0, `queued_jobs` = 0. - Rate is steady at roughly 2–7 occurrences per hour over 24 h, with different volume **and** server IDs each time — so it is not one stuck volume retried forever. The volumes evidently do get released, presumably because deleting the server cascades to its root device. The detach attempt that precedes it can never succeed: OpenStack refuses to detach a root device volume by design. ## Why it is still worth fixing A prior CI outage was caused by orphaned runner boot volumes (227 of them), and the orphan sweep is the control that stops that recurring. This message is emitted at `WARNING` in that exact code path, several times an hour, permanently. A genuine reap failure would arrive in the middle of a steady stream of identical-looking warnings that everyone has learned to ignore. That is the condition under which the next volume leak goes unnoticed, and the fix is cheap. ## Suggested Do not attempt to detach a volume that is the server's root device — delete the server and let the root volume cascade, which is what already happens. If an explicit detach is wanted for non-root volumes, skip root devices rather than calling and catching the 400. Either way the reaper should stop logging a `WARNING` for an outcome that is normal and expected. ## Open question, not investigated The baseline of ~640 GB used with **zero** active runner VMs is larger than the known persistent volumes obviously account for. It may be entirely legitimate (host root volumes, the monitoring host, images or snapshots counted against the same project quota) — but it was not chased down, and if it is not fully explained it deserves its own look. Recorded so the number is not mistaken for verified.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#305
Ingen beskrivning angiven.