An undeclared 20 GiB boot volume is provisioned once per CI run — matching neither the runner's 40 GB nor the drill's 80 GB #372

Öppen
öppnade 2026-08-04 14:40:01 +00:00 av supernaut · 0 kommentarer
Ägare

What is observed

The orphan sweep reclaims one unnamed, unattached 20 GiB volume per CI run, continuously. From
gitborg-prod on 2026-08-04 (each deleted at the sweep's 1-hour grace, so the sweep is healthy):

11:02:23 sweep: deleted orphan boot volume 694bb730… (20 GiB, age=3603s)
11:05:13 sweep: deleted orphan boot volume e1f885ef… (20 GiB, age=3603s)
11:05:23 sweep: deleted orphan boot volume 5d236561… (20 GiB, age=3603s)
11:14:23 sweep: deleted 2 orphan boot volume(s)      (age=3603s, 3601s)
13:03:41 sweep: deleted orphan boot volume 8bed814a… (20 GiB, age=3605s)
13:16:11 sweep: deleted orphan boot volume 6bac70d7… (20 GiB, age=3605s)

20 GiB matches nothing we declare

Producer Image Declared boot volume
runner-controller gitborg-runner 40 GB (runner_controller_os_boot_volume_size)
backup-drill gitborg-runner 80 GB (backup_drill_os_boot_volume_size)

Confirmed against the live config on the host, not just the role defaults:

$ grep boot_volume_size /home/gitborg/.config/runner-controller/config.yaml
  boot_volume_size: 40

So the volumes being reclaimed are not the declared boot volume of either producer. #320 recorded the
eight historical orphans the same way — 20 GB, bootable, sourced from Debian13 — which is not
the image either producer boots from.

Something is provisioning an undeclared 20 GiB volume roughly once per CI run.

Why file it when the sweep already reclaims them

Not cost: the sweep takes each one an hour later and Cinder usage is stable at 440 GB. Two other
reasons:

  1. Undeclared provisioning is unexplained behaviour, and the volumes are bootable and unnamed —
    the same shape as the class that exhausted the volume quota and stopped CI in #133. If the sweep
    ever regresses, this refills the quota on its own.
  2. This exact mismatch already misdirected an investigation. #320 reasoned the orphans could not
    be runner volumes because the controller declares 40 GB and the orphans were 20 GB, and named the
    weekly backup-drill VM as the likeliest producer on cadence grounds. The per-CI-run rhythm above
    rules the drill out. A declared-vs-actual gap that survives is a trap for the next person.

Suggested diagnostic

The volumes exist for a full hour before the sweep takes them, so there is a comfortable window to
catch one live. While gitborg_runner_controller_os_volumes_available{class="sweepable"} is 1:

openstack volume list --status available -f value -c ID
openstack volume show <id> -f json   # volume_image_metadata.image_name, bootable, size, created_at

Then correlate against the runner VM active at that timestamp — specifically whether its own
declared boot volume (40 GB, delete_on_termination=true) cascaded correctly, i.e. whether this is a
second volume rather than the VM's own.

#320's disclosure hypothesised exactly that and is worth testing first: "That VM's declared boot
volume cascaded correctly on delete_on_termination=true, so how a second one appeared is
unexplained — possibly a create that Nova rejected after Cinder had already provisioned."
If Nova
rejects a create after Cinder has provisioned, the abandoned volume would be image-sized rather than
request-sized, which would explain 20 GiB against a 40 GB request.

Refs #320, #133.

## What is observed The orphan sweep reclaims **one unnamed, unattached 20 GiB volume per CI run**, continuously. From `gitborg-prod` on 2026-08-04 (each deleted at the sweep's 1-hour grace, so the sweep is healthy): ``` 11:02:23 sweep: deleted orphan boot volume 694bb730… (20 GiB, age=3603s) 11:05:13 sweep: deleted orphan boot volume e1f885ef… (20 GiB, age=3603s) 11:05:23 sweep: deleted orphan boot volume 5d236561… (20 GiB, age=3603s) 11:14:23 sweep: deleted 2 orphan boot volume(s) (age=3603s, 3601s) 13:03:41 sweep: deleted orphan boot volume 8bed814a… (20 GiB, age=3605s) 13:16:11 sweep: deleted orphan boot volume 6bac70d7… (20 GiB, age=3605s) ``` ## 20 GiB matches nothing we declare | Producer | Image | Declared boot volume | | --- | --- | --- | | runner-controller | `gitborg-runner` | **40 GB** (`runner_controller_os_boot_volume_size`) | | backup-drill | `gitborg-runner` | **80 GB** (`backup_drill_os_boot_volume_size`) | Confirmed against the *live* config on the host, not just the role defaults: ``` $ grep boot_volume_size /home/gitborg/.config/runner-controller/config.yaml boot_volume_size: 40 ``` So the volumes being reclaimed are not the declared boot volume of either producer. #320 recorded the eight historical orphans the same way — **20 GB, bootable, sourced from `Debian13`** — which is not the image either producer boots from. Something is provisioning an undeclared 20 GiB volume roughly once per CI run. ## Why file it when the sweep already reclaims them Not cost: the sweep takes each one an hour later and Cinder usage is stable at 440 GB. Two other reasons: 1. **Undeclared provisioning is unexplained behaviour**, and the volumes are bootable and unnamed — the same shape as the class that exhausted the volume quota and stopped CI in #133. If the sweep ever regresses, this refills the quota on its own. 2. **This exact mismatch already misdirected an investigation.** #320 reasoned the orphans could not be runner volumes *because* the controller declares 40 GB and the orphans were 20 GB, and named the weekly backup-drill VM as the likeliest producer on cadence grounds. The per-CI-run rhythm above rules the drill out. A declared-vs-actual gap that survives is a trap for the next person. ## Suggested diagnostic The volumes exist for a full hour before the sweep takes them, so there is a comfortable window to catch one live. While `gitborg_runner_controller_os_volumes_available{class="sweepable"}` is 1: ```bash openstack volume list --status available -f value -c ID openstack volume show <id> -f json # volume_image_metadata.image_name, bootable, size, created_at ``` Then correlate against the runner VM active at that timestamp — specifically whether its *own* declared boot volume (40 GB, `delete_on_termination=true`) cascaded correctly, i.e. whether this is a **second** volume rather than the VM's own. #320's disclosure hypothesised exactly that and is worth testing first: *"That VM's declared boot volume cascaded correctly on `delete_on_termination=true`, so how a second one appeared is unexplained — possibly a create that Nova rejected after Cinder had already provisioned."* If Nova rejects a create after Cinder has provisioned, the abandoned volume would be image-sized rather than request-sized, which would explain 20 GiB against a 40 GB request. Refs #320, #133.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#372
Ingen beskrivning angiven.