Each runner image rebake strands a 20 GiB boot volume and its snapshot #340

Stängd
öppnade 2026-08-02 12:49:47 +00:00 av supernaut · 1 kommentar
Ägare

Problem

Every rebake of the gitborg-runner Glance image permanently strands a 20 GiB Cinder volume plus
its 20 GiB snapshot
. Neither is ever reclaimed, and the pair is invisible to the orphan sweep by
design.

Baking the image snapshots the staging VM's boot volume. The snapshot keeps a reference to that
volume, so Cinder refuses to delete it — and the sweep in
ansible/roles/runner-controller/files/controller.py correctly skips any volume that has snapshots,
because that heuristic is what stops it destroying image and backup infrastructure (added in #133).

The sweep is behaving exactly as designed. Nothing here is a sweep defect. The gap is that
nothing owns the other end: the bake process never cleans up after itself.

Evidence

Nine volumes are currently available (detached), all 20 GiB and bootable. Three were created
today and are simply inside the sweep's age window — they will be reclaimed normally. The other six
are permanent, and each pairs with an image-bake snapshot:

Volume Created Paired snapshot
f1ce224d 2026-07-29 21:09 snapshot for gitborg-runner-staging-20260729-230936
7b4588d3 2026-07-22 14:12 snapshot for gitborg-runner-staging-20260722-161157
e8dd5800 2026-07-20 17:58 snapshot for gitborg-runner-staging-20260720-195756
4d4a7cd2 2026-07-20 17:40 snapshot for gitborg-runner-staging-20260720-194019
40781d11 2026-06-28 22:04 snapshot for gitborg-runner
e610f170 2026-06-28 14:41 snapshot for gitborg-runner

The controller logs the refusal on every restart, and has done since at least 2026-07-29:

sweep: skipping volume e610f170-… — it has snapshot(s), so it is image/backup infra,
not an ephemeral leak (will not retry this session)

Current block-storage consumption is 500 GiB across 14 volumes and 280 GiB across 8 snapshots.
The six stranded pairs account for roughly 200 GiB of that — about a quarter of all allocated block
storage, holding nothing.

Why it matters

This is the same quota that caused the CI outage in #133: once volume-gigabytes is exhausted, every
runner boot fails for lack of a boot volume, self-reinforcingly and silently. The growth rate here
is slow — one pair per rebake — but it is unbounded and it consumes the exact resource that has
already taken CI down once.

Careful — one pair is live

The current Glance image boots from snapshot_id 4daabf1b-…, which is
snapshot for gitborg-runner-staging-20260729-230936 — the 2026-07-29 row above. That snapshot
and its volume f1ce224d are load-bearing and must not be deleted. Deleting them would leave the
image unbootable and break all CI.

The five older pairs are safe to reclaim.

Proposed work

  1. Reclaim the five stale pairs, snapshot first then volume, verifying against the live image's
    block_device_mapping immediately beforehand rather than trusting the table above. Roughly
    200 GiB returned.
  2. Make the bake clean up after itself. After the Glance image is created and verified, delete
    the staging VM's volume and the intermediate snapshot, keeping only what the image references.
    Whichever step bakes the image should own this — a leak that needs a human to notice is not
    fixed.
  3. Retain the previous image's pair deliberately, not accidentally. Keeping exactly one
    generation back is a reasonable rollback story; keeping five is an accident. Make the retention
    count explicit.
  4. Surface it. gitborg_runner_controller_os_volume_gb_* is already exported and dashboarded.
    Add a gauge for detached-but-unsweepable volumes, so the next thing that accumulates behind the
    snapshot heuristic is visible rather than discovered by hand.

Also worth a look while in here

Two unrelated snapshots, predrill-backup-20260709 (100 GiB) and predrill-data-20260709
(60 GiB), have been retained since 2026-07-09. If the restore drill they belong to is long finished,
that is another 160 GiB. Out of scope for this issue, but it is the same failure shape: a safety
artefact that nothing is responsible for removing.

Acceptance

  • The five stale volume/snapshot pairs are gone and block-storage consumption drops by ~200 GiB.
  • A rebake leaves behind only the pair the new image references, plus at most one retained
    generation.
  • Detached volumes the sweep cannot reclaim are visible on the runners dashboard.
  • The runbook says which snapshot the live image depends on and how to check before deleting.
## Problem Every rebake of the `gitborg-runner` Glance image permanently strands a **20 GiB Cinder volume plus its 20 GiB snapshot**. Neither is ever reclaimed, and the pair is invisible to the orphan sweep by design. Baking the image snapshots the staging VM's boot volume. The snapshot keeps a reference to that volume, so Cinder refuses to delete it — and the sweep in `ansible/roles/runner-controller/files/controller.py` correctly skips any volume that has snapshots, because that heuristic is what stops it destroying image and backup infrastructure (added in #133). **The sweep is behaving exactly as designed. Nothing here is a sweep defect.** The gap is that nothing owns the *other* end: the bake process never cleans up after itself. ## Evidence Nine volumes are currently `available` (detached), all 20 GiB and bootable. Three were created today and are simply inside the sweep's age window — they will be reclaimed normally. The other six are permanent, and each pairs with an image-bake snapshot: | Volume | Created | Paired snapshot | | --- | --- | --- | | `f1ce224d` | 2026-07-29 21:09 | `snapshot for gitborg-runner-staging-20260729-230936` | | `7b4588d3` | 2026-07-22 14:12 | `snapshot for gitborg-runner-staging-20260722-161157` | | `e8dd5800` | 2026-07-20 17:58 | `snapshot for gitborg-runner-staging-20260720-195756` | | `4d4a7cd2` | 2026-07-20 17:40 | `snapshot for gitborg-runner-staging-20260720-194019` | | `40781d11` | 2026-06-28 22:04 | `snapshot for gitborg-runner` | | `e610f170` | 2026-06-28 14:41 | `snapshot for gitborg-runner` | The controller logs the refusal on every restart, and has done since at least 2026-07-29: ```text sweep: skipping volume e610f170-… — it has snapshot(s), so it is image/backup infra, not an ephemeral leak (will not retry this session) ``` Current block-storage consumption is **500 GiB across 14 volumes and 280 GiB across 8 snapshots**. The six stranded pairs account for roughly 200 GiB of that — about a quarter of all allocated block storage, holding nothing. ## Why it matters This is the same quota that caused the CI outage in #133: once volume-gigabytes is exhausted, every runner boot fails for lack of a boot volume, self-reinforcingly and silently. The growth rate here is slow — one pair per rebake — but it is unbounded and it consumes the exact resource that has already taken CI down once. ## Careful — one pair is live The current Glance image boots from `snapshot_id 4daabf1b-…`, which is `snapshot for gitborg-runner-staging-20260729-230936` — the **2026-07-29** row above. That snapshot and its volume `f1ce224d` are load-bearing and must not be deleted. Deleting them would leave the image unbootable and break all CI. The five older pairs are safe to reclaim. ## Proposed work 1. **Reclaim the five stale pairs**, snapshot first then volume, verifying against the live image's `block_device_mapping` immediately beforehand rather than trusting the table above. Roughly 200 GiB returned. 2. **Make the bake clean up after itself.** After the Glance image is created and verified, delete the staging VM's volume and the intermediate snapshot, keeping only what the image references. Whichever step bakes the image should own this — a leak that needs a human to notice is not fixed. 3. **Retain the previous image's pair deliberately**, not accidentally. Keeping exactly one generation back is a reasonable rollback story; keeping five is an accident. Make the retention count explicit. 4. **Surface it.** `gitborg_runner_controller_os_volume_gb_*` is already exported and dashboarded. Add a gauge for detached-but-unsweepable volumes, so the next thing that accumulates behind the snapshot heuristic is visible rather than discovered by hand. ## Also worth a look while in here Two unrelated snapshots, `predrill-backup-20260709` (100 GiB) and `predrill-data-20260709` (60 GiB), have been retained since 2026-07-09. If the restore drill they belong to is long finished, that is another 160 GiB. Out of scope for this issue, but it is the same failure shape: a safety artefact that nothing is responsible for removing. ## Acceptance - [ ] The five stale volume/snapshot pairs are gone and block-storage consumption drops by ~200 GiB. - [ ] A rebake leaves behind only the pair the new image references, plus at most one retained generation. - [ ] Detached volumes the sweep cannot reclaim are visible on the runners dashboard. - [ ] The runbook says which snapshot the live image depends on and how to check before deleting.
Upphovsperson
Ägare

Duplicate of #320, which predates this and already carries the accumulation evidence and the metric-sensitivity analysis. The root-cause analysis from this issue has been moved to #320, where it answers that issue's open "identify the producer" step. Closing in favour of #320.

Duplicate of #320, which predates this and already carries the accumulation evidence and the metric-sensitivity analysis. The root-cause analysis from this issue has been moved to #320, where it answers that issue's open "identify the producer" step. Closing in favour of #320.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#340
Ingen beskrivning angiven.