fix(runner-controller): sweep leaked boot volumes + boot-failure/quota alerts (#133) #134

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/runner-controller-volume-leak in i main 2026-07-19 08:16:31 +00:00
Ägare

Permanent fix for the CI outage in #133 (leaked ephemeral boot volumes exhausted the Cinder quota → every runner booted to ERROR, silently).

Root cause

create_server(boot_from_volume=True, terminate_volume=True) auto-creates the boot volume; terminate_volume only cascades it on a normal termination. A VM that dies in ERROR during build never has the volume attached, so _delete_server_and_resources (which enumerates server.volumes) can't see it → the volume orphans as detached/available. 227 leaked (≈4540 GiB) and filled the quota; after that every boot ERRORed for lack of a boot volume — self-reinforcing, and the reconcile loop still reported last_loop_ok=1, so nothing alerted.

Changes (controller.py)

  • _sweep_orphan_volumes() — each cycle, delete detached boot-volume leaks. Three independent guards make it impossible to touch a live disk: status == available (detached), blank name (Nova boot volumes are unnamed; every persistent bitborg volume is named), and age > orphan_volume_max_age_seconds (default 1h, so a volume mid-attach is never caught). Respects DRY_RUN.
  • New metrics: boot_error_vms, orphan_volumes_swept, and best-effort Cinder os_volume_gb_used / os_volume_gb_quota.
  • Config knob orphan_volume_max_age_seconds.

Alerts (monitoring) — so this never goes silent again

  • RunnerBootFailing — boot_error_vms > 0 for 15m (a healthy pool reaps a transient ERROR within a cycle; sustained = CI not running).
  • RunnerVolumeQuotaHigh — Cinder volume quota > 80% (alert_runner_volume_quota_pct), guarded against the -1 unknown sentinel.

Validation

python3 -m py_compile clean; alert template renders to valid YAML (27 alerts, both new ones present); ansible-lint roles/runner-controller roles/monitoring passes (production profile). Not applied — needs the controller image rebuilt + site.yml apply (yours to run). Does not clean the existing 227 legacy orphans (they're unnamed and pre-date this) — those are handled by the manual runbook on #133; the sweep prevents recurrence.

Future hardening (noted in #133): tag boot volumes gitborg_ephemeral=true at create and match the sweep on the tag instead of the blank-name heuristic.

Permanent fix for the CI outage in #133 (leaked ephemeral boot volumes exhausted the Cinder quota → every runner booted to ERROR, silently). ### Root cause `create_server(boot_from_volume=True, terminate_volume=True)` auto-creates the boot volume; `terminate_volume` only cascades it on a **normal** termination. A VM that dies in `ERROR` during build never has the volume attached, so `_delete_server_and_resources` (which enumerates `server.volumes`) can't see it → the volume orphans as detached/`available`. 227 leaked (≈4540 GiB) and filled the quota; after that every boot ERRORed for lack of a boot volume — self-reinforcing, and the reconcile loop still reported `last_loop_ok=1`, so nothing alerted. ### Changes (`controller.py`) - **`_sweep_orphan_volumes()`** — each cycle, delete detached boot-volume leaks. Three independent guards make it impossible to touch a live disk: `status == available` (detached), **blank name** (Nova boot volumes are unnamed; every persistent bitborg volume is named), and **age > `orphan_volume_max_age_seconds`** (default 1h, so a volume mid-attach is never caught). Respects `DRY_RUN`. - New metrics: `boot_error_vms`, `orphan_volumes_swept`, and best-effort Cinder `os_volume_gb_used` / `os_volume_gb_quota`. - Config knob `orphan_volume_max_age_seconds`. ### Alerts (`monitoring`) — so this never goes silent again - **RunnerBootFailing** — `boot_error_vms > 0` for 15m (a healthy pool reaps a transient ERROR within a cycle; sustained = CI not running). - **RunnerVolumeQuotaHigh** — Cinder volume quota > 80% (`alert_runner_volume_quota_pct`), guarded against the `-1` unknown sentinel. ### Validation `python3 -m py_compile` clean; alert template renders to valid YAML (27 alerts, both new ones present); `ansible-lint roles/runner-controller roles/monitoring` passes (production profile). Not applied — needs the controller image rebuilt + `site.yml` apply (yours to run). Does **not** clean the existing 227 legacy orphans (they're unnamed and pre-date this) — those are handled by the manual runbook on #133; the sweep prevents recurrence. **Future hardening (noted in #133):** tag boot volumes `gitborg_ephemeral=true` at create and match the sweep on the tag instead of the blank-name heuristic.
supernaut lade till 1 incheckning 2026-07-19 08:07:45 +00:00
fix(runner-controller): sweep leaked boot volumes + alert on boot failures (#133)
Alla kontroller lyckades
ci / ci (pull_request) Successful in 2m35s
190af0bf3e
Ephemeral runner VMs that die in ERROR during build leave an unattached boot
volume behind: create_server(terminate_volume=True) only cascades the volume on
a NORMAL termination, and the reap enumerates server.volumes (empty for an
ERROR-boot server), so the volume orphans as detached/available. Enough leaked to
exhaust the Cinder volume quota, after which every runner boot ERRORed for lack of
a boot volume — CI down, and silent (the loop still reported last_loop_ok=1).

- add _sweep_orphan_volumes(): each cycle, delete detached (available), unnamed,
  aged-out boot volumes. Three guards (available + blank name + age) make it
  impossible to touch a persistent volume (those are in-use and named).
- export boot_error_vms, orphan_volumes_swept, and best-effort Cinder volume
  quota used/limit gauges.
- alerts RunnerBootFailing (boot_error_vms>0 for 15m) + RunnerVolumeQuotaHigh
  (>80%) so this pages instead of failing silently.
- config knob orphan_volume_max_age_seconds (default 1h).
supernaut sammanfogade incheckning 91f40f7b32 till main 2026-07-19 08:16:31 +00:00
supernaut tog bort grenen fix/runner-controller-volume-leak 2026-07-19 08:16:31 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!134
Ingen beskrivning angiven.