fix(runner-controller): sweep leaked boot volumes + boot-failure/quota alerts (#133) #134
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!134
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/runner-controller-volume-leak"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Permanent fix for the CI outage in #133 (leaked ephemeral boot volumes exhausted the Cinder quota → every runner booted to ERROR, silently).
Root cause
create_server(boot_from_volume=True, terminate_volume=True)auto-creates the boot volume;terminate_volumeonly cascades it on a normal termination. A VM that dies inERRORduring build never has the volume attached, so_delete_server_and_resources(which enumeratesserver.volumes) can't see it → the volume orphans as detached/available. 227 leaked (≈4540 GiB) and filled the quota; after that every boot ERRORed for lack of a boot volume — self-reinforcing, and the reconcile loop still reportedlast_loop_ok=1, so nothing alerted.Changes (
controller.py)_sweep_orphan_volumes()— each cycle, delete detached boot-volume leaks. Three independent guards make it impossible to touch a live disk:status == available(detached), blank name (Nova boot volumes are unnamed; every persistent bitborg volume is named), and age >orphan_volume_max_age_seconds(default 1h, so a volume mid-attach is never caught). RespectsDRY_RUN.boot_error_vms,orphan_volumes_swept, and best-effort Cinderos_volume_gb_used/os_volume_gb_quota.orphan_volume_max_age_seconds.Alerts (
monitoring) — so this never goes silent againboot_error_vms > 0for 15m (a healthy pool reaps a transient ERROR within a cycle; sustained = CI not running).alert_runner_volume_quota_pct), guarded against the-1unknown sentinel.Validation
python3 -m py_compileclean; alert template renders to valid YAML (27 alerts, both new ones present);ansible-lint roles/runner-controller roles/monitoringpasses (production profile). Not applied — needs the controller image rebuilt +site.ymlapply (yours to run). Does not clean the existing 227 legacy orphans (they're unnamed and pre-date this) — those are handled by the manual runbook on #133; the sweep prevents recurrence.Future hardening (noted in #133): tag boot volumes
gitborg_ephemeral=trueat create and match the sweep on the tag instead of the blank-name heuristic.