Ephemeral runners boot-loop to OpenStack ERROR → CI down, silent (no alert), burning volumes #133
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#133
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Severity: HIGH (CI fully down) + silent + burning OpenStack resources. Found investigating a stuck Actions run (run 142 queued, never starts).
Symptom
Every queued Actions job (
labels=ci) waits forever. Forgejo has runner registrations churning (RegisterRunner201 →DeleteRunner204 seconds later, orphaned).Root cause
The runner-controller is healthy and NOT in dry-run — it boots ephemeral VMs correctly, but every ephemeral runner VM goes to OpenStack
status=ERRORwithin seconds of boot and is reaped before it can register+run. Controller log pattern (live,container="runner-controller"):This repeats continuously — dozens of volume-backed VM boots per minute, all ERROR → real, ongoing Cinder/Nova churn (and likely orphaned boot volumes) with zero successful jobs.
The controller doesn't log the OpenStack fault, so the reason the VMs ERROR is not in Loki. Leading hypotheses (verify on host): Cinder volume quota / ceph-pool capacity exhausted (note the concurrent disk pressure — #126; and a boot-loop leaking ERROR'd volumes is self-reinforcing), OR Nova "No valid host"/capacity, OR a bad/missing
gitborg-runnerGlance image.Why it's silent (monitoring gap)
gitborg_runner_controller_last_loop_ok = 1andactive_vms= 1–4 — the controller counts a boot+reap cycle as a successful loop, so no alert fires even though no runner ever becomes healthy. There is no alert on "VMs booted but 0 reached a healthy/registered state" or on a boot-error rate.Diagnose (needs OpenStack API on the host)
Remediate
systemctl --user stop gitborg-runner-controller.timer(as bitborg) so it stops booting failing VMs / leaking volumes while diagnosing.server show … fault(raise quota / clean leaked volumes / repair image / capacity).ephemeral-runner-*volumes left by ERROR'd boots.active_healthy == 0over N minutes, or a boot-error-rate metric — so a 100%-ERROR loop pages instead of reading as "loop OK".Root cause CONFIRMED (via read-only OpenStack API,
--os-cloud bahnhof)Cinder volume quota is exhausted:
max_total_volume_gigabytes = 5000,total_gigabytes_used = 4990→ 10 GiB free.gitborg-prod-root40 GiB (from the 2026-07-09 root downsize).With <10 GiB free, no new runner boot volume can be created → every ephemeral runner boots to
status=ERROR→ CI down.Why the volumes leak (controller.py): boot uses
boot_from_volume + terminate_volume=True(L516-526), and reap enumeratesserver.volumesto delete them (L339-356). But a server that dies in ERROR during boot has a boot volume that was created-from-image but never fully attached, so it's absent fromserver.volumesandterminate_volume's delete-on-terminate never fires → the volume is orphaned as detached/"available". Self-reinforcing: each ERROR boot that got as far as creating its volume leaked one, until the quota filled.Remediation
--status availablefilter cannot touch the 4 in-use persistent volumes.server.volumes(empty for ERROR-boot servers); (c) explicitly delete the created volume whencreate_serverraises/ERRORs.last_loop_ok=1).Permanent fix drafted — PR #134 (not applied; needs the controller image rebuild +
site.ymlapply)._sweep_orphan_volumes()deletes detached/unnamed/aged-out boot-volume leaks each cycle (three guards → cannot touch a live disk); handles the ERROR-boot orphan the server-reap misses.boot_error_vms/orphan_volumes_swept/ Cinderos_volume_gb_used+_quota.last_loop_ok=1.py_compile + alert-template YAML render +
ansible-lint(production) all clean. PR #134 does not delete the existing 227 legacy orphans (unnamed, pre-date the tag/heuristic) — the manual runbook still needs running to reclaim the ~4540 GiB now; the sweep prevents recurrence going forward.Checklist:
site.ymlapply (prevents recurrence + adds alerts).Post-apply verification of #134 — found + fixed a defect (PR #135).
After
site.yml, the new controller image is live and healthy: metrics present (boot_error_vms=0,os_volume_gb_used=630,os_volume_gb_quota=5000,last_loop_ok=1), CI recovered (web Run 142 → success), quota reclaimed to 630/5000 GiB.BUT
orphan_volumes_sweptstayed 0 with 11 orphans present. Root cause: OpenStackcreated_atis timezone-naive (2026-07-19T08:07:59.000000);_iso_age_secondsraised on the naive-vs-aware subtraction and theexceptreturned0.0, so every volume read as 0s old → the sweep (and thereap_max_ageserver backstop) never fired. PR #135 stamps a missing tz as UTC. Verified against the real format (0s → 1917s).Remaining: merge #135 → rebuild image →
site.ymlapply. Then the sweep clears the 11 leftover orphans once they cross the 1h age threshold. (Harmless meanwhile — 630/5000 GiB used.)Validation after #135 apply — sweep confirmed working in prod ✅
boot_error_vms=0,last_loop_ok=1, quota gauges accurate; CI green (Run 142 success).One edge case found + fixed → PR #136
The sweep was hitting a 400 every 10s on two unnamed detached volumes: Cinder refuses to delete a volume that
must not have snapshots or belong to a group. Those two aren't leaks — they carry thegitborg-runnerimage source snapshot and thepredrill-*2026-07-09 safety snapshots (image/backup infra). The blank-name heuristic was too broad. PR #136 makes the sweep skip any volume with snapshots and skip-list any delete that's refused (no per-cycle retry/spam). Unit-tested with a stub connection.⚠️ Do NOT manually delete those two volumes or their
snapshot for gitborg-runnersnapshots — they back the runner image. (Thepredrill-*2026-07-09 snapshots are stale and optionally reclaimable, your call.)Status
Resolved + deployed + verified — closing.
The full remediation chain is merged and applied to prod (
site.yml, health-gate green):RunnerBootFailing/RunnerVolumeQuotaHighalerts (applied earlier).939026a2).Verified live after the apply (controller logs):
sweep: skipping volume …e610f170 / …40781d11 — it has snapshot(s), image/backup infra→ the twogitborg-runner/pre-drill volumes are skip-listed; the every-10s 400 spam stopped.sweep: deleted orphan boot volume …0bf861f4 (20 GiB, age=3604s)→ real leaks are cleaned as they age past 1h.boot_error_vms=0,last_loop_ok=1; detached volumes down 228 → 3 (the 2 infra + 1 recent), Cinder quota 4990 → 470 GiB.Original incident (ephemeral runners boot-looping to ERROR → CI down from Cinder-quota exhaustion) is fully resolved, recurrence is prevented, and it's now alertable instead of silent.