docs(runbook): record the #320 prune and the snapshot-quota constraint #367
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!367
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "docs/runner-image-prune-320"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Documentation only — the prune itself has already run. Records what #320 actually cost and the
constraint that made it urgent.
The alert was framed wrong
RunnerSnapshotPinnedVolumesHighreads as a disk-waste warning, which is why it was tolerable at 6generations. Disk was never the constraint:
gigabytesquota is 5000 with ~460 in use.The binding limit is the snapshot COUNT quota of 10, and every bake consumes one. The project sat
at 8/10 — six bake generations plus the two
predrill-*snapshots — so it was two bakes awayfrom
bake-runner-image.shfailing outright at snapshot creation. Nothing alerts on that, and therunbook did not mention it. Now recorded, with the commands to check headroom before a bake.
The prune
6 generations → 2 (live
4daabf1b, which the Glance image boots from, plus rollbacka500b034).Reclaimed 4 snapshots and 80 GiB; snapshot usage 8 → 4. Verified afterwards: the image is still
activewith its snapshotavailable, all five named production volumes untouched and attached, andboth
RunnerSnapshotPinnedVolumesHighandRunnerOrphanVolumesUnsweptcleared on the nextevaluation — only
Watchdogremains, as intended.Two leftovers the prune can never reach
Its filter is
snapshot for gitborg-runner*, sopredrill-backup-20260709(100 GiB) andpredrill-data-20260709(60 GiB) from the 2026-07-09 drill rehearsal are invisible to it — 160 GiBand 2 of the 10 slots. No Glance image references either. Left in place pending a decision on
whether they are a retained recovery point; recorded so they are not rediscovered from scratch.
A trap worth knowing next time
The runner-controller's orphan sweep deletes hour-old CI boot volumes concurrently with a prune,
so the volume count moves for two independent reasons. During this run the count went 12 → 7 while the
prune accounted for only 4 of the 5 deletions — the fifth was the sweep reclaiming a transient CI boot
volume. Reconcile against the named production volumes, not against a total.
Scope grew by one commit: the two
predrill-*snapshots have now been deleted as well, so the note saying they were left in place is superseded by2fcb107.Final state: snapshot usage 2/10 (was 8/10) — only the two retained
gitborg-runnerbake generations remain. 240 GiB of snapshots reclaimed in total across both steps.Before deleting them I established that the recovery point was superseded rather than assuming it: resolved each
volume_id(they were snapshots ofgitborg-prod-backupandgitborg-prod-data), confirmed no Glance image boots from either, and required all five backup signals to be0and recent — nightly archive 5.2 h, restic tier 3.8 h, both off-site destinations OK, weekly structural verify 2.2 d, and the DR drill 3.1 d. The drill is the one that counts, since it boots a VM, restores the latest off-site archive and runsforgejo doctor.Verified after: both source volumes still
in-useand attached,/srv/gitborg-dataand/srv/gitborg-backupmounted at expected sizes, 10 containers running, 5 core units active, 6 probes green, onlyWatchdogfiring.The added commit also records that these snapshots were not blocking anything —
gitborg-prod-backupwas grown 100 → 150 GiB while its snapshot existed, so Cinder permitsvolume extendon a snapshotted volume here.