docs(runbook): record the #320 prune and the snapshot-quota constraint #367

Sammanfogat
supernaut sammanfogade 2 incheckningar från docs/runner-image-prune-320 in i main 2026-08-04 08:53:52 +00:00
Ägare

Documentation only — the prune itself has already run. Records what #320 actually cost and the
constraint that made it urgent.

The alert was framed wrong

RunnerSnapshotPinnedVolumesHigh reads as a disk-waste warning, which is why it was tolerable at 6
generations. Disk was never the constraint: gigabytes quota is 5000 with ~460 in use.

The binding limit is the snapshot COUNT quota of 10, and every bake consumes one. The project sat
at 8/10 — six bake generations plus the two predrill-* snapshots — so it was two bakes away
from bake-runner-image.sh failing outright
at snapshot creation. Nothing alerts on that, and the
runbook did not mention it. Now recorded, with the commands to check headroom before a bake.

The prune

6 generations → 2 (live 4daabf1b, which the Glance image boots from, plus rollback a500b034).
Reclaimed 4 snapshots and 80 GiB; snapshot usage 8 → 4. Verified afterwards: the image is still
active with its snapshot available, all five named production volumes untouched and attached, and
both RunnerSnapshotPinnedVolumesHigh and RunnerOrphanVolumesUnswept cleared on the next
evaluation — only Watchdog remains, as intended.

Two leftovers the prune can never reach

Its filter is snapshot for gitborg-runner*, so predrill-backup-20260709 (100 GiB) and
predrill-data-20260709 (60 GiB) from the 2026-07-09 drill rehearsal are invisible to it — 160 GiB
and 2 of the 10 slots
. No Glance image references either. Left in place pending a decision on
whether they are a retained recovery point; recorded so they are not rediscovered from scratch.

A trap worth knowing next time

The runner-controller's orphan sweep deletes hour-old CI boot volumes concurrently with a prune,
so the volume count moves for two independent reasons. During this run the count went 12 → 7 while the
prune accounted for only 4 of the 5 deletions — the fifth was the sweep reclaiming a transient CI boot
volume. Reconcile against the named production volumes, not against a total.

Documentation only — the prune itself has already run. Records what #320 actually cost and the constraint that made it urgent. ### The alert was framed wrong `RunnerSnapshotPinnedVolumesHigh` reads as a disk-waste warning, which is why it was tolerable at 6 generations. Disk was never the constraint: `gigabytes` quota is 5000 with ~460 in use. The binding limit is the **snapshot COUNT quota of 10**, and every bake consumes one. The project sat at **8/10** — six bake generations plus the two `predrill-*` snapshots — so it was **two bakes away from `bake-runner-image.sh` failing outright** at snapshot creation. Nothing alerts on that, and the runbook did not mention it. Now recorded, with the commands to check headroom before a bake. ### The prune 6 generations → 2 (live `4daabf1b`, which the Glance image boots from, plus rollback `a500b034`). Reclaimed 4 snapshots and 80 GiB; snapshot usage 8 → 4. Verified afterwards: the image is still `active` with its snapshot `available`, all five named production volumes untouched and attached, and both `RunnerSnapshotPinnedVolumesHigh` and `RunnerOrphanVolumesUnswept` cleared on the next evaluation — only `Watchdog` remains, as intended. ### Two leftovers the prune can never reach Its filter is `snapshot for gitborg-runner*`, so `predrill-backup-20260709` (100 GiB) and `predrill-data-20260709` (60 GiB) from the 2026-07-09 drill rehearsal are invisible to it — **160 GiB and 2 of the 10 slots**. No Glance image references either. Left in place pending a decision on whether they are a retained recovery point; recorded so they are not rediscovered from scratch. ### A trap worth knowing next time The runner-controller's orphan sweep deletes hour-old CI boot volumes **concurrently** with a prune, so the volume count moves for two independent reasons. During this run the count went 12 → 7 while the prune accounted for only 4 of the 5 deletions — the fifth was the sweep reclaiming a transient CI boot volume. Reconcile against the named production volumes, not against a total.
supernaut lade till 1 incheckning 2026-08-04 08:27:07 +00:00
docs(runbook): record the #320 prune and the snapshot-quota constraint
Alla kontroller lyckades
ci / ci (pull_request) Successful in 17s
8f0fffc5d4
Pruned 2026-08-04: 6 `gitborg-runner` bake generations → 2, reclaiming 4 Cinder snapshots and 80 GiB.
`RunnerSnapshotPinnedVolumesHigh` and `RunnerOrphanVolumesUnswept` both cleared on the next
evaluation.

The alert that surfaced this reads as a disk-waste warning, which is the wrong mental model and the
reason it sat at 6 generations. Disk was never the constraint — `gigabytes` is 5000 with ~460 in use.
The binding limit is the **snapshot COUNT quota of 10**, and every bake consumes one: the project was
at 8/10, i.e. two bakes away from `bake-runner-image.sh` failing at snapshot creation, with no metric
or alert covering that at all. Recorded with the commands to check headroom.

Also recorded two snapshots the prune can never reach, because its filter is
`snapshot for gitborg-runner*`: the `predrill-*` pair from the 2026-07-09 drill rehearsal, holding
160 GiB and 2 of the 10 slots. No image references either. Left in place pending a decision on
whether they are a retained recovery point.

And a trap for next time: the runner-controller's orphan sweep deletes hour-old CI boot volumes
concurrently with a prune, so the volume count moves for two independent reasons — reconcile against
the named production volumes, not against a total. During this prune the count went 12 → 7 while the
prune itself accounted for only 4 of the 5 deletions.
supernaut lade till 1 incheckning 2026-08-04 08:49:48 +00:00
docs(runbook): the predrill snapshots are deleted, with the check that justified it
Alla kontroller lyckades
ci / ci (pull_request) Successful in 17s
2fcb10728b
Both `predrill-*` snapshots from the 2026-07-09 drill rehearsal are gone (160 GiB, 2 snapshot slots);
snapshot usage is now 2/10. Supersedes the "left in place pending a decision" note earlier on this
branch.

Records the diligence rather than just the outcome, because the next person deleting a point-in-time
snapshot of a production volume needs the procedure, not the verdict: resolve `volume_id` to find what
it is a snapshot OF, check no Glance image boots from it, then require all five backup signals to be
`0` and recent — with the DR drill weighted highest, since it boots a VM, restores the latest off-site
archive and runs `forgejo doctor`, making it proof of restorability rather than of presence. Here that
was ~26 nights of age-encrypted archives on two off-site providers plus four passing drills since the
snapshots were taken.

Also corrects a plausible misconception worth having on record: these snapshots were not blocking
anything. `gitborg-prod-backup` was grown 100 -> 150 GiB while its snapshot existed, so Cinder permits
`volume extend` on a snapshotted volume here, and deleting a snapshot leaves its source volume alone —
both stayed `in-use` and attached throughout.
Upphovsperson
Ägare

Scope grew by one commit: the two predrill-* snapshots have now been deleted as well, so the note saying they were left in place is superseded by 2fcb107.

Final state: snapshot usage 2/10 (was 8/10) — only the two retained gitborg-runner bake generations remain. 240 GiB of snapshots reclaimed in total across both steps.

Before deleting them I established that the recovery point was superseded rather than assuming it: resolved each volume_id (they were snapshots of gitborg-prod-backup and gitborg-prod-data), confirmed no Glance image boots from either, and required all five backup signals to be 0 and recent — nightly archive 5.2 h, restic tier 3.8 h, both off-site destinations OK, weekly structural verify 2.2 d, and the DR drill 3.1 d. The drill is the one that counts, since it boots a VM, restores the latest off-site archive and runs forgejo doctor.

Verified after: both source volumes still in-use and attached, /srv/gitborg-data and /srv/gitborg-backup mounted at expected sizes, 10 containers running, 5 core units active, 6 probes green, only Watchdog firing.

The added commit also records that these snapshots were not blocking anything — gitborg-prod-backup was grown 100 → 150 GiB while its snapshot existed, so Cinder permits volume extend on a snapshotted volume here.

Scope grew by one commit: the two `predrill-*` snapshots have now been **deleted** as well, so the note saying they were left in place is superseded by 2fcb107. Final state: snapshot usage **2/10** (was 8/10) — only the two retained `gitborg-runner` bake generations remain. 240 GiB of snapshots reclaimed in total across both steps. Before deleting them I established that the recovery point was superseded rather than assuming it: resolved each `volume_id` (they were snapshots of `gitborg-prod-backup` and `gitborg-prod-data`), confirmed no Glance image boots from either, and required all five backup signals to be `0` and recent — nightly archive 5.2 h, restic tier 3.8 h, both off-site destinations OK, weekly structural verify 2.2 d, and the DR drill 3.1 d. The drill is the one that counts, since it boots a VM, restores the latest off-site archive and runs `forgejo doctor`. Verified after: both source volumes still `in-use` and attached, `/srv/gitborg-data` and `/srv/gitborg-backup` mounted at expected sizes, 10 containers running, 5 core units active, 6 probes green, only `Watchdog` firing. The added commit also records that these snapshots were **not** blocking anything — `gitborg-prod-backup` was grown 100 → 150 GiB while its snapshot existed, so Cinder permits `volume extend` on a snapshotted volume here.
supernaut sammanfogade incheckning 3308b4f0e0 till main 2026-08-04 08:53:52 +00:00
supernaut tog bort grenen docs/runner-image-prune-320 2026-08-04 08:53:52 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!367
Ingen beskrivning angiven.