Restore drill fails with ENOSPC — drill VM boot volume (20 GB) < restored dataset (~25 GB) #131
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#131
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Severity: HIGH — surfaced by the OpenTofu infra audit (2026-07-19).
BackupDrillFailed(critical) is firing.The weekly restore drill has been failing every run:
The throwaway drill VM boots on a 20 GB Cinder volume (
backup_drill_os_boot_volume_size), but the restored dataset now unpacks to ~25 GB in container storage (/srv/gitborg-datais ~24.6 GB and growing). The restore hits ENOSPC beforeforgejo doctorruns — so there is currently no verified proof backups can be restored, right before onboarding (ties to the audit's H6/M10 DR-proof gaps).Done / remaining
backup_drill_os_boot_volume_size20 → 40 GB — PR #130.forgejo doctorgreen,BackupDrillFailedclears). — verified 2026-07-19; needed PR #143 (guard + 60 GB) and PR #144 (proportional staging + 80 GB) beyond #130. See the verification comment below.Blocks confidence in the DR path before onboarding.
Reopening scope — the 40 GB fix (#130) is INSUFFICIENT, and running the drill caused a transient prod disk-critical
Verified by triggering a drill run after applying #130 (2026-07-19):
no space left on devicewriting packages on the drill VM (exit 125), even at 40 GB. The VM holds the transferred bundle and the extracted ~21 GB dataset simultaneously, so peak usage exceeds 40 GB.BackupDrillFailedstill firing; DR path still unproven./srv/gitborg-backup/drill— on a volume already at 81% that crossed the 90%DiskUsageCriticalthreshold (81→89→92→96.8%), then recovered to 81% on teardown. So the drill is actively dangerous to run while the backup volume is near-full.Stopgap applied:
bitborg-backup-drill.timerstopped + disabled so it can't auto-fire and re-spike. (Asite.ymlapply would re-enable it — gate it in group_vars if it must survive an apply.) Do not run the drill until redesigned.Redesign needed (both must be fixed)
backup_volume_size100→150 first (ties to #126 / PR #26 item 2).Root cause of the disk-critical event was my triggering the drill without accounting for its decrypt-staging on the near-full backup volume — my mistake.
Redesign up as PR #143 (review-only; not applied, drill timer stays disabled):
WORKDIR's volume has <backup_drill_min_free_gb(25 GB) free, so it can never fill a shared prod volume again (fixes the incident class).WORKDIRpeak.Operator sequence to get a passing drill (in PR body): merge #143 → apply
--tags backup-drill→ growbackup_volume_size100→150 (so the 25 GB guard clears with margin; = PR #26 backup-headroom item) → re-enable the timer → run + verify. Box 3 (derive VM disk from live data size / stream the decrypt to drop the local bundle) noted as future hardening.✅ Restore drill now PASSES — DR path verified end-to-end (2026-07-19 14:58Z).
Timeline after the original 20 GB ENOSPC:
restore-on-scratch.shkeepingbundle.tar+ the extractedforgejo-data.tar.gz+ the decompressed data volume alive at once (peak ≈ 2× compressed + uncompressed). Now frees each intermediate as it's consumed (peak proportional to the dataset) + VM boot volume 60 → 80 GB.Verified run (post-#144 apply):
gitborg_backup_drill_last_run_status = 0, fresh success timestampBackupDrillFailedcleared; backup volume back to 54% (66 GB free) — 7.6 GB staging cleaned up, no leakBox 2 (confirm next drill passes) ✅. Box 3 (auto-size the drill VM from live data growth) deferred — 80 GB now gives ample headroom over a 7.6 GB compressed bundle / ~21 GB live dataset with proportional staging; the
defaults/main.ymlcomment flags deriving from the data volume should it ever be outgrown again.