fix(backup-drill): guard shared-volume + size drill VM for bundle+extract (#131) #143

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/131-drill-redesign in i main 2026-07-19 13:23:45 +00:00
Ägare

Fixes #131 properly after #130 (40 GB) proved insufficient and triggering the drill spiked the shared prod backup volume 81%→97% (DiskUsageCritical) — root cause: the orchestrator fetches the archive and decrypts a ~3× larger bundle into WORKDIR on /srv/gitborg-backup.

Changes (orchestrator + defaults)

  • Pre-flight free-space guard — abort before fetching if WORKDIR's volume has < backup_drill_min_free_gb (25) free. A drill can now never fill a shared prod volume; it fails safe instead. (This is the guarantee that the 2026-07-19 incident can't recur.)
  • Delete the encrypted archive right after decrypt — only the bundle ships to the VM; keeping both ~doubled the WORKDIR peak.
  • Drill VM boot volume 40 → 60 GB — the VM holds the ~15 GB bundle and the ~21 GB extraction at once (>40 GB peak → the ENOSPC).

Validated: rendered orchestrator py_compiles; ansible-lint roles/backup-drill passes (production).

⚠️ Operator sequence (this PR alone is not enough to make the drill pass)

The guard needs headroom to clear — with the backup volume at 81% (~19 GB free), the guard will (correctly) abort the drill. To get a passing drill:

  1. Merge this PR, ansible-playbook site.yml --tags backup-drill (installs guard + 60 GB VM).
  2. Grow the backup volume backup_volume_size 100 → 150 in terraform.tfvars (gitignored — local edit), tofu apply, then resize2fs /dev/vdc — this is PR #26's backup-headroom item too. Now WORKDIR clears the 25 GB floor with margin.
  3. Re-enable the drill timer (systemctl --user enable --now bitborg-backup-drill.timer) — I stopped+disabled it as the incident stopgap.
  4. Trigger one drill; confirm forgejo doctor passes + BackupDrillFailed clears.

I won't run the drill or apply anything here — this is review-only, and the timer stays disabled until you've done the above.

Future hardening (not here): stream the decrypt straight to the VM (age -d | ssh 'cat > bundle') to eliminate the local bundle entirely, and derive the VM disk size from the live data volume (#131 box 3).

Fixes #131 properly after #130 (40 GB) proved insufficient **and** triggering the drill spiked the shared prod backup volume 81%→97% (DiskUsageCritical) — root cause: the orchestrator fetches the archive and decrypts a ~3× larger bundle into `WORKDIR` on `/srv/gitborg-backup`. ### Changes (orchestrator + defaults) - **Pre-flight free-space guard** — abort *before* fetching if `WORKDIR`'s volume has `< backup_drill_min_free_gb` (25) free. A drill can now **never** fill a shared prod volume; it fails safe instead. (This is the guarantee that the 2026-07-19 incident can't recur.) - **Delete the encrypted archive right after decrypt** — only the bundle ships to the VM; keeping both ~doubled the `WORKDIR` peak. - **Drill VM boot volume 40 → 60 GB** — the VM holds the ~15 GB bundle *and* the ~21 GB extraction at once (>40 GB peak → the ENOSPC). Validated: rendered orchestrator `py_compile`s; `ansible-lint roles/backup-drill` passes (production). ### ⚠️ Operator sequence (this PR alone is not enough to make the drill pass) The guard needs headroom to clear — with the backup volume at 81% (~19 GB free), the guard will (correctly) **abort** the drill. To get a passing drill: 1. Merge this PR, `ansible-playbook site.yml --tags backup-drill` (installs guard + 60 GB VM). 2. **Grow the backup volume** `backup_volume_size` 100 → 150 in `terraform.tfvars` (gitignored — local edit), `tofu apply`, then `resize2fs /dev/vdc` — this is PR #26's backup-headroom item too. Now `WORKDIR` clears the 25 GB floor with margin. 3. Re-enable the drill timer (`systemctl --user enable --now bitborg-backup-drill.timer`) — I stopped+disabled it as the incident stopgap. 4. Trigger one drill; confirm `forgejo doctor` passes + `BackupDrillFailed` clears. **I won't run the drill or apply anything here** — this is review-only, and the timer stays disabled until you've done the above. Future hardening (not here): stream the decrypt straight to the VM (`age -d | ssh 'cat > bundle'`) to eliminate the local bundle entirely, and derive the VM disk size from the live data volume (#131 box 3).
supernaut lade till 1 incheckning 2026-07-19 13:20:17 +00:00
The 40 GB VM bump (#130) was insufficient and, worse, running the drill spiked the
shared prod backup volume 81%->97% (DiskUsageCritical): the orchestrator fetches the
archive and decrypts a ~3x-larger bundle into WORKDIR on /srv/gitborg-backup.

- Pre-flight free-space guard: abort BEFORE fetching if WORKDIR's volume has
  < backup_drill_min_free_gb (25) free, so a drill can never push a shared prod
  volume to critical — a failed drill is acceptable, a disk-full prod is not.
- Delete the encrypted archive right after decrypt (only the bundle ships to the VM)
  to roughly halve the WORKDIR peak footprint.
- Drill VM boot volume 40 -> 60 GB: it holds the ~15 GB bundle AND the ~21 GB
  extraction at once (>40 GB peak → the ENOSPC we saw).

Rendered orchestrator py_compiles; ansible-lint (production) clean. NOT applied; drill
timer stays disabled until this lands + the backup volume is grown (see PR body).
supernaut sammanfogade incheckning 44260eb81e till main 2026-07-19 13:23:45 +00:00
supernaut tog bort grenen fix/131-drill-redesign 2026-07-19 13:23:46 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!143
Ingen beskrivning angiven.