Restore drill fails with ENOSPC — drill VM boot volume (20 GB) < restored dataset (~25 GB) #131

Stängd
öppnade 2026-07-19 07:28:01 +00:00 av supernaut · 3 kommentarer
Ägare

Severity: HIGH — surfaced by the OpenTofu infra audit (2026-07-19). BackupDrillFailed (critical) is firing.

The weekly restore drill has been failing every run:

[drill] restore/forgejo-doctor FAILED on the drill VM (exit 125)
Error: write .../drill-forgejo-data/_data/gitea/data/packages/...: no space left on device

The throwaway drill VM boots on a 20 GB Cinder volume (backup_drill_os_boot_volume_size), but the restored dataset now unpacks to ~25 GB in container storage (/srv/gitborg-data is ~24.6 GB and growing). The restore hits ENOSPC before forgejo doctor runs — so there is currently no verified proof backups can be restored, right before onboarding (ties to the audit's H6/M10 DR-proof gaps).

Done / remaining

  • Bump backup_drill_os_boot_volume_size 20 → 40 GB — PR #130.
  • Apply the backup-drill role and confirm the next drill passes (forgejo doctor green, BackupDrillFailed clears). — verified 2026-07-19; needed PR #143 (guard + 60 GB) and PR #144 (proportional staging + 80 GB) beyond #130. See the verification comment below.
  • Make the drill boot-volume size track data growth (derive from data-volume usage, or alert when the drill VM disk headroom < live data size) so it doesn't silently outgrow again. — deferred hardening; 80 GB now leaves ample headroom (7.6 GB compressed bundle / ~21 GB live dataset with proportional staging).

Blocks confidence in the DR path before onboarding.

**Severity: HIGH** — surfaced by the OpenTofu infra audit (2026-07-19). `BackupDrillFailed` (critical) is firing. The weekly restore drill has been failing every run: ``` [drill] restore/forgejo-doctor FAILED on the drill VM (exit 125) Error: write .../drill-forgejo-data/_data/gitea/data/packages/...: no space left on device ``` The throwaway drill VM boots on a **20 GB** Cinder volume (`backup_drill_os_boot_volume_size`), but the restored dataset now unpacks to ~25 GB in container storage (`/srv/gitborg-data` is ~24.6 GB and growing). The restore hits ENOSPC before `forgejo doctor` runs — so there is currently **no verified proof backups can be restored**, right before onboarding (ties to the audit's H6/M10 DR-proof gaps). ### Done / remaining - [x] Bump `backup_drill_os_boot_volume_size` 20 → 40 GB — **PR #130**. - [x] Apply the backup-drill role and confirm the **next drill passes** (`forgejo doctor` green, `BackupDrillFailed` clears). — verified 2026-07-19; needed **PR #143** (guard + 60 GB) and **PR #144** (proportional staging + 80 GB) beyond #130. See the verification comment below. - [ ] Make the drill boot-volume size track data growth (derive from data-volume usage, or alert when the drill VM disk headroom < live data size) so it doesn't silently outgrow again. — *deferred hardening; 80 GB now leaves ample headroom (7.6 GB compressed bundle / ~21 GB live dataset with proportional staging).* Blocks confidence in the DR path before onboarding.
Upphovsperson
Ägare

Reopening scope — the 40 GB fix (#130) is INSUFFICIENT, and running the drill caused a transient prod disk-critical

Verified by triggering a drill run after applying #130 (2026-07-19):

  1. Drill still ENOSPC'd — no space left on device writing packages on the drill VM (exit 125), even at 40 GB. The VM holds the transferred bundle and the extracted ~21 GB dataset simultaneously, so peak usage exceeds 40 GB. BackupDrillFailed still firing; DR path still unproven.
  2. Running the drill spiked the PROD backup volume to 96.8% (critical). The drill decrypts the off-site archive into /srv/gitborg-backup/drill — on a volume already at 81% that crossed the 90% DiskUsageCritical threshold (81→89→92→96.8%), then recovered to 81% on teardown. So the drill is actively dangerous to run while the backup volume is near-full.

Stopgap applied: bitborg-backup-drill.timer stopped + disabled so it can't auto-fire and re-spike. (A site.yml apply would re-enable it — gate it in group_vars if it must survive an apply.) Do not run the drill until redesigned.

Redesign needed (both must be fixed)

  • Drill VM disk: size for bundle + extracted data (≥60 GB), or have the restore delete the transferred bundle before extracting so it doesn't hold both.
  • Staging location: decrypt the archive somewhere with headroom — the data volume (36 G free) or a dedicated scratch — NOT the near-full backup volume; and/or grow backup_volume_size 100→150 first (ties to #126 / PR #26 item 2).

Root cause of the disk-critical event was my triggering the drill without accounting for its decrypt-staging on the near-full backup volume — my mistake.

## Reopening scope — the 40 GB fix (#130) is INSUFFICIENT, and running the drill caused a transient prod disk-critical Verified by triggering a drill run after applying #130 (2026-07-19): 1. **Drill still ENOSPC'd** — `no space left on device` writing packages on the drill VM (exit 125), even at 40 GB. The VM holds the transferred bundle **and** the extracted ~21 GB dataset simultaneously, so peak usage exceeds 40 GB. `BackupDrillFailed` still firing; DR path still unproven. 2. **Running the drill spiked the PROD backup volume to 96.8% (critical).** The drill decrypts the off-site archive into `/srv/gitborg-backup/drill` — on a volume already at 81% that crossed the 90% `DiskUsageCritical` threshold (`81→89→92→96.8%`), then recovered to 81% on teardown. So the drill is actively **dangerous** to run while the backup volume is near-full. **Stopgap applied:** `bitborg-backup-drill.timer` **stopped + disabled** so it can't auto-fire and re-spike. (A `site.yml` apply would re-enable it — gate it in group_vars if it must survive an apply.) Do **not** run the drill until redesigned. ### Redesign needed (both must be fixed) - **Drill VM disk:** size for bundle + extracted data (≥60 GB), or have the restore delete the transferred bundle before extracting so it doesn't hold both. - **Staging location:** decrypt the archive somewhere with headroom — the data volume (36 G free) or a dedicated scratch — NOT the near-full backup volume; and/or grow `backup_volume_size` 100→150 first (ties to #126 / PR #26 item 2). Root cause of the disk-critical event was my triggering the drill without accounting for its decrypt-staging on the near-full backup volume — my mistake.
Upphovsperson
Ägare

Redesign up as PR #143 (review-only; not applied, drill timer stays disabled):

  • Pre-flight free-space guard — the drill aborts before fetching if WORKDIR's volume has < backup_drill_min_free_gb (25 GB) free, so it can never fill a shared prod volume again (fixes the incident class).
  • Delete the fetched archive right after decrypt — halves the WORKDIR peak.
  • Drill VM boot volume 40 → 60 GB — holds the ~15 GB bundle + ~21 GB extraction at once.

Operator sequence to get a passing drill (in PR body): merge #143 → apply --tags backup-drill → grow backup_volume_size 100→150 (so the 25 GB guard clears with margin; = PR #26 backup-headroom item) → re-enable the timer → run + verify. Box 3 (derive VM disk from live data size / stream the decrypt to drop the local bundle) noted as future hardening.

Redesign up as **PR #143** (review-only; not applied, drill timer stays disabled): - **Pre-flight free-space guard** — the drill aborts before fetching if `WORKDIR`'s volume has < `backup_drill_min_free_gb` (25 GB) free, so it can **never** fill a shared prod volume again (fixes the incident class). - **Delete the fetched archive right after decrypt** — halves the `WORKDIR` peak. - **Drill VM boot volume 40 → 60 GB** — holds the ~15 GB bundle + ~21 GB extraction at once. Operator sequence to get a passing drill (in PR body): merge #143 → apply `--tags backup-drill` → grow `backup_volume_size` 100→150 (so the 25 GB guard clears with margin; = PR #26 backup-headroom item) → re-enable the timer → run + verify. Box 3 (derive VM disk from live data size / stream the decrypt to drop the local bundle) noted as future hardening.
Upphovsperson
Ägare

✅ Restore drill now PASSES — DR path verified end-to-end (2026-07-19 14:58Z).

Timeline after the original 20 GB ENOSPC:

  • PR #130: 20 → 40 GB (still ENOSPC'd — the dataset had grown via the registry mirror #137).
  • PR #143: → 60 GB + pre-flight free-space guard + host-side archive cleanup (this fixed a separate host disk-critical incident; the VM still ENOSPC'd importing the forgejo package registry).
  • PR #144: root cause was restore-on-scratch.sh keeping bundle.tar + the extracted forgejo-data.tar.gz + the decompressed data volume alive at once (peak ≈ 2× compressed + uncompressed). Now frees each intermediate as it's consumed (peak proportional to the dataset) + VM boot volume 60 → 80 GB.

Verified run (post-#144 apply):

[restore] restored user rows: 10
[restore] running forgejo doctor check (default suite)
[restore] OK — off-site archive restored and forgejo doctor passed with no errors
[drill]   restore-drill PASSED — the off-site archive restores and forgejo doctor is clean
[drill]   teardown: deleted drill VM (boot volume cascades)
  • gitborg_backup_drill_last_run_status = 0, fresh success timestamp
  • BackupDrillFailed cleared; backup volume back to 54% (66 GB free) — 7.6 GB staging cleaned up, no leak

Box 2 (confirm next drill passes) ✅. Box 3 (auto-size the drill VM from live data growth) deferred — 80 GB now gives ample headroom over a 7.6 GB compressed bundle / ~21 GB live dataset with proportional staging; the defaults/main.yml comment flags deriving from the data volume should it ever be outgrown again.

✅ **Restore drill now PASSES — DR path verified end-to-end** (2026-07-19 14:58Z). Timeline after the original 20 GB ENOSPC: - **PR #130**: 20 → 40 GB (still ENOSPC'd — the dataset had grown via the registry mirror #137). - **PR #143**: → 60 GB + pre-flight free-space guard + host-side archive cleanup (this fixed a separate host disk-critical incident; the VM still ENOSPC'd importing the forgejo package registry). - **PR #144**: root cause was `restore-on-scratch.sh` keeping `bundle.tar` + the extracted `forgejo-data.tar.gz` + the decompressed data volume alive at once (peak ≈ 2× compressed + uncompressed). Now frees each intermediate as it's consumed (peak proportional to the dataset) + VM boot volume 60 → 80 GB. **Verified run (post-#144 apply):** ``` [restore] restored user rows: 10 [restore] running forgejo doctor check (default suite) [restore] OK — off-site archive restored and forgejo doctor passed with no errors [drill] restore-drill PASSED — the off-site archive restores and forgejo doctor is clean [drill] teardown: deleted drill VM (boot volume cascades) ``` - `gitborg_backup_drill_last_run_status = 0`, fresh success timestamp - `BackupDrillFailed` cleared; backup volume back to 54% (66 GB free) — 7.6 GB staging cleaned up, no leak **Box 2** (confirm next drill passes) ✅. **Box 3** (auto-size the drill VM from live data growth) deferred — 80 GB now gives ample headroom over a 7.6 GB compressed bundle / ~21 GB live dataset with proportional staging; the `defaults/main.yml` comment flags deriving from the data volume should it ever be outgrown again.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#131
Ingen beskrivning angiven.