fix(backup-drill): free restore intermediates + bump drill VM to 80G #144

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/backup-drill-vm-disk-cleanup in i main 2026-07-19 14:30:23 +00:00
Ägare

Problem

The re-enabled restore drill (2026-07-19) failed: the drill VM ENOSPC'd importing the Forgejo data volume —

Error: write .../drill-forgejo-data/_data/gitea/data/packages/22/ba/… : no space left on device
[drill] restore/forgejo-doctor FAILED on the drill VM (exit 125)

Host side worked as designed (#143): the pre-flight guard held, teardown cleaned up, /srv/gitborg-backup returned to 54% — no leak, the disk incident cannot recur. The failure is purely the drill VM boot volume.

Root cause

restore-on-scratch.sh kept three copies of the dataset alive at once: bundle.tar (kept the whole run) + the extracted forgejo-data.tar.gz + the decompressed data volume in container storage. Peak ≈ 2× the compressed dataset + the uncompressed volume. The registry mirror (#137) grew the package registry with barely-compressible container blobs (gitea/data/packages/…) and pushed a 60 GB VM over.

Fix

  • Free each intermediate as it's consumed — bundle.tar after extract, db.dump after pg_restore, forgejo-data.tar.gz after volume import — so peak disk is proportional to the dataset (same principle as the host-side os.remove(archive) in #143).
  • Bump the throwaway VM boot volume 60→80 GB for headroom vs a ~21 GB live dataset (~3.5×). +20 GB transient Cinder for the ~20-min run, well within quota now the runner-vol leak is fixed.

Verification plan (post-merge/apply)

Apply --tags backup-drill, trigger one run, confirm forgejo doctor passes + BackupDrillFailed clears + backup vol stays flat → then close #131.

Refs #131.

## Problem The re-enabled restore drill (2026-07-19) failed: the drill VM ENOSPC'd importing the Forgejo data volume — ``` Error: write .../drill-forgejo-data/_data/gitea/data/packages/22/ba/… : no space left on device [drill] restore/forgejo-doctor FAILED on the drill VM (exit 125) ``` Host side worked as designed (#143): the pre-flight guard held, teardown cleaned up, `/srv/gitborg-backup` returned to 54% — no leak, the disk incident cannot recur. The failure is purely the **drill VM boot volume**. ## Root cause `restore-on-scratch.sh` kept three copies of the dataset alive at once: `bundle.tar` (kept the whole run) + the extracted `forgejo-data.tar.gz` + the decompressed data volume in container storage. Peak ≈ 2× the compressed dataset + the uncompressed volume. The registry mirror (#137) grew the package registry with barely-compressible container blobs (`gitea/data/packages/…`) and pushed a 60 GB VM over. ## Fix - **Free each intermediate as it's consumed** — `bundle.tar` after extract, `db.dump` after `pg_restore`, `forgejo-data.tar.gz` after volume import — so peak disk is proportional to the dataset (same principle as the host-side `os.remove(archive)` in #143). - **Bump the throwaway VM boot volume 60→80 GB** for headroom vs a ~21 GB live dataset (~3.5×). +20 GB transient Cinder for the ~20-min run, well within quota now the runner-vol leak is fixed. ## Verification plan (post-merge/apply) Apply `--tags backup-drill`, trigger one run, confirm `forgejo doctor` passes + `BackupDrillFailed` clears + backup vol stays flat → then close #131. Refs #131.
supernaut lade till 1 incheckning 2026-07-19 14:20:02 +00:00
fix(backup-drill): free restore intermediates + bump drill VM to 80G
Alla kontroller lyckades
ci / ci (pull_request) Successful in 2m54s
a701d64490
The restore-drill VM ENOSPC'd importing the Forgejo data volume
(gitea/data/packages): restore-on-scratch.sh kept bundle.tar, the
extracted forgejo-data.tar.gz, and the decompressed data volume alive
at once, so peak disk was ~2x the compressed dataset plus the
uncompressed volume. The registry mirror (#137) grew the package
registry with barely-compressible blobs and pushed a 60G VM over.

Free each intermediate as soon as it's consumed (bundle.tar after
extract, db.dump after pg_restore, forgejo-data.tar.gz after volume
import) so peak disk is proportional to the dataset, mirroring the
host-side archive cleanup in #143. Bump the throwaway VM boot volume
60->80G for headroom against a ~21G live dataset (~3.5x).

Refs #131.
supernaut sammanfogade incheckning 922d6271a1 till main 2026-07-19 14:30:23 +00:00
supernaut tog bort grenen fix/backup-drill-vm-disk-cleanup 2026-07-19 14:30:23 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!144
Ingen beskrivning angiven.