fix(backup-drill): grow drill VM boot volume 20→40 GB (restore ENOSPC) #130

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/backup-drill-boot-volume-size in i main 2026-07-19 12:38:46 +00:00
Ägare

The weekly restore drill has been failing (BackupDrillFailed critical, firing now). Root cause from the drill logs:

[drill] restore/forgejo-doctor FAILED on the drill VM (exit 125)
Error: write .../drill-forgejo-data/_data/gitea/data/packages/...: no space left on device

The throwaway drill VM boots on a 20 GB Cinder volume (backup_drill_os_boot_volume_size), but the restored dataset now unpacks to ~25 GB in container storage (the live /srv/gitborg-data volume is at ~24.6 GB and growing). So the restore runs out of space before forgejo doctor — meaning there is currently no verified proof backups can be restored, right before onboarding.

Bumps the boot volume to 40 GB (headroom above the live data volume) and rewrites the now-false "20 GiB is enough" comment to explain the sizing rule. The volume is a throwaway scratch disk on a VM that's created and destroyed per drill — no data-loss risk, no prod-service impact; it takes effect on the next drill run after the role is applied.

Validated: ansible-lint roles/backup-drill/ passes (production profile), site.yml --syntax-check OK. Found by the OpenTofu infra audit (2026-07-19).

The weekly restore drill has been failing (`BackupDrillFailed` critical, firing now). Root cause from the drill logs: ``` [drill] restore/forgejo-doctor FAILED on the drill VM (exit 125) Error: write .../drill-forgejo-data/_data/gitea/data/packages/...: no space left on device ``` The throwaway drill VM boots on a **20 GB** Cinder volume (`backup_drill_os_boot_volume_size`), but the restored dataset now unpacks to ~25 GB in container storage (the live `/srv/gitborg-data` volume is at ~24.6 GB and growing). So the restore runs out of space before `forgejo doctor` — meaning there is currently **no verified proof backups can be restored**, right before onboarding. Bumps the boot volume to **40 GB** (headroom above the live data volume) and rewrites the now-false "20 GiB is enough" comment to explain the sizing rule. The volume is a throwaway scratch disk on a VM that's created and destroyed per drill — no data-loss risk, no prod-service impact; it takes effect on the next drill run after the role is applied. Validated: `ansible-lint roles/backup-drill/` passes (production profile), `site.yml --syntax-check` OK. Found by the OpenTofu infra audit (2026-07-19).
supernaut lade till 1 incheckning 2026-07-19 07:27:36 +00:00
fix(backup-drill): grow drill VM boot volume 20->40 GB (restore ENOSPC)
Alla kontroller lyckades
ci / ci (pull_request) Successful in 3m51s
2c15f244b7
The throwaway restore-drill VM boots on a 20 GB Cinder volume, but the restored
dataset now unpacks to ~25 GB in container storage, so the drill fails with
'no space left on device' writing restored packages and forgejo doctor never
runs (BackupDrillFailed firing). Size it comfortably above the live data volume;
40 GB gives headroom. Throwaway volume — no data-loss risk.
supernaut sammanfogade incheckning d0661d491d till main 2026-07-19 12:38:46 +00:00
supernaut tog bort grenen fix/backup-drill-boot-volume-size 2026-07-19 12:38:47 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!130
Ingen beskrivning angiven.