fix(backup): harden retention + off-site failure handling (M6/M7/L2/L3) #147
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!147
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/126-backup-robustness"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Scope (#126 — genuine data-safety bugs)
The bug subset of the backup-robustness audit. Design items (H6 DR-key escrow, H7 PITR/WAL, M5 hot-skew, M9 multipart abort, M10 full RTO) are not in this PR — they'll be tracked as separate follow-ups. All changes are in
roles/backup/templates/bitborg-backup.sh.j2.M6 — orphaned
*.tar.age.partialleakA SIGKILL/OOM mid-
ageskips the EXIT trap, leaving a multi-GBbitborg-*.tar.age.partialthat pruning (which matches*.tar.age, not*.partial) never removes → compounds disk-fill run after run. Fix: sweep stale partials at run start (any present then are from an aborted prior run — this run's ENC name is timestamped and not created yet).M7 — off-site failure escalated to "total failure" + halted local prune
An off-site upload failure
exit 1'd before the local prune and before recording local success, so a Glesys outage both masked a good local backup (run went red) and halted local pruning (disk fill). Fix: off-site failure no longer fails the run or skips the local prune — it's a WARNING; local success is recorded independently, and the per-destinationBackupOffsiteFailedalert (already wired) surfaces the off-site problem.L2 — no keep-last-N floor
Local prune was time-only, so >14 days of failed/degraded backups would prune every good copy. Fix:
backup_keep_min(default 3) — always retain the newest N regardless of age. Functionally tested: with all archives older than retention, the newest 3 still survive.L3 — off-site prune had no name filter
rclone delete --min-agematched any object in the bucket. Fix:--include "bitborg-*.tar.age" --include "bitborg-*.tar.enc"so it only ever deletes our own archives.Validation
ansible-lint (production profile) ·
--syntax-check·--check --diffrenders clean (0 failed) ·bash -non the rendered script · functional test of the keep-N prune (normal + all-expired edge case).Apply notes
--tags backup. Re-renders the script only; effect is on the next scheduled/manual backup run. New knob:backup_keep_miningroup_vars/all/vars.yml.Refs #126.