Backup robustness for production (DR key drill, RPO/PITR, hot-skew, partial leak, off-site failure, RTO) #126
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#126
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Severity: HIGH — pre-onboarding infra audit (2026-07-19). Backup robustness for production; groups several related findings. (The web-DB gap is tracked separately as a CRITICAL.)
All references
ansible/roles/backup/templates/bitborg-backup.sh.j2unless noted.backup-drill/...py.j2:276) and verify (bitborg-backup-verify.sh.j2:54) both decrypt with the on-host verify key; the real DR keybackup_age_recipient(group_vars/all/vars.yml:286) lives only in 1Password. Nothing proves the archive decrypts with the DR key, and it's a single-vault SPOF. → Periodic manual DR-key restore test; store a second sealed copy of the DR identity independent of 1Password.vars.yml:25903:30), no WAL archiving → up to 24 h of pushes/issues/PRs lost on host loss; never set as a target. → Document RPO/RTO; more-frequent DB dumps and/or Postgres WAL archiving off-site.pg_dumpthenpodman volume exportrun live seconds apart (:84-97); a push in between leaves the volume with objects the DB dump doesn't reference. → Briefly quiesce Forgejo for the volume-export window, or evaluateforgejo dump; at minimum document.*.tar.age.partialleak. OOM/SIGKILL mid-encrypt skips the EXIT trap (:72-78); prune matches onlybitborg-*.tar.age, not.partial(:196-198) → partials accumulate. → Sweep stale partials at run start.exit 1s before local prune (:196) and the success timestamp (:201-207) → a Glesys outage makes a good local archive look like total failure AND stops pruning (disk fill). → Separate local-success from off-site-success; surface off-site via a distinctBackupOffsiteFailed.:169-173). → S3 lifecycle rule to abort incomplete multipart after N days; periodicrclone cleanup.site.ymlstand-up, no measured RTO. → Periodically run the full manual restore on a scratch instance; quantify RTO.:172-198): 14 days of silently-degraded backups prunes every good copy. → Add a min-count floor.:172-173):rclone delete --min-agedeletes any object >14 d. → Add--include "bitborg-*.tar.age".Live confirmation from the OpenTofu infra audit (2026-07-19) — this is not hypothetical, it's happening now.
/srv/gitborg-backup(100 GB volume) is at ~82% (85.9 GB used) and bothBackupVolumeLow(70%) andDiskUsageWarning(80%) are firing. Trend: it climbed 40% → 82% over ~10 days (~4%/day) with no pruning-driven drop, so at the current rate it hitsDiskUsageCritical(90%) in ~2 days and ENOSPC (backups start failing) in ~4–5 days.86 GB used is far more than 14-day retention of the ~24.6 GB dataset should need — which points straight at the root causes tracked here:
*.tar.age.partialnever pruned (an OOM/SIGKILL mid-encrypt skips the trap; prune matches only*.tar.age).exit 1s before local prune, so a transient S3 outage halts local pruning and the volume fills.Immediate (separate from this cleanup): grow the volume
backup_volume_size100 → 150 (online Cinder extend +resize2fs, non-destructive — same as the 40→100 bump on 2026-07-09) to buy headroom while the pruning fixes land. On the host, check what's actually consuming it first:du -sh /srv/gitborg-backup/*andfind /srv/gitborg-backup -name '*.tar.age.partial'— if it's leaked partials, the fix here reclaims the space without growing.✅ Bug subset (M6/M7/L2/L3) done — PR #147, applied + verified live (2026-07-19).
Verified with a real end-to-end backup run:
gitborg_backup_last_run_status 0; fresh success timestamp.gitborg_backup_offsite_last_status{destination="glesys"} 0,{destination="hetzner"} 0— local and off-site are now tracked independently (M7).*.tar.age.partial; fresh archive written cleanly.pruning local archives older than 14 days (always keeping newest 3)ran (pruned nothing this run — all archives within retention; the keep-N floor was functionally tested for the all-expired case).--include "bitborg-*.tar.age"/.enc).BackupFailed/BackupOffsiteFailed/BackupStale/BackupVolumeLowfiring.Remaining under this issue (design-level, not quick bugs) — keeping #126 open to track:
pg_dumpand thepodman volume export.rclone cleanup).site.ymlstand-up, measured RTO).These can stay here or be split into individual issues — say which you prefer.
Splitting the remaining design items into their own tracked issues so this umbrella can close.
Done under #126 (shipped + live):
.partialleak sweep), M7 (off-site failure decoupled from local prune + success), L2 (keep-last-N floor), L3 (off-site prune name filter) — PR #147, applied + verified via a live backup run.Remaining design items → split out:
Closing this umbrella; tracking continues in #157–#161.