Backup robustness for production (DR key drill, RPO/PITR, hot-skew, partial leak, off-site failure, RTO) #126

Stängd
öppnade 2026-07-19 06:58:03 +00:00 av supernaut · 3 kommentarer
Ägare

Severity: HIGH — pre-onboarding infra audit (2026-07-19). Backup robustness for production; groups several related findings. (The web-DB gap is tracked separately as a CRITICAL.)

All references ansible/roles/backup/templates/bitborg-backup.sh.j2 unless noted.

  • H6 — DR decryption path never exercised/escrowed. The drill (backup-drill/...py.j2:276) and verify (bitborg-backup-verify.sh.j2:54) both decrypt with the on-host verify key; the real DR key backup_age_recipient (group_vars/all/vars.yml:286) lives only in 1Password. Nothing proves the archive decrypts with the DR key, and it's a single-vault SPOF. → Periodic manual DR-key restore test; store a second sealed copy of the DR identity independent of 1Password.
  • H7 — RPO up to ~24 h, no PITR. Daily dump (vars.yml:259 03:30), no WAL archiving → up to 24 h of pushes/issues/PRs lost on host loss; never set as a target. → Document RPO/RTO; more-frequent DB dumps and/or Postgres WAL archiving off-site.
  • M5 — hot, non-quiesced backup skew. pg_dump then podman volume export run live seconds apart (:84-97); a push in between leaves the volume with objects the DB dump doesn't reference. → Briefly quiesce Forgejo for the volume-export window, or evaluate forgejo dump; at minimum document.
  • M6 — orphaned *.tar.age.partial leak. OOM/SIGKILL mid-encrypt skips the EXIT trap (:72-78); prune matches only bitborg-*.tar.age, not .partial (:196-198) → partials accumulate. → Sweep stale partials at run start.
  • M7 — off-site failure escalates to "no backups" + halts local prune. Off-site failure exit 1s before local prune (:196) and the success timestamp (:201-207) → a Glesys outage makes a good local archive look like total failure AND stops pruning (disk fill). → Separate local-success from off-site-success; surface off-site via a distinct BackupOffsiteFailed.
  • M9 — interrupted S3 multipart uploads not aborted (:169-173). → S3 lifecycle rule to abort incomplete multipart after N days; periodic rclone cleanup.
  • M10 — RTO never proven end-to-end. Drill restores forgejo only (not Kanidm, not web DB), no full site.yml stand-up, no measured RTO. → Periodically run the full manual restore on a scratch instance; quantify RTO.
  • L2 — retention is time-only, no keep-last-N floor (:172-198): 14 days of silently-degraded backups prunes every good copy. → Add a min-count floor.
  • L3 — off-site prune has no name filter (:172-173): rclone delete --min-age deletes any object >14 d. → Add --include "bitborg-*.tar.age".
**Severity: HIGH** — pre-onboarding infra audit (2026-07-19). Backup robustness for production; groups several related findings. (The web-DB gap is tracked separately as a CRITICAL.) All references `ansible/roles/backup/templates/bitborg-backup.sh.j2` unless noted. - **H6 — DR decryption path never exercised/escrowed.** The drill (`backup-drill/...py.j2:276`) and verify (`bitborg-backup-verify.sh.j2:54`) both decrypt with the on-host *verify* key; the real DR key `backup_age_recipient` (`group_vars/all/vars.yml:286`) lives only in 1Password. Nothing proves the archive decrypts with the DR key, and it's a single-vault SPOF. → Periodic manual DR-key restore test; store a second sealed copy of the DR identity independent of 1Password. - **H7 — RPO up to ~24 h, no PITR.** Daily dump (`vars.yml:259` `03:30`), no WAL archiving → up to 24 h of pushes/issues/PRs lost on host loss; never set as a target. → Document RPO/RTO; more-frequent DB dumps and/or Postgres WAL archiving off-site. - **M5 — hot, non-quiesced backup skew.** `pg_dump` then `podman volume export` run live seconds apart (`:84-97`); a push in between leaves the volume with objects the DB dump doesn't reference. → Briefly quiesce Forgejo for the volume-export window, or evaluate `forgejo dump`; at minimum document. - **M6 — orphaned `*.tar.age.partial` leak.** OOM/SIGKILL mid-encrypt skips the EXIT trap (`:72-78`); prune matches only `bitborg-*.tar.age`, not `.partial` (`:196-198`) → partials accumulate. → Sweep stale partials at run start. - **M7 — off-site failure escalates to "no backups" + halts local prune.** Off-site failure `exit 1`s before local prune (`:196`) and the success timestamp (`:201-207`) → a Glesys outage makes a good local archive look like total failure AND stops pruning (disk fill). → Separate local-success from off-site-success; surface off-site via a distinct `BackupOffsiteFailed`. - **M9 — interrupted S3 multipart uploads not aborted** (`:169-173`). → S3 lifecycle rule to abort incomplete multipart after N days; periodic `rclone cleanup`. - **M10 — RTO never proven end-to-end.** Drill restores forgejo only (not Kanidm, not web DB), no full `site.yml` stand-up, no measured RTO. → Periodically run the full manual restore on a scratch instance; quantify RTO. - **L2 — retention is time-only, no keep-last-N floor** (`:172-198`): 14 days of silently-degraded backups prunes every good copy. → Add a min-count floor. - **L3 — off-site prune has no name filter** (`:172-173`): `rclone delete --min-age` deletes any object >14 d. → Add `--include "bitborg-*.tar.age"`.
Upphovsperson
Ägare

Live confirmation from the OpenTofu infra audit (2026-07-19) — this is not hypothetical, it's happening now.

/srv/gitborg-backup (100 GB volume) is at ~82% (85.9 GB used) and both BackupVolumeLow (70%) and DiskUsageWarning (80%) are firing. Trend: it climbed 40% → 82% over ~10 days (~4%/day) with no pruning-driven drop, so at the current rate it hits DiskUsageCritical (90%) in ~2 days and ENOSPC (backups start failing) in ~4–5 days.

86 GB used is far more than 14-day retention of the ~24.6 GB dataset should need — which points straight at the root causes tracked here:

  • M6 — orphaned *.tar.age.partial never pruned (an OOM/SIGKILL mid-encrypt skips the trap; prune matches only *.tar.age).
  • M7 — off-site (Glesys) upload failure exit 1s before local prune, so a transient S3 outage halts local pruning and the volume fills.
  • L2 — no keep-last-N floor.

Immediate (separate from this cleanup): grow the volume backup_volume_size 100 → 150 (online Cinder extend + resize2fs, non-destructive — same as the 40→100 bump on 2026-07-09) to buy headroom while the pruning fixes land. On the host, check what's actually consuming it first: du -sh /srv/gitborg-backup/* and find /srv/gitborg-backup -name '*.tar.age.partial' — if it's leaked partials, the fix here reclaims the space without growing.

**Live confirmation from the OpenTofu infra audit (2026-07-19) — this is not hypothetical, it's happening now.** `/srv/gitborg-backup` (100 GB volume) is at **~82% (85.9 GB used)** and both `BackupVolumeLow` (70%) and `DiskUsageWarning` (80%) are **firing**. Trend: it climbed **40% → 82% over ~10 days (~4%/day)** with no pruning-driven drop, so at the current rate it hits `DiskUsageCritical` (90%) in **~2 days** and ENOSPC (backups start failing) in ~4–5 days. 86 GB used is far more than 14-day retention of the ~24.6 GB dataset should need — which points straight at the root causes tracked here: - **M6** — orphaned `*.tar.age.partial` never pruned (an OOM/SIGKILL mid-encrypt skips the trap; prune matches only `*.tar.age`). - **M7** — off-site (Glesys) upload failure `exit 1`s before local prune, so a transient S3 outage halts local pruning and the volume fills. - **L2** — no keep-last-N floor. **Immediate (separate from this cleanup):** grow the volume `backup_volume_size` 100 → 150 (online Cinder extend + `resize2fs`, non-destructive — same as the 40→100 bump on 2026-07-09) to buy headroom while the pruning fixes land. On the host, check what's actually consuming it first: `du -sh /srv/gitborg-backup/*` and `find /srv/gitborg-backup -name '*.tar.age.partial'` — if it's leaked partials, the fix here reclaims the space without growing.
Upphovsperson
Ägare

✅ Bug subset (M6/M7/L2/L3) done — PR #147, applied + verified live (2026-07-19).

Verified with a real end-to-end backup run:

  • gitborg_backup_last_run_status 0; fresh success timestamp.
  • Both off-site destinations green: gitborg_backup_offsite_last_status{destination="glesys"} 0, {destination="hetzner"} 0 — local and off-site are now tracked independently (M7).
  • M6: 0 orphan *.tar.age.partial; fresh archive written cleanly.
  • L2: pruning local archives older than 14 days (always keeping newest 3) ran (pruned nothing this run — all archives within retention; the keep-N floor was functionally tested for the all-expired case).
  • L3: per-destination off-site prunes ran name-filtered (--include "bitborg-*.tar.age" / .enc).
  • No BackupFailed / BackupOffsiteFailed / BackupStale / BackupVolumeLow firing.

Remaining under this issue (design-level, not quick bugs) — keeping #126 open to track:

  • H6 — exercise the real DR-key decryption path + store a second sealed copy of the DR identity independent of 1Password.
  • H7 — RPO/PITR: document RPO/RTO targets; more-frequent DB dumps and/or Postgres WAL archiving off-site.
  • M5 — hot, non-quiesced skew between pg_dump and the podman volume export.
  • M9 — abort incomplete S3 multipart uploads (bucket lifecycle rule / periodic rclone cleanup).
  • M10 — prove full RTO end-to-end (Kanidm + web DB, full site.yml stand-up, measured RTO).

These can stay here or be split into individual issues — say which you prefer.

✅ **Bug subset (M6/M7/L2/L3) done — PR #147, applied + verified live** (2026-07-19). Verified with a real end-to-end backup run: - `gitborg_backup_last_run_status 0`; fresh success timestamp. - Both off-site destinations green: `gitborg_backup_offsite_last_status{destination="glesys"} 0`, `{destination="hetzner"} 0` — local and off-site are now tracked independently (M7). - **M6:** 0 orphan `*.tar.age.partial`; fresh archive written cleanly. - **L2:** `pruning local archives older than 14 days (always keeping newest 3)` ran (pruned nothing this run — all archives within retention; the keep-N floor was functionally tested for the all-expired case). - **L3:** per-destination off-site prunes ran name-filtered (`--include "bitborg-*.tar.age"` / `.enc`). - No `BackupFailed` / `BackupOffsiteFailed` / `BackupStale` / `BackupVolumeLow` firing. **Remaining under this issue** (design-level, not quick bugs) — keeping #126 open to track: - **H6** — exercise the real DR-key decryption path + store a second sealed copy of the DR identity independent of 1Password. - **H7** — RPO/PITR: document RPO/RTO targets; more-frequent DB dumps and/or Postgres WAL archiving off-site. - **M5** — hot, non-quiesced skew between `pg_dump` and the `podman volume export`. - **M9** — abort incomplete S3 multipart uploads (bucket lifecycle rule / periodic `rclone cleanup`). - **M10** — prove full RTO end-to-end (Kanidm + web DB, full `site.yml` stand-up, measured RTO). These can stay here or be split into individual issues — say which you prefer.
Upphovsperson
Ägare

Splitting the remaining design items into their own tracked issues so this umbrella can close.

Done under #126 (shipped + live):

  • Immediate pressure relieved — backup volume grown 100 → 150 GB.
  • M6 (.partial leak sweep), M7 (off-site failure decoupled from local prune + success), L2 (keep-last-N floor), L3 (off-site prune name filter) — PR #147, applied + verified via a live backup run.

Remaining design items → split out:

  • #157 — H6: exercise + escrow the disaster-recovery age key
  • #158 — H7: RPO/RTO target + Postgres PITR / WAL archiving
  • #159 — M5: hot pg_dump/volume-export skew
  • #160 — M9: abort incomplete S3 multipart uploads
  • #161 — M10: prove full end-to-end RTO

Closing this umbrella; tracking continues in #157–#161.

Splitting the remaining design items into their own tracked issues so this umbrella can close. **Done under #126 (shipped + live):** - Immediate pressure relieved — backup volume grown 100 → 150 GB. - **M6** (`.partial` leak sweep), **M7** (off-site failure decoupled from local prune + success), **L2** (keep-last-N floor), **L3** (off-site prune name filter) — PR #147, applied + verified via a live backup run. **Remaining design items → split out:** - #157 — **H6**: exercise + escrow the disaster-recovery age key - #158 — **H7**: RPO/RTO target + Postgres PITR / WAL archiving - #159 — **M5**: hot pg_dump/volume-export skew - #160 — **M9**: abort incomplete S3 multipart uploads - #161 — **M10**: prove full end-to-end RTO Closing this umbrella; tracking continues in #157–#161.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#126
Ingen beskrivning angiven.