feat(backup): weekly end-to-end restore drill #66
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!66
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "backup-restore-drill"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
What
Adds a weekly end-to-end restore drill (
backup-drillrole +backup-drill.tofu), closing the "tested backups" half of guiding principle #4. It upgrades the existing weekly structural verify (pg_restore --list/tar -t) into a true recovery proof:create_serverpattern (ADR 0021).forgejo doctor.gitborg_backup_drill_last_run_status/…_success_timestamp_seconds(alerts:BackupDrillFailedcritical,BackupDrillStale>10 d).Orchestrated by a
gitborg-user systemd timer on the services host — keeping the OpenStack credential + backup-decrypt capability off the shared CI runner pool (least privilege). Opt-in (backup_drill_enabled, default off); weekly cadence stays inside the 14-day retention window.See ADR 0027 (bitborg-docs) and the runbook section.
Verification (deployed + verified live on prod)
Took 8 live runs to go green — each surfaced a real integration issue (exactly what a DR drill is for):
-Pvs ssh-p; wait for SSH auth readiness (cloud-init enables root after the port opens)forgejo doctoras uid 1000 (refuses root); pre-create + chown restored volumeforgejo-repositoriesunder/data/gitea/data) instead of guessing--network host(runner image has no aardvark-dns or catatonit)forgejo doctor's[E]markers, not just exit code (it can exit 0 with check errors — this caught a false green)Final run: restored the latest off-site archive,
forgejo doctorconfirmed 5 repos + DB consistency + version 305, VM torn down with no leak, metric= 0, scraped into VictoriaMetrics, alerts loaded (only the intentional Watchdog fires).Files
opentofu/backup-drill.tofu— opt-in drill security groupansible/roles/backup-drill/— orchestrator, on-VM restore, weekly timer/serviceansible/roles/monitoring/—BackupDrillFailed/BackupDrillStaleansible/site.yml,group_vars/all/vars.yml,docs/runbook.mdEpic: gitborg/gitborg-docs#2
Closes #40