feat(backup): weekly end-to-end restore drill #66

Sammanfogat
supernaut sammanfogade 2 incheckningar från backup-restore-drill in i main 2026-07-11 20:40:34 +00:00
Ägare

What

Adds a weekly end-to-end restore drill (backup-drill role + backup-drill.tofu), closing the "tested backups" half of guiding principle #4. It upgrades the existing weekly structural verify (pg_restore --list / tar -t) into a true recovery proof:

  1. Fetch the latest off-site archive from S3 and decrypt it with the on-host verify age key (the offline DR key is never used).
  2. Boot a throwaway, volume-backed OpenStack VM — no floating IP, private subnet, reusing the ephemeral-runner create_server pattern (ADR 0021).
  3. Restore Postgres + the Forgejo data volume into a clean stack over SSH and run forgejo doctor.
  4. Record gitborg_backup_drill_last_run_status / …_success_timestamp_seconds (alerts: BackupDrillFailed critical, BackupDrillStale >10 d).
  5. Always tear the VM down (boot volume cascades; never touches floating IPs).

Orchestrated by a gitborg-user systemd timer on the services host — keeping the OpenStack credential + backup-decrypt capability off the shared CI runner pool (least privilege). Opt-in (backup_drill_enabled, default off); weekly cadence stays inside the 14-day retention window.

See ADR 0027 (bitborg-docs) and the runbook section.

Verification (deployed + verified live on prod)

Took 8 live runs to go green — each surfaced a real integration issue (exactly what a DR drill is for):

  • scp -P vs ssh -p; wait for SSH auth readiness (cloud-init enables root after the port opens)
  • run forgejo doctor as uid 1000 (refuses root); pre-create + chown restored volume
  • discover the real repo root (forgejo-repositories under /data/gitea/data) instead of guessing
  • --network host (runner image has no aardvark-dns or catatonit)
  • assert on forgejo doctor's [E] markers, not just exit code (it can exit 0 with check errors — this caught a false green)

Final run: restored the latest off-site archive, forgejo doctor confirmed 5 repos + DB consistency + version 305, VM torn down with no leak, metric = 0, scraped into VictoriaMetrics, alerts loaded (only the intentional Watchdog fires).

Files

  • opentofu/backup-drill.tofu — opt-in drill security group
  • ansible/roles/backup-drill/ — orchestrator, on-VM restore, weekly timer/service
  • ansible/roles/monitoring/ — BackupDrillFailed / BackupDrillStale
  • ansible/site.yml, group_vars/all/vars.yml, docs/runbook.md

Epic: gitborg/gitborg-docs#2
Closes #40

## What Adds a **weekly end-to-end restore drill** (`backup-drill` role + `backup-drill.tofu`), closing the "tested backups" half of guiding principle #4. It upgrades the existing weekly *structural* verify (`pg_restore --list` / `tar -t`) into a true recovery proof: 1. Fetch the latest **off-site** archive from S3 and decrypt it with the on-host **verify** age key (the offline DR key is never used). 2. Boot a throwaway, volume-backed OpenStack VM — no floating IP, private subnet, reusing the ephemeral-runner `create_server` pattern (ADR 0021). 3. Restore Postgres + the Forgejo data volume into a clean stack over SSH and run `forgejo doctor`. 4. Record `gitborg_backup_drill_last_run_status` / `…_success_timestamp_seconds` (alerts: `BackupDrillFailed` critical, `BackupDrillStale` >10 d). 5. Always tear the VM down (boot volume cascades; never touches floating IPs). Orchestrated by a `gitborg`-user systemd timer on the services host — keeping the OpenStack credential + backup-decrypt capability off the shared CI runner pool (least privilege). Opt-in (`backup_drill_enabled`, default off); weekly cadence stays inside the 14-day retention window. See ADR 0027 (bitborg-docs) and the runbook section. ## Verification (deployed + verified live on prod) Took 8 live runs to go green — each surfaced a real integration issue (exactly what a DR drill is for): - scp `-P` vs ssh `-p`; wait for SSH **auth** readiness (cloud-init enables root after the port opens) - run `forgejo doctor` as uid 1000 (refuses root); pre-create + chown restored volume - discover the real repo root (`forgejo-repositories` under `/data/gitea/data`) instead of guessing - `--network host` (runner image has no aardvark-dns or catatonit) - assert on `forgejo doctor`'s `[E]` markers, not just exit code (it can exit 0 with check errors — this caught a false green) Final run: restored the latest off-site archive, `forgejo doctor` confirmed **5 repos + DB consistency + version 305**, VM torn down with **no leak**, metric `= 0`, scraped into VictoriaMetrics, alerts loaded (only the intentional Watchdog fires). ## Files - `opentofu/backup-drill.tofu` — opt-in drill security group - `ansible/roles/backup-drill/` — orchestrator, on-VM restore, weekly timer/service - `ansible/roles/monitoring/` — `BackupDrillFailed` / `BackupDrillStale` - `ansible/site.yml`, `group_vars/all/vars.yml`, `docs/runbook.md` Epic: gitborg/gitborg-docs#2 Closes #40
supernaut lade till 2 incheckningar 2026-07-11 20:18:25 +00:00
Upgrade the structural verify (pg_restore --list / tar -t) into a true
recovery proof. A weekly gitborg-user timer boots a throwaway, volume-backed
OpenStack VM (no floating IP), restores the latest off-site archive into a
clean Postgres + Forgejo stack, runs `forgejo doctor`, records a metric, and
tears the VM down.

- opentofu/backup-drill.tofu: opt-in security group for the drill VM
  (ingress SSH from the private subnet; egress 443 + 53 only).
- roles/backup-drill: orchestrator (openstacksdk boot + SSH restore +
  teardown), on-VM restore script, weekly timer/service; reuses the
  runner-controller OpenStack credential and the decrypt-only verify age
  identity (the offline DR key is never used).
- monitoring: BackupDrillFailed (critical) + BackupDrillStale (>10d) alerts,
  kept inside the 14-day retention window.
- site.yml/group_vars: wired in, inert until opted in and secrets present.

Epic: gitborg/gitborg-docs#2
fix(backup): make the restore drill work end-to-end (#40)
Alla kontroller lyckades
ci / ci (pull_request) Successful in 2m5s
7e99323fe9
Fixes found by running the drill live against the off-site archive:

- orchestrator: scp needs -P (not ssh's -p) for the port; wait for SSH AUTH
  readiness (cloud-init enables root login after the port opens); handle
  SIGTERM so a timeout still runs teardown; drop the redundant root-volume
  detach (terminate_volume already cascades the boot volume); raise the
  timeout to 3600s (the off-site fetch alone is ~6 min).
- on-VM restore: run forgejo as uid 1000 (it refuses to run as root);
  pre-create + chown the restored volume so it can write its work dirs;
  discover the real repository root (forgejo-repositories under
  /data/gitea/data) instead of guessing; use --network host so Postgres is
  reachable on 127.0.0.1 (the runner image has no aardvark-dns or catatonit);
  assert on forgejo doctor's [E] markers, not just its exit code (it can exit
  0 with check-level errors).

Verified on prod: restores the latest off-site archive into a throwaway VM,
`forgejo doctor` confirms 5 repos + DB consistency, VM torn down with no leak,
gitborg_backup_drill_last_run_status=0.

Epic: gitborg/gitborg-docs#2
supernaut sammanfogade incheckning 3b53f9250f till main 2026-07-11 20:40:34 +00:00
supernaut tog bort grenen backup-restore-drill 2026-07-11 20:40:34 +00:00
supernaut refererade denna ändringsförfrågan från en incheckning 2026-07-11 20:40:35 +00:00
supernaut refererade denna ändringsförfrågan från en incheckning 2026-08-03 09:41:33 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!66
Ingen beskrivning angiven.