Backups: CI/scratch-instance test-restore job #40

Stängd
öppnade 2026-07-09 23:03:57 +00:00 av supernaut · 1 kommentar
Ägare

A weekly on-host restore-verification timer already runs. Add the CI/scratch-instance variant: restore into a scratch instance and run forgejo doctor.

Epic: gitborg/gitborg-docs#2

A weekly on-host restore-verification timer already runs. Add the CI/scratch-instance variant: restore into a scratch instance and run `forgejo doctor`. Epic: gitborg/gitborg-docs#2
supernaut refererade till detta ärende från en incheckning 2026-07-11 20:17:38 +00:00
Upphovsperson
Ägare

Implemented, deployed, and verified live on prod. 🎉

The CI/scratch-instance variant now runs as a weekly bitborg-backup-drill systemd timer on the services host: it fetches the latest off-site archive, decrypts it with the on-host verify key, boots a throwaway OpenStack VM (ADR 0021 pattern, no FIP), restores Postgres + the Forgejo data volume into a clean stack, runs forgejo doctor, records a metric, and tears the VM down. Orchestration lives on the services host (not the CI runner pool) to keep the OpenStack credential + backup-decrypt capability least-privilege.

Verified end-to-end: a live run restored the latest off-site archive into a throwaway VM; forgejo doctor confirmed 5 repos + DB consistency + version 305; the VM was torn down with no leaked server/volume; gitborg_backup_drill_last_run_status = 0 is scraped into VictoriaMetrics; BackupDrillFailed / BackupDrillStale alerts are loaded. (It took 8 runs to go green — each surfaced a real integration bug, which is exactly the point of a restore drill.)

Sibling tasks #39 (Bahnhof S3 migration, blocked on the DPA) and #41 (second off-provider destination) remain open under epic gitborg/gitborg-docs#2.

Implemented, deployed, and **verified live on prod**. 🎉 The CI/scratch-instance variant now runs as a weekly `bitborg-backup-drill` systemd timer on the services host: it fetches the latest **off-site** archive, decrypts it with the on-host verify key, boots a throwaway OpenStack VM (ADR 0021 pattern, no FIP), restores Postgres + the Forgejo data volume into a clean stack, runs `forgejo doctor`, records a metric, and tears the VM down. Orchestration lives on the services host (not the CI runner pool) to keep the OpenStack credential + backup-decrypt capability least-privilege. **Verified end-to-end:** a live run restored the latest off-site archive into a throwaway VM; `forgejo doctor` confirmed **5 repos + DB consistency + version 305**; the VM was torn down with no leaked server/volume; `gitborg_backup_drill_last_run_status = 0` is scraped into VictoriaMetrics; `BackupDrillFailed` / `BackupDrillStale` alerts are loaded. (It took 8 runs to go green — each surfaced a real integration bug, which is exactly the point of a restore drill.) - Implementation + verification: gitborg/gitborg-infra#66 - Decision record (ADR 0027) + architecture: gitborg/gitborg-docs#17 Sibling tasks #39 (Bahnhof S3 migration, blocked on the DPA) and #41 (second off-provider destination) remain open under epic gitborg/gitborg-docs#2.
supernaut refererade till detta ärende från en incheckning 2026-07-11 20:39:30 +00:00
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#40
Ingen beskrivning angiven.