Backups: define RPO/RTO + Postgres PITR / WAL archiving #158
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#158
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Severity: HIGH —
area/backups.RPO up to ~24 h, no PITR
The DB backup is a single daily
pg_dump(03:30,group_vars/all/vars.yml) with no WAL archiving → up to 24 h of pushes/issues/PRs lost on host loss. This was never set as an explicit target.Ask
Refs #126.
Epic: bitborg/bitborg-docs#107
Research summary, 2026-10-02. Needs operator sign-off before any PR.
Today
forgejo,gitborg_web, and the billing DB (dumped only whenbilling_enabled). No WAL archiving. Kanidm is SQLite, dumped withkanidmd database backup.pg_dump -Fc+ Forgejo repo volume export, age-encrypted, local 14 d, off-site to Glesys and Hetzner.[storage]tier via restic at 05:00.Constraint
Forgejo's DB and its git repos must restore to the same point. WAL PITR on the DB alone does not improve Forgejo's RPO: the repo volume needs the same cadence. WAL archiving is per cluster, so it is all three DBs or none.
Proposed targets
[storage](packages, LFS)RTO 4 h is a guess until one full rebuild is timed.
Recommendation
Hourly dumps (DBs and Kanidm,
-Z0) followed by a restic snapshot of the repo volume, into a new restic repo on Glesys, reusing the[storage]restic machinery. Keep 48 hourly + 14 daily. A 15-min billing dump gated onbilling_enabled. The daily archive stays: it is the only copy that reaches Hetzner.Rejected for now: raw
archive_command(no restic/rclone in the image, a stalled archive fillspg_waland stops Postgres), pgBackRest (custom image, more config than one small cluster needs). WAL-G is the upgrade path if seconds-level RPO is ever required.Drill
Restore from the hourly repo and assert: snapshot at most 65 min old; a canary row written just before the dump comes back;
forgejo doctorpasses on the matched DB and repo pair. Export drill duration and restored-age metrics so RTO and RPO become measured values.Rollout
BackupHotStalealert (apply: backup, monitoring).Unverified: DB and repo-volume sizes (so storage cost), and whether Forgejo tolerates repos slightly newer than its DB. The job dumps the DB first so repos are never older than it.
The plan is approved.
Approved by the operator 2026-10-02: targets and the hourly-dump design as summarised above. Rollout starting with the ADR and the hourly job.
Reopened: #521 closed this on merge, but the first production drill run (2026-10-02 17:19 UTC) failed in the hot leg at
restic restore(exit 1, stderr not captured). The daily leg passed, and a manual restore of the same snapshot works (201.6 MiB, 7 s). Fix in progress. This closes when a drill passes both legs.Done. Drill on 2026-10-02 19:26 UTC passed both legs: daily archive restore with forgejo doctor and Kanidm (persons 9, groups 18), and the hourly hot leg (snapshot 6 s old, 18345 files restored, canary present, git fsck clean on 14 repos, 1 orphan dir reported). Metrics:
bitborg_backup_drill_last_run_status 0,bitborg_backup_drill_hot_status 0. The orphan issupernaut/migration-verify-scratch.git(an empty leftover dir); remove it under Site Administration, Unadopted Repositories.