Backups: define RPO/RTO + Postgres PITR / WAL archiving #158

Stängd
öppnade 2026-07-19 21:51:34 +00:00 av supernaut · 6 kommentarer
Ägare

Split from #126 (backup-robustness umbrella).

Severity: HIGH — area/backups.

RPO up to ~24 h, no PITR

The DB backup is a single daily pg_dump (03:30, group_vars/all/vars.yml) with no WAL archiving → up to 24 h of pushes/issues/PRs lost on host loss. This was never set as an explicit target.

Ask

  • Document an explicit RPO/RTO target.
  • More-frequent DB dumps and/or Postgres WAL archiving off-site (continuous archiving → point-in-time recovery).

Refs #126.

> Split from #126 (backup-robustness umbrella). **Severity: HIGH** — `area/backups`. ## RPO up to ~24 h, no PITR The DB backup is a single daily `pg_dump` (03:30, `group_vars/all/vars.yml`) with no WAL archiving → up to **24 h** of pushes/issues/PRs lost on host loss. This was never set as an explicit target. ### Ask - Document an explicit **RPO/RTO** target. - More-frequent DB dumps and/or **Postgres WAL archiving off-site** (continuous archiving → point-in-time recovery). Refs #126.
Upphovsperson
Ägare
Epic: bitborg/bitborg-docs#107
Upphovsperson
Ägare

Research summary, 2026-10-02. Needs operator sign-off before any PR.

Today

  • One Postgres 17 cluster: forgejo, gitborg_web, and the billing DB (dumped only when billing_enabled). No WAL archiving. Kanidm is SQLite, dumped with kanidmd database backup.
  • Daily 03:30 pg_dump -Fc + Forgejo repo volume export, age-encrypted, local 14 d, off-site to Glesys and Hetzner. [storage] tier via restic at 05:00.
  • RPO about 24 h everywhere (BackupStale at 25 h). No RTO target. The 2026-09-30 drill restored in 555 s on a scratch VM; a full host rebuild has never been timed.

Constraint

Forgejo's DB and its git repos must restore to the same point. WAL PITR on the DB alone does not improve Forgejo's RPO: the repo volume needs the same cadence. WAL archiving is per cluster, so it is all three DBs or none.

Proposed targets

Tier RPO RTO
Billing (when enabled) 15 min 4 h
Portal, Kanidm 1 h 4 h
Forgejo DB + repo volume 1 h 4 h
[storage] (packages, LFS) 24 h 8 h

RTO 4 h is a guess until one full rebuild is timed.

Recommendation

Hourly dumps (DBs and Kanidm, -Z0) followed by a restic snapshot of the repo volume, into a new restic repo on Glesys, reusing the [storage] restic machinery. Keep 48 hourly + 14 daily. A 15-min billing dump gated on billing_enabled. The daily archive stays: it is the only copy that reaches Hetzner.

Rejected for now: raw archive_command (no restic/rclone in the image, a stalled archive fills pg_wal and stops Postgres), pgBackRest (custom image, more config than one small cluster needs). WAL-G is the upgrade path if seconds-level RPO is ever required.

Drill

Restore from the hourly repo and assert: snapshot at most 65 min old; a canary row written just before the dump comes back; forgejo doctor passes on the matched DB and repo pair. Export drill duration and restored-age metrics so RTO and RPO become measured values.

Rollout

  1. ADR + runbook targets (docs only).
  2. Hourly job + BackupHotStale alert (apply: backup, monitoring).
  3. Drill from the hourly repo with the new assertions (apply: backup-drill).
  4. 15-min billing dump, gated (apply, inert until billing is on).
  5. Time one full rebuild rehearsal (operator).

Unverified: DB and repo-volume sizes (so storage cost), and whether Forgejo tolerates repos slightly newer than its DB. The job dumps the DB first so repos are never older than it.

Research summary, 2026-10-02. Needs operator sign-off before any PR. ## Today - One Postgres 17 cluster: `forgejo`, `gitborg_web`, and the billing DB (dumped only when `billing_enabled`). No WAL archiving. Kanidm is SQLite, dumped with `kanidmd database backup`. - Daily 03:30 `pg_dump -Fc` + Forgejo repo volume export, age-encrypted, local 14 d, off-site to Glesys and Hetzner. `[storage]` tier via restic at 05:00. - RPO about 24 h everywhere (BackupStale at 25 h). No RTO target. The 2026-09-30 drill restored in 555 s on a scratch VM; a full host rebuild has never been timed. ## Constraint Forgejo's DB and its git repos must restore to the same point. WAL PITR on the DB alone does not improve Forgejo's RPO: the repo volume needs the same cadence. WAL archiving is per cluster, so it is all three DBs or none. ## Proposed targets | Tier | RPO | RTO | | --- | --- | --- | | Billing (when enabled) | 15 min | 4 h | | Portal, Kanidm | 1 h | 4 h | | Forgejo DB + repo volume | 1 h | 4 h | | `[storage]` (packages, LFS) | 24 h | 8 h | RTO 4 h is a guess until one full rebuild is timed. ## Recommendation Hourly dumps (DBs and Kanidm, `-Z0`) followed by a restic snapshot of the repo volume, into a new restic repo on Glesys, reusing the `[storage]` restic machinery. Keep 48 hourly + 14 daily. A 15-min billing dump gated on `billing_enabled`. The daily archive stays: it is the only copy that reaches Hetzner. Rejected for now: raw `archive_command` (no restic/rclone in the image, a stalled archive fills `pg_wal` and stops Postgres), pgBackRest (custom image, more config than one small cluster needs). WAL-G is the upgrade path if seconds-level RPO is ever required. ## Drill Restore from the hourly repo and assert: snapshot at most 65 min old; a canary row written just before the dump comes back; `forgejo doctor` passes on the matched DB and repo pair. Export drill duration and restored-age metrics so RTO and RPO become measured values. ## Rollout 1. ADR + runbook targets (docs only). 2. Hourly job + `BackupHotStale` alert (apply: backup, monitoring). 3. Drill from the hourly repo with the new assertions (apply: backup-drill). 4. 15-min billing dump, gated (apply, inert until billing is on). 5. Time one full rebuild rehearsal (operator). Unverified: DB and repo-volume sizes (so storage cost), and whether Forgejo tolerates repos slightly newer than its DB. The job dumps the DB first so repos are never older than it.
Upphovsperson
Ägare

The plan is approved.

The plan is approved.
Upphovsperson
Ägare

Approved by the operator 2026-10-02: targets and the hourly-dump design as summarised above. Rollout starting with the ADR and the hourly job.

Approved by the operator 2026-10-02: targets and the hourly-dump design as summarised above. Rollout starting with the ADR and the hourly job.
supernaut refererade till detta ärende från en incheckning 2026-10-02 11:59:36 +00:00
supernaut refererade till detta ärende från en incheckning 2026-10-02 11:59:36 +00:00
Upphovsperson
Ägare

Reopened: #521 closed this on merge, but the first production drill run (2026-10-02 17:19 UTC) failed in the hot leg at restic restore (exit 1, stderr not captured). The daily leg passed, and a manual restore of the same snapshot works (201.6 MiB, 7 s). Fix in progress. This closes when a drill passes both legs.

Reopened: #521 closed this on merge, but the first production drill run (2026-10-02 17:19 UTC) failed in the hot leg at `restic restore` (exit 1, stderr not captured). The daily leg passed, and a manual restore of the same snapshot works (201.6 MiB, 7 s). Fix in progress. This closes when a drill passes both legs.
supernaut återöppnade detta ärende 2026-10-02 17:41:52 +00:00
Upphovsperson
Ägare

Done. Drill on 2026-10-02 19:26 UTC passed both legs: daily archive restore with forgejo doctor and Kanidm (persons 9, groups 18), and the hourly hot leg (snapshot 6 s old, 18345 files restored, canary present, git fsck clean on 14 repos, 1 orphan dir reported). Metrics: bitborg_backup_drill_last_run_status 0, bitborg_backup_drill_hot_status 0. The orphan is supernaut/migration-verify-scratch.git (an empty leftover dir); remove it under Site Administration, Unadopted Repositories.

Done. Drill on 2026-10-02 19:26 UTC passed both legs: daily archive restore with forgejo doctor and Kanidm (persons 9, groups 18), and the hourly hot leg (snapshot 6 s old, 18345 files restored, canary present, git fsck clean on 14 repos, 1 orphan dir reported). Metrics: `bitborg_backup_drill_last_run_status 0`, `bitborg_backup_drill_hot_status 0`. The orphan is `supernaut/migration-verify-scratch.git` (an empty leftover dir); remove it under Site Administration, Unadopted Repositories.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#158
Ingen beskrivning angiven.