fix(backup): back up + verify + drill the portal DB (gitborg_web) #149

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/122-portal-db-backup in i main 2026-07-19 19:23:57 +00:00
Ägare

Problem (#122 — CRITICAL, pre-onboarding)

The portal DB gitborg_web (same Postgres as forgejo; created by the web role) holds signup accounts, consent/GDPR records, invites, subscriptions, sessions — and the backup only dumps -d forgejo, so it's captured by nothing. On host loss, portal state is unrecoverable. Worse, the weekly verify and the restore drill are also blind to it (both only touch forgejo), so forgejo doctor stays green while gitborg_web is silently absent from every archive.

Fix (close the loop across all three)

  1. Backup (bitborg-backup.sh.j2) — add a custom-format pg_dump -d gitborg_web → gitborg_web.dump into the bundle (fail-closed: empty dump aborts the run). Targeted per-DB dump keeps the existing pg_restore flow; roles/globals are re-provisioned by Ansible on restore.
  2. Verify (bitborg-backup-verify.sh.j2) — pg_restore --list gitborg_web.dump (TOC intact / restorable).
  3. Drill (restore-on-scratch.sh.j2) — restore gitborg_web into the scratch Postgres and assert it has ≥1 public table (schema + data round-trip), so the drill actually proves portal state is recoverable.

Backward-compatible: the verify + drill checks are guarded on gitborg_web.dump being present, so a transition-period run against a pre-#122 archive skips cleanly; every archive written after this becomes an always-on check.

Design note: targeted pg_dump -d gitborg_web (custom format) rather than pg_dumpall — keeps the custom-format/pg_restore path the drill was just stabilized on; Ansible re-provisions roles on restore.

Validation

ansible-lint (production profile) · --check --diff renders clean (0 failed) · bash -n on all three rendered scripts.

Verify after apply

--tags backup,backup-drill, then one backup run (archive contains gitborg_web.dump) + one drill run (asserts the portal DB round-trips).

Closes #122.

## Problem (#122 — CRITICAL, pre-onboarding) The portal DB `gitborg_web` (same Postgres as forgejo; created by the web role) holds **signup accounts, consent/GDPR records, invites, subscriptions, sessions** — and the backup only dumps `-d forgejo`, so it's captured by **nothing**. On host loss, portal state is unrecoverable. Worse, the weekly verify and the restore drill are also blind to it (both only touch forgejo), so `forgejo doctor` stays green while `gitborg_web` is silently absent from every archive. ## Fix (close the loop across all three) 1. **Backup** (`bitborg-backup.sh.j2`) — add a custom-format `pg_dump -d gitborg_web` → `gitborg_web.dump` into the bundle (fail-closed: empty dump aborts the run). Targeted per-DB dump keeps the existing `pg_restore` flow; roles/globals are re-provisioned by Ansible on restore. 2. **Verify** (`bitborg-backup-verify.sh.j2`) — `pg_restore --list gitborg_web.dump` (TOC intact / restorable). 3. **Drill** (`restore-on-scratch.sh.j2`) — restore `gitborg_web` into the scratch Postgres and assert it has ≥1 public table (schema + data round-trip), so the drill actually **proves** portal state is recoverable. **Backward-compatible:** the verify + drill checks are guarded on `gitborg_web.dump` being present, so a transition-period run against a pre-#122 archive skips cleanly; every archive written after this becomes an always-on check. **Design note:** targeted `pg_dump -d gitborg_web` (custom format) rather than `pg_dumpall` — keeps the custom-format/`pg_restore` path the drill was just stabilized on; Ansible re-provisions roles on restore. ## Validation ansible-lint (production profile) · `--check --diff` renders clean (0 failed) · `bash -n` on all three rendered scripts. ## Verify after apply `--tags backup,backup-drill`, then one backup run (archive contains `gitborg_web.dump`) + one drill run (asserts the portal DB round-trips). Closes #122.
supernaut lade till 1 incheckning 2026-07-19 19:18:32 +00:00
fix(backup): back up + verify + drill the portal DB (gitborg_web)
Alla kontroller lyckades
ci / ci (pull_request) Successful in 2m51s
2dffe8d14c
The portal DB gitborg_web (accounts, consent/GDPR, invites,
subscriptions, sessions) shared the forgejo Postgres but was dumped by
nothing — and the verify + drill were blind to it, so forgejo doctor
stayed green while portal state was silently absent from every archive
(#122, CRITICAL, pre-onboarding).

- backup: add a custom-format pg_dump of gitborg_web to the bundle
  (fail-closed on empty).
- verify: pg_restore --list gitborg_web.dump.
- drill: restore gitborg_web into the scratch Postgres and assert >=1
  public table (proves the portal DB round-trips, not just forgejo).

Verify + drill checks are guarded on the dump's presence so a run
against a pre-#122 archive skips cleanly; new archives make it always-on.

Validated: ansible-lint (production), --check renders clean, bash -n on
all three rendered scripts.

Closes #122.
supernaut sammanfogade incheckning adcf297ab6 till main 2026-07-19 19:23:57 +00:00
supernaut tog bort grenen fix/122-portal-db-backup 2026-07-19 19:23:57 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!149
Ingen beskrivning angiven.