#188 follow-up: reclaim duplicate storage copy + object-level cutover verification #197

Stängd
öppnade 2026-07-21 22:39:11 +00:00 av supernaut · 1 kommentar
Ägare

Follow-up to the 2026-07-21 storage incident (#188 cutover left Forgejo [storage] on an empty ceph volume → avatars/packages/LFS 404'd instance-wide; restored by migrating the 9 GB of blobs with rsync).

1. Reclaim the duplicate old copy (after backup verification)

rsync copied (not moved), so the ~9 GB of blobs now exist on both the data volume (old default path) and the ceph volume. Once a post-cutover backup is confirmed to contain forgejo-storage.tar.gz (decrypt + tar -t), reclaim the old copy from the data volume to free ~9 GB and stop double-archiving:

OLD=/srv/gitborg-data/containers/storage/volumes/gitborg-forgejo-data/_data/gitea/data
for t in avatars packages actions_log …; do sudo rm -rf "$OLD/$t"/*; done

Destructive — do only after a clean soak + verified backup.

2. Object-level post-cutover verification (prevention)

The cutover was verified with --check and /api/healthz (200) — neither reads a stored object, so an empty target was invisible. This is the same shallow-verification failure mode as the #174 runner rebake (looked healthy, never ran a real job → #195/#196 territory).

Add a real post-change check that exercises the changed path:

  • a storage assertion (fetch a known avatar/package → expect 200, not 404), runnable after any storage-affecting apply;
  • consider a controller/monitoring probe for 'storage configured but target empty' (dir exists, 0 files, but the instance has users/packages).

Runbook now documents the migrate-data + fetch-a-real-object procedure (PR on docs/storage-cutover-migration).

area/ops, area/backup, type/enhancement

Follow-up to the 2026-07-21 storage incident (#188 cutover left Forgejo `[storage]` on an empty ceph volume → avatars/packages/LFS 404'd instance-wide; restored by migrating the 9 GB of blobs with rsync). ## 1. Reclaim the duplicate old copy (after backup verification) `rsync` *copied* (not moved), so the ~9 GB of blobs now exist on **both** the data volume (old default path) and the ceph volume. Once a post-cutover backup is confirmed to contain `forgejo-storage.tar.gz` (decrypt + `tar -t`), reclaim the old copy from the data volume to free ~9 GB and stop double-archiving: ``` OLD=/srv/gitborg-data/containers/storage/volumes/gitborg-forgejo-data/_data/gitea/data for t in avatars packages actions_log …; do sudo rm -rf "$OLD/$t"/*; done ``` Destructive — do only after a clean soak + verified backup. ## 2. Object-level post-cutover verification (prevention) The cutover was verified with `--check` and `/api/healthz` (200) — **neither reads a stored object**, so an empty target was invisible. This is the same shallow-verification failure mode as the #174 runner rebake (looked healthy, never ran a real job → #195/#196 territory). Add a real post-change check that exercises the changed path: - a storage assertion (fetch a known avatar/package → expect 200, not 404), runnable after any storage-affecting apply; - consider a controller/monitoring probe for 'storage configured but target empty' (dir exists, 0 files, but the instance has users/packages). Runbook now documents the migrate-data + fetch-a-real-object procedure (PR on `docs/storage-cutover-migration`). area/ops, area/backup, type/enhancement
Upphovsperson
Ägare

Both halves complete.

  1. Object-level verification — PR #205 (merged + applied): the ADR-0030 health gate pulls a custom-avatar hash from the DB and fetches /avatars/, failing the apply on non-200. Verified on prod (executes, 200). Would have caught the #188 cutover.
  2. Reclaim — the ~9 GB duplicate (old packages/avatars/actions_log) removed from the forgejo-data volume after a byte-exact safety gate. This also unblocked backups: the rsync double-copy (data.tar.gz + storage.tar.gz) had caused ENOSPC. Post-reclaim backup succeeded (bitborg-20260722T062816Z, 8.9 GB, back to pre-#188 footprint) and restore-verify confirmed forgejo-storage.tar.gz + every component restorable.

Also fixed the lost+found storage-tar bug found along the way (#199). Closing — the #188 storage chain is fully closed.

Both halves complete. 1. **Object-level verification** — PR #205 (merged + applied): the ADR-0030 health gate pulls a custom-avatar hash from the DB and fetches /avatars/<hash>, failing the apply on non-200. Verified on prod (executes, 200). Would have caught the #188 cutover. 2. **Reclaim** — the ~9 GB duplicate (old packages/avatars/actions_log) removed from the forgejo-data volume after a byte-exact safety gate. This also *unblocked backups*: the rsync double-copy (data.tar.gz + storage.tar.gz) had caused ENOSPC. Post-reclaim backup succeeded (bitborg-20260722T062816Z, 8.9 GB, back to pre-#188 footprint) and **restore-verify confirmed forgejo-storage.tar.gz + every component restorable**. Also fixed the lost+found storage-tar bug found along the way (#199). Closing — the #188 storage chain is fully closed.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#197
Ingen beskrivning angiven.