Route LFS/asset storage to a cheaper ceph tier (keep repos + Postgres on ceph-ssd) #188

Stängd
öppnade 2026-07-21 19:09:16 +00:00 av supernaut · 3 kommentarer
Ägare

Route Forgejo's bulk object storage (LFS + packages + attachments — Forgejo's [storage] PATH) off the ceph-ssd data volume onto a cheaper standard-ceph Cinder volume, keeping repos + the Postgres graphroot (ADR 0025) on ceph-ssd for git/DB latency. LFS/asset bytes are large, cold and sequential — a good fit for the slower, cheaper tier (standard ceph ~0,54 vs ceph-ssd ~1,08 kr/GB·mo, Bahnhof published unit prices). Forgejo quota still counts these bytes wherever they live, so tier gating is unaffected.

Measured 2026-07-21: LFS is empty today, so this is a near-zero-migration change — best done before data accumulates.

Design: new standard-ceph Cinder volume → format + mount /srv/gitborg-lfs → bind into Forgejo → [storage] PATH. Repos + Postgres stay on ceph-ssd. Required companion change: extend the backup to archive the new volume (moving assets to a bind mount removes them from podman volume export), and the restore-drill path.

Relates: #45 (HA native S3-backed repo/LFS storage), #184/#185 (quota model), data-tier decoupling.

Route Forgejo's bulk object storage (LFS + packages + attachments — Forgejo's `[storage] PATH`) off the ceph-ssd data volume onto a cheaper standard-`ceph` Cinder volume, keeping repos + the Postgres graphroot (ADR 0025) on ceph-ssd for git/DB latency. LFS/asset bytes are large, cold and sequential — a good fit for the slower, cheaper tier (standard `ceph` ~0,54 vs ceph-ssd ~1,08 kr/GB·mo, Bahnhof published unit prices). Forgejo quota still counts these bytes wherever they live, so tier gating is unaffected. Measured 2026-07-21: LFS is empty today, so this is a **near-zero-migration** change — best done before data accumulates. Design: new standard-`ceph` Cinder volume → format + mount `/srv/gitborg-lfs` → bind into Forgejo → `[storage] PATH`. Repos + Postgres stay on ceph-ssd. **Required companion change:** extend the backup to archive the new volume (moving assets to a bind mount removes them from `podman volume export`), and the restore-drill path. Relates: #45 (HA native S3-backed repo/LFS storage), #184/#185 (quota model), data-tier decoupling.
Upphovsperson
Ägare

Scoped + measured (2026-07-21). Current LFS footprint on prod: data/lfs does not exist yet (0 bytes), git 16K, whole gitborg-forgejo-data volume 9 G, data volume 39% full. So this is a near-zero-migration change — no maintenance-window rsync needed; Forgejo creates the LFS dir on the new volume on first use. Best done now, before LFS data accumulates.

Design (dedicated-volume route): new openstack_blockstorage_volume_v3 "lfs" (standard ceph) + attach (mirror the backup volume pattern in opentofu/compute.tofu) → format+mount /srv/gitborg-lfs (copy backup role's mount tasks; podman unshare chown -R 1000:1000) → bind into Forgejo (Volume=/srv/gitborg-lfs:/srv/lfs) → [lfs] PATH = /srv/lfs in app.ini. Repos + Postgres stay on ceph-ssd.

Required companion change: the backup currently captures LFS only because it lives inside the gitborg-forgejo-data named volume (podman volume export). Moving LFS to a bind mount removes it from the archive — so bitborg-backup.sh.j2 + the restore-drill path MUST be extended to also archive /srv/gitborg-lfs, or LFS silently drops out of DR.

Open decisions: (1) scope — LFS-only ([lfs] PATH) vs all bulk assets incl. packages/attachments ([storage] PATH); (2) volume type — ceph (known price) vs ceph-ec (cheaper, unpriced); (3) size (growable online). Relates #45.

**Scoped + measured (2026-07-21).** Current LFS footprint on prod: `data/lfs` does not exist yet (0 bytes), `git` 16K, whole `gitborg-forgejo-data` volume 9 G, data volume 39% full. So this is a **near-zero-migration** change — no maintenance-window rsync needed; Forgejo creates the LFS dir on the new volume on first use. Best done now, before LFS data accumulates. Design (dedicated-volume route): new `openstack_blockstorage_volume_v3 "lfs"` (standard `ceph`) + attach (mirror the `backup` volume pattern in opentofu/compute.tofu) → format+mount `/srv/gitborg-lfs` (copy backup role's mount tasks; `podman unshare chown -R 1000:1000`) → bind into Forgejo (`Volume=/srv/gitborg-lfs:/srv/lfs`) → `[lfs] PATH = /srv/lfs` in app.ini. Repos + Postgres stay on `ceph-ssd`. **Required companion change:** the backup currently captures LFS only because it lives inside the `gitborg-forgejo-data` named volume (`podman volume export`). Moving LFS to a bind mount removes it from the archive — so `bitborg-backup.sh.j2` + the restore-drill path MUST be extended to also archive `/srv/gitborg-lfs`, or LFS silently drops out of DR. Open decisions: (1) scope — LFS-only (`[lfs] PATH`) vs all bulk assets incl. packages/attachments (`[storage] PATH`); (2) volume type — `ceph` (known price) vs `ceph-ec` (cheaper, unpriced); (3) size (growable online). Relates #45.
Upphovsperson
Ägare

OpenTofu layer done (branch feat/188-lfs-storage-tier): dedicated standard-ceph volume + attach + lfs_volume_id output; tofu validate + fmt clean.

Ansible layer (next), verified against the live setup — rootful image, data volume at /data (WORK_PATH=/data/gitea), backup captures assets via podman volume export of the forgejo named volume:

  • group_vars/all/vars.yml: lfs_volume_id (from tofu output -raw lfs_volume_id after the tofu apply), mirroring backup_volume_id.
  • roles/forgejo/defaults: …_lfs_volume_device: /dev/disk/by-id/virtio-{{ lfs_volume_id[:20] }}, mount dir /srv/gitborg-lfs, container path /srv/storage.
  • roles/forgejo/tasks: format + mount /srv/gitborg-lfs by Cinder serial (mirror the backup role, fail-closed), then podman unshare chown to the container git uid so Forgejo can write.
  • forgejo.container.j2: add Volume=/srv/gitborg-lfs:/srv/storage:Z.
  • app.ini.j2: add [storage] PATH = /srv/storage (relocates LFS + packages + attachments; repos stay in /data on ceph-ssd).
  • bitborg-backup.sh.j2 + backup-drill restore: also archive /srv/gitborg-lfs — REQUIRED, else assets drop out of DR (they leave the named volume the export captures).

Two-phase gated apply (infra-apply): (1) tofu apply → set lfs_volume_id; (2) ansible-playbook site.yml --tags forgejo,backup --limit bitborg → mount, relocate storage, restart Forgejo (brief blip), extend backup. Verify LFS push/pull + a package push + backup includes /srv/gitborg-lfs. Near-zero migration (assets empty today).

**OpenTofu layer done** (branch `feat/188-lfs-storage-tier`): dedicated standard-`ceph` volume + attach + `lfs_volume_id` output; `tofu validate` + fmt clean. **Ansible layer (next), verified against the live setup** — rootful image, data volume at `/data` (`WORK_PATH=/data/gitea`), backup captures assets via `podman volume export` of the forgejo named volume: - `group_vars/all/vars.yml`: `lfs_volume_id` (from `tofu output -raw lfs_volume_id` after the tofu apply), mirroring `backup_volume_id`. - `roles/forgejo/defaults`: `…_lfs_volume_device: /dev/disk/by-id/virtio-{{ lfs_volume_id[:20] }}`, mount dir `/srv/gitborg-lfs`, container path `/srv/storage`. - `roles/forgejo/tasks`: format + mount `/srv/gitborg-lfs` by Cinder serial (mirror the backup role, fail-closed), then `podman unshare chown` to the container git uid so Forgejo can write. - `forgejo.container.j2`: add `Volume=/srv/gitborg-lfs:/srv/storage:Z`. - `app.ini.j2`: add `[storage] PATH = /srv/storage` (relocates LFS + packages + attachments; repos stay in `/data` on ceph-ssd). - **`bitborg-backup.sh.j2` + backup-drill restore: also archive `/srv/gitborg-lfs`** — REQUIRED, else assets drop out of DR (they leave the named volume the export captures). Two-phase gated apply (infra-apply): (1) `tofu apply` → set `lfs_volume_id`; (2) `ansible-playbook site.yml --tags forgejo,backup --limit bitborg` → mount, relocate storage, restart Forgejo (brief blip), extend backup. Verify LFS push/pull + a package push + backup includes `/srv/gitborg-lfs`. Near-zero migration (assets empty today).
Upphovsperson
Ägare

Delivered + live. Shipped the dedicated cheaper-ceph volume for Forgejo [storage] (LFS/packages/attachments/avatars/actions) in #192/#193 and cut over to prod in #194; backups extended to archive/restore/verify forgejo-storage.tar.gz.

Post-cutover incident (2026-07-21): the cutover repointed [storage] at the fresh volume without migrating existing blobs, so avatars/packages 404'd instance-wide (healthz stayed 200). Resolved by rsync-migrating the 9 GB byte-exact + verifying by fetching a real avatar (200). Runbook now documents the mandatory migrate-data + fetch-a-real-object procedure (#198).

Residual work tracked in #197 (reclaim the now-duplicate old copy after backup verification; add object-level post-cutover verification). Closing — the storage-tier split is complete.

Delivered + live. Shipped the dedicated cheaper-ceph volume for Forgejo [storage] (LFS/packages/attachments/avatars/actions) in #192/#193 and cut over to prod in #194; backups extended to archive/restore/verify forgejo-storage.tar.gz. Post-cutover incident (2026-07-21): the cutover repointed [storage] at the fresh volume without migrating existing blobs, so avatars/packages 404'd instance-wide (healthz stayed 200). Resolved by rsync-migrating the 9 GB byte-exact + verifying by fetching a real avatar (200). Runbook now documents the mandatory migrate-data + fetch-a-real-object procedure (#198). Residual work tracked in #197 (reclaim the now-duplicate old copy after backup verification; add object-level post-cutover verification). Closing — the storage-tier split is complete.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#188
Ingen beskrivning angiven.