Route LFS/asset storage to a cheaper ceph tier (keep repos + Postgres on ceph-ssd) #188
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#188
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Route Forgejo's bulk object storage (LFS + packages + attachments — Forgejo's
[storage] PATH) off the ceph-ssd data volume onto a cheaper standard-cephCinder volume, keeping repos + the Postgres graphroot (ADR 0025) on ceph-ssd for git/DB latency. LFS/asset bytes are large, cold and sequential — a good fit for the slower, cheaper tier (standardceph~0,54 vs ceph-ssd ~1,08 kr/GB·mo, Bahnhof published unit prices). Forgejo quota still counts these bytes wherever they live, so tier gating is unaffected.Measured 2026-07-21: LFS is empty today, so this is a near-zero-migration change — best done before data accumulates.
Design: new standard-
cephCinder volume → format + mount/srv/gitborg-lfs→ bind into Forgejo →[storage] PATH. Repos + Postgres stay on ceph-ssd. Required companion change: extend the backup to archive the new volume (moving assets to a bind mount removes them frompodman volume export), and the restore-drill path.Relates: #45 (HA native S3-backed repo/LFS storage), #184/#185 (quota model), data-tier decoupling.
Scoped + measured (2026-07-21). Current LFS footprint on prod:
data/lfsdoes not exist yet (0 bytes),git16K, wholegitborg-forgejo-datavolume 9 G, data volume 39% full. So this is a near-zero-migration change — no maintenance-window rsync needed; Forgejo creates the LFS dir on the new volume on first use. Best done now, before LFS data accumulates.Design (dedicated-volume route): new
openstack_blockstorage_volume_v3 "lfs"(standardceph) + attach (mirror thebackupvolume pattern in opentofu/compute.tofu) → format+mount/srv/gitborg-lfs(copy backup role's mount tasks;podman unshare chown -R 1000:1000) → bind into Forgejo (Volume=/srv/gitborg-lfs:/srv/lfs) →[lfs] PATH = /srv/lfsin app.ini. Repos + Postgres stay onceph-ssd.Required companion change: the backup currently captures LFS only because it lives inside the
gitborg-forgejo-datanamed volume (podman volume export). Moving LFS to a bind mount removes it from the archive — sobitborg-backup.sh.j2+ the restore-drill path MUST be extended to also archive/srv/gitborg-lfs, or LFS silently drops out of DR.Open decisions: (1) scope — LFS-only (
[lfs] PATH) vs all bulk assets incl. packages/attachments ([storage] PATH); (2) volume type —ceph(known price) vsceph-ec(cheaper, unpriced); (3) size (growable online). Relates #45.OpenTofu layer done (branch
feat/188-lfs-storage-tier): dedicated standard-cephvolume + attach +lfs_volume_idoutput;tofu validate+ fmt clean.Ansible layer (next), verified against the live setup — rootful image, data volume at
/data(WORK_PATH=/data/gitea), backup captures assets viapodman volume exportof the forgejo named volume:group_vars/all/vars.yml:lfs_volume_id(fromtofu output -raw lfs_volume_idafter the tofu apply), mirroringbackup_volume_id.roles/forgejo/defaults:…_lfs_volume_device: /dev/disk/by-id/virtio-{{ lfs_volume_id[:20] }}, mount dir/srv/gitborg-lfs, container path/srv/storage.roles/forgejo/tasks: format + mount/srv/gitborg-lfsby Cinder serial (mirror the backup role, fail-closed), thenpodman unshare chownto the container git uid so Forgejo can write.forgejo.container.j2: addVolume=/srv/gitborg-lfs:/srv/storage:Z.app.ini.j2: add[storage] PATH = /srv/storage(relocates LFS + packages + attachments; repos stay in/dataon ceph-ssd).bitborg-backup.sh.j2+ backup-drill restore: also archive/srv/gitborg-lfs— REQUIRED, else assets drop out of DR (they leave the named volume the export captures).Two-phase gated apply (infra-apply): (1)
tofu apply→ setlfs_volume_id; (2)ansible-playbook site.yml --tags forgejo,backup --limit bitborg→ mount, relocate storage, restart Forgejo (brief blip), extend backup. Verify LFS push/pull + a package push + backup includes/srv/gitborg-lfs. Near-zero migration (assets empty today).Delivered + live. Shipped the dedicated cheaper-ceph volume for Forgejo [storage] (LFS/packages/attachments/avatars/actions) in #192/#193 and cut over to prod in #194; backups extended to archive/restore/verify forgejo-storage.tar.gz.
Post-cutover incident (2026-07-21): the cutover repointed [storage] at the fresh volume without migrating existing blobs, so avatars/packages 404'd instance-wide (healthz stayed 200). Resolved by rsync-migrating the 9 GB byte-exact + verifying by fetching a real avatar (200). Runbook now documents the mandatory migrate-data + fetch-a-real-object procedure (#198).
Residual work tracked in #197 (reclaim the now-duplicate old copy after backup verification; add object-level post-cutover verification). Closing — the storage-tier split is complete.