feat(registry): retention policy, and grow the filesystem when a volume grows #310
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!310
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "feat/registry-retention"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Closes #297, closes #298.
#298 — growing a volume did not grow its filesystem
resizefs: trueadded to thecommunity.general.filesystemtasks for the data, LFS and backupvolumes. The backup volume was not named in the issue, but it has the identical one-line gap and the
runbook documented the identical manual
resize2fsfor it — leaving it as the only volume still carryingthe footgun would have been arbitrary.
Idempotent and a no-op once a filesystem already fills its device, so it costs nothing on a normal
converge and removes a trap that only ever appears during a capacity emergency — which is exactly when
nobody wants to discover it.
The root volume is deliberately NOT automated, and the issue asked the question. Its ext4 is created
by the cloud image, so there is no
community.general.filesystemtask to attachresizefsto; andbecause that disk is partitioned,
resize2fsalone cannot help —growparthas to move the partition endfirst. Automating that means writing the boot disk's partition table on every converge, to serve a
roughly annual operation that always follows a
tofu applythis playbook does not drive. The runbook nowcarries that reasoning next to the manual commands.
#297 — registry retention
A new
registry-retentionrole on a daily 22:10 timer, in the same shape asregistry-mirrorandtoken-audit. The policy is data, not code:sha-*tags — solatest, whichpodman auto-updatefollows, can never bedeleted;
bitborg-web/cacheversions older than 2 days. Age-based rather than count-based because asingle push writes ~10–12 cache versions, which is exactly why the measured count was 44 against 6 image
tags;
Reuses gitborg-ci's existing
write:packagePAT — no new secret and no admin scope.--dry-runwrites nometrics. There is a per-run cap of 500, and it is not run on deploy, because it deletes.
[cron.cleanup_packages]is pinned inapp.inibecause it is the load-bearing half. Deleting aversion only unlinks it; that cron is what frees the blobs. Retention without it would report success and
free nothing. This is also why applying #297 restarts Forgejo.
The trend alert would have caught this nine days earlier — the level threshold never fired at all
Replayed against production VictoriaMetrics, read-only.
/srv/gitborg-lfspeaked at 79.17% used(2026-07-31 18:00Z), so 80% was never crossed before the volume was grown. The level alert was not
late; it was never going to fire.
DiskWillFill(predict_linear) is true from 2026-07-22 00:00Z —nine days of lead — and returns nothing against today's post-grow state.
A 3-day regression window is used rather than the common 6 h, because writes here arrive in lumps.
Verified
--syntax-checkclean.ansible-lint: 17 failures on the three edited roles both with the change and onmain(pre-existingvar-naming), and the new role produces only the repo-wide hyphenatedrole-namefinding that
registry-mirrorandtoken-auditalso produce.shellcheckclean on the rendered script.The retention logic was exercised offline against a stubbed API (57 synthetic versions across 2 pages): it
kept the newest 20
sha-tags, deleted the 5 oldest, leftlatestand a mirroredgitborg/forgejountouched, deleted the 18 cache versions older than 48 h while keeping 12, and encoded the slash as
%2F. A dry run issued 0 DELETEs and wrote 0 metric files. With the cap lowered to 4 it stopped at 4 andset
..._capped 1. Metrics parse as valid exposition format with each family contiguous, mode 0644.Applying
--tags forgejo—resizefs: trueis a no-op on already-full filesystems;app.inigains[cron.cleanup_packages], so Forgejo restarts. Do it in a quiet window.--tags podman,backup— expect no changes;resizefsis a no-op today.--tags registry-retention— installs the script, env and units and starts the timer; it doesnot run the sweep. Rehearse first:
sudo -u gitborg /home/gitborg/bin/bitborg-registry-retention.sh --dry-run, and read the list beforeletting the timer fire. The first real run will hit the 500 cap and set
gitborg_registry_retention_capped— expected; the next run continues.df -h /srv/gitborg-lfswillnot move until after the following midnight, when blob GC runs.
--tags monitoringon the monitoring host — vmalert reloads the rule file.DiskWillFillshouldbe silent, which was verified against live data.
Note on
--check: it fabricatescommand/shellresults, so a dry run proves nothing about theretention job itself. Nothing here depends on command output in check mode, and the
systemd_servicetasks are gated on
not ansible_check_mode.#178 covers Actions artifacts, which are a different store and are not addressed here.
supernaut refererade till denna ändringsförfrågan2026-08-01 15:11:56 +00:00
a0e9309882c86bbde4f6