feat(registry): retention policy, and grow the filesystem when a volume grows #310

Sammanfogat
supernaut sammanfogade 2 incheckningar från feat/registry-retention in i main 2026-08-01 17:15:44 +00:00
Ägare

Closes #297, closes #298.

#298 — growing a volume did not grow its filesystem

resizefs: true added to the community.general.filesystem tasks for the data, LFS and backup
volumes. The backup volume was not named in the issue, but it has the identical one-line gap and the
runbook documented the identical manual resize2fs for it — leaving it as the only volume still carrying
the footgun would have been arbitrary.

Idempotent and a no-op once a filesystem already fills its device, so it costs nothing on a normal
converge and removes a trap that only ever appears during a capacity emergency — which is exactly when
nobody wants to discover it.

The root volume is deliberately NOT automated, and the issue asked the question. Its ext4 is created
by the cloud image, so there is no community.general.filesystem task to attach resizefs to; and
because that disk is partitioned, resize2fs alone cannot help — growpart has to move the partition end
first. Automating that means writing the boot disk's partition table on every converge, to serve a
roughly annual operation that always follows a tofu apply this playbook does not drive. The runbook now
carries that reasoning next to the manual commands.

#297 — registry retention

A new registry-retention role on a daily 22:10 timer, in the same shape as registry-mirror and
token-audit. The policy is data, not code:

  • keep the newest 20 sha-* tags — so latest, which podman auto-update follows, can never be
    deleted;
  • delete bitborg-web/cache versions older than 2 days. Age-based rather than count-based because a
    single push writes ~10–12 cache versions, which is exactly why the measured count was 44 against 6 image
    tags;
  • packages not named in a rule are never touched, which is what protects the mirrored images.

Reuses gitborg-ci's existing write:package PAT — no new secret and no admin scope. --dry-run writes no
metrics. There is a per-run cap of 500, and it is not run on deploy, because it deletes.

[cron.cleanup_packages] is pinned in app.ini because it is the load-bearing half. Deleting a
version only unlinks it; that cron is what frees the blobs. Retention without it would report success and
free nothing. This is also why applying #297 restarts Forgejo.

The trend alert would have caught this nine days earlier — the level threshold never fired at all

Replayed against production VictoriaMetrics, read-only. /srv/gitborg-lfs peaked at 79.17% used
(2026-07-31 18:00Z), so 80% was never crossed before the volume was grown. The level alert was not
late; it was never going to fire. DiskWillFill (predict_linear) is true from 2026-07-22 00:00Z —
nine days of lead — and returns nothing against today's post-grow state.

A 3-day regression window is used rather than the common 6 h, because writes here arrive in lumps.

Verified

--syntax-check clean. ansible-lint: 17 failures on the three edited roles both with the change and on
main (pre-existing var-naming), and the new role produces only the repo-wide hyphenated role-name
finding that registry-mirror and token-audit also produce. shellcheck clean on the rendered script.

The retention logic was exercised offline against a stubbed API (57 synthetic versions across 2 pages): it
kept the newest 20 sha- tags, deleted the 5 oldest, left latest and a mirrored gitborg/forgejo
untouched, deleted the 18 cache versions older than 48 h while keeping 12, and encoded the slash as
%2F. A dry run issued 0 DELETEs and wrote 0 metric files. With the cap lowered to 4 it stopped at 4 and
set ..._capped 1. Metrics parse as valid exposition format with each family contiguous, mode 0644.

Applying

  1. --tags forgejo — resizefs: true is a no-op on already-full filesystems; app.ini gains
    [cron.cleanup_packages], so Forgejo restarts. Do it in a quiet window.
  2. --tags podman,backup — expect no changes; resizefs is a no-op today.
  3. --tags registry-retention — installs the script, env and units and starts the timer; it does
    not run the sweep. Rehearse first:
    sudo -u gitborg /home/gitborg/bin/bitborg-registry-retention.sh --dry-run, and read the list before
    letting the timer fire. The first real run will hit the 500 cap and set
    gitborg_registry_retention_capped — expected; the next run continues. df -h /srv/gitborg-lfs will
    not move until after the following midnight, when blob GC runs.
  4. --tags monitoring on the monitoring host — vmalert reloads the rule file. DiskWillFill should
    be silent, which was verified against live data.

Note on --check: it fabricates command/shell results, so a dry run proves nothing about the
retention job itself. Nothing here depends on command output in check mode, and the systemd_service
tasks are gated on not ansible_check_mode.

#178 covers Actions artifacts, which are a different store and are not addressed here.

Closes #297, closes #298. ## #298 — growing a volume did not grow its filesystem `resizefs: true` added to the `community.general.filesystem` tasks for the data, LFS **and backup** volumes. The backup volume was not named in the issue, but it has the identical one-line gap and the runbook documented the identical manual `resize2fs` for it — leaving it as the only volume still carrying the footgun would have been arbitrary. Idempotent and a no-op once a filesystem already fills its device, so it costs nothing on a normal converge and removes a trap that only ever appears during a capacity emergency — which is exactly when nobody wants to discover it. **The root volume is deliberately NOT automated**, and the issue asked the question. Its ext4 is created by the cloud image, so there is no `community.general.filesystem` task to attach `resizefs` to; and because that disk is partitioned, `resize2fs` alone cannot help — `growpart` has to move the partition end first. Automating that means writing the boot disk's partition table on every converge, to serve a roughly annual operation that always follows a `tofu apply` this playbook does not drive. The runbook now carries that reasoning next to the manual commands. ## #297 — registry retention A new `registry-retention` role on a daily 22:10 timer, in the same shape as `registry-mirror` and `token-audit`. The policy is data, not code: - keep the newest **20** `sha-*` tags — so `latest`, which `podman auto-update` follows, can never be deleted; - delete `bitborg-web/cache` versions older than **2 days**. Age-based rather than count-based because a single push writes ~10–12 cache versions, which is exactly why the measured count was 44 against 6 image tags; - packages not named in a rule are **never** touched, which is what protects the mirrored images. Reuses gitborg-ci's existing `write:package` PAT — no new secret and no admin scope. `--dry-run` writes no metrics. There is a per-run cap of 500, and it is not run on deploy, because it deletes. **`[cron.cleanup_packages]` is pinned in `app.ini` because it is the load-bearing half.** Deleting a version only unlinks it; that cron is what frees the blobs. Retention without it would report success and free nothing. **This is also why applying #297 restarts Forgejo.** ## The trend alert would have caught this nine days earlier — the level threshold never fired at all Replayed against production VictoriaMetrics, read-only. `/srv/gitborg-lfs` peaked at **79.17%** used (2026-07-31 18:00Z), so **80% was never crossed** before the volume was grown. The level alert was not late; it was never going to fire. `DiskWillFill` (`predict_linear`) is true from **2026-07-22 00:00Z** — nine days of lead — and returns nothing against today's post-grow state. A 3-day regression window is used rather than the common 6 h, because writes here arrive in lumps. ## Verified `--syntax-check` clean. `ansible-lint`: 17 failures on the three edited roles both with the change and on `main` (pre-existing `var-naming`), and the new role produces only the repo-wide hyphenated `role-name` finding that `registry-mirror` and `token-audit` also produce. `shellcheck` clean on the rendered script. The retention logic was exercised offline against a stubbed API (57 synthetic versions across 2 pages): it kept the newest 20 `sha-` tags, deleted the 5 oldest, left `latest` and a mirrored `gitborg/forgejo` untouched, deleted the 18 cache versions older than 48 h while keeping 12, and encoded the slash as `%2F`. A dry run issued 0 DELETEs and wrote 0 metric files. With the cap lowered to 4 it stopped at 4 and set `..._capped 1`. Metrics parse as valid exposition format with each family contiguous, mode 0644. ## Applying 1. **`--tags forgejo`** — `resizefs: true` is a no-op on already-full filesystems; `app.ini` gains `[cron.cleanup_packages]`, so **Forgejo restarts**. Do it in a quiet window. 2. **`--tags podman,backup`** — expect **no changes**; `resizefs` is a no-op today. 3. **`--tags registry-retention`** — installs the script, env and units and starts the timer; it does **not** run the sweep. **Rehearse first:** `sudo -u gitborg /home/gitborg/bin/bitborg-registry-retention.sh --dry-run`, and read the list before letting the timer fire. The first real run will hit the 500 cap and set `gitborg_registry_retention_capped` — expected; the next run continues. `df -h /srv/gitborg-lfs` will not move until after the following midnight, when blob GC runs. 4. **`--tags monitoring`** on the monitoring host — vmalert reloads the rule file. `DiskWillFill` should be silent, which was verified against live data. **Note on `--check`:** it fabricates `command`/`shell` results, so a dry run proves nothing about the retention job itself. Nothing here depends on command output in check mode, and the `systemd_service` tasks are gated on `not ansible_check_mode`. #178 covers Actions artifacts, which are a different store and are not addressed here.
supernaut lade till 2 incheckningar 2026-08-01 14:35:24 +00:00
Growing a Cinder volume does not grow its filesystem, and nothing in the
playbook closed the gap. `community.general.filesystem` created a filesystem
but never resized one, so bumping `lfs_volume_size` / `data_volume_size` and
applying left the block device larger and the filesystem exactly as it was.
On 2026-08-01 that meant `lsblk` reporting 60G while `df` still showed 20G
with 4.1G free, until `resize2fs` was run by hand.

The failure is invisible in the places you look: the plan says `~ size`, the
apply is green, and monitoring keeps showing a nearly full disk with nothing
linking the two. It only ever bites during a capacity emergency.

Add `resizefs: true` to the `filesystem` task for the LFS/storage volume
(forgejo), the data volume (podman) and — same footgun, same one-line fix, and
the runbook documented the same manual step for it — the backup volume. ext4
grows online and the option is a no-op once the filesystem already fills the
device, so a normal converge is unaffected.

Root is deliberately left manual: its ext4 is created by the cloud image (no
Ansible task owns it), and because the disk is partitioned `resize2fs` alone is
insufficient — `growpart` must move the partition end first. Automating that
would write the boot disk's partition table on every converge to serve a
roughly annual operation that always follows a `tofu apply` the playbook does
not drive. The runbook now states that reasoning where the manual commands are.

Verified: `ansible-playbook site.yml --syntax-check` clean; `ansible-lint` on
the three roles reports the same 17 pre-existing var-naming findings as `main`
(no new ones).
feat(registry): retention policy for the package registry (#297)
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m28s
a0e9309882
`/srv/gitborg-lfs` (Forgejo's `[storage]` — LFS, packages, attachments,
avatars) reached 79% with 4 GB free on 2026-08-01, roughly a week from full. It
was grown 20→60 GB as an emergency measure, but that bought runway, it did not
fix the leak.

Cause: gitborg-web's deploy workflow pushes three things on every merge to
`main` — `latest`, an immutable `sha-XXXX` tag, and a BuildKit registry layer
cache via `--cache-to` — and nothing pruned the registry. Measured the same day,
`gitborg-web` had 6 versions while `gitborg-web/cache` had 44: the layer cache
is the dominant consumer, not the image tags, which is worth stating because
the obvious guess is wrong. Drift was ~0.6 GB/day. (`ARTIFACT_RETENTION_DAYS`
covers Actions artifacts — a different store — and does not help here.)

New `registry-retention` role: a gitborg-user oneshot service plus a daily
timer, the same shape as `registry-mirror` and `token-audit` (host script +
EnvironmentFile + systemd user timer, textfile metrics, inert until its
credential is vaulted). Policy is data in `registry_retention_rules`: keep the
newest 20 `sha-*` tags of gitborg-web, delete `gitborg-web/cache` versions
older than 2 days. Notable choices:

- `latest` cannot be deleted (the rule's `^sha-` match excludes it) — that tag
  is what `podman auto-update` follows.
- Packages not named in a rule are never touched, which is what keeps the
  registry-mirror images safe from a job that shares gitborg-ci's token.
- The cache rule is age-based, not count-based: one push writes ~10-12 cache
  versions, so any `keep_newest` would be a magic number that shreds the
  current cache the moment the layer count changes.
- `--dry-run` lists what a run would delete and deliberately does not write
  metrics, so a rehearsal cannot refresh `last_run_timestamp` and hide a dead
  timer. There is no undo for a deleted version.
- A per-run deletion cap (500) stops a mistyped rule emptying the registry in
  one pass; hitting it sets a metric and the next run continues.
- Deliberately NOT run on deploy, unlike the mirror. A converge is the wrong
  moment to fire a job that deletes.
- Reuses gitborg-ci's existing `write:package` PAT. Package deletion needs
  write on the owner's packages and nothing more — no new secret, no admin
  scope.

Also pin `[cron.cleanup_packages]` in app.ini. This is the load-bearing half:
deleting a version only unlinks it, and that cron is what sweeps the
unreferenced blobs and actually returns the bytes. Pinned rather than inherited
because an upstream default change would silently stop reclaiming space, and
the symptom — retention runs clean, `df` never moves — is not one anybody would
trace back to a default. @midnight follows the 22:10 sweep by ~2 h.

Alerting, per the issue's "alert on the trend, not just the level":

- `DiskWillFill` — `predict_linear` over a 3-day window, projected
  `alert_disk_predict_days` (14) ahead, gated on >40% used. Replayed against
  production VictoriaMetrics the level thresholds are worse than late: they
  never fired, because the volume peaked at 79.17% used and 80% was never
  crossed before it was grown. The same expression is true from 2026-07-22
  00:00Z, nine days of lead. Against today's post-grow state it returns nothing.
  The 3-day window (not the mixin's 6h) is deliberate — writes here arrive in
  lumps, and one big LFS push must not read as a trend.
- `RegistryRetentionFailed`, `RegistryRetentionStale` (deadman — a dead timer
  never sets a non-zero status, it just stops updating) and
  `RegistryRetentionCapped`.

Verified:
- `ansible-playbook site.yml --syntax-check` clean.
- `ansible-lint` on the new role: only the repo-wide hyphenated `role-name`
  finding that `registry-mirror` and `token-audit` also produce.
- `shellcheck` clean on the rendered script.
- Script logic exercised offline against a stubbed API (57 synthetic versions
  over 2 pages): kept the newest 20 `sha-` tags and deleted the 5 oldest, left
  `latest` and a mirrored `gitborg/forgejo` image untouched, deleted the 18
  cache versions older than 48 h and kept 12, and encoded the slash in
  `gitborg-web/cache` as `%2F`. Dry run issued 0 DELETEs and wrote no metrics.
  With the cap set to 4 it stopped at 4 and set `..._capped 1`.
- Emitted metrics parse as valid exposition format with each family contiguous
  (the trap that previously made a whole textfile unreadable) and land 0644.
- Alert rules render to valid YAML (38 rules); `DiskWillFill` executed as a
  read-only instant and range query against production VictoriaMetrics.
supernaut tvångsskickade feat/registry-retention från a0e9309882
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m28s
till c86bbde4f6
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m43s
2026-08-01 17:10:41 +00:00
Jämför
supernaut sammanfogade incheckning 7664b50190 till main 2026-08-01 17:15:44 +00:00
supernaut tog bort grenen feat/registry-retention 2026-08-01 17:15:44 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!310
Ingen beskrivning angiven.