fix(runner-controller): stop the reaper warning about a detach that cannot succeed #311

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/reaper-root-volume-noise in i main 2026-08-01 17:09:21 +00:00
Ägare

Closes #305.

What it is

Boot volumes are created with delete_on_termination=True, so the root volume cascades when the server
is deleted. The detach loop in _delete_ephemeral_server is belt-and-suspenders for extra attached
volumes — but on every ordinary reap it also tried the root volume, which OpenStack always refuses:

WARNING [controller] reap: could not delete volume ade7c40a-…: BadRequestException: 400: … Cannot
detach a root device volume

Measured at 2–7 occurrences per hour, indefinitely, with different volume and server IDs each time.

Volumes are not leaking — checked before changing anything

gitborg_runner_controller_os_volume_gb_used oscillates between 640 and 1020 GB over 7 days and returns
to a 640–720 baseline against a 5000 GB quota. It does not climb. last_loop_ok is 1, active_vms and
boot_error_vms are 0. The volumes do get released, because deleting the server cascades the root device.

Why fix it anyway

This is the code path whose failures are the signal for orphaned boot volumes — and a prior CI outage was
caused by 227 of them. A genuine reap failure arriving in the middle of a permanent stream of
identical-looking warnings, which everyone has learned to ignore, is exactly how the next leak goes
unnoticed. Removing a false alarm from an alerting path is worth more than the tidiness suggests.

How

A pure predicate, _is_expected_root_volume_refusal(status_code, message), splits "expected, cascades
anyway" from "actually failed". Only a genuine failure stays a WARNING; the root-device refusal logs at
INFO and says that the server delete will cascade it.

Kept pure and separate from the SDK exception so it can be tested on a bare interpreter, matching the
other decision predicates in this file. Matching on the message text is deliberate and noted inline:
OpenStack has no distinct error code for this, and 400 on its own is far too broad to treat as benign.

Verified

8 new checks in test_controller_logic.py, following the existing dependency-free assert pattern —
17/17 checks passed.

Mutation-checked rather than merely passing. Replacing the predicate body with return True — the
dangerous over-broad direction, which would silence real failures — fails exactly the five cases that
matter:

[FAIL] 400 without the message -> genuine: got=True expected=False
[FAIL] 409 conflict -> genuine: got=True expected=False
[FAIL] 500 -> genuine: got=True expected=False
[FAIL] no status code -> genuine: got=True expected=False
[FAIL] empty message -> genuine: got=True expected=False
12/17 checks passed

Both files parse, and every line is inside ruff.toml's 120-character limit. ruff itself is baked into
the runner image rather than available locally, so CI is the first place it actually runs — flagging that
rather than claiming a lint pass I did not perform.

Disclosure on ordering: the implementation was written before its tests here, unlike the test-first
work elsewhere today. This changes a log level rather than behaviour, and the mutation check is the
evidence that the tests constrain it. Said plainly rather than implied otherwise.

Applying

--tags runner-controller. The controller image is rebuilt from files/controller.py, so the unit
restarts. No behavioural change to provisioning, reaping or deletion — only which severity a known
outcome is logged at.

Open question carried over from the issue

The ~640 GB baseline with zero active runner VMs is larger than the obviously-known persistent volumes
account for. It may be entirely legitimate (host root volumes, the monitoring host, images or snapshots
against the same project quota) but it was not chased down, and this PR does not address it. Recorded in
#305 so the number is not mistaken for verified.

Closes #305. ## What it is Boot volumes are created with `delete_on_termination=True`, so the root volume cascades when the server is deleted. The detach loop in `_delete_ephemeral_server` is belt-and-suspenders for *extra* attached volumes — but on every ordinary reap it also tried the root volume, which OpenStack always refuses: ```text WARNING [controller] reap: could not delete volume ade7c40a-…: BadRequestException: 400: … Cannot detach a root device volume ``` Measured at 2–7 occurrences per hour, indefinitely, with different volume **and** server IDs each time. ## Volumes are not leaking — checked before changing anything `gitborg_runner_controller_os_volume_gb_used` oscillates between 640 and 1020 GB over 7 days and returns to a 640–720 baseline against a 5000 GB quota. It does not climb. `last_loop_ok` is 1, `active_vms` and `boot_error_vms` are 0. The volumes do get released, because deleting the server cascades the root device. ## Why fix it anyway This is the code path whose failures are the signal for orphaned boot volumes — and a prior CI outage was caused by 227 of them. A genuine reap failure arriving in the middle of a permanent stream of identical-looking warnings, which everyone has learned to ignore, is exactly how the next leak goes unnoticed. Removing a false alarm from an alerting path is worth more than the tidiness suggests. ## How A pure predicate, `_is_expected_root_volume_refusal(status_code, message)`, splits "expected, cascades anyway" from "actually failed". Only a genuine failure stays a `WARNING`; the root-device refusal logs at `INFO` and says that the server delete will cascade it. Kept pure and separate from the SDK exception so it can be tested on a bare interpreter, matching the other decision predicates in this file. Matching on the message text is deliberate and noted inline: OpenStack has no distinct error code for this, and 400 on its own is far too broad to treat as benign. ## Verified 8 new checks in `test_controller_logic.py`, following the existing dependency-free assert pattern — `17/17 checks passed`. **Mutation-checked rather than merely passing.** Replacing the predicate body with `return True` — the dangerous over-broad direction, which would silence real failures — fails exactly the five cases that matter: ```text [FAIL] 400 without the message -> genuine: got=True expected=False [FAIL] 409 conflict -> genuine: got=True expected=False [FAIL] 500 -> genuine: got=True expected=False [FAIL] no status code -> genuine: got=True expected=False [FAIL] empty message -> genuine: got=True expected=False 12/17 checks passed ``` Both files parse, and every line is inside `ruff.toml`'s 120-character limit. `ruff` itself is baked into the runner image rather than available locally, so CI is the first place it actually runs — flagging that rather than claiming a lint pass I did not perform. **Disclosure on ordering:** the implementation was written before its tests here, unlike the test-first work elsewhere today. This changes a log level rather than behaviour, and the mutation check is the evidence that the tests constrain it. Said plainly rather than implied otherwise. ## Applying `--tags runner-controller`. The controller image is rebuilt from `files/controller.py`, so the unit restarts. No behavioural change to provisioning, reaping or deletion — only which severity a known outcome is logged at. ## Open question carried over from the issue The ~640 GB baseline with **zero** active runner VMs is larger than the obviously-known persistent volumes account for. It may be entirely legitimate (host root volumes, the monitoring host, images or snapshots against the same project quota) but it was not chased down, and this PR does not address it. Recorded in #305 so the number is not mistaken for verified.
supernaut lade till 3 incheckningar 2026-08-01 14:53:37 +00:00
Growing a Cinder volume does not grow its filesystem, and nothing in the
playbook closed the gap. `community.general.filesystem` created a filesystem
but never resized one, so bumping `lfs_volume_size` / `data_volume_size` and
applying left the block device larger and the filesystem exactly as it was.
On 2026-08-01 that meant `lsblk` reporting 60G while `df` still showed 20G
with 4.1G free, until `resize2fs` was run by hand.

The failure is invisible in the places you look: the plan says `~ size`, the
apply is green, and monitoring keeps showing a nearly full disk with nothing
linking the two. It only ever bites during a capacity emergency.

Add `resizefs: true` to the `filesystem` task for the LFS/storage volume
(forgejo), the data volume (podman) and — same footgun, same one-line fix, and
the runbook documented the same manual step for it — the backup volume. ext4
grows online and the option is a no-op once the filesystem already fills the
device, so a normal converge is unaffected.

Root is deliberately left manual: its ext4 is created by the cloud image (no
Ansible task owns it), and because the disk is partitioned `resize2fs` alone is
insufficient — `growpart` must move the partition end first. Automating that
would write the boot disk's partition table on every converge to serve a
roughly annual operation that always follows a `tofu apply` the playbook does
not drive. The runbook now states that reasoning where the manual commands are.

Verified: `ansible-playbook site.yml --syntax-check` clean; `ansible-lint` on
the three roles reports the same 17 pre-existing var-naming findings as `main`
(no new ones).
feat(registry): retention policy for the package registry (#297)
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m28s
a0e9309882
`/srv/gitborg-lfs` (Forgejo's `[storage]` — LFS, packages, attachments,
avatars) reached 79% with 4 GB free on 2026-08-01, roughly a week from full. It
was grown 20→60 GB as an emergency measure, but that bought runway, it did not
fix the leak.

Cause: gitborg-web's deploy workflow pushes three things on every merge to
`main` — `latest`, an immutable `sha-XXXX` tag, and a BuildKit registry layer
cache via `--cache-to` — and nothing pruned the registry. Measured the same day,
`gitborg-web` had 6 versions while `gitborg-web/cache` had 44: the layer cache
is the dominant consumer, not the image tags, which is worth stating because
the obvious guess is wrong. Drift was ~0.6 GB/day. (`ARTIFACT_RETENTION_DAYS`
covers Actions artifacts — a different store — and does not help here.)

New `registry-retention` role: a gitborg-user oneshot service plus a daily
timer, the same shape as `registry-mirror` and `token-audit` (host script +
EnvironmentFile + systemd user timer, textfile metrics, inert until its
credential is vaulted). Policy is data in `registry_retention_rules`: keep the
newest 20 `sha-*` tags of gitborg-web, delete `gitborg-web/cache` versions
older than 2 days. Notable choices:

- `latest` cannot be deleted (the rule's `^sha-` match excludes it) — that tag
  is what `podman auto-update` follows.
- Packages not named in a rule are never touched, which is what keeps the
  registry-mirror images safe from a job that shares gitborg-ci's token.
- The cache rule is age-based, not count-based: one push writes ~10-12 cache
  versions, so any `keep_newest` would be a magic number that shreds the
  current cache the moment the layer count changes.
- `--dry-run` lists what a run would delete and deliberately does not write
  metrics, so a rehearsal cannot refresh `last_run_timestamp` and hide a dead
  timer. There is no undo for a deleted version.
- A per-run deletion cap (500) stops a mistyped rule emptying the registry in
  one pass; hitting it sets a metric and the next run continues.
- Deliberately NOT run on deploy, unlike the mirror. A converge is the wrong
  moment to fire a job that deletes.
- Reuses gitborg-ci's existing `write:package` PAT. Package deletion needs
  write on the owner's packages and nothing more — no new secret, no admin
  scope.

Also pin `[cron.cleanup_packages]` in app.ini. This is the load-bearing half:
deleting a version only unlinks it, and that cron is what sweeps the
unreferenced blobs and actually returns the bytes. Pinned rather than inherited
because an upstream default change would silently stop reclaiming space, and
the symptom — retention runs clean, `df` never moves — is not one anybody would
trace back to a default. @midnight follows the 22:10 sweep by ~2 h.

Alerting, per the issue's "alert on the trend, not just the level":

- `DiskWillFill` — `predict_linear` over a 3-day window, projected
  `alert_disk_predict_days` (14) ahead, gated on >40% used. Replayed against
  production VictoriaMetrics the level thresholds are worse than late: they
  never fired, because the volume peaked at 79.17% used and 80% was never
  crossed before it was grown. The same expression is true from 2026-07-22
  00:00Z, nine days of lead. Against today's post-grow state it returns nothing.
  The 3-day window (not the mixin's 6h) is deliberate — writes here arrive in
  lumps, and one big LFS push must not read as a trend.
- `RegistryRetentionFailed`, `RegistryRetentionStale` (deadman — a dead timer
  never sets a non-zero status, it just stops updating) and
  `RegistryRetentionCapped`.

Verified:
- `ansible-playbook site.yml --syntax-check` clean.
- `ansible-lint` on the new role: only the repo-wide hyphenated `role-name`
  finding that `registry-mirror` and `token-audit` also produce.
- `shellcheck` clean on the rendered script.
- Script logic exercised offline against a stubbed API (57 synthetic versions
  over 2 pages): kept the newest 20 `sha-` tags and deleted the 5 oldest, left
  `latest` and a mirrored `gitborg/forgejo` image untouched, deleted the 18
  cache versions older than 48 h and kept 12, and encoded the slash in
  `gitborg-web/cache` as `%2F`. Dry run issued 0 DELETEs and wrote no metrics.
  With the cap set to 4 it stopped at 4 and set `..._capped 1`.
- Emitted metrics parse as valid exposition format with each family contiguous
  (the trap that previously made a whole textfile unreadable) and land 0644.
- Alert rules render to valid YAML (38 rules); `DiskWillFill` executed as a
  read-only instant and range query against production VictoriaMetrics.
fix(runner-controller): stop the reaper warning about a detach that cannot succeed
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m40s
a4a9d7489d
Boot volumes are created with delete_on_termination=True, so the root volume
cascades when the server is deleted. The detach loop is belt-and-suspenders for
extra attached volumes, but on every ordinary reap it also tried the root volume,
which OpenStack always refuses with 400 'Cannot detach a root device volume'.
Logged as a warning, that produced 2-7 alarming-but-harmless lines an hour,
indefinitely.

Volumes were verified not to be leaking: os_volume_gb_used oscillates around a
640-720 GB baseline over 7 days against a 5000 GB quota, and never climbs.

The reason to fix it anyway is that this is the code path whose failures are the
signal for orphaned boot volumes, which once caused a CI outage with 227 of them.
A real reap failure arriving among identical benign warnings is how the next leak
goes unmissed. The classification is a pure predicate so only a GENUINE failure
stays a warning, with dependency-free tests in the existing pattern; a mutation
check confirmed the tests catch an over-broad predicate rather than merely
passing.
supernaut tvångsskickade fix/reaper-root-volume-noise från a4a9d7489d
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m40s
till f7916e8e9a
Väntande kontroller
ci / ci (pull_request) Has started running
2026-08-01 16:58:15 +00:00
Jämför
supernaut tvångsskickade fix/reaper-root-volume-noise från f7916e8e9a
Väntande kontroller
ci / ci (pull_request) Has started running
till 0b14f76e0f
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m22s
2026-08-01 17:02:54 +00:00
Jämför
supernaut sammanfogade incheckning eab5c8af64 till main 2026-08-01 17:09:21 +00:00
supernaut tog bort grenen fix/reaper-root-volume-noise 2026-08-01 17:09:22 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!311
Ingen beskrivning angiven.