runbook: the registry-retention dry-run rehearsal command fails (FORGEJO_TOKEN not set) #316
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#316
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
The runbook tells the operator to rehearse the registry-retention sweep before letting the timer fire.
That command does not work. Found while applying #310 to production on 2026-08-01.
docs/runbook.md(Registry retention section) says:Run exactly as written it exits immediately:
Cause
The token reaches the script through the systemd unit's
EnvironmentFile=(
~/.config/gitborg/registry-retention.env, set fromvault_forgejo_ci_token). systemd sources thatfile; a manual shell invocation does not, so
FORGEJO_TOKENis unset and the script's own guard online 22 stops it. The script is right to fail closed — the documentation is what is wrong.
Why this is worth fixing rather than filing under "obvious once you hit it"
This is the only safety step between installing the role and an irreversible job. The sweep deletes
package versions and there is no undo; the role deliberately does not run on deploy so that the operator
can look at a dry run first. A rehearsal command that cannot be executed removes that gate at exactly
the moment it is supposed to be used — and it fails with a message that reads like a broken deployment
(missing credential) rather than a wrong command, so the natural next move is to go hunting through the
vault instead of adding one line.
What works
Verified on production the same day: it lists what would be deleted, writes no metrics, and issues no
DELETEs.
Suggested
Correct the command in
docs/runbook.md. Two other options exist and are probably worse:systemd-run --user --scope -p EnvironmentFile=…is harder to read for a documented procedure, andteaching the script to source its own env file would duplicate what the unit already declares and
diverge from the
registry-mirror/token-auditshape.Worth a scan of the other three env-file-driven jobs for the same documentation shape —
registry-mirror,token-auditand the backup jobs all useEnvironmentFile=and all have runbookentries describing a manual run.
Also observed on the same apply, recorded here so it is not mistaken for a fault
The first real sweep will hit the 500-deletion cap and stay capped for several nights.
#310 measured 44
bitborg-web/cacheversions; the live dry run found 862 (434 kept, 428 deletable)plus 92
bitborg-webversions (20 kept, 72 deletable). Convergence therefore takes several runs, eachfollowed by the
@midnightblob GC beforedfmoves. A capped run and a flatdfare both expected inthat window.
Refs #297, #310.
Scanned the other
EnvironmentFile=-driven jobs as this issue suggests. Nothing to fix there — thescope is one line, and the suggested scan can be dropped.
Why it cannot affect the others
A job is only exposed to this if it has a rehearsal mode that its unit's
ExecStartdoes not expose.Where the only manual path is
systemctl --user start <unit>, systemd sources theEnvironmentFileandthe failure cannot occur.
Exactly two jobs have such a mode:
backup-drillDRILL_HOLD=1backup.envexplicitlyregistry-retention--dry-runbackup,backup-storage,backup-verify,registry-mirror,renovateandtoken-audithave norehearsal mode at all — every one is a bare binary in
ExecStart, and the runbook documentssystemctl --user startfor each. They are structurally immune.There are only two direct script invocations in the whole of
docs/, which is the entire exposuresurface:
The fix is a pattern this repo already has
docs/runbook.md:592-608does not merely get it right by luck — it states the reason ("a bareDRILL_HOLD=1 python3 …from another shell exits withincomplete config") and warns thatDRILL_HOLD=1must sit inside the
envaftersudo, or the normal drill runs and tears the VM down. That entry is themodel for this one, ~1400 lines up in the same file.
Note also that
registry-retentionalready documents the workingsystemctl --user start bitborg-registry-retention.serviceat line 2076. It is specifically and only thedry run — the one thing systemd cannot provide, and the reason a direct invocation was needed at all —
that is broken.
One adjacent gap, not part of this issue
token-audithas aFORGEJO_TOKENguard but no documented manual run at all: no direct invocation andno
systemctl --user startline, unlike every other timer job. Nothing there is broken, so it is not thisbug, but there is no entry for running it out of cycle. Worth its own issue rather than widening this one.