runbook: the registry-retention dry-run rehearsal command fails (FORGEJO_TOKEN not set) #316

Stängd
öppnade 2026-08-01 17:41:49 +00:00 av supernaut · 1 kommentar
Ägare

The runbook tells the operator to rehearse the registry-retention sweep before letting the timer fire.
That command does not work. Found while applying #310 to production on 2026-08-01.

docs/runbook.md (Registry retention section) says:

ssh -p 2222 admin@<host> \
  'sudo -u gitborg /home/gitborg/bin/bitborg-registry-retention.sh --dry-run'

Run exactly as written it exits immediately:

/home/gitborg/bin/bitborg-registry-retention.sh: line 22: FORGEJO_TOKEN: FORGEJO_TOKEN not set

Cause

The token reaches the script through the systemd unit's EnvironmentFile=
(~/.config/gitborg/registry-retention.env, set from vault_forgejo_ci_token). systemd sources that
file; a manual shell invocation does not, so FORGEJO_TOKEN is unset and the script's own guard on
line 22 stops it. The script is right to fail closed — the documentation is what is wrong.

Why this is worth fixing rather than filing under "obvious once you hit it"

This is the only safety step between installing the role and an irreversible job. The sweep deletes
package versions and there is no undo; the role deliberately does not run on deploy so that the operator
can look at a dry run first. A rehearsal command that cannot be executed removes that gate at exactly
the moment it is supposed to be used — and it fails with a message that reads like a broken deployment
(missing credential) rather than a wrong command, so the natural next move is to go hunting through the
vault instead of adding one line.

What works

set -a; . /home/gitborg/.config/gitborg/registry-retention.env; set +a
/home/gitborg/bin/bitborg-registry-retention.sh --dry-run

Verified on production the same day: it lists what would be deleted, writes no metrics, and issues no
DELETEs.

Suggested

Correct the command in docs/runbook.md. Two other options exist and are probably worse:
systemd-run --user --scope -p EnvironmentFile=… is harder to read for a documented procedure, and
teaching the script to source its own env file would duplicate what the unit already declares and
diverge from the registry-mirror / token-audit shape.

Worth a scan of the other three env-file-driven jobs for the same documentation shape —
registry-mirror, token-audit and the backup jobs all use EnvironmentFile= and all have runbook
entries describing a manual run.

Also observed on the same apply, recorded here so it is not mistaken for a fault

The first real sweep will hit the 500-deletion cap and stay capped for several nights.
#310 measured 44 bitborg-web/cache versions; the live dry run found 862 (434 kept, 428 deletable)
plus 92 bitborg-web versions (20 kept, 72 deletable). Convergence therefore takes several runs, each
followed by the @midnight blob GC before df moves. A capped run and a flat df are both expected in
that window.

Refs #297, #310.

The runbook tells the operator to rehearse the registry-retention sweep before letting the timer fire. That command does not work. Found while applying #310 to production on 2026-08-01. `docs/runbook.md` (Registry retention section) says: ```bash ssh -p 2222 admin@<host> \ 'sudo -u gitborg /home/gitborg/bin/bitborg-registry-retention.sh --dry-run' ``` Run exactly as written it exits immediately: ```text /home/gitborg/bin/bitborg-registry-retention.sh: line 22: FORGEJO_TOKEN: FORGEJO_TOKEN not set ``` ## Cause The token reaches the script through the systemd unit's `EnvironmentFile=` (`~/.config/gitborg/registry-retention.env`, set from `vault_forgejo_ci_token`). systemd sources that file; a manual shell invocation does not, so `FORGEJO_TOKEN` is unset and the script's own guard on line 22 stops it. The script is right to fail closed — the documentation is what is wrong. ## Why this is worth fixing rather than filing under "obvious once you hit it" This is the **only safety step between installing the role and an irreversible job**. The sweep deletes package versions and there is no undo; the role deliberately does not run on deploy so that the operator can look at a dry run first. A rehearsal command that cannot be executed removes that gate at exactly the moment it is supposed to be used — and it fails with a message that reads like a broken deployment (missing credential) rather than a wrong command, so the natural next move is to go hunting through the vault instead of adding one line. ## What works ```bash set -a; . /home/gitborg/.config/gitborg/registry-retention.env; set +a /home/gitborg/bin/bitborg-registry-retention.sh --dry-run ``` Verified on production the same day: it lists what would be deleted, writes no metrics, and issues no DELETEs. ## Suggested Correct the command in `docs/runbook.md`. Two other options exist and are probably worse: `systemd-run --user --scope -p EnvironmentFile=…` is harder to read for a documented procedure, and teaching the script to source its own env file would duplicate what the unit already declares and diverge from the `registry-mirror` / `token-audit` shape. Worth a scan of the other three env-file-driven jobs for the same documentation shape — `registry-mirror`, `token-audit` and the backup jobs all use `EnvironmentFile=` and all have runbook entries describing a manual run. ## Also observed on the same apply, recorded here so it is not mistaken for a fault The first real sweep will hit the 500-deletion cap and stay capped for several nights. #310 measured 44 `bitborg-web/cache` versions; the live dry run found **862** (434 kept, 428 deletable) plus 92 `bitborg-web` versions (20 kept, 72 deletable). Convergence therefore takes several runs, each followed by the `@midnight` blob GC before `df` moves. A capped run and a flat `df` are both expected in that window. Refs #297, #310.
Upphovsperson
Ägare

Scanned the other EnvironmentFile=-driven jobs as this issue suggests. Nothing to fix there — the
scope is one line, and the suggested scan can be dropped.

Why it cannot affect the others

A job is only exposed to this if it has a rehearsal mode that its unit's ExecStart does not expose.
Where the only manual path is systemctl --user start <unit>, systemd sources the EnvironmentFile and
the failure cannot occur.

Exactly two jobs have such a mode:

Job Rehearsal mode How the runbook documents it Verdict
backup-drill DRILL_HOLD=1 sources backup.env explicitly correct
registry-retention --dry-run bare invocation this issue

backup, backup-storage, backup-verify, registry-mirror, renovate and token-audit have no
rehearsal mode at all — every one is a bare binary in ExecStart, and the runbook documents
systemctl --user start for each. They are structurally immune.

There are only two direct script invocations in the whole of docs/, which is the entire exposure
surface:

docs/runbook.md:601   /home/gitborg/bin/bitborg-backup-drill.py'      ← sources backup.env, correct
docs/runbook.md:2069  …/bitborg-registry-retention.sh --dry-run'      ← broken

The fix is a pattern this repo already has

docs/runbook.md:592-608 does not merely get it right by luck — it states the reason ("a bare
DRILL_HOLD=1 python3 … from another shell exits with incomplete config"
) and warns that DRILL_HOLD=1
must sit inside the env after sudo, or the normal drill runs and tears the VM down. That entry is the
model for this one, ~1400 lines up in the same file.

Note also that registry-retention already documents the working
systemctl --user start bitborg-registry-retention.service at line 2076. It is specifically and only the
dry run — the one thing systemd cannot provide, and the reason a direct invocation was needed at all —
that is broken.

One adjacent gap, not part of this issue

token-audit has a FORGEJO_TOKEN guard but no documented manual run at all: no direct invocation and
no systemctl --user start line, unlike every other timer job. Nothing there is broken, so it is not this
bug, but there is no entry for running it out of cycle. Worth its own issue rather than widening this one.

Scanned the other `EnvironmentFile=`-driven jobs as this issue suggests. **Nothing to fix there — the scope is one line, and the suggested scan can be dropped.** ## Why it cannot affect the others A job is only exposed to this if it has a **rehearsal mode that its unit's `ExecStart` does not expose**. Where the only manual path is `systemctl --user start <unit>`, systemd sources the `EnvironmentFile` and the failure cannot occur. Exactly two jobs have such a mode: | Job | Rehearsal mode | How the runbook documents it | Verdict | | --- | --- | --- | --- | | `backup-drill` | `DRILL_HOLD=1` | sources `backup.env` explicitly | correct | | `registry-retention` | `--dry-run` | bare invocation | **this issue** | `backup`, `backup-storage`, `backup-verify`, `registry-mirror`, `renovate` and `token-audit` have no rehearsal mode at all — every one is a bare binary in `ExecStart`, and the runbook documents `systemctl --user start` for each. They are structurally immune. There are only two direct script invocations in the whole of `docs/`, which is the entire exposure surface: ```text docs/runbook.md:601 /home/gitborg/bin/bitborg-backup-drill.py' ← sources backup.env, correct docs/runbook.md:2069 …/bitborg-registry-retention.sh --dry-run' ← broken ``` ## The fix is a pattern this repo already has `docs/runbook.md:592-608` does not merely get it right by luck — it states the reason (*"a bare `DRILL_HOLD=1 python3 …` from another shell exits with `incomplete config`"*) and warns that `DRILL_HOLD=1` must sit inside the `env` after `sudo`, or the normal drill runs and tears the VM down. That entry is the model for this one, ~1400 lines up in the same file. Note also that `registry-retention` already documents the working `systemctl --user start bitborg-registry-retention.service` at line 2076. It is specifically and only the dry run — the one thing systemd cannot provide, and the reason a direct invocation was needed at all — that is broken. ## One adjacent gap, not part of this issue `token-audit` has a `FORGEJO_TOKEN` guard but **no documented manual run at all**: no direct invocation and no `systemctl --user start` line, unlike every other timer job. Nothing there is broken, so it is not this bug, but there is no entry for running it out of cycle. Worth its own issue rather than widening this one.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#316
Ingen beskrivning angiven.