docs(runbook): fix the registry-retention rehearsal command and record the auto-update warning #341

Sammanfogat
supernaut sammanfogade 2 incheckningar från docs/runbook-corrections in i main 2026-08-02 16:07:32 +00:00
Ägare

Closes #316, closes #318.

Two places where the documentation described something that was not true.

#316 — the registry-retention dry-run rehearsal command in the runbook fails, which means it has
probably never been run as written. The token path was traced through the unit rather than guessed:
EnvironmentFile points at a 0600 file owned by gitborg, the script is 0750, and the guard
that fires is the script's own ${FORGEJO_TOKEN:?…}. So the invocation has to both run as
gitborg and source the file inside that sudo. The corrected form matches an idiom already proven
elsewhere in the same runbook.

Reproduced locally with byte-shaped stand-ins — same env-file layout, same guard, same file modes:
the documented form fails with the exact reported message and exit 1; the corrected one exits 0,
including the backslash continuation inside single quotes, which a copy-paste reproduces verbatim.
Not run against production.

A scan for the same defect elsewhere found this was the only affected entry: every other
env-file job documents systemctl --user start, where systemd reads the EnvironmentFile itself, so
the problem cannot arise. That distinction is now written down so it does not get re-lost.

#318 — the MONITOR_* propagation warning is genuine noise, not a masked bug, and is now
documented as expected. Measured rather than estimated: 385 occurrences in 24 h, spaced
~157–205 s, i.e. one per auto-update timer tick rather than per deploy.

The mechanism holds up: the verify unit is named on both OnSuccess= and OnFailure=, and systemd
only propagates MONITOR_* when there is exactly one trigger source. Decisively, grep -rn "MONITOR_" across the whole repo returns nothing — the verify script never reads those variables,
determining outcomes from podman inspect plus a journal scan, precisely because a rollback exits 0.
So the propagation was never load-bearing.

Recorded alongside it: the one real constraint this creates (the verifier cannot learn which hook
invoked it, so a future failed-versus-succeeded branch needs two units), and an explicit note not
to drop OnFailure= just to silence the warning.

Verification

markdownlint 0 issues, Prettier clean, shellcheck exit 0. pnpm ansible:check could not run in the
worktree (needs the vault password file); the only Ansible change is a comment block, verified to
contain no Jinja delimiters and to be entirely #-prefixed.

Closes #316, closes #318. Two places where the documentation described something that was not true. **#316** — the registry-retention dry-run rehearsal command in the runbook fails, which means it has probably never been run as written. The token path was traced through the unit rather than guessed: `EnvironmentFile` points at a `0600` file owned by `gitborg`, the script is `0750`, and the guard that fires is the script's own `${FORGEJO_TOKEN:?…}`. So the invocation has to both run as `gitborg` and source the file inside that `sudo`. The corrected form matches an idiom already proven elsewhere in the same runbook. Reproduced locally with byte-shaped stand-ins — same env-file layout, same guard, same file modes: the documented form fails with the exact reported message and exit 1; the corrected one exits 0, including the backslash continuation inside single quotes, which a copy-paste reproduces verbatim. Not run against production. A scan for the same defect elsewhere found this was the **only** affected entry: every other env-file job documents `systemctl --user start`, where systemd reads the `EnvironmentFile` itself, so the problem cannot arise. That distinction is now written down so it does not get re-lost. **#318** — the `MONITOR_*` propagation warning is genuine noise, not a masked bug, and is now documented as expected. Measured rather than estimated: **385 occurrences in 24 h**, spaced ~157–205 s, i.e. one per auto-update timer tick rather than per deploy. The mechanism holds up: the verify unit is named on both `OnSuccess=` and `OnFailure=`, and systemd only propagates `MONITOR_*` when there is exactly one trigger source. Decisively, `grep -rn "MONITOR_"` across the whole repo returns nothing — the verify script never reads those variables, determining outcomes from `podman inspect` plus a journal scan, precisely because a rollback exits 0. So the propagation was never load-bearing. Recorded alongside it: the one real constraint this creates (the verifier cannot learn which hook invoked it, so a future failed-versus-succeeded branch needs two units), and an explicit note **not** to drop `OnFailure=` just to silence the warning. ### Verification markdownlint 0 issues, Prettier clean, shellcheck exit 0. `pnpm ansible:check` could not run in the worktree (needs the vault password file); the only Ansible change is a comment block, verified to contain no Jinja delimiters and to be entirely `#`-prefixed.
supernaut lade till 2 incheckningar 2026-08-02 15:56:03 +00:00
The documented rehearsal invoked the script bare, so it exited immediately on
its own guard with `FORGEJO_TOKEN: FORGEJO_TOKEN not set`. The token reaches the
script only through the systemd unit's EnvironmentFile= (registry-retention.env);
systemd sources that file, a manual shell does not.

This mattered more than a typo: the dry run is the only safety step between
installing the role and an irreversible delete, and it failed with a message that
reads like a missing credential rather than a wrong command — pointing the
operator at the vault instead of at one missing line.

Document the sourcing prefix, matching the shape the backup-drill entry already
uses, and note why `systemctl --user start` needs none of it. Scanned the other
env-file-driven jobs (image mirror, token audit, backup): all document a
`systemctl --user start`, which systemd sources the EnvironmentFile for, so the
retention rehearsal was the only documented bare-script invocation affected.

Closes #316
docs(web): record the MONITOR_* propagation cost of the deploy-verify hooks
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m58s
25dc2721aa
Naming podman-auto-update.service on both OnSuccess= and OnFailure= gives systemd
two trigger-source candidates, so it propagates no MONITOR_* variables and logs
"multiple trigger source candidates for exit status propagation … skipping" on
every auto-update tick. Confirmed against production logs: the exact line, 385
occurrences in 24 h, spaced ~3 min — one per timer tick, not one per deploy.

It is noise, not a defect. gitborg-deploy-verify.sh reads no MONITOR_* variable
anywhere; it determines what happened from `podman inspect` plus the rollback
journal scan, precisely because a rollback exits 0 and is invisible to systemd's
exit status. So the propagated status was never load-bearing.

Documented rather than silenced, in both places an operator would look: the
drop-in template (why the cost is accepted, and why dropping OnFailure= to quieten
it would trade a real signal for a quiet journal) and the runbook deploy section
(so the line is identifiable as expected when grepping the journal mid-incident).
Also records the one real constraint — the verifier can never learn which hook
invoked it, so any future "failed vs succeeded" branch needs two units.

No behaviour change.

Closes #318
supernaut sammanfogade incheckning 095315fbfc till main 2026-08-02 16:07:32 +00:00
supernaut tog bort grenen docs/runbook-corrections 2026-08-02 16:07:32 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!341
Ingen beskrivning angiven.