deploy-verify: systemd logs a MONITOR_* propagation warning on every auto-update tick (~480/day) #318
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#318
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Since #315 was applied to production on 2026-08-01, every start of
bitborg-deploy-verify.serviceis preceded by a systemd warning:
It is harmless today, but it is permanent, it is not rare, and it silently forecloses a design option.
Worth writing down rather than leaving for whoever next reads this journal during an incident.
Cause
The drop-in names the same unit on both hooks:
That is deliberate and correct —
podman auto-updateexits 0 after a rollback, soOnFailure=alone would never see one. But systemd populates
$MONITOR_UNIT/$MONITOR_EXIT_STATUS/$MONITOR_SERVICE_RESULTfor the triggering unit, and with two candidate sources it cannot decidewhich to propagate. So it propagates neither and says so.
Frequency: every auto-update tick, not every deploy
OnSuccess=fires wheneverpodman-auto-update.servicesucceeds, which is every timer tick (~3 min),not only when an image actually changed — the verifier's own early exit is what makes a quiet tick
cheap. Measured on the services host: 5 occurrences in the ~25 minutes since the drop-in was
installed, i.e. roughly 480 lines/day shipped to Loki, indefinitely.
Consequence
Nothing is broken.
bitborg-deploy-verify.shreads noMONITOR_*variable anywhere — it determineswhat happened from
podman inspectplus ajournalctl … grep 'start operation timed out'scan over awindow, precisely because the rollback case is invisible to systemd's exit status. So the propagated
status was never load-bearing.
What the warning does mean is that the verifier can never learn which of the two hooks invoked it.
Any future logic wanting to branch on "auto-update failed" versus "auto-update succeeded" is not
available from the trigger, and would need two separate units (or a wrapper) rather than a
MONITOR_*check. That is a real constraint and it is currently recorded nowhere.
Suggested
Documentation, not a behaviour change:
roles/web/templates/podman-auto-update-verify.conf.j2stating that naming one unit onboth hooks is intentional, that it costs
MONITOR_*propagation, and that the verifier does not relyon it. The template already explains why both hooks are needed; it does not mention this cost.
docs/runbook.mdunder the deploy-verification section so the warning is identifiable asexpected when someone greps the journal during an incident.
Deliberately not suggested: dropping
OnFailure=to silence it. That hook is there because a failedpull or registry outage should still trigger verification, and trading a real signal for a quiet journal
is the wrong way round.
If the volume is judged too high on its own, the alternative is a Loki drop rule for that exact line —
but that hides a systemd diagnostic wholesale, so the comment plus runbook entry is the cheaper fix.
Refs #315.