deploy-verify: systemd logs a MONITOR_* propagation warning on every auto-update tick (~480/day) #318

Stängd
öppnade 2026-08-01 18:46:14 +00:00 av supernaut · 0 kommentarer
Ägare

Since #315 was applied to production on 2026-08-01, every start of bitborg-deploy-verify.service
is preceded by a systemd warning:

bitborg-deploy-verify.service: multiple trigger source candidates for exit status propagation
(podman-auto-update.service, podman-auto-update.service), skipping.

It is harmless today, but it is permanent, it is not rare, and it silently forecloses a design option.
Worth writing down rather than leaving for whoever next reads this journal during an incident.

Cause

The drop-in names the same unit on both hooks:

OnSuccess=bitborg-deploy-verify.service
OnFailure=bitborg-deploy-verify.service

That is deliberate and correct — podman auto-update exits 0 after a rollback, so OnFailure=
alone would never see one. But systemd populates $MONITOR_UNIT / $MONITOR_EXIT_STATUS /
$MONITOR_SERVICE_RESULT for the triggering unit, and with two candidate sources it cannot decide
which to propagate. So it propagates neither and says so.

Frequency: every auto-update tick, not every deploy

OnSuccess= fires whenever podman-auto-update.service succeeds, which is every timer tick (~3 min),
not only when an image actually changed — the verifier's own early exit is what makes a quiet tick
cheap. Measured on the services host: 5 occurrences in the ~25 minutes since the drop-in was
installed
, i.e. roughly 480 lines/day shipped to Loki, indefinitely.

Consequence

Nothing is broken. bitborg-deploy-verify.sh reads no MONITOR_* variable anywhere — it determines
what happened from podman inspect plus a journalctl … grep 'start operation timed out' scan over a
window, precisely because the rollback case is invisible to systemd's exit status. So the propagated
status was never load-bearing.

What the warning does mean is that the verifier can never learn which of the two hooks invoked it.
Any future logic wanting to branch on "auto-update failed" versus "auto-update succeeded" is not
available from the trigger, and would need two separate units (or a wrapper) rather than a MONITOR_*
check. That is a real constraint and it is currently recorded nowhere.

Suggested

Documentation, not a behaviour change:

  • A comment in roles/web/templates/podman-auto-update-verify.conf.j2 stating that naming one unit on
    both hooks is intentional, that it costs MONITOR_* propagation, and that the verifier does not rely
    on it. The template already explains why both hooks are needed; it does not mention this cost.
  • A line in docs/runbook.md under the deploy-verification section so the warning is identifiable as
    expected when someone greps the journal during an incident.

Deliberately not suggested: dropping OnFailure= to silence it. That hook is there because a failed
pull or registry outage should still trigger verification, and trading a real signal for a quiet journal
is the wrong way round.

If the volume is judged too high on its own, the alternative is a Loki drop rule for that exact line —
but that hides a systemd diagnostic wholesale, so the comment plus runbook entry is the cheaper fix.

Refs #315.

Since #315 was applied to production on 2026-08-01, every start of `bitborg-deploy-verify.service` is preceded by a systemd warning: ```text bitborg-deploy-verify.service: multiple trigger source candidates for exit status propagation (podman-auto-update.service, podman-auto-update.service), skipping. ``` It is harmless today, but it is permanent, it is not rare, and it silently forecloses a design option. Worth writing down rather than leaving for whoever next reads this journal during an incident. ## Cause The drop-in names the same unit on both hooks: ```ini OnSuccess=bitborg-deploy-verify.service OnFailure=bitborg-deploy-verify.service ``` That is deliberate and correct — `podman auto-update` **exits 0 after a rollback**, so `OnFailure=` alone would never see one. But systemd populates `$MONITOR_UNIT` / `$MONITOR_EXIT_STATUS` / `$MONITOR_SERVICE_RESULT` for the *triggering* unit, and with two candidate sources it cannot decide which to propagate. So it propagates neither and says so. ## Frequency: every auto-update tick, not every deploy `OnSuccess=` fires whenever `podman-auto-update.service` succeeds, which is every timer tick (~3 min), not only when an image actually changed — the verifier's own early exit is what makes a quiet tick cheap. Measured on the services host: **5 occurrences in the ~25 minutes since the drop-in was installed**, i.e. roughly **480 lines/day** shipped to Loki, indefinitely. ## Consequence Nothing is broken. `bitborg-deploy-verify.sh` reads no `MONITOR_*` variable anywhere — it determines what happened from `podman inspect` plus a `journalctl … grep 'start operation timed out'` scan over a window, precisely because the rollback case is invisible to systemd's exit status. So the propagated status was never load-bearing. What the warning does mean is that the verifier **can never learn which of the two hooks invoked it**. Any future logic wanting to branch on "auto-update failed" versus "auto-update succeeded" is not available from the trigger, and would need two separate units (or a wrapper) rather than a `MONITOR_*` check. That is a real constraint and it is currently recorded nowhere. ## Suggested Documentation, not a behaviour change: - A comment in `roles/web/templates/podman-auto-update-verify.conf.j2` stating that naming one unit on both hooks is intentional, that it costs `MONITOR_*` propagation, and that the verifier does not rely on it. The template already explains *why* both hooks are needed; it does not mention this cost. - A line in `docs/runbook.md` under the deploy-verification section so the warning is identifiable as expected when someone greps the journal during an incident. Deliberately **not** suggested: dropping `OnFailure=` to silence it. That hook is there because a failed pull or registry outage should still trigger verification, and trading a real signal for a quiet journal is the wrong way round. If the volume is judged too high on its own, the alternative is a Loki drop rule for that exact line — but that hides a systemd diagnostic wholesale, so the comment plus runbook entry is the cheaper fix. Refs #315.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#318
Ingen beskrivning angiven.