Image-tag bumps silently no-op on prod (postgres/forgejo/caddy roles don't pull+restart) #63
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#63
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
On gitborg-prod, bumping a pinned container image tag (e.g. via #62: forgejo/postgres/caddy)
re-renders the
.containerQuadlet unit but does not bring the new image live. Applying thetag bump reports
changedand success, yet the running containers stay on the old image untilsomeone restarts them by hand. Discovered while applying #62 — prod had to be converged manually
(
podman pullthe three new tags +systemctl --user restart postgres forgejo caddy).Two distinct defects
postgresandcaddyroles never restart on an image-only change. Their start step isstart-if-stopped — a no-op when the container is already running — so a re-rendered unit with a
new
Image=tag is never picked up. (Only the Caddyfile-changed path triggers a Caddy restart.)The
forgejorole does restart on unit change, but after #62 it came back on15.0.3because15.0.4was never pulled to the host. There is nopodman_imagepull task forforgejo/postgres/caddy anywhere.
The
monitoringrole already does this correctly — its "Enable and start monitoring services(restart any whose unit changed)" task converged ntfy
v2.26.0+ caddy2.11.4-alpineautomatically during the same apply, no manual step. So the fix pattern already exists in-repo.
Impact
Any future image-tag bump applied through the
postgres/forgejo/caddyroles silently no-ops onprod — the security update is not actually deployed despite a green apply. Easy to believe you're
patched when you aren't.
Proposed fix
containers.podman.podman_imageor equivalent) for the pinnedimages, gated on the tag/unit having changed, before the start step — mirroring the
webrole's"Pull the image so a fresh deploy has it" task.
.containerunit changed, copying themonitoringrole's"restart any whose unit changed" pattern into
postgres/forgejo/caddy.--check-mode verification that the running image digest matches the unit's tag, so ano-op apply is visible.
Surfaced by: #62 · Related: #61 (these image tags aren't Renovate-tracked)
Labels: area/infra, type/bug
Reproduced live during the v16 upgrade (#73), with the mechanism pinned down:
site.yml --tags forgejoupdated the Quadlet.container(Image 15.0.4 → 16.0.0) and reportedRestart Forgejochanged — but the container came back up on 15.0.4 (no 16.0.0 image was even pulled).forgejo: Restart Forgejobefore the systemd user daemon-reload handler (which came from the runner-controller role), so the restart used the stale generated unit still pointing at 15.0.4. After the reload, the unit said 16.0.0 (NeedDaemonReload=no) but nothing restarted it again.roles/forgejo/tasks/main.yml(~L68: "Reload (podman role) runs before Restart at flush") does not hold under--tags forgejo— the podman role's reload handler isn't the one notified in that scope.podman pull+systemctl --user restart forgejoby hand.Fix direction: make the forgejo role's Quadlet-install task notify a role-local daemon-reload handler ordered before
Restart Forgejo(handler listen/ordering that survives--tagsscoping), or fold reload+restart into one handler. Same pattern presumably affects postgres/caddy roles.Reproduced live during the v16 upgrade (#73), with the mechanism pinned down:
site.yml --tags forgejoupdated the Quadlet.container(Image 15.0.4 → 16.0.0) and reportedRestart Forgejochanged — but the container came back up on 15.0.4 (no 16.0.0 image was even pulled).forgejo: Restart Forgejobefore the systemd user daemon-reload handler (which came from the runner-controller role), so the restart used the stale generated unit still pointing at 15.0.4. After the reload, the unit said 16.0.0 (NeedDaemonReload=no) but nothing restarted it again.roles/forgejo/tasks/main.yml(~L68: "Reload (podman role) runs before Restart at flush") does not hold under--tags forgejo— the podman role's reload handler isn't the one notified in that scope.podman pull+systemctl --user restart forgejoby hand.Fix direction: make the forgejo role's Quadlet-install task notify a role-local daemon-reload handler ordered before
Restart Forgejo(handler listen/ordering that survives--tagsscoping), or fold reload+restart into one handler. Same pattern presumably affects postgres/caddy roles.