Image-tag bumps silently no-op on prod (postgres/forgejo/caddy roles don't pull+restart) #63

Stängd
öppnade 2026-07-10 13:03:26 +00:00 av supernaut · 2 kommentarer
Ägare

On gitborg-prod, bumping a pinned container image tag (e.g. via #62: forgejo/postgres/caddy)
re-renders the .container Quadlet unit but does not bring the new image live. Applying the
tag bump reports changed and success, yet the running containers stay on the old image until
someone restarts them by hand. Discovered while applying #62 — prod had to be converged manually
(podman pull the three new tags + systemctl --user restart postgres forgejo caddy).

Two distinct defects

  1. postgres and caddy roles never restart on an image-only change. Their start step is
    start-if-stopped — a no-op when the container is already running — so a re-rendered unit with a
    new Image= tag is never picked up. (Only the Caddyfile-changed path triggers a Caddy restart.)
  2. No pull step for the pinned images, and the restart that does happen doesn't reliably pull.
    The forgejo role does restart on unit change, but after #62 it came back on 15.0.3 because
    15.0.4 was never pulled to the host. There is no podman_image pull task for
    forgejo/postgres/caddy anywhere.

The monitoring role already does this correctly — its "Enable and start monitoring services
(restart any whose unit changed)"
task converged ntfy v2.26.0 + caddy 2.11.4-alpine
automatically during the same apply, no manual step. So the fix pattern already exists in-repo.

Impact

Any future image-tag bump applied through the postgres/forgejo/caddy roles silently no-ops on
prod — the security update is not actually deployed despite a green apply. Easy to believe you're
patched when you aren't.

Proposed fix

  • Add an explicit image-pull task (containers.podman.podman_image or equivalent) for the pinned
    images, gated on the tag/unit having changed, before the start step — mirroring the web role's
    "Pull the image so a fresh deploy has it" task.
  • Restart the container when its .container unit changed, copying the monitoring role's
    "restart any whose unit changed" pattern into postgres/forgejo/caddy.
  • Consider a --check-mode verification that the running image digest matches the unit's tag, so a
    no-op apply is visible.

Surfaced by: #62 · Related: #61 (these image tags aren't Renovate-tracked)

Labels: area/infra, type/bug

On **gitborg-prod**, bumping a pinned container image tag (e.g. via #62: forgejo/postgres/caddy) re-renders the `.container` Quadlet unit but **does not bring the new image live**. Applying the tag bump reports `changed` and success, yet the running containers stay on the old image until someone restarts them by hand. Discovered while applying #62 — prod had to be converged manually (`podman pull` the three new tags + `systemctl --user restart postgres forgejo caddy`). ## Two distinct defects 1. **`postgres` and `caddy` roles never restart on an image-only change.** Their start step is start-if-stopped — a no-op when the container is already running — so a re-rendered unit with a new `Image=` tag is never picked up. (Only the Caddyfile-changed path triggers a Caddy restart.) 2. **No pull step for the pinned images, and the restart that does happen doesn't reliably pull.** The `forgejo` role *does* restart on unit change, but after #62 it came back on `15.0.3` because `15.0.4` was never pulled to the host. There is no `podman_image` pull task for forgejo/postgres/caddy anywhere. The `monitoring` role already does this correctly — its *"Enable and start monitoring services (restart any whose unit changed)"* task converged ntfy `v2.26.0` + caddy `2.11.4-alpine` automatically during the same apply, no manual step. So the fix pattern already exists in-repo. ## Impact Any future image-tag bump applied through the `postgres`/`forgejo`/`caddy` roles silently no-ops on prod — the security update is *not* actually deployed despite a green apply. Easy to believe you're patched when you aren't. ## Proposed fix - Add an explicit image-pull task (`containers.podman.podman_image` or equivalent) for the pinned images, gated on the tag/unit having changed, before the start step — mirroring the `web` role's *"Pull the image so a fresh deploy has it"* task. - Restart the container when its `.container` unit changed, copying the `monitoring` role's *"restart any whose unit changed"* pattern into `postgres`/`forgejo`/`caddy`. - Consider a `--check`-mode verification that the running image digest matches the unit's tag, so a no-op apply is visible. Surfaced by: #62 · Related: #61 (these image tags aren't Renovate-tracked) Labels: area/infra, type/bug
Upphovsperson
Ägare

Reproduced live during the v16 upgrade (#73), with the mechanism pinned down:

  • site.yml --tags forgejo updated the Quadlet .container (Image 15.0.4 → 16.0.0) and reported Restart Forgejo changed — but the container came back up on 15.0.4 (no 16.0.0 image was even pulled).
  • Cause: handler ordering. The run's flush executed forgejo: Restart Forgejo before the systemd user daemon-reload handler (which came from the runner-controller role), so the restart used the stale generated unit still pointing at 15.0.4. After the reload, the unit said 16.0.0 (NeedDaemonReload=no) but nothing restarted it again.
  • The comment in roles/forgejo/tasks/main.yml (~L68: "Reload (podman role) runs before Restart at flush") does not hold under --tags forgejo — the podman role's reload handler isn't the one notified in that scope.
  • Recovery: podman pull + systemctl --user restart forgejo by hand.

Fix direction: make the forgejo role's Quadlet-install task notify a role-local daemon-reload handler ordered before Restart Forgejo (handler listen/ordering that survives --tags scoping), or fold reload+restart into one handler. Same pattern presumably affects postgres/caddy roles.

Reproduced live during the v16 upgrade (#73), with the mechanism pinned down: - `site.yml --tags forgejo` updated the Quadlet `.container` (Image 15.0.4 → 16.0.0) and reported `Restart Forgejo` **changed** — but the container came back up on **15.0.4** (no 16.0.0 image was even pulled). - Cause: handler ordering. The run's flush executed `forgejo: Restart Forgejo` **before** the systemd user daemon-reload handler (which came from the runner-controller role), so the restart used the **stale generated unit** still pointing at 15.0.4. After the reload, the unit said 16.0.0 (`NeedDaemonReload=no`) but nothing restarted it again. - The comment in `roles/forgejo/tasks/main.yml` (~L68: "Reload (podman role) runs before Restart at flush") does not hold under `--tags forgejo` — the podman role's reload handler isn't the one notified in that scope. - Recovery: `podman pull` + `systemctl --user restart forgejo` by hand. Fix direction: make the forgejo role's Quadlet-install task notify a role-local daemon-reload handler ordered before `Restart Forgejo` (handler listen/ordering that survives `--tags` scoping), or fold reload+restart into one handler. Same pattern presumably affects postgres/caddy roles.
Upphovsperson
Ägare

Reproduced live during the v16 upgrade (#73), with the mechanism pinned down:

  • site.yml --tags forgejo updated the Quadlet .container (Image 15.0.4 → 16.0.0) and reported Restart Forgejo changed — but the container came back up on 15.0.4 (no 16.0.0 image was even pulled).
  • Cause: handler ordering. The run's flush executed forgejo: Restart Forgejo before the systemd user daemon-reload handler (which came from the runner-controller role), so the restart used the stale generated unit still pointing at 15.0.4. After the reload, the unit said 16.0.0 (NeedDaemonReload=no) but nothing restarted it again.
  • The comment in roles/forgejo/tasks/main.yml (~L68: "Reload (podman role) runs before Restart at flush") does not hold under --tags forgejo — the podman role's reload handler isn't the one notified in that scope.
  • Recovery: podman pull + systemctl --user restart forgejo by hand.

Fix direction: make the forgejo role's Quadlet-install task notify a role-local daemon-reload handler ordered before Restart Forgejo (handler listen/ordering that survives --tags scoping), or fold reload+restart into one handler. Same pattern presumably affects postgres/caddy roles.

Reproduced live during the v16 upgrade (#73), with the mechanism pinned down: - `site.yml --tags forgejo` updated the Quadlet `.container` (Image 15.0.4 → 16.0.0) and reported `Restart Forgejo` **changed** — but the container came back up on **15.0.4** (no 16.0.0 image was even pulled). - Cause: handler ordering. The run's flush executed `forgejo: Restart Forgejo` **before** the systemd user daemon-reload handler (which came from the runner-controller role), so the restart used the **stale generated unit** still pointing at 15.0.4. After the reload, the unit said 16.0.0 (`NeedDaemonReload=no`) but nothing restarted it again. - The comment in `roles/forgejo/tasks/main.yml` (~L68: "Reload (podman role) runs before Restart at flush") does not hold under `--tags forgejo` — the podman role's reload handler isn't the one notified in that scope. - Recovery: `podman pull` + `systemctl --user restart forgejo` by hand. Fix direction: make the forgejo role's Quadlet-install task notify a role-local daemon-reload handler ordered before `Restart Forgejo` (handler listen/ordering that survives `--tags` scoping), or fold reload+restart into one handler. Same pattern presumably affects postgres/caddy roles.
supernaut refererade till detta ärende från en incheckning 2026-07-28 18:00:44 +00:00
supernaut refererade till detta ärende från en incheckning 2026-08-03 09:41:34 +00:00
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#63
Ingen beskrivning angiven.