fix(ansible): stop Ansible and podman fighting over directory ownership/modes #258

Sammanfogat
supernaut sammanfogade 2 incheckningar från fix/idempotent-dir-ownership in i main 2026-07-30 16:20:13 +00:00
Ägare

Every apply reported 8 spurious changed tasks on a fully converged host. Three separate cases,
all the same shape: Ansible declares state that a container mechanism immediately overrides, so
the two fight on every run.

Stacked on #257 — that PR makes ansible-playbook --check complete at all, without which none of
this was visible. Merge #257 first and this PR's diff reduces to the churn fixes alone.

1. ~/.config/systemd/user — the mode was ordering-dependent

Nine roles create this one shared path. Three (caddy, forgejo, monitoring) declare 0755.
The other six had it mixed into a loop whose mode is 0700 — because that loop also covers
~/.config/gitborg, which holds the 0600 token env files. So the final mode depended on which role
ran last, and whichever way it landed the other set reported changed.

Split into its own task at 0755 in all six. That is the value to converge on: it is what the host
already has, and 0700 protected nothing here — the unit files inside are 0644, while the real
secrets stay in ~/.config/gitborg, which remains 0700.

2. /srv/gitborg-lfs — Ansible vs podman unshare chown

Give the storage mountpoint to the bitborg user set owner=bitborg (2000), and the very next
task
handed it to the container's uid via podman unshare chown 1000:1000. So every apply chowned a
converged host to 2000 and straight back to the mapped subuid — reporting changed each time, and
briefly leaving the directory owned by the uid whose use here caused a production outage during the
ADR 0031 cutover.

podman unshare chown is authoritative, so ownership is no longer declared; only the mode is. Renamed
to Ensure the storage mountpoint exists, since it no longer gives anything to the bitborg user.

3. ~/.cache/renovate — Ansible vs the :U mount

gitborg-renovate.sh.j2 mounts the cache as :U, which makes podman chown it to the container uid —
renovate's image runs as uid 12021, landing on a mapped subuid on the host (177556 = 165536 + 12021 − 1,
which is exactly what the dry-run showed). Ansible declared owner=bitborg, so each apply chowned it
back and the next renovate run chowned it forward. :U is authoritative; we now only ensure the
directory exists as the mount source.

Effect on production: none

Every case already converged to the container-owned or 0755 state. Measured against real prod:

before after
changed on a converged host 13 6
of which spurious 7 0

The remaining six are the genuinely pending changes — #254's Quadlet and sentinel dir, the web restart,
#256's renovate onboarding config, and #250's Caddyfile (merged but never applied).

That is the real win: --check becomes a usable signal instead of 8 permanent false positives that
mask actual drift. #250 sitting unapplied is exactly the kind of thing they were hiding.

Verification

  • ansible-playbook site.yml --check --diff --limit gitborg-prod → ok=222 changed=6 failed=0
  • --syntax-check passes; ansible-lint clean at the production profile across all seven touched roles.

One caught in review: the first attempt at the backup role orphaned {{ backup_restic_cache_dir }}
out of its loop, which --syntax-check caught as a YAML parse error. It is back in the 0700 loop
where it belongs.

Every apply reported **8 spurious `changed` tasks on a fully converged host**. Three separate cases, all the same shape: **Ansible declares state that a container mechanism immediately overrides**, so the two fight on every run. > Stacked on #257 — that PR makes `ansible-playbook --check` complete at all, without which none of > this was visible. Merge #257 first and this PR's diff reduces to the churn fixes alone. ## 1. `~/.config/systemd/user` — the mode was ordering-dependent **Nine** roles create this one shared path. Three (`caddy`, `forgejo`, `monitoring`) declare `0755`. The other six had it mixed into a loop whose mode is `0700` — because that loop *also* covers `~/.config/gitborg`, which holds the 0600 token env files. So the final mode depended on which role ran last, and whichever way it landed the other set reported `changed`. Split into its own task at `0755` in all six. That is the value to converge on: it is what the host already has, and `0700` protected nothing here — the unit files inside are `0644`, while the real secrets stay in `~/.config/gitborg`, which remains `0700`. ## 2. `/srv/gitborg-lfs` — Ansible vs `podman unshare chown` *Give the storage mountpoint to the bitborg user* set `owner=bitborg` (2000), and the **very next task** handed it to the container's uid via `podman unshare chown 1000:1000`. So every apply chowned a converged host to 2000 and straight back to the mapped subuid — reporting `changed` each time, and briefly leaving the directory owned by the uid whose use here caused a production outage during the ADR 0031 cutover. `podman unshare chown` is authoritative, so ownership is no longer declared; only the mode is. Renamed to *Ensure the storage mountpoint exists*, since it no longer gives anything to the bitborg user. ## 3. `~/.cache/renovate` — Ansible vs the `:U` mount `gitborg-renovate.sh.j2` mounts the cache as `:U`, which makes podman chown it to the container uid — renovate's image runs as uid 12021, landing on a mapped subuid on the host (177556 = 165536 + 12021 − 1, which is exactly what the dry-run showed). Ansible declared `owner=bitborg`, so each apply chowned it back and the next renovate run chowned it forward. `:U` is authoritative; we now only ensure the directory exists as the mount source. ## Effect on production: none Every case already converged to the container-owned or `0755` state. Measured against real prod: | | before | after | | --- | --- | --- | | `changed` on a converged host | **13** | **6** | | of which spurious | **7** | **0** | The remaining six are the genuinely pending changes — #254's Quadlet and sentinel dir, the web restart, #256's renovate onboarding config, and #250's Caddyfile (merged but never applied). That is the real win: `--check` becomes a usable signal instead of 8 permanent false positives that mask actual drift. #250 sitting unapplied is exactly the kind of thing they were hiding. ## Verification - `ansible-playbook site.yml --check --diff --limit gitborg-prod` → `ok=222 changed=6 failed=0` - `--syntax-check` passes; `ansible-lint` clean at the `production` profile across all seven touched roles. One caught in review: the first attempt at the `backup` role orphaned `{{ backup_restic_cache_dir }}` out of its loop, which `--syntax-check` caught as a YAML parse error. It is back in the `0700` loop where it belongs.
supernaut lade till 2 incheckningar 2026-07-30 15:47:43 +00:00
fix(forgejo): let the playbook survive --check past the system-webhook task
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m22s
386135a918
`ansible-playbook site.yml --check` aborted at the forgejo role with:

  A 'when' expression failed: object of type 'dict' has no attribute 'json'
  roles/forgejo/tasks/system-webhook.yml:27

ansible.builtin.uri does not run under --check by default, even for a read-only
GET. So "Check for the reconcile-trigger system webhook" was skipped,
`_fj_system_hooks` registered as a skip-dict with no `.json`, and the `when` on
the following warn task died on it — taking the whole play down (failed=1) before
the web, caddy, renovate and health-check roles were ever evaluated.

A real run was unaffected, which is why this went unnoticed: the GET executes,
`.json` exists, the verification works. But it made `--check` unusable past this
role, and every production apply is supposed to be gated on a clean dry-run — so
in practice the gate could not be satisfied.

Two changes:

- `check_mode: false` on the GET. It is read-only, so running it under --check is
  safe and is what lets the verification below evaluate at all.
- `| default([])` on the `when`. Belt-and-braces: if this GET is ever skipped
  again or made failure-tolerant, a missing `.json` should degrade to "warn"
  rather than abort the play.

With this, `--check --diff` completes: ok=215 changed=13 failed=0.
fix(ansible): stop Ansible and podman fighting over directory ownership/modes
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m29s
830e79fccb
Every apply reported 8 spurious `changed` tasks on a fully converged host. Three
separate cases, all the same shape: Ansible declares state that a container
mechanism immediately overrides, so the two fight on every run.

1. ~/.config/systemd/user — mode was ordering-dependent

NINE roles create this one shared path. Three (caddy, forgejo, monitoring)
declare 0755; six had it mixed into a loop whose mode is 0700 because the loop
ALSO covers ~/.config/gitborg, which holds the 0600 token env files. So the final
mode depended on which role ran last, and whichever way it landed the other set
reported `changed`.

Split out of those loops into its own task at 0755 in all six. 0755 is the value
to converge on: it is what the host already has, and 0700 protected nothing here —
the unit files inside are 0644, while the real secrets stay in ~/.config/gitborg,
which remains 0700.

2. /srv/gitborg-lfs — Ansible vs `podman unshare chown`

"Give the storage mountpoint to the gitborg user" set owner=gitborg (2000), and
the very next task handed it to the container's uid via `podman unshare chown
1000:1000`. Every apply therefore chowned a converged host to 2000 and straight
back to the mapped subuid, reporting `changed` each time — and briefly leaving the
dir owned by the uid whose use here caused a production outage during the ADR 0031
cutover.

The `podman unshare chown` is authoritative, so ownership is no longer declared;
only the mode is. Renamed to "Ensure the storage mountpoint exists", since it no
longer gives anything to the gitborg user.

3. ~/.cache/renovate — Ansible vs the `:U` mount

The wrapper mounts this dir as `:U` (gitborg-renovate.sh.j2), which makes podman
chown it to the container uid; renovate's image runs as uid 12021, landing on a
mapped subuid on the host. Ansible declared owner=gitborg, so each apply chowned
it back and the next renovate run chowned it forward. `:U` is authoritative — we
now only ensure the directory exists as the mount source.

Net effect on production: none. Every case already converged to the
container-owned or 0755 state; this makes the playbook agree with it instead of
flapping, so `--check` becomes a usable signal rather than 8 permanent false
positives that mask real drift.

Verified: --syntax-check passes, ansible-lint clean at the production profile
across all seven touched roles.
supernaut tvångsskickade fix/idempotent-dir-ownership från 830e79fccb
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m29s
till b796025985
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m30s
2026-07-30 16:09:31 +00:00
Jämför
supernaut sammanfogade incheckning de300bd7ab till main 2026-07-30 16:20:13 +00:00
supernaut tog bort grenen fix/idempotent-dir-ownership 2026-07-30 16:20:13 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!258
Ingen beskrivning angiven.