fix(ansible): stop Ansible and podman fighting over directory ownership/modes #258
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!258
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/idempotent-dir-ownership"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Every apply reported 8 spurious
changedtasks on a fully converged host. Three separate cases,all the same shape: Ansible declares state that a container mechanism immediately overrides, so
the two fight on every run.
1.
~/.config/systemd/user— the mode was ordering-dependentNine roles create this one shared path. Three (
caddy,forgejo,monitoring) declare0755.The other six had it mixed into a loop whose mode is
0700— because that loop also covers~/.config/gitborg, which holds the 0600 token env files. So the final mode depended on which roleran last, and whichever way it landed the other set reported
changed.Split into its own task at
0755in all six. That is the value to converge on: it is what the hostalready has, and
0700protected nothing here — the unit files inside are0644, while the realsecrets stay in
~/.config/gitborg, which remains0700.2.
/srv/gitborg-lfs— Ansible vspodman unshare chownGive the storage mountpoint to the bitborg user set
owner=bitborg(2000), and the very nexttask handed it to the container's uid via
podman unshare chown 1000:1000. So every apply chowned aconverged host to 2000 and straight back to the mapped subuid — reporting
changedeach time, andbriefly leaving the directory owned by the uid whose use here caused a production outage during the
ADR 0031 cutover.
podman unshare chownis authoritative, so ownership is no longer declared; only the mode is. Renamedto Ensure the storage mountpoint exists, since it no longer gives anything to the bitborg user.
3.
~/.cache/renovate— Ansible vs the:Umountgitborg-renovate.sh.j2mounts the cache as:U, which makes podman chown it to the container uid —renovate's image runs as uid 12021, landing on a mapped subuid on the host (177556 = 165536 + 12021 − 1,
which is exactly what the dry-run showed). Ansible declared
owner=bitborg, so each apply chowned itback and the next renovate run chowned it forward.
:Uis authoritative; we now only ensure thedirectory exists as the mount source.
Effect on production: none
Every case already converged to the container-owned or
0755state. Measured against real prod:changedon a converged hostThe remaining six are the genuinely pending changes — #254's Quadlet and sentinel dir, the web restart,
#256's renovate onboarding config, and #250's Caddyfile (merged but never applied).
That is the real win:
--checkbecomes a usable signal instead of 8 permanent false positives thatmask actual drift. #250 sitting unapplied is exactly the kind of thing they were hiding.
Verification
ansible-playbook site.yml --check --diff --limit gitborg-prod→ok=222 changed=6 failed=0--syntax-checkpasses;ansible-lintclean at theproductionprofile across all seven touched roles.One caught in review: the first attempt at the
backuprole orphaned{{ backup_restic_cache_dir }}out of its loop, which
--syntax-checkcaught as a YAML parse error. It is back in the0700loopwhere it belongs.
830e79fccbb796025985