deploy: UserNS=keep-id makes every web deploy a 100s outage (regression from #238) #253
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#253
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Every
bitborg-webdeploy has taken the portal down for 83–139 seconds since 2026-07-28, up fromabout 13 seconds before that. The cause is
UserNS=keep-idon the web Quadlet unit, added by#238 (ADR 0037 Phase 1). This is what the 502s behind #249 / #250 were actually a symptom of.
Measurement
Podman stamps each journald event with a monotonic offset (
m=+N) from the start of the command thatcaused it, so the phases of a deploy can be priced exactly rather than inferred from probe samples.
Breakdown of the 2026-07-29 22:27 UTC deploy:
podman auto-updatepullpodman createTwo things worth stating plainly, because both contradict the assumption made while triaging #249:
scripts/migrate.mjs, Astrolistening — is 2.8 s. Migrations are 2 % of the outage.
podman create.It is a regression, not a standing property
container createduration forbitborg-web, every deploy since 07-24:Cleanly bimodal: eight deploys at ~0.06 s, then every deploy since at 83–139 s, a ~1500× regression.
The step lands between 07-28 21:42 and 22:16 — which brackets the apply of #238
(
adb2fd8, merged 07-28 21:39 UTC), the only change toansible/roles/web/templates/bitborg-web.container.j2in that window.Cause
That commit added, to give the container's
nodeuser (uid 1000) the hostgitborguid so theADR 0037 sentinel it writes is owned by
gitborg:Rootless Podman with
keep-idhas to produce a rootfs with shifted ownership on every containercreation. For the web image —
node:24-trixie-slimplus a productionnode_modules— that is arecursive chown across tens of thousands of files.
The cost scales with file count, which is why only this container shows it:
gitborg-reconcilercarries the identical
UserNS=keep-id:uid=1000,gid=1000line and creates in 0.05–0.2 s, because itsimage is small. So
keep-idis not wrong in general — it is wrong on our largest image.Proposed fix
Idmap only the bind mount that needs it, instead of remapping the whole rootfs:
and drop
UserNS=keep-idfrom this unit (the reconciler keeps it — small image, andnode_exporter must be able to read its textfile metrics).
The host side of ADR 0037 only ever performs metadata operations on the sentinel —
PathExists=inbitborg-reconcile.path, andstat -c %Yplusrm -finbitborg-reconcile-kick.sh. None of thoseread file contents, and
rmneeds write permission on the directory, not the file. So thesentinel's own ownership does not have to match
gitborg; only the directory's does, and it isalready
bitborg:bitborg 0755.Prerequisite to verify before committing to this:
idmapon a bind mount requires the sourcefilesystem to support idmapped mounts (kernel ≥ 5.12, ext4/xfs — not overlayfs).
/home/gitborg/reconcile-triggerneeds checking on the host. If it is unsupported, the fallback is tomake the sentinel directory writable by the container's mapped subuid — a group or a mode change —
which keeps the default userns and the fast create path.
Acceptance
idmapbind-mount support ongitborg-prodfor the sentinel directory's filesystem.UserNS=keep-idwith an idmapped volume on the web unit (or apply the fallback).container createforbitborg-webis back to sub-second, from the podmanevent
m=+offset — not from probe samples, which lack the resolution to attribute this..pathunit kicks,
gitborg_reconcile_trigger_lag_secondsis emitted.HealthStartPeriod=30smargin against the measured 2.8 s in-container startup.Notes
figure at ~117 s, #250's 2 s window is honest about not bridging a deploy — that stays correct
regardless of this fix, and becomes an adequate mitigation once creation is sub-second again.
both fixes, expected deploy unavailability is roughly 3 s.
Correction to the proposed fix above — plain
,idmapis not sufficient.The diagnosis stands, but the one-line fix I suggested does not work as written.
Volume=…:rw,idmapidmaps the bind mount using the container's default user-namespace mapping.For rootless
gitborg(uid 2000) that mapping is: container uid 0 → host 2000, container uids 1–65536→ the subuid range. So a sentinel directory owned by host uid 2000 appears inside the container as
owned by uid 0, and the app runs as
node(uid 1000) — which still cannot write it. Same failurekeep-idwas added to avoid.The candidate that should work is an explicit mapping, scoped to that one mount:
@2000-1000-1maps host uid 2000 to container uid 1000 for a range of 1 — the mount-scopedequivalent of
keep-id:uid=1000, without remapping the rootfs. Two caveats before relying on it:;separator and@syntax have to survive Quadlet'sVolume=generation. This unit alreadycarries a comment documenting that Quadlet/podman 5.4.2 corrupt embedded double quotes in
HealthCmd, so the generator's escaping is not to be taken on trust — verify withquadlet -dryrunand confirm the resulting--mountinpodman inspect.Hardcoding uid 2000 also duplicates
gitborg_user_uid; it should be templated asuids=@{{ gitborg_user_uid }}-1000-1, with 1000 matching thenodeuid in bitborg-web'sContainerfile.
Not committing to this until it is measured on the host. A diagnostic runs six variants — default
userns,
keep-id, plainidmap, explicitidmap— timingpodman createfor each and probingwhether uid 1000 can actually write the sentinel. Options if the explicit mapping does not hold up:
userns and the fast create path.
keep-idcan use idmapped mounts instead of a recursive chown on this host — thatwould fix the cost globally with no unit change. The ~100 s strongly suggests it is currently
falling back to chown;
podman infostorage options will say.Phase 2's shim already does for the webhook path.
Option 2 is worth checking first — it is the only one that fixes the cause rather than working around
it, and it would also protect any future container that needs
keep-id.Measured on the host. The cause is confirmed; both fixes proposed above are wrong. Details below.
The mechanism, now established
The first attempt to reproduce this failed misleadingly:
podman create --userns keep-id:uid=1000,gid=1000against the running image took 113 ms, which looked like a refutation. It was not — the
shifted-ownership copy for that image and that mapping was already cached in the store.
Re-running with a mapping never used before (
uid=1001) on the same image forces a cold cache:keep-idkeep-id:uid=1001— cold cachekeep-id:uid=1001— repeat, now cachedkeep-id:uid=1000— prod's mapping, cached108.8 s reproduces the ~103 s seen in prod deploys. So:
keep-idtriggers a full recursive chown of the rootfs — 849 MB across 11 layers — cached per(image digest, mapping). Every deploy ships a new digest, so every deploy pays it cold. A warm-cache
run is 538× faster, which is why this is invisible to any test that reuses the running image.
The reason Podman has to chown at all is in
podman info:Supports shifting: false— the overlay driver cannot shift UIDs on the fly, so a non-identitymapping can only be satisfied by physically rewriting ownership. This is also why
gitborg-reconcilercarries the identical directive at ~0.1 s: the cost is proportional to imagesize, and its image is small. The directive is not wrong in general; it is wrong on our largest image.
Both proposed fixes are dead
Idmapped bind mounts are unavailable on this host for a rootless container:
The relative
@-prefixed form fails earlier still, on mapping resolution:So strike the
rw,idmapsuggestion in the issue body and the explicit-mapping variant in the firstcomment.
Supports shifting: falsepredicted this and I should have read it before proposing either.Surviving options
Drop
UserNS=keep-idfrom the web unit. Recovers the full ~109 s. ADR 0037's web-side triggerdegrades to the 5-min timer backstop — which is by design:
triggerReconcile(
bitborg-web/src/lib/reconcile-trigger.ts) catches every write failure and never throws, and thetimer is documented as the fallback. Net effect is pre-#238 behaviour: entitlement changes apply on
the next tick instead of instantly.
Own the sentinel directory as the container's mapped uid. Keeps both the instant trigger and
the fast create path. With the default rootless userns, container uid 1000 lands in the subuid range
(
bitborg:165536:65536). Setting the directoryowner=<mapped uid>, group=bitborg, mode=0775letsthe container write as owner while the host side keeps the group bit it needs — and the host side
only ever does metadata operations (
PathExists=,stat -c %Y,rm -f), never a content read.Express it as
podman unshare chown 1000:1000 <dir>, which performs the mapping arithmetic itselfrather than hardcoding 166535 — that number is a function of
/etc/subuidand should not beembedded in a playbook.
Fails safe if the mapping ever changes: the write fails, is logged, and the timer reconciles.
Shrink the image. Mitigation only — the chown stays O(image), just smaller.
Replace the file sentinel with a network call, as ADR 0037 Phase 2's shim already does for the
webhook path. The largest change; worth considering if Phase 2 lands anyway.
Option 2 is the recommendation, pending a host test of the ownership scheme end to end. Option 1 is
available as a one-line mitigation at any time and is worth taking immediately if option 2 needs more
than a moment — a recurring ~2-minute outage on every deploy is worse than 5-minute entitlement latency.
Correction to the acceptance criteria
The first item above ("confirm idmap bind-mount support") is settled: not supported. Replace with
verifying whichever ownership scheme is chosen, and keep the requirement to measure
container createfrom the podman event
m=+offset after the fix — noting that a warm cache makes any post-deployre-test of the same digest meaningless. It has to be measured on a genuinely new image.