fix(web): drop UserNS=keep-id — it cost ~109s of downtime per deploy #254
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!254
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/web-keepid-deploy-outage"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Fixes #253. Every
bitborg-webdeploy has taken the portal down for 83–139 s since 2026-07-28,up from ~13 s before. This removes ~109 s of that.
Cause
UserNS=keep-id:uid=1000,gid=1000, added to the web Quadlet by #238 (ADR 0037 Phase 1) so thesentinel the container writes would be owned by
gitborg.The overlay store on this host reports
Supports shifting: false, so a non-identity user-namespacemapping can only be satisfied by physically chowning the entire rootfs — 849 MB across 11 layers.
That result is cached per
(image digest, mapping), and every deploy ships a new digest, so everydeploy paid it cold.
Measured on the host, same image:
keep-idkeep-id, cold mappingkeep-id, same mapping cachedkeep-id, prod's mapping (cached)108.8 s reproduces the ~103 s seen in prod. The cache is why this hid so well — the first attempt to
reproduce it used prod's own mapping and came back at 113 ms, which looked like a refutation. Forcing
a mapping never used before (
uid=1001) exposes it.The wider deploy breakdown, from podman's journald event offsets:
podman auto-updatepullpodman createWorth stating plainly because it contradicts the assumption made while triaging #249: migrations are
2 % of the outage. In-container startup is 2.8 s.
Fix
Drop
keep-idand own the sentinel directory as the container's mapped uid instead —owner=<mapped uid>, group=bitborg, mode=0775:.pathunit and kick script keep the group bit, which covers every operation theyperform —
PathExists=,stat -c %Y, andrm(write on the directory, not the file). Nothingreads the sentinel's contents.
The owner is derived from
/etc/subuid(container uid N →subuid_start + N - 1) rather thanhardcoded, with an
assertso a missing range fails loudly instead of silently handing the directoryto uid 999.
web_container_uidis a new default documenting the coupling toUSER nodein bitborg-web'sContainerfile — if the image's user changes, the trigger stops firing and falls back to the timer.
Verification
Ownership scheme tested end to end on the host before this was written — container write, host
stat/rm, and re-write after removal all pass, with the directory at166535:2000 drwxrwxr-x:The
/etc/subuidparsing and uid arithmetic are unit-tested offline for both the happy path(
165536→ owner166535) and the missing-range path (yields 0, so the assert fires).ansible-playbook --syntax-checkpasses andansible-lintis clean at theproductionprofile.Not yet verified, and it cannot be until this is applied: cold-cache create time in prod. A
re-test against the same image digest will look fast whether or not this worked, because the cache is
already warm. The real proof is the first genuinely new digest deployed after the apply, measured from
the podman event
m=+offset — not from blackbox probe samples, whose 30 s spacing is what got thismisattributed in the first place.
What was tried and rejected
Idmapped bind mounts (
Volume=…:rw,idmap) were the first choice — they would have keptkeep-id'sownership guarantee while leaving the rootfs alone. They are unavailable here:
and the relative
@form fails earlier still, on mapping resolution. Both are consistent with thesame
Supports shifting: false. This also means the change carries no new Quadlet quoting risk — itremoves a line rather than adding exotic
Volume=syntax, so theHealthCmdquote-corruption hazarddocumented in this unit does not apply.
Notes
the backstop as designed.
keep-idthe container's uid maps to an unprivilegedsubuid with no host presence, rather than to
gitborg— better least privilege (principle 4).keep-id. Its image is small and it pays ~0.1 s. The directive is not wrongin general, only on our largest image.
.pathunit firesand the kick script (which owns the directory) removes it, so no migration task is needed.
SIGTERMhandler, filed separately againstbitborg-web. After both, expected deploy unavailability is ~3 s.
honest about not bridging a deploy, and becomes an adequate mitigation once creation is sub-second.