fix(web): drop UserNS=keep-id — it cost ~109s of downtime per deploy #254

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/web-keepid-deploy-outage in i main 2026-07-30 14:22:22 +00:00
Ägare

Fixes #253. Every bitborg-web deploy has taken the portal down for 83–139 s since 2026-07-28,
up from ~13 s before. This removes ~109 s of that.

Cause

UserNS=keep-id:uid=1000,gid=1000, added to the web Quadlet by #238 (ADR 0037 Phase 1) so the
sentinel the container writes would be owned by gitborg.

The overlay store on this host reports Supports shifting: false, so a non-identity user-namespace
mapping can only be satisfied by physically chowning the entire rootfs — 849 MB across 11 layers.
That result is cached per (image digest, mapping), and every deploy ships a new digest, so every
deploy paid it cold.

Measured on the host, same image:

Variant Create time
default userns, no keep-id 129 ms
keep-id, cold mapping 108 786 ms
keep-id, same mapping cached 202 ms
keep-id, prod's mapping (cached) 149 ms

108.8 s reproduces the ~103 s seen in prod. The cache is why this hid so well — the first attempt to
reproduce it used prod's own mapping and came back at 113 ms, which looked like a refutation. Forcing
a mapping never used before (uid=1001) exposes it.

The wider deploy breakdown, from podman's journald event offsets:

Phase Duration
podman auto-update pull 0.7 s
SIGTERM ignored → SIGKILL 10.1 s
podman create 103.2 s
node boot + migrations 2.3 s
Total ~117 s

Worth stating plainly because it contradicts the assumption made while triaging #249: migrations are
2 % of the outage.
In-container startup is 2.8 s.

Fix

Drop keep-id and own the sentinel directory as the container's mapped uid instead —
owner=<mapped uid>, group=bitborg, mode=0775:

  • the container writes the sentinel as the directory's owner;
  • the host-side .path unit and kick script keep the group bit, which covers every operation they
    perform — PathExists=, stat -c %Y, and rm (write on the directory, not the file). Nothing
    reads the sentinel's contents.

The owner is derived from /etc/subuid (container uid N → subuid_start + N - 1) rather than
hardcoded, with an assert so a missing range fails loudly instead of silently handing the directory
to uid 999.

web_container_uid is a new default documenting the coupling to USER node in bitborg-web's
Containerfile — if the image's user changes, the trigger stops firing and falls back to the timer.

Verification

Ownership scheme tested end to end on the host before this was written — container write, host
stat/rm, and re-write after removal all pass, with the directory at 166535:2000 drwxrwxr-x:

uid_map:  0 → 2000 (len 1)    1 → 165536 (len 65536)
container writes sentinel as 166535 → WRITE-OK
bitborg:  PathExists YES · stat OK · kick stamp OK · rm OK · re-write OK
create:   127 ms

The /etc/subuid parsing and uid arithmetic are unit-tested offline for both the happy path
(165536 → owner 166535) and the missing-range path (yields 0, so the assert fires).
ansible-playbook --syntax-check passes and ansible-lint is clean at the production profile.

Not yet verified, and it cannot be until this is applied: cold-cache create time in prod. A
re-test against the same image digest will look fast whether or not this worked, because the cache is
already warm. The real proof is the first genuinely new digest deployed after the apply, measured from
the podman event m=+ offset — not from blackbox probe samples, whose 30 s spacing is what got this
misattributed in the first place.

What was tried and rejected

Idmapped bind mounts (Volume=…:rw,idmap) were the first choice — they would have kept keep-id's
ownership guarantee while leaving the rootfs alone. They are unavailable here:

Error: crun: mount_setattr `/mnt/rt`: Operation not permitted: OCI permission denied

and the relative @ form fails earlier still, on mapping resolution. Both are consistent with the
same Supports shifting: false. This also means the change carries no new Quadlet quoting risk — it
removes a line rather than adding exotic Volume= syntax, so the HealthCmd quote-corruption hazard
documented in this unit does not apply.

Notes

  • Trade-off: none in behaviour. ADR 0037's instant trigger is preserved; the 5-min timer remains
    the backstop as designed.
  • Security: a small improvement. Without keep-id the container's uid maps to an unprivileged
    subuid with no host presence, rather than to gitborg — better least privilege (principle 4).
  • The reconciler keeps keep-id. Its image is small and it pays ~0.1 s. The directive is not wrong
    in general, only on our largest image.
  • Transition: a sentinel left over from the old ownership is self-clearing — the .path unit fires
    and the kick script (which owns the directory) removes it, so no migration task is needed.
  • The remaining 10.1 s is an app-side defect with no SIGTERM handler, filed separately against
    bitborg-web. After both, expected deploy unavailability is ~3 s.
  • #249 and #250 tuned Caddy's retry window against an assumed "few seconds". #250's 2 s window is
    honest about not bridging a deploy, and becomes an adequate mitigation once creation is sub-second.
Fixes #253. Every `bitborg-web` deploy has taken the portal down for **83–139 s** since 2026-07-28, up from ~13 s before. This removes ~109 s of that. ## Cause `UserNS=keep-id:uid=1000,gid=1000`, added to the web Quadlet by #238 (ADR 0037 Phase 1) so the sentinel the container writes would be owned by `gitborg`. The overlay store on this host reports `Supports shifting: false`, so a non-identity user-namespace mapping can only be satisfied by **physically chowning the entire rootfs** — 849 MB across 11 layers. That result is cached per `(image digest, mapping)`, and every deploy ships a new digest, so **every deploy paid it cold.** Measured on the host, same image: | Variant | Create time | | --- | --- | | default userns, no `keep-id` | **129 ms** | | `keep-id`, cold mapping | **108 786 ms** | | `keep-id`, same mapping cached | 202 ms | | `keep-id`, prod's mapping (cached) | 149 ms | 108.8 s reproduces the ~103 s seen in prod. The cache is why this hid so well — the first attempt to reproduce it used prod's own mapping and came back at 113 ms, which looked like a refutation. Forcing a mapping never used before (`uid=1001`) exposes it. The wider deploy breakdown, from podman's journald event offsets: | Phase | Duration | | --- | --- | | `podman auto-update` pull | 0.7 s | | SIGTERM ignored → SIGKILL | 10.1 s | | **`podman create`** | **103.2 s** | | node boot + migrations | 2.3 s | | **Total** | **~117 s** | Worth stating plainly because it contradicts the assumption made while triaging #249: **migrations are 2 % of the outage.** In-container startup is 2.8 s. ## Fix Drop `keep-id` and own the sentinel **directory** as the container's *mapped* uid instead — `owner=<mapped uid>, group=bitborg, mode=0775`: - the container writes the sentinel as the directory's **owner**; - the host-side `.path` unit and kick script keep the **group** bit, which covers every operation they perform — `PathExists=`, `stat -c %Y`, and `rm` (write on the *directory*, not the file). Nothing reads the sentinel's contents. The owner is derived from `/etc/subuid` (container uid N → `subuid_start + N - 1`) rather than hardcoded, with an `assert` so a missing range fails loudly instead of silently handing the directory to uid 999. `web_container_uid` is a new default documenting the coupling to `USER node` in bitborg-web's Containerfile — if the image's user changes, the trigger stops firing and falls back to the timer. ## Verification Ownership scheme tested end to end on the host before this was written — container write, host `stat`/`rm`, and re-write after removal all pass, with the directory at `166535:2000 drwxrwxr-x`: ``` uid_map: 0 → 2000 (len 1) 1 → 165536 (len 65536) container writes sentinel as 166535 → WRITE-OK bitborg: PathExists YES · stat OK · kick stamp OK · rm OK · re-write OK create: 127 ms ``` The `/etc/subuid` parsing and uid arithmetic are unit-tested offline for both the happy path (`165536` → owner `166535`) and the missing-range path (yields 0, so the assert fires). `ansible-playbook --syntax-check` passes and `ansible-lint` is clean at the `production` profile. **Not yet verified, and it cannot be until this is applied:** cold-cache create time in prod. A re-test against the same image digest will look fast whether or not this worked, because the cache is already warm. The real proof is the first genuinely new digest deployed after the apply, measured from the podman event `m=+` offset — not from blackbox probe samples, whose 30 s spacing is what got this misattributed in the first place. ## What was tried and rejected Idmapped bind mounts (`Volume=…:rw,idmap`) were the first choice — they would have kept `keep-id`'s ownership guarantee while leaving the rootfs alone. They are unavailable here: ``` Error: crun: mount_setattr `/mnt/rt`: Operation not permitted: OCI permission denied ``` and the relative `@` form fails earlier still, on mapping resolution. Both are consistent with the same `Supports shifting: false`. This also means the change carries no new Quadlet quoting risk — it removes a line rather than adding exotic `Volume=` syntax, so the `HealthCmd` quote-corruption hazard documented in this unit does not apply. ## Notes - **Trade-off:** none in behaviour. ADR 0037's instant trigger is preserved; the 5-min timer remains the backstop as designed. - **Security:** a small improvement. Without `keep-id` the container's uid maps to an unprivileged subuid with no host presence, rather than to `gitborg` — better least privilege (principle 4). - **The reconciler keeps `keep-id`.** Its image is small and it pays ~0.1 s. The directive is not wrong in general, only on our largest image. - **Transition:** a sentinel left over from the old ownership is self-clearing — the `.path` unit fires and the kick script (which owns the directory) removes it, so no migration task is needed. - The remaining 10.1 s is an app-side defect with no `SIGTERM` handler, filed separately against bitborg-web. After both, expected deploy unavailability is ~3 s. - #249 and #250 tuned Caddy's retry window against an assumed "few seconds". #250's 2 s window is honest about not bridging a deploy, and becomes an adequate mitigation once creation is sub-second.
supernaut lade till 1 incheckning 2026-07-30 12:42:52 +00:00
fix(web): drop UserNS=keep-id — it cost ~109s of downtime per deploy (#253)
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m27s
7a02e29e21
Every gitborg-web deploy took the portal down for 83-139s, up from ~13s before
2026-07-28. `UserNS=keep-id` (added by #238 for the ADR 0037 sentinel) was the
cause: the overlay store reports `Supports shifting: false`, so a non-identity
userns mapping can only be satisfied by physically chowning the whole rootfs —
849 MB across 11 layers. The result is cached per (image digest, mapping), and
every deploy ships a new digest, so every deploy paid it cold.

Measured on the host, same image:

  default userns, no keep-id          129 ms
  keep-id, cold mapping           108 786 ms
  keep-id, same mapping cached        202 ms
  keep-id, prod's mapping (cached)    149 ms

The cache is why this hid so well: any test reusing the running image sees
~150 ms. Reproducing it needs a mapping never used before (uid=1001).

Instead of remapping the rootfs, the sentinel DIRECTORY is now owned by the
container's mapped uid with group=gitborg and mode 0775. The container writes as
owner; the host-side .path unit and kick script keep the group bit, which covers
every operation they perform — PathExists, `stat -c %Y`, `rm` (write on the
directory, not the file). Nothing reads the sentinel's contents. Verified
end-to-end on the host: container write, host stat/rm, and re-write after removal.

The owner is derived from /etc/subuid rather than hardcoded, with an assert so a
missing range fails loudly instead of silently handing the directory to uid 999.

Idmapped bind mounts were the first choice and are not available here — crun
rejects mount_setattr for rootless containers, consistent with the same
`Supports shifting: false`.

Side benefit: without keep-id the container's uid maps to an unprivileged subuid
with no host presence, rather than to gitborg — better least privilege
(principle 4).

The reconciler keeps keep-id: its image is small and it pays ~0.1 s.
supernaut sammanfogade incheckning a2a1dace75 till main 2026-07-30 14:22:22 +00:00
supernaut tog bort grenen fix/web-keepid-deploy-outage 2026-07-30 14:22:22 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!254
Ingen beskrivning angiven.