caddy: raise lb_try_duration to bridge a deploy — #250's premise no longer holds #259

Stängd
öppnade 2026-07-30 16:38:26 +00:00 av supernaut · 1 kommentar
Ägare

#250 concluded that no sane Caddy retry window could bridge a web deploy, and scoped
lb_try_duration down to 2 s accordingly. That reasoning was sound for the numbers available at the
time — but those numbers were wrong, and the two fixes since have invalidated the conclusion.

What #250 assumed, and what is actually true

The Caddyfile currently says:

Explicitly NOT what this does: it does not bridge a deploy. The web container's restart window was
measured at ~60-90s (auto-update pull + teardown + startup migrations), so no sane retry window
covers it — a longer one would just make visitors wait a minute before failing.

Every factual claim in that paragraph is now false:

  • It was never "startup migrations". Migrations take ~1.8 s. The ~60–90 s was podman create,
    because UserNS=keep-id forced a full rootfs chown on every new image digest (#253).
  • ~60–90 s was itself an underestimate — the real figure was ~117 s. The 30 s-spaced blackbox
    probes could not resolve it.
  • It is now ~2.6 s.
Phase Before After #254 After bitborg-web #109
podman auto-update pull 0.7 s 0.7 s 0.7 s
SIGTERM → SIGKILL 10.1 s 10.07 s 0.19 s
podman create 103.2 s 0.072 s 0.072 s
container init + start 0.5 s 0.17 s 0.17 s
node boot + migrations 2.3 s 1.79 s 1.79 s
migrations → listening 0.5 s 0.37 s 0.37 s
Total ~117 s ~12.7 s ~2.6 s

#254 is applied and verified in production. bitborg-web #109 is measured end-to-end against the real
image as pid 1 (podman stop 190 ms, exit 0) but not yet deployed.

Proposal

Once #109 is deployed and the ~2.6 s figure is confirmed from the podman event offsets on a real
deploy:

  1. Raise lb_try_duration to roughly 5 s — about 2× the measured restart, so a deploy is
    bridged with margin rather than exactly.
  2. Rewrite the comment. It is load-bearing documentation that currently asserts three false things,
    and it is the reason nobody has revisited this. Replace the "no sane retry window covers it"
    paragraph with the measured breakdown and the actual reason the window works.

lb_try_interval 250ms is fine as-is.

The retry semantics need no change and are already correct for this: a failed connection is retried
regardless of method, which is exactly the restart case, so POST /api/signup is covered without
risking a double submission. A failure after the round-trip started is retried only for GET. The
Caddyfile already documents this and #250 got it right.

The trade-off to state plainly

A 5 s window means that when the upstream is genuinely down — not restarting — every request hangs 5 s
before erroring, instead of 2 s. That is the cost #250 was managing, and it is real. It is worth paying
now only because the window would actually succeed for the case it is there to cover, which was not
true at 117 s. EndpointDown (for: 5m) remains the signal for a real outage; this only changes how a
visitor experiences the first few seconds of one.

Why this matters beyond the 2.6 s

This is the cheaper half of a bigger reversal. bitborg-web's process-local captcha and rate-limit state
was on the critical path to zero-downtime deploys, because the only way to remove a
minute-plus of downtime looked like running two instances — which needs that state in Postgres first,
plus a standing additive-only constraint on every future schema migration.

At ~2.6 s a retry window closes the gap on its own. So this issue plus #109 may deliver zero visible
downtime without any of that, and the state-externalisation work reverts to being worth doing on its own
merits at normal priority.

Acceptance

  • bitborg-web #109 deployed; restart re-measured from the podman event m=+ offsets on a real deploy
  • lb_try_duration raised to cover it with margin
  • The stale ~60–90 s / "startup migrations" rationale replaced with the measured breakdown
  • Confirm a deploy produces no probe_success dip — the actual test of whether it bridges
#250 concluded that no sane Caddy retry window could bridge a web deploy, and scoped `lb_try_duration` down to 2 s accordingly. That reasoning was sound for the numbers available at the time — but those numbers were wrong, and the two fixes since have invalidated the conclusion. ## What #250 assumed, and what is actually true The Caddyfile currently says: > Explicitly NOT what this does: it does not bridge a deploy. The web container's restart window was > measured at ~60-90s (auto-update pull + teardown + startup migrations), so no sane retry window > covers it — a longer one would just make visitors wait a minute before failing. Every factual claim in that paragraph is now false: - **It was never "startup migrations".** Migrations take ~1.8 s. The ~60–90 s was `podman create`, because `UserNS=keep-id` forced a full rootfs chown on every new image digest (#253). - **~60–90 s was itself an underestimate** — the real figure was ~117 s. The 30 s-spaced blackbox probes could not resolve it. - **It is now ~2.6 s.** | Phase | Before | After #254 | After bitborg-web #109 | | --- | --- | --- | --- | | `podman auto-update` pull | 0.7 s | 0.7 s | 0.7 s | | SIGTERM → SIGKILL | 10.1 s | 10.07 s | **0.19 s** | | `podman create` | 103.2 s | **0.072 s** | 0.072 s | | container init + start | 0.5 s | 0.17 s | 0.17 s | | node boot + migrations | 2.3 s | 1.79 s | 1.79 s | | migrations → listening | 0.5 s | 0.37 s | 0.37 s | | **Total** | **~117 s** | **~12.7 s** | **~2.6 s** | #254 is applied and verified in production. bitborg-web #109 is measured end-to-end against the real image as pid 1 (`podman stop` 190 ms, exit 0) but not yet deployed. ## Proposal Once #109 is deployed and the ~2.6 s figure is confirmed from the podman event offsets on a real deploy: 1. **Raise `lb_try_duration`** to roughly **5 s** — about 2× the measured restart, so a deploy is bridged with margin rather than exactly. 2. **Rewrite the comment.** It is load-bearing documentation that currently asserts three false things, and it is the reason nobody has revisited this. Replace the "no sane retry window covers it" paragraph with the measured breakdown and the actual reason the window works. `lb_try_interval 250ms` is fine as-is. The retry semantics need no change and are already correct for this: a failed **connection** is retried regardless of method, which is exactly the restart case, so `POST /api/signup` is covered without risking a double submission. A failure *after* the round-trip started is retried only for GET. The Caddyfile already documents this and #250 got it right. ## The trade-off to state plainly A 5 s window means that when the upstream is genuinely down — not restarting — every request hangs 5 s before erroring, instead of 2 s. That is the cost #250 was managing, and it is real. It is worth paying now only because the window would actually *succeed* for the case it is there to cover, which was not true at 117 s. `EndpointDown` (`for: 5m`) remains the signal for a real outage; this only changes how a visitor experiences the first few seconds of one. ## Why this matters beyond the 2.6 s This is the cheaper half of a bigger reversal. bitborg-web's process-local captcha and rate-limit state was on the critical path to zero-downtime deploys, because the only way to remove a minute-plus of downtime looked like running two instances — which needs that state in Postgres first, plus a standing additive-only constraint on every future schema migration. At ~2.6 s a retry window closes the gap on its own. So this issue plus #109 may deliver zero *visible* downtime without any of that, and the state-externalisation work reverts to being worth doing on its own merits at normal priority. ## Acceptance - [ ] bitborg-web #109 deployed; restart re-measured from the podman event `m=+` offsets on a real deploy - [ ] `lb_try_duration` raised to cover it with margin - [ ] The stale ~60–90 s / "startup migrations" rationale replaced with the measured breakdown - [ ] Confirm a deploy produces no `probe_success` dip — the actual test of whether it bridges
Upphovsperson
Ägare

Measured. bitborg-web #109 is deployed and the restart is now 1.694 s — so this is unblocked, and the
number is smaller than the estimate in the issue body.

The deploy at 2026-07-30 16:57:51Z — first with both fixes live

Event Time Δ
Stopping bitborg-web.service 16:57:51.270 —
[shutdown] SIGTERM received — draining 16:57:51.343 73 ms
[shutdown] drained cleanly 16:57:51.344 0.6 ms
container died (m=+0.064) 16:57:51.366
container create (m=+0.052) 16:57:51.673
container start 16:57:51.869
[migrate] migrations applied 16:57:52.773 904 ms
Server listening 16:57:52.964

Total unavailability: 1.694 s (the issue body estimated ~2.6 s; migrations came in at 904 ms rather
than 1.79 s). No resorting to SIGKILL in the journal.

Two things this also settles:

  • container create at m=+0.052 on a genuinely NEW image digest. Every prior measurement after
    #254 reused the existing digest, so the cold-cache path was untested. It holds: the chown is gone, not
    merely warmed.
  • The SIGTERM stop went from 10.08 s to 0.064 s — visible in the same log pair, because the deploy at
    16:42 stopped a pre-#109 container (died … m=+10.08) and the one at 16:57 stopped a post-#109 one
    (died … m=+0.064).

Revised recommendation

lb_try_duration 5s gives roughly 3× margin over the measured 1.694 s. Still comfortably short
enough that a genuinely-down upstream fails fast rather than hanging a visitor.

A caveat on how to verify this

min_over_time(probe_success[20m]) is 1 across both deploys — no dip. That is encouraging but
not proof the deploy is invisible: the blackbox probe interval also missed the 16:42 deploy's ~12 s
gap completely. A 1.7 s gap will essentially never be sampled.

So the acceptance item "confirm a deploy produces no probe_success dip" is too weak to be the test —
absence of a dip is the expected result whether or not the retry window works. Replace it with something
that actually exercises the window:

  • During a deploy, issue continuous requests (e.g. a 1-per-100ms loop against https://www.gitborg.se)
    and confirm zero non-200 responses across the restart. That is the only way to observe a 1.7 s
    gap being bridged.

Updated acceptance:

  • bitborg-web #109 deployed; restart re-measured from the podman event m=+ offsets → 1.694 s
  • lb_try_duration raised to 5 s
  • The stale ~60–90 s / "startup migrations" rationale in the Caddyfile replaced with the measured breakdown
  • Continuous-request loop across a deploy shows no non-200 (replaces the probe-dip check, which cannot resolve 1.7 s)
**Measured. bitborg-web #109 is deployed and the restart is now 1.694 s — so this is unblocked, and the number is smaller than the estimate in the issue body.** ## The deploy at 2026-07-30 16:57:51Z — first with both fixes live | Event | Time | Δ | | --- | --- | --- | | `Stopping bitborg-web.service` | 16:57:51.270 | — | | `[shutdown] SIGTERM received — draining` | 16:57:51.343 | 73 ms | | `[shutdown] drained cleanly` | 16:57:51.344 | **0.6 ms** | | `container died` (`m=+0.064`) | 16:57:51.366 | | | `container create` (**`m=+0.052`**) | 16:57:51.673 | | | `container start` | 16:57:51.869 | | | `[migrate] migrations applied` | 16:57:52.773 | 904 ms | | **`Server listening`** | **16:57:52.964** | | **Total unavailability: 1.694 s** (the issue body estimated ~2.6 s; migrations came in at 904 ms rather than 1.79 s). No `resorting to SIGKILL` in the journal. Two things this also settles: - **`container create` at `m=+0.052` on a genuinely NEW image digest.** Every prior measurement after #254 reused the existing digest, so the cold-cache path was untested. It holds: the chown is gone, not merely warmed. - **The SIGTERM stop went from 10.08 s to 0.064 s** — visible in the same log pair, because the deploy at 16:42 stopped a pre-#109 container (`died … m=+10.08`) and the one at 16:57 stopped a post-#109 one (`died … m=+0.064`). ## Revised recommendation `lb_try_duration 5s` gives roughly **3×** margin over the measured 1.694 s. Still comfortably short enough that a genuinely-down upstream fails fast rather than hanging a visitor. ## A caveat on how to verify this `min_over_time(probe_success[20m])` is **1** across both deploys — no dip. That is encouraging but **not** proof the deploy is invisible: the blackbox probe interval also missed the 16:42 deploy's ~12 s gap completely. A 1.7 s gap will essentially never be sampled. So the acceptance item "confirm a deploy produces no `probe_success` dip" is too weak to be the test — absence of a dip is the expected result whether or not the retry window works. Replace it with something that actually exercises the window: - [ ] During a deploy, issue continuous requests (e.g. a 1-per-100ms loop against `https://www.gitborg.se`) and confirm **zero** non-200 responses across the restart. That is the only way to observe a 1.7 s gap being bridged. Updated acceptance: - [x] bitborg-web #109 deployed; restart re-measured from the podman event `m=+` offsets → **1.694 s** - [ ] `lb_try_duration` raised to 5 s - [ ] The stale ~60–90 s / "startup migrations" rationale in the Caddyfile replaced with the measured breakdown - [ ] Continuous-request loop across a deploy shows no non-200 (replaces the probe-dip check, which cannot resolve 1.7 s)
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#259
Ingen beskrivning angiven.