caddy: raise lb_try_duration to bridge a deploy — #250's premise no longer holds #259
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#259
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
#250 concluded that no sane Caddy retry window could bridge a web deploy, and scoped
lb_try_durationdown to 2 s accordingly. That reasoning was sound for the numbers available at thetime — but those numbers were wrong, and the two fixes since have invalidated the conclusion.
What #250 assumed, and what is actually true
The Caddyfile currently says:
Every factual claim in that paragraph is now false:
podman create,because
UserNS=keep-idforced a full rootfs chown on every new image digest (#253).probes could not resolve it.
podman auto-updatepullpodman create#254 is applied and verified in production. bitborg-web #109 is measured end-to-end against the real
image as pid 1 (
podman stop190 ms, exit 0) but not yet deployed.Proposal
Once #109 is deployed and the ~2.6 s figure is confirmed from the podman event offsets on a real
deploy:
lb_try_durationto roughly 5 s — about 2× the measured restart, so a deploy isbridged with margin rather than exactly.
and it is the reason nobody has revisited this. Replace the "no sane retry window covers it"
paragraph with the measured breakdown and the actual reason the window works.
lb_try_interval 250msis fine as-is.The retry semantics need no change and are already correct for this: a failed connection is retried
regardless of method, which is exactly the restart case, so
POST /api/signupis covered withoutrisking a double submission. A failure after the round-trip started is retried only for GET. The
Caddyfile already documents this and #250 got it right.
The trade-off to state plainly
A 5 s window means that when the upstream is genuinely down — not restarting — every request hangs 5 s
before erroring, instead of 2 s. That is the cost #250 was managing, and it is real. It is worth paying
now only because the window would actually succeed for the case it is there to cover, which was not
true at 117 s.
EndpointDown(for: 5m) remains the signal for a real outage; this only changes how avisitor experiences the first few seconds of one.
Why this matters beyond the 2.6 s
This is the cheaper half of a bigger reversal. bitborg-web's process-local captcha and rate-limit state
was on the critical path to zero-downtime deploys, because the only way to remove a
minute-plus of downtime looked like running two instances — which needs that state in Postgres first,
plus a standing additive-only constraint on every future schema migration.
At ~2.6 s a retry window closes the gap on its own. So this issue plus #109 may deliver zero visible
downtime without any of that, and the state-externalisation work reverts to being worth doing on its own
merits at normal priority.
Acceptance
m=+offsets on a real deploylb_try_durationraised to cover it with marginprobe_successdip — the actual test of whether it bridgesMeasured. bitborg-web #109 is deployed and the restart is now 1.694 s — so this is unblocked, and the
number is smaller than the estimate in the issue body.
The deploy at 2026-07-30 16:57:51Z — first with both fixes live
Stopping bitborg-web.service[shutdown] SIGTERM received — draining[shutdown] drained cleanlycontainer died(m=+0.064)container create(m=+0.052)container start[migrate] migrations appliedServer listeningTotal unavailability: 1.694 s (the issue body estimated ~2.6 s; migrations came in at 904 ms rather
than 1.79 s). No
resorting to SIGKILLin the journal.Two things this also settles:
container createatm=+0.052on a genuinely NEW image digest. Every prior measurement after#254 reused the existing digest, so the cold-cache path was untested. It holds: the chown is gone, not
merely warmed.
16:42 stopped a pre-#109 container (
died … m=+10.08) and the one at 16:57 stopped a post-#109 one(
died … m=+0.064).Revised recommendation
lb_try_duration 5sgives roughly 3× margin over the measured 1.694 s. Still comfortably shortenough that a genuinely-down upstream fails fast rather than hanging a visitor.
A caveat on how to verify this
min_over_time(probe_success[20m])is 1 across both deploys — no dip. That is encouraging butnot proof the deploy is invisible: the blackbox probe interval also missed the 16:42 deploy's ~12 s
gap completely. A 1.7 s gap will essentially never be sampled.
So the acceptance item "confirm a deploy produces no
probe_successdip" is too weak to be the test —absence of a dip is the expected result whether or not the retry window works. Replace it with something
that actually exercises the window:
https://www.gitborg.se)and confirm zero non-200 responses across the restart. That is the only way to observe a 1.7 s
gap being bridged.
Updated acceptance:
m=+offsets → 1.694 slb_try_durationraised to 5 s