fix(caddy): raise lb_try_duration to 5s — a deploy is now bridgeable #261
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!261
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/caddy-bridge-deploy"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Closes #259. Raises
caddy_web_lb_try_durationfrom 2s → 5s, so a web deploy costs an arrivingrequest a short wait instead of a 502.
This value has been wrong twice, in opposite directions
Both times because it was sized from blackbox probe samples rather than a measurement.
#250's reasoning followed correctly from what it had, but ~60–90s came from three consecutive
30s-spaced probes failing — which can neither resolve nor attribute a gap. The real figure was ~117s,
and it was not "startup runs migrations" (those were ~2s). It was
podman create, becauseUserNS=keep-idforced a full rootfs chown on every new image digest (#253, fixed in #254).The measurement this is sized from
Deploy at 2026-07-30 16:57:51Z — the first with both #254 and bitborg-web #109 live, priced from
podman's monotonic journald event offsets:
5s is ~3× that — bridging with margin rather than exactly, because startup is dominated by
migrations and a schema change can lengthen them.
The trade-off, stated plainly
A genuinely down upstream now hangs a request 5s before erroring instead of 2s. That is worth
paying only because the window actually succeeds for the case it exists to cover — which was not true
at 117s.
EndpointDown(probe_success == 0 for: 5m) remains the signal for a real outage; this onlychanges how a visitor experiences its first few seconds.
Retry semantics need no change and were already right: a failed connection is retried regardless of
method, which is exactly the restart case, so
POST /api/signupis covered without risking a doublesubmission.
Why the comments are rewritten too
The old text in both the defaults and the Caddyfile asserted three things that are now false — 60–90s,
"startup migrations", and "no sane retry window covers it". That stale reasoning is precisely why nobody
revisited this, so leaving it while changing the number would be worse than leaving both alone.
One consequence worth recording, and now recorded in the Caddyfile: this makes externalising
bitborg-web's process-local captcha and rate-limit state NOT a prerequisite for hiding deploy
downtime. The previous reasoning had that state on the critical path, because the only way to remove a
minute-plus of downtime looked like two concurrent instances. It still isn't a rolling deploy — this
hides the restart from arriving requests, it does not run two versions.
How to verify — and how NOT to
Both files now carry this warning: do not verify with
probe_success. The blackbox interval missedeven a 12s deploy gap, so a 1.7s one will never be sampled, and "no dip" is the expected result whether
or not retrying works.
The real test is a continuous request loop across a deploy — confirming zero non-200 responses
through the restart. That is the only thing that can observe a 1.7s gap being bridged, and it is what
I'd run right after applying.
--syntax-checkpasses;ansible-lintclean at theproductionprofile.