fix(caddy): scope the web retry window to brief blips, not deploys #250
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!250
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/scope-caddy-retry-to-blips"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Corrects #249. I sized that retry window against an assumed restart duration rather than a
measured one, and the assumption was wrong.
What the next deploy measured
/healthz(PR bitborg-web #105) deployed at ~00:27 and gave a clean natural experiment:probe_duration_secondsprobe_http_status_codeprobe_successThree consecutive probes, 30s apart, all failed. So the restart window is ~60-90 seconds — not
"a few seconds". Status code
0means no HTTP response at all: the probe waited 9.5s and got nothing.The window is that long because
podman auto-updatepulls the image, tears the container down, andstartup runs
scripts/migrate.mjsbefore Astro begins listening.Why #249 as merged was a regression
lb_try_duration 10s: requests hang ~10s and then fail, for ~60-90s.Ten seconds spans about a seventh of the gap, so the retry never completes — it just delays the
error. A visitor who waits ten seconds to be told it failed is worse served than one told
immediately. My claim on #249 that "visitors get a slightly slower request instead of an error" was
wrong.
Why not simply raise it
Covering 60-90s would mean visitors hanging for a minute, browsers timing out anyway, and — worse —
a genuine upstream outage hanging instead of failing fast, which also delays every alert keyed on
response failure. Deploy downtime is an architecture problem, not something a retry window can hide.
What this PR does
Scopes the window to what a retry legitimately fixes: 2s, for a sub-second blip — a Caddy config
reload, a momentary network hiccup, a single dropped connection.
lb_try_interval 250msgives ~8attempts inside it.
The measurements and the reasoning are written into the caddy role's defaults, so the next person to
consider raising this number finds out why it is small before doing so.
Follow-on
This makes blue-green (internal #68) the actual fix rather than a nice-to-have — with a health-checked
second instance there is no gap to bridge and window-sizing becomes moot. It also sharpens #66: 60-90s
of downtime per deploy is 60-90s in which a broken image is indistinguishable from a slow one.
Both issues are being updated with these numbers, plus a note to measure restart duration explicitly
— every decision here depends on it and nobody had it until tonight.
Verified:
caddy validate→Valid configuration;format:check,mdlint,ansible-lintpass.Refs #66, #68