fix(caddy): scope the web retry window to brief blips, not deploys #250

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/scope-caddy-retry-to-blips in i main 2026-07-29 23:07:24 +00:00
Ägare

Corrects #249. I sized that retry window against an assumed restart duration rather than a
measured one, and the assumption was wrong.

What the next deploy measured

/healthz (PR bitborg-web #105) deployed at ~00:27 and gave a clean natural experiment:

Metric Before During restart After
probe_duration_seconds 0.02s 9.50s ×3 0.02s
probe_http_status_code 200 0 ×3 200
probe_success 1 0 1

Three consecutive probes, 30s apart, all failed. So the restart window is ~60-90 seconds — not
"a few seconds". Status code 0 means no HTTP response at all: the probe waited 9.5s and got nothing.

The window is that long because podman auto-update pulls the image, tears the container down, and
startup runs scripts/migrate.mjs before Astro begins listening.

Why #249 as merged was a regression

  • Before: an immediate 502 for ~60-90s.
  • With lb_try_duration 10s: requests hang ~10s and then fail, for ~60-90s.

Ten seconds spans about a seventh of the gap, so the retry never completes — it just delays the
error. A visitor who waits ten seconds to be told it failed is worse served than one told
immediately. My claim on #249 that "visitors get a slightly slower request instead of an error" was
wrong.

Why not simply raise it

Covering 60-90s would mean visitors hanging for a minute, browsers timing out anyway, and — worse —
a genuine upstream outage hanging instead of failing fast, which also delays every alert keyed on
response failure. Deploy downtime is an architecture problem, not something a retry window can hide.

What this PR does

Scopes the window to what a retry legitimately fixes: 2s, for a sub-second blip — a Caddy config
reload, a momentary network hiccup, a single dropped connection. lb_try_interval 250ms gives ~8
attempts inside it.

The measurements and the reasoning are written into the caddy role's defaults, so the next person to
consider raising this number finds out why it is small before doing so.

Follow-on

This makes blue-green (internal #68) the actual fix rather than a nice-to-have — with a health-checked
second instance there is no gap to bridge and window-sizing becomes moot. It also sharpens #66: 60-90s
of downtime per deploy is 60-90s in which a broken image is indistinguishable from a slow one.

Both issues are being updated with these numbers, plus a note to measure restart duration explicitly
— every decision here depends on it and nobody had it until tonight.

Verified: caddy validate → Valid configuration; format:check, mdlint, ansible-lint pass.

Refs #66, #68

Corrects #249. I sized that retry window against an **assumed** restart duration rather than a measured one, and the assumption was wrong. ## What the next deploy measured `/healthz` (PR bitborg-web #105) deployed at ~00:27 and gave a clean natural experiment: | Metric | Before | During restart | After | | --- | --- | --- | --- | | `probe_duration_seconds` | 0.02s | **9.50s** ×3 | 0.02s | | `probe_http_status_code` | 200 | **0** ×3 | 200 | | `probe_success` | 1 | **0** | 1 | Three consecutive probes, 30s apart, all failed. So the restart window is **~60-90 seconds** — not "a few seconds". Status code `0` means no HTTP response at all: the probe waited 9.5s and got nothing. The window is that long because `podman auto-update` pulls the image, tears the container down, and startup runs `scripts/migrate.mjs` before Astro begins listening. ## Why #249 as merged was a regression - **Before:** an immediate 502 for ~60-90s. - **With `lb_try_duration 10s`:** requests hang ~10s and *then* fail, for ~60-90s. Ten seconds spans about a seventh of the gap, so the retry never completes — it just delays the error. A visitor who waits ten seconds to be told it failed is worse served than one told immediately. My claim on #249 that "visitors get a slightly slower request instead of an error" was wrong. ## Why not simply raise it Covering 60-90s would mean visitors hanging for a minute, browsers timing out anyway, and — worse — a *genuine* upstream outage hanging instead of failing fast, which also delays every alert keyed on response failure. Deploy downtime is an architecture problem, not something a retry window can hide. ## What this PR does Scopes the window to what a retry legitimately fixes: **2s**, for a sub-second blip — a Caddy config reload, a momentary network hiccup, a single dropped connection. `lb_try_interval 250ms` gives ~8 attempts inside it. The measurements and the reasoning are written into the caddy role's defaults, so the next person to consider raising this number finds out why it is small before doing so. ## Follow-on This makes blue-green (internal #68) the actual fix rather than a nice-to-have — with a health-checked second instance there is no gap to bridge and window-sizing becomes moot. It also sharpens #66: 60-90s of downtime per deploy is 60-90s in which a *broken* image is indistinguishable from a slow one. Both issues are being updated with these numbers, plus a note to measure restart duration explicitly — every decision here depends on it and nobody had it until tonight. Verified: `caddy validate` → `Valid configuration`; `format:check`, `mdlint`, `ansible-lint` pass. Refs #66, #68
supernaut lade till 1 incheckning 2026-07-29 22:47:43 +00:00
fix(caddy): scope the web retry window to brief blips, not deploys
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m27s
e450f0051e
supernaut sammanfogade incheckning da14a08157 till main 2026-07-29 23:07:24 +00:00
supernaut tog bort grenen fix/scope-caddy-retry-to-blips 2026-07-29 23:07:24 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!250
Ingen beskrivning angiven.