fix(caddy): ride out the web container restart instead of 502ing #249
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!249
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/web-deploy-restart-502"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
www.gitborg.sereturned 502 twice tonight, at 23:32 and 23:42. Both were deploy restarts, and bothself-healed — but a visitor loading the page inside either window saw an error.
What happened
Under ADR 0019 a merge builds and pushes an image, the host pulls it with
podman auto-update, andsystemd restarts the unit. bitborg-web is a single container, so for those few seconds there is
no upstream at all and Caddy answers 502.
Three web deploys ran tonight (#104, #98, #102).
probe_successcaught two of the resulting gaps:11Only
www, one 120s sample each, recovering unaided. Nothing was broken — no bad image, no crashloop.
EndpointDown(probe_success == 0 for: 5m) correctly did not fire: a two-minute deploy blipshould not page anyone. That threshold is right and is left alone.
The fix
Two directives on the web upstream, so a request arriving mid-restart waits for the new container
instead of erroring:
No
lb_retry_matchis needed, and adding one would have been a mistake. Caddy's defaults arealready the safe semantics: a failed connection is always retried regardless of method — correct,
because the request never reached the app, so replaying it cannot double-apply anything — while a
failure after the round-trip began is retried only for
GET. A restart is the connection-failedcase, so
POST /api/signupis covered without any risk of a double submission.Verified
caddy validateon the changed directives →Valid configuration(Caddy v2.11.4). Worth notingbecause my first draft declared
@idempotent method GET HEADinside thereverse_proxyblock,which is invalid — request matchers cannot be defined there — and would have broken the Caddyfile
and taken down all four sites. Validating caught it.
format:check,mdlintandansible-lintall pass.What this is NOT
This hides the restart from arriving requests; it does not run two versions, and it is not a rolling
deploy. That distinction is load-bearing: two concurrent bitborg-web instances would break its
process-local captcha and rate-limit state — both
src/lib/captcha.tsandsrc/lib/rate-limit.tsstate outright that the portal "runs as a SINGLE Node process … so process-local state is
authoritative". Blue-green would trade an occasional 502 for occasional broken sign-ups and halved
rate limiting, which is a worse deal. True zero-downtime is therefore sequenced behind moving that
state to Postgres — filed separately.
Deployment
This changes the Caddy template, so it needs an Ansible apply. Not applied — please run it
through the
infra-applyflow (validate →--check→ review → apply) when you want it live. Caddyreloads config gracefully, so the apply itself should not interrupt traffic.