fix(caddy): ride out the web container restart instead of 502ing #249

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/web-deploy-restart-502 in i main 2026-07-29 22:02:29 +00:00
Ägare

www.gitborg.se returned 502 twice tonight, at 23:32 and 23:42. Both were deploy restarts, and both
self-healed — but a visitor loading the page inside either window saw an error.

What happened

Under ADR 0019 a merge builds and pushes an image, the host pulls it with podman auto-update, and
systemd restarts the unit. bitborg-web is a single container, so for those few seconds there is
no upstream at all and Caddy answers 502.

Three web deploys ran tonight (#104, #98, #102). probe_success caught two of the resulting gaps:

Time www git / auth / grafana / ntfy / stats
23:32 0 all 1
23:42 0 all 1

Only www, one 120s sample each, recovering unaided. Nothing was broken — no bad image, no crash
loop.

EndpointDown (probe_success == 0 for: 5m) correctly did not fire: a two-minute deploy blip
should not page anyone. That threshold is right and is left alone.

The fix

Two directives on the web upstream, so a request arriving mid-restart waits for the new container
instead of erroring:

lb_try_duration 10s
lb_try_interval 250ms

No lb_retry_match is needed, and adding one would have been a mistake. Caddy's defaults are
already the safe semantics: a failed connection is always retried regardless of method — correct,
because the request never reached the app, so replaying it cannot double-apply anything — while a
failure after the round-trip began is retried only for GET. A restart is the connection-failed
case, so POST /api/signup is covered without any risk of a double submission.

Verified

caddy validate on the changed directives → Valid configuration (Caddy v2.11.4). Worth noting
because my first draft declared @idempotent method GET HEAD inside the reverse_proxy block,
which is invalid — request matchers cannot be defined there — and would have broken the Caddyfile
and taken down all four sites. Validating caught it.

format:check, mdlint and ansible-lint all pass.

What this is NOT

This hides the restart from arriving requests; it does not run two versions, and it is not a rolling
deploy. That distinction is load-bearing: two concurrent bitborg-web instances would break its
process-local captcha and rate-limit state — both src/lib/captcha.ts and src/lib/rate-limit.ts
state outright that the portal "runs as a SINGLE Node process … so process-local state is
authoritative". Blue-green would trade an occasional 502 for occasional broken sign-ups and halved
rate limiting, which is a worse deal. True zero-downtime is therefore sequenced behind moving that
state to Postgres — filed separately.

Deployment

This changes the Caddy template, so it needs an Ansible apply. Not applied — please run it
through the infra-apply flow (validate → --check → review → apply) when you want it live. Caddy
reloads config gracefully, so the apply itself should not interrupt traffic.

`www.gitborg.se` returned 502 twice tonight, at 23:32 and 23:42. Both were deploy restarts, and both self-healed — but a visitor loading the page inside either window saw an error. ## What happened Under ADR 0019 a merge builds and pushes an image, the host pulls it with `podman auto-update`, and systemd restarts the unit. bitborg-web is a **single** container, so for those few seconds there is no upstream at all and Caddy answers 502. Three web deploys ran tonight (#104, #98, #102). `probe_success` caught two of the resulting gaps: | Time | www | git / auth / grafana / ntfy / stats | | --- | --- | --- | | 23:32 | **0** | all `1` | | 23:42 | **0** | all `1` | Only `www`, one 120s sample each, recovering unaided. Nothing was broken — no bad image, no crash loop. `EndpointDown` (`probe_success == 0 for: 5m`) correctly did **not** fire: a two-minute deploy blip should not page anyone. That threshold is right and is left alone. ## The fix Two directives on the web upstream, so a request arriving mid-restart waits for the new container instead of erroring: ``` lb_try_duration 10s lb_try_interval 250ms ``` **No `lb_retry_match` is needed, and adding one would have been a mistake.** Caddy's defaults are already the safe semantics: a failed *connection* is always retried regardless of method — correct, because the request never reached the app, so replaying it cannot double-apply anything — while a failure *after* the round-trip began is retried only for `GET`. A restart is the connection-failed case, so `POST /api/signup` is covered without any risk of a double submission. ## Verified `caddy validate` on the changed directives → **`Valid configuration`** (Caddy v2.11.4). Worth noting because my first draft declared `@idempotent method GET HEAD` *inside* the `reverse_proxy` block, which is invalid — request matchers cannot be defined there — and would have broken the Caddyfile and taken down all four sites. Validating caught it. `format:check`, `mdlint` and `ansible-lint` all pass. ## What this is NOT This hides the restart from arriving requests; it does not run two versions, and it is not a rolling deploy. That distinction is load-bearing: two concurrent bitborg-web instances would break its **process-local** captcha and rate-limit state — both `src/lib/captcha.ts` and `src/lib/rate-limit.ts` state outright that the portal "runs as a SINGLE Node process … so process-local state is authoritative". Blue-green would trade an occasional 502 for occasional *broken sign-ups* and halved rate limiting, which is a worse deal. True zero-downtime is therefore sequenced behind moving that state to Postgres — filed separately. ## Deployment This changes the Caddy template, so it needs an Ansible apply. **Not applied** — please run it through the `infra-apply` flow (validate → `--check` → review → apply) when you want it live. Caddy reloads config gracefully, so the apply itself should not interrupt traffic.
supernaut lade till 1 incheckning 2026-07-29 21:59:05 +00:00
fix(caddy): ride out the web container restart instead of 502ing
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m30s
d86b106380
supernaut sammanfogade incheckning 95a83de589 till main 2026-07-29 22:02:29 +00:00
supernaut tog bort grenen fix/web-deploy-restart-502 2026-07-29 22:02:29 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!249
Ingen beskrivning angiven.