fix(caddy): raise lb_try_duration to 5s — a deploy is now bridgeable #261

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/caddy-bridge-deploy in i main 2026-07-30 17:21:42 +00:00
Ägare

Closes #259. Raises caddy_web_lb_try_duration from 2s → 5s, so a web deploy costs an arriving
request a short wait instead of a 502.

This value has been wrong twice, in opposite directions

Both times because it was sized from blackbox probe samples rather than a measurement.

value assumption reality
#249 10s restart takes "a few seconds" ~117s
#250 2s restart is ~60–90s, unbridgeable ~117s, and the cause was misattributed
this 5s restart is 1.694s, measured —

#250's reasoning followed correctly from what it had, but ~60–90s came from three consecutive
30s-spaced probes failing — which can neither resolve nor attribute a gap. The real figure was ~117s,
and it was not "startup runs migrations" (those were ~2s). It was podman create, because
UserNS=keep-id forced a full rootfs chown on every new image digest (#253, fixed in #254).

The measurement this is sized from

Deploy at 2026-07-30 16:57:51Z — the first with both #254 and bitborg-web #109 live, priced from
podman's monotonic journald event offsets:

Stopping bitborg-web.service     16:57:51.270
[shutdown] drained cleanly       16:57:51.344   (0.6ms drain; container died m=+0.064)
container create                 16:57:51.673   (m=+0.052 — on a NEW image digest)
[migrate] migrations applied     16:57:52.773   (904ms)
Server listening                 16:57:52.964
─────────────────────────────────────────────
total unavailability                  1.694s

5s is ~3× that — bridging with margin rather than exactly, because startup is dominated by
migrations and a schema change can lengthen them.

The trade-off, stated plainly

A genuinely down upstream now hangs a request 5s before erroring instead of 2s. That is worth
paying only because the window actually succeeds for the case it exists to cover — which was not true
at 117s. EndpointDown (probe_success == 0 for: 5m) remains the signal for a real outage; this only
changes how a visitor experiences its first few seconds.

Retry semantics need no change and were already right: a failed connection is retried regardless of
method, which is exactly the restart case, so POST /api/signup is covered without risking a double
submission.

Why the comments are rewritten too

The old text in both the defaults and the Caddyfile asserted three things that are now false — 60–90s,
"startup migrations", and "no sane retry window covers it". That stale reasoning is precisely why nobody
revisited this, so leaving it while changing the number would be worse than leaving both alone.

One consequence worth recording, and now recorded in the Caddyfile: this makes externalising
bitborg-web's process-local captcha and rate-limit state NOT a prerequisite for hiding deploy
downtime.
The previous reasoning had that state on the critical path, because the only way to remove a
minute-plus of downtime looked like two concurrent instances. It still isn't a rolling deploy — this
hides the restart from arriving requests, it does not run two versions.

How to verify — and how NOT to

Both files now carry this warning: do not verify with probe_success. The blackbox interval missed
even a 12s deploy gap, so a 1.7s one will never be sampled, and "no dip" is the expected result whether
or not retrying works.

The real test is a continuous request loop across a deploy — confirming zero non-200 responses
through the restart. That is the only thing that can observe a 1.7s gap being bridged, and it is what
I'd run right after applying.

--syntax-check passes; ansible-lint clean at the production profile.

Closes #259. Raises `caddy_web_lb_try_duration` from **2s → 5s**, so a web deploy costs an arriving request a short wait instead of a 502. ## This value has been wrong twice, in opposite directions Both times because it was sized from blackbox probe samples rather than a measurement. | | value | assumption | reality | | --- | --- | --- | --- | | #249 | 10s | restart takes "a few seconds" | ~117s | | #250 | 2s | restart is ~60–90s, unbridgeable | ~117s, and the cause was misattributed | | this | **5s** | restart is **1.694s**, measured | — | #250's reasoning followed correctly from what it had, but ~60–90s came from three consecutive 30s-spaced probes failing — which can neither resolve nor attribute a gap. The real figure was ~117s, and it was **not** "startup runs migrations" (those were ~2s). It was `podman create`, because `UserNS=keep-id` forced a full rootfs chown on every new image digest (#253, fixed in #254). ## The measurement this is sized from Deploy at 2026-07-30 16:57:51Z — the first with both #254 and bitborg-web #109 live, priced from podman's monotonic journald event offsets: ``` Stopping bitborg-web.service 16:57:51.270 [shutdown] drained cleanly 16:57:51.344 (0.6ms drain; container died m=+0.064) container create 16:57:51.673 (m=+0.052 — on a NEW image digest) [migrate] migrations applied 16:57:52.773 (904ms) Server listening 16:57:52.964 ───────────────────────────────────────────── total unavailability 1.694s ``` **5s is ~3× that** — bridging with margin rather than exactly, because startup is dominated by migrations and a schema change can lengthen them. ## The trade-off, stated plainly A genuinely **down** upstream now hangs a request 5s before erroring instead of 2s. That is worth paying only because the window actually *succeeds* for the case it exists to cover — which was not true at 117s. `EndpointDown` (`probe_success == 0 for: 5m`) remains the signal for a real outage; this only changes how a visitor experiences its first few seconds. Retry semantics need no change and were already right: a failed **connection** is retried regardless of method, which is exactly the restart case, so `POST /api/signup` is covered without risking a double submission. ## Why the comments are rewritten too The old text in both the defaults and the Caddyfile asserted three things that are now false — 60–90s, "startup migrations", and "no sane retry window covers it". That stale reasoning is precisely why nobody revisited this, so leaving it while changing the number would be worse than leaving both alone. One consequence worth recording, and now recorded in the Caddyfile: **this makes externalising bitborg-web's process-local captcha and rate-limit state NOT a prerequisite for hiding deploy downtime.** The previous reasoning had that state on the critical path, because the only way to remove a minute-plus of downtime looked like two concurrent instances. It still isn't a rolling deploy — this hides the restart from arriving requests, it does not run two versions. ## How to verify — and how NOT to Both files now carry this warning: **do not verify with `probe_success`.** The blackbox interval missed even a 12s deploy gap, so a 1.7s one will never be sampled, and "no dip" is the expected result whether or not retrying works. The real test is a continuous request loop across a deploy — confirming **zero** non-200 responses through the restart. That is the only thing that can observe a 1.7s gap being bridged, and it is what I'd run right after applying. `--syntax-check` passes; `ansible-lint` clean at the `production` profile.
supernaut lade till 1 incheckning 2026-07-30 17:18:39 +00:00
fix(caddy): raise lb_try_duration to 5s — a deploy is now bridgeable
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m32s
b2eb42bedf
Closes #259. This value has been wrong twice, in opposite directions, and both
times because it was sized from blackbox probe samples rather than a measurement.

#249 set 10s assuming a restart took "a few seconds". #250 then cut it to 2s after
concluding the restart was ~60-90s and therefore unbridgeable. That conclusion
followed correctly from the numbers available, but the numbers were wrong: ~60-90s
came from three consecutive 30s-spaced probes failing, which can neither resolve
nor attribute a gap. The real figure was ~117s, and it was not "startup runs
migrations" — migrations were ~2s. It was `podman create`, because UserNS=keep-id
forced a full rootfs chown on every new image digest (#253, fixed in #254).

With that fixed, and gitborg-web #109 adding the SIGTERM handler the server never
had, the restart is 1.694s — priced from podman's monotonic journald event offsets
on the 2026-07-30 16:57:51Z deploy:

  Stopping gitborg-web.service     16:57:51.270
  [shutdown] drained cleanly       16:57:51.344   (0.6ms drain, died m=+0.064)
  container create                 16:57:51.673   (m=+0.052, NEW image digest)
  [migrate] migrations applied     16:57:52.773   (904ms)
  Server listening                 16:57:52.964
  ─────────────────────────────────────────────
  total                                 1.694s

5s is ~3x that: it bridges a deploy with margin rather than exactly, which matters
because startup is dominated by migrations and a schema change can lengthen them.

The trade-off, stated in the defaults: a genuinely DOWN upstream now hangs a
request 5s instead of 2s before erroring. Worth paying only because the window
actually succeeds for the case it exists to cover, which was not true at 117s.
EndpointDown remains the signal for a real outage.

Retry semantics need no change and were already right: a failed CONNECTION is
retried regardless of method, which is the restart case, so POST /api/signup is
covered without risking a double submission.

Also rewrites the rationale in both the defaults and the Caddyfile comment. The
old text asserted three things that are no longer true (60-90s, "startup
migrations", "no sane retry window covers it"), and that stale reasoning is why
nobody revisited this.

Notably this makes externalising gitborg-web's process-local captcha and
rate-limit state NOT a prerequisite for hiding deploy downtime — the reversal is
recorded in the Caddyfile comment so the next reader does not re-derive the old
conclusion.

Verification note in both files: do NOT verify with probe_success. The blackbox
interval missed even a 12s deploy gap, so a 1.7s one will never be sampled and
"no dip" is the expected result whether or not retrying works. It needs a
continuous request loop across a deploy.
supernaut sammanfogade incheckning ee096e1ab5 till main 2026-07-30 17:21:42 +00:00
supernaut tog bort grenen fix/caddy-bridge-deploy 2026-07-30 17:21:42 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!261
Ingen beskrivning angiven.