server: no SIGTERM handler, so every deploy burns 10s and drops in-flight requests #107

Stängd
öppnade 2026-07-30 11:43:00 +00:00 av supernaut · 0 kommentarer
Ägare

The server does not handle SIGTERM, so every deploy spends a fixed 10 seconds waiting for a
graceful stop that never happens, then gets SIGKILLed. In-flight requests are dropped abruptly
rather than drained.

Evidence

From the services host journal, on the 2026-07-29 22:27 UTC deploy:

22:27:11.6  Stopping bitborg-web.service - bitborg public web portal (Astro)...
22:27:21.7  StopSignal SIGTERM failed to stop container bitborg-web in 10 seconds,
            resorting to SIGKILL
22:27:21.7  container died

Every deploy shows the same 10.1 s gap. It is the second-largest component of deploy downtime.

Why it happens

Containerfile ends with:

CMD ["sh", "-c", "node scripts/migrate.mjs || echo '...'; exec node ./dist/server/entry.mjs"]

The exec correctly hands PID 1 to the server, so the process does receive SIGTERM — the
comment in the Containerfile is right that this is what makes signal handling possible. But the Astro
node adapter's standalone server installs no SIGTERM handler, and Node's default action for a signal
with no handler and no default disposition is to do nothing. So the process sits there until Podman's
10 s StopTimeout expires.

This is a stop-path defect, not a startup one: nothing is broken on the way up.

Consequences

  1. 10 s added to every deploy's unavailability. Deploys are frequent (AutoUpdate=registry), so
    this is a recurring, entirely avoidable cost.
  2. In-flight requests are killed, not drained. A visitor mid-request during a deploy gets a
    dropped connection. The severity is currently masked by a much larger infra-side delay, but it is
    independent of it.
  3. It blocks any future graceful rollout. A rolling or blue-green deploy depends on an instance
    being able to stop draining-first; an instance that must be SIGKILLed cannot participate in one.

Fix

Install a signal handler that stops accepting new connections, closes idle ones, and exits once
in-flight requests finish or a deadline passes:

  • Handle both SIGTERM and SIGINT.
  • server.close() to stop accepting, plus server.closeIdleConnections() so keep-alive sockets do
    not hold the process open for the full timeout.
  • A hard deadline shorter than Podman's StopTimeout (default 10 s) — e.g. 5 s — then process.exit(0),
    so a stuck request cannot reintroduce the SIGKILL path.
  • Idempotent: a second signal during shutdown should not restart the sequence.

The Astro node adapter in standalone mode does not expose its http.Server directly, so the handler
needs a small wrapper entry point around dist/server/entry.mjs, or the adapter's middleware mode
with our own server. Worth checking which the current adapter version supports before choosing.

Acceptance

  • SIGTERM to the running container exits cleanly in well under 10 s.
  • No resorting to SIGKILL warning in the journal on a deploy.
  • A request in flight when the signal arrives completes rather than being dropped.
  • The hard deadline is exercised by a test, so a hung request still exits cleanly.

Context

Found while measuring the deploy phase breakdown after the 502s in bitborg-infra #249 / #250. The
dominant cost there is a separate infra-side regression filed against bitborg-infra; this 10 s is the
app's own share and is fixable independently of it.

The server does not handle `SIGTERM`, so every deploy spends a fixed 10 seconds waiting for a graceful stop that never happens, then gets `SIGKILL`ed. In-flight requests are dropped abruptly rather than drained. ## Evidence From the services host journal, on the 2026-07-29 22:27 UTC deploy: ``` 22:27:11.6 Stopping bitborg-web.service - bitborg public web portal (Astro)... 22:27:21.7 StopSignal SIGTERM failed to stop container bitborg-web in 10 seconds, resorting to SIGKILL 22:27:21.7 container died ``` Every deploy shows the same 10.1 s gap. It is the second-largest component of deploy downtime. ## Why it happens `Containerfile` ends with: ``` CMD ["sh", "-c", "node scripts/migrate.mjs || echo '...'; exec node ./dist/server/entry.mjs"] ``` The `exec` correctly hands PID 1 to the server, so the process **does** receive `SIGTERM` — the comment in the Containerfile is right that this is what makes signal handling possible. But the Astro node adapter's standalone server installs no `SIGTERM` handler, and Node's default action for a signal with no handler and no default disposition is to do nothing. So the process sits there until Podman's 10 s `StopTimeout` expires. This is a stop-path defect, not a startup one: nothing is broken on the way up. ## Consequences 1. **10 s added to every deploy's unavailability.** Deploys are frequent (`AutoUpdate=registry`), so this is a recurring, entirely avoidable cost. 2. **In-flight requests are killed, not drained.** A visitor mid-request during a deploy gets a dropped connection. The severity is currently masked by a much larger infra-side delay, but it is independent of it. 3. **It blocks any future graceful rollout.** A rolling or blue-green deploy depends on an instance being able to stop draining-first; an instance that must be `SIGKILL`ed cannot participate in one. ## Fix Install a signal handler that stops accepting new connections, closes idle ones, and exits once in-flight requests finish or a deadline passes: - Handle both `SIGTERM` and `SIGINT`. - `server.close()` to stop accepting, plus `server.closeIdleConnections()` so keep-alive sockets do not hold the process open for the full timeout. - A hard deadline shorter than Podman's `StopTimeout` (default 10 s) — e.g. 5 s — then `process.exit(0)`, so a stuck request cannot reintroduce the `SIGKILL` path. - Idempotent: a second signal during shutdown should not restart the sequence. The Astro node adapter in standalone mode does not expose its `http.Server` directly, so the handler needs a small wrapper entry point around `dist/server/entry.mjs`, or the adapter's middleware mode with our own server. Worth checking which the current adapter version supports before choosing. ## Acceptance - [ ] `SIGTERM` to the running container exits cleanly in well under 10 s. - [ ] No `resorting to SIGKILL` warning in the journal on a deploy. - [ ] A request in flight when the signal arrives completes rather than being dropped. - [ ] The hard deadline is exercised by a test, so a hung request still exits cleanly. ## Context Found while measuring the deploy phase breakdown after the 502s in bitborg-infra #249 / #250. The dominant cost there is a separate infra-side regression filed against bitborg-infra; this 10 s is the app's own share and is fixable independently of it.
supernaut lade till detta till projektet Bitborg Web 2026-07-30 17:01:12 +00:00
Logga in för att delta i denna konversation.
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-web#107
Ingen beskrivning angiven.