feat(web): gate deploys on readiness, and verify they actually served #315

Sammanfogat
supernaut sammanfogade 1 incheckning från feat/deploy-verification in i main 2026-08-01 18:21:36 +00:00
Ägare

Closes the deploy-safety follow-ups to the 2026-07-29 502 incident: a readiness gate with automatic
rollback, plus a post-deploy probe and a deploy record so a failure is known rather than merely
recovered from.

Corrected premise: the health endpoint already existed

The plan assumed bitborg-web exposed nothing to probe. That is stale — /healthz and src/lib/health.ts
exist, are a genuine readiness check (select 1 against Postgres, 503 on failure, timeout-raced, no
driver detail leaked), and are live: curl https://www.gitborg.se/healthz → 200 {"db":"ok","ok":true}.
No bitborg-web change was needed. /healthz was already exempt from rate limiting; the missing piece
was access-log exclusion, which is added here.

What changed

Readiness gate + rollback. The Quadlet healthcheck now targets /healthz instead of /, the interval
drops 30s→10s, and the unit gains Notify=healthy so systemd does not consider the container started
until it is actually healthy. TimeoutStartSec 300s→90s, plus TimeoutStopSec=30s. The post-apply gate in
site.yml now asserts /healthz rather than just /. web_healthcheck_path is defined once, because
three consumers have to agree on it.

Post-deploy probe + deploy record. A verify unit hooked to podman-auto-update.service via an
OnSuccess=/OnFailure= drop-in (the same pattern as the existing .timer.d override), three alert rules
(WebDeployFailed at for: 0m — the whole point is to be impatient where EndpointDown's 5m is
deliberately not), and a Loki-sourced deploy annotation on the three dashboards we control, including
the one where probe_success lives, which is the correlation that took a range query and a CI-log dig on
the night of the incident.

Four things measurement changed, that reasoning would not have

  1. podman auto-update exits 0 after a rollback. So OnFailure= never fires for it, and a
    successful rollback is otherwise completely silent — the old image serves, every probe passes, and
    the change you published simply is not live. The verifier detects it from start operation timed out in
    the web unit's journal. The first draft of the drop-in comment asserted the opposite; it is corrected.
  2. A broken tag is re-pulled and re-rolled-back every interval (~3 min). Reproduced, and documented
    with the remediation, because otherwise it looks like flapping.
  3. The failed-start window ≈ TimeoutStartSec + TimeoutStopSec, because podman run --sdnotify=healthy sits out the full 90s default stop timeout. Hence capping TimeoutStopSec at 30s.
    Actual downtime inside that window was ~11s — the new container serves until the final replace.
  4. A bug caught before shipping: a 90s probe window inside a Type=oneshot whose default
    TimeoutStartSec is also 90s means systemd SIGTERMs the check mid-probe and no metric is written —
    silent failure in precisely the case the check exists for. Now TimeoutStartSec = deadline + 60.

⚠ The rollback is NOT yet verified on production's podman

Stated plainly because an unexercised rollback tends not to work when it is finally needed, and that was
the part of the plan flagged as mattering most.

  • Rehearsed end to end on podman 5.8.3 with throwaway units and registry: registry policy +
    Notify=healthy + a broken image pushed over the tracked tag → UPDATED: rolled back, container back on
    the good image and healthy. Repeated, and also done with the local policy.
  • Production runs podman 5.4.2. I confirmed the mechanism exists there rather than assuming it:
    podman 5.4.2's own Quadlet man page documents it — "setting Notify to healthy will postpone startup
    notifications until such time as the container is marked healthy"
    . So the capability is present on
    prod's version.
  • What is still unproven is the end-to-end behaviour on 5.4.2, which needs the rehearsal in
    docs/runbook.md § "Rehearse the automatic rollback". The runbook carries a results table with the
    production row marked not yet rehearsed. Please do not treat this as a working rollback until that row
    is filled in.

Verified

quadlet -dryrun on the real rendered unit → Type=notify, --sdnotify=healthy, health-cmd quoting
intact, the new timeouts present. OnSuccess= drop-in confirmed end-to-end on a user oneshot in a VM. The
verify script driven through a 10-case stubbed state machine (deploy / idle / broken / recovery / restart /
thin body / absent container / rollback-while-healthy / no-flap / clear). shellcheck clean under both
0.11 and the CI-pinned 0.10. promtool check rules → 37 rules. caddy validate with real Caddy 2.11.4 →
Valid configuration. The LogQL annotation query runs against production Loki without error (0 hits — the
line does not exist yet) and job="journald" provably carries user-unit lines. ansible-lint passed on
the production profile
; site.yml --syntax-check OK; pnpm smoke:textfile 11 ok; prettier and
markdownlint clean. No production mutation.

Deliberate omissions

  • No Grafana API annotation / service-account token. A Loki-sourced annotation gives the same
    correlation with no new credential. Trade-off: it appears only on dashboards whose JSON we control — all
    three now do — not org-wide.
  • /healthz not added to the blackbox targets, because that would change EndpointDown's meaning for
    www: a Postgres blip would page after 5m. The post-deploy probe covers readiness instead.
  • StartLimitIntervalSec/Burst untouched. A latch would keep the site down pending a manual
    reset-failed — a worse trade, and not one to make silently. CoreUnitDown / ContainerFlapping cover
    the loop.

Two hazards that cannot be engineered away, both documented

A rollback reverts the image but not the database migration. And a readiness-gated start couples web
startup to Postgres being up, so a restart during a Postgres outage will fail and retry.

Follow-up

This materially amends ADR 0019's deploy mechanism (readiness gate, automatic rollback), so the ADR needs
an amendment. That lives in bitborg-docs — filed there rather than papered over here.

Closes the deploy-safety follow-ups to the 2026-07-29 502 incident: a readiness gate with automatic rollback, plus a post-deploy probe and a deploy record so a failure is *known* rather than merely recovered from. ## Corrected premise: the health endpoint already existed The plan assumed bitborg-web exposed nothing to probe. That is stale — `/healthz` and `src/lib/health.ts` exist, are a genuine readiness check (`select 1` against Postgres, 503 on failure, timeout-raced, no driver detail leaked), and are live: `curl https://www.gitborg.se/healthz` → `200 {"db":"ok","ok":true}`. **No bitborg-web change was needed.** `/healthz` was already exempt from rate limiting; the missing piece was access-log exclusion, which is added here. ## What changed **Readiness gate + rollback.** The Quadlet healthcheck now targets `/healthz` instead of `/`, the interval drops 30s→10s, and the unit gains **`Notify=healthy`** so systemd does not consider the container started until it is actually healthy. `TimeoutStartSec` 300s→90s, plus `TimeoutStopSec=30s`. The post-apply gate in `site.yml` now asserts `/healthz` rather than just `/`. `web_healthcheck_path` is defined once, because three consumers have to agree on it. **Post-deploy probe + deploy record.** A verify unit hooked to `podman-auto-update.service` via an `OnSuccess=`/`OnFailure=` drop-in (the same pattern as the existing `.timer.d` override), three alert rules (`WebDeployFailed` at **`for: 0m`** — the whole point is to be impatient where `EndpointDown`'s 5m is deliberately not), and a **Loki-sourced** deploy annotation on the three dashboards we control, including the one where `probe_success` lives, which is the correlation that took a range query and a CI-log dig on the night of the incident. ## Four things measurement changed, that reasoning would not have 1. **`podman auto-update` exits 0 after a rollback.** So `OnFailure=` never fires for it, and a *successful* rollback is otherwise **completely silent** — the old image serves, every probe passes, and the change you published simply is not live. The verifier detects it from `start operation timed out` in the web unit's journal. The first draft of the drop-in comment asserted the opposite; it is corrected. 2. **A broken tag is re-pulled and re-rolled-back every interval** (~3 min). Reproduced, and documented with the remediation, because otherwise it looks like flapping. 3. **The failed-start window ≈ `TimeoutStartSec` + `TimeoutStopSec`**, because `podman run --sdnotify=healthy` sits out the full 90s default stop timeout. Hence capping `TimeoutStopSec` at 30s. Actual downtime inside that window was ~11s — the new container serves until the final replace. 4. **A bug caught before shipping:** a 90s probe window inside a `Type=oneshot` whose default `TimeoutStartSec` is also 90s means systemd SIGTERMs the check mid-probe and *no metric is written* — silent failure in precisely the case the check exists for. Now `TimeoutStartSec = deadline + 60`. ## ⚠ The rollback is NOT yet verified on production's podman Stated plainly because an unexercised rollback tends not to work when it is finally needed, and that was the part of the plan flagged as mattering most. - **Rehearsed end to end on podman 5.8.3** with throwaway units and registry: `registry` policy + `Notify=healthy` + a broken image pushed over the tracked tag → `UPDATED: rolled back`, container back on the good image and `healthy`. Repeated, and also done with the `local` policy. - **Production runs podman 5.4.2.** I confirmed the mechanism *exists* there rather than assuming it: podman 5.4.2's own Quadlet man page documents it — *"setting `Notify` to `healthy` will postpone startup notifications until such time as the container is marked healthy"*. So the capability is present on prod's version. - **What is still unproven is the end-to-end behaviour on 5.4.2**, which needs the rehearsal in `docs/runbook.md § "Rehearse the automatic rollback"`. The runbook carries a results table with the production row marked *not yet rehearsed*. Please do not treat this as a working rollback until that row is filled in. ## Verified `quadlet -dryrun` on the **real rendered** unit → `Type=notify`, `--sdnotify=healthy`, health-cmd quoting intact, the new timeouts present. `OnSuccess=` drop-in confirmed end-to-end on a user oneshot in a VM. The verify script driven through a 10-case stubbed state machine (deploy / idle / broken / recovery / restart / thin body / absent container / rollback-while-healthy / no-flap / clear). `shellcheck` clean under both 0.11 and the CI-pinned 0.10. `promtool check rules` → 37 rules. `caddy validate` with real Caddy 2.11.4 → Valid configuration. The LogQL annotation query runs against production Loki without error (0 hits — the line does not exist yet) and `job="journald"` provably carries user-unit lines. `ansible-lint` **passed on the production profile**; `site.yml --syntax-check` OK; `pnpm smoke:textfile` 11 ok; prettier and markdownlint clean. No production mutation. ## Deliberate omissions - **No Grafana API annotation / service-account token.** A Loki-sourced annotation gives the same correlation with no new credential. Trade-off: it appears only on dashboards whose JSON we control — all three now do — not org-wide. - **`/healthz` not added to the blackbox targets**, because that would change `EndpointDown`'s meaning for www: a Postgres blip would page after 5m. The post-deploy probe covers readiness instead. - **`StartLimitIntervalSec`/`Burst` untouched.** A latch would keep the site down pending a manual `reset-failed` — a worse trade, and not one to make silently. `CoreUnitDown` / `ContainerFlapping` cover the loop. ## Two hazards that cannot be engineered away, both documented A rollback reverts the **image** but not the **database migration**. And a readiness-gated start couples web startup to Postgres being up, so a restart during a Postgres outage will fail and retry. ## Follow-up This materially amends ADR 0019's deploy mechanism (readiness gate, automatic rollback), so the ADR needs an amendment. That lives in `bitborg-docs` — filed there rather than papered over here.
supernaut lade till 1 incheckning 2026-08-01 15:03:00 +00:00
feat(web): gate deploys on readiness, and verify they actually served
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m34s
9a851618b6
Follow-up to the 2026-07-29 502. A merge to main reaches production with no
operator present (AutoUpdate=registry on a moving tag), and nothing checked the
result — the outage was found by a human loading the page. Two independent
mechanisms now stand in for that person.

The gate: the container healthcheck moves from `/` to the portal's /healthz,
which is 200 only when the server serves AND Postgres answers, so an image with
a wrong DATABASE_URL or a failed migration can no longer pass by rendering a
home page. `Notify=healthy` makes podman withhold systemd's READY message until
that probe passes, which is what lets `podman auto-update` (rollback is on by
default) see a failed start and restore the previous image. Still no
`HealthOnFailure=` — the anti-kill-loop stance is unchanged; this consults the
probe once per deploy, and the consequence is a bounded one-shot rollback.
TimeoutStartSec drops 300s -> 90s and TimeoutStopSec is capped at 30s, because
together they bound how long a broken image holds the unit before the rollback
lands.

The check: a drop-in starts gitborg-deploy-verify.service when
podman-auto-update.service finishes. It exits in two `podman inspect` calls when
nothing changed; on a new image or container it logs one `gitborg-deploy:` line
(the dashboards annotate deploys from it, so a probe_success dip is correlatable
at a glance) and probes /healthz plus a real page through the local Caddy,
retrying briefly so an ordinary restart is never called an outage. The result
becomes textfile metrics: WebDeployFailed fires with `for: 0m` rather than
waiting out EndpointDown's five minutes, and WebDeployVerifyStale catches the
verifier — or the auto-update timer — going quiet.

A successful rollback would otherwise be silent: the previous image serves, every
probe passes, and the change CI published simply is not live. `podman
auto-update` exits 0 after a rollback, so no unit fails either. The verifier
therefore looks for the trace the gate does leave in the web unit's journal and
raises WebDeployRolledBack, which also warns that a still-broken tag is re-pulled
and re-rejected every interval.

Also: /healthz is excluded from the Caddy access log (machine traffic that would
otherwise be the most-logged path and distort the request-rate and latency
panels), the post-apply health gate now asserts /healthz and not just `/`, and
the runbook gains the deploy-verification section plus a rehearsal procedure.

The rollback mechanism is rehearsed only on a developer machine with throwaway
units and a throwaway registry (podman 5.8.3): a deliberately-broken image over
the tracked tag produced `UPDATED: rolled back` and came back healthy on the
previous image. It is UNVERIFIED on the services host. The runbook records that
plainly and tracks where it has been exercised.
supernaut tvångsskickade feat/deploy-verification från 9a851618b6
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m34s
till fdabb9684b
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m39s
2026-08-01 17:12:25 +00:00
Jämför
supernaut tvångsskickade feat/deploy-verification från fdabb9684b
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m39s
till fe3f14ee96
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m26s
2026-08-01 18:17:19 +00:00
Jämför
supernaut sammanfogade incheckning 3f7f31b6f4 till main 2026-08-01 18:21:36 +00:00
supernaut tog bort grenen feat/deploy-verification 2026-08-01 18:21:36 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!315
Ingen beskrivning angiven.