ADR 0019 needs amending: deploys are now readiness-gated and roll back automatically #62

Stängd
öppnade 2026-08-01 15:03:22 +00:00 av supernaut · 0 kommentarer
Ägare

ADR 0019 records the bitborg-web deploy mechanism: CI builds and pushes an image, the services host pulls
it via podman auto-update, and CI never touches production. That remains true, but the mechanism has just
gained two properties the ADR does not describe, and both change what a deploy means.

What changed in bitborg-infra

  1. The start is readiness-gated. The Quadlet unit now sets Notify=healthy with a healthcheck against
    bitborg-web's /healthz (a real readiness check — process up and Postgres reachable). systemd does not
    consider the container started until it is healthy, so a start that never becomes ready now fails
    rather than silently serving a broken container.
  2. A failed new image rolls back automatically, via podman auto-update's rollback on healthcheck
    failure. ADR 0019 as written implies a deploy is one-directional: pull, restart, done.

There is also a post-deploy probe and a Grafana deploy annotation, which are observability rather than
mechanism and probably do not need ADR text.

Why it needs recording rather than just a runbook entry

Three consequences a reader of ADR 0019 would currently get wrong:

  • A "successful" podman auto-update run may mean a rollback happened. It exits 0 either way, so the
    published change can silently not be live. That is surprising enough to belong in the decision record, not
    only in a runbook procedure.
  • A rollback reverts the image but not the database migration. So the guarantee is narrower than
    "reverts the deploy", and anyone reasoning about safety from the ADR alone would over-trust it.
  • Web startup is now coupled to Postgres availability. A restart during a Postgres outage fails and
    retries. That is a deliberate trade — fail closed rather than serve a broken portal — but it is a real
    change to the deploy's failure modes and it should be a recorded decision rather than an emergent property
    of a healthcheck.

Suggested

An amendment to ADR 0019 rather than a new ADR: the decision (pull-based deploy, CI never touches
production) has not changed, only its mechanism and its failure modes. If the amendment convention here
prefers a superseding ADR, follow that instead.

Note when writing it: as of filing, the automatic rollback is verified on podman 5.8.3 but not yet
rehearsed on production's 5.4.2
. The capability is documented as present in 5.4.2's Quadlet man page, but
the end-to-end rehearsal is outstanding and the infra runbook tracks it in a results table. The ADR should
not describe the rollback as proven until that row is filled in.

ADR 0019 records the bitborg-web deploy mechanism: CI builds and pushes an image, the services host pulls it via `podman auto-update`, and CI never touches production. That remains true, but the mechanism has just gained two properties the ADR does not describe, and both change what a deploy *means*. ## What changed in bitborg-infra 1. **The start is readiness-gated.** The Quadlet unit now sets `Notify=healthy` with a healthcheck against bitborg-web's `/healthz` (a real readiness check — process up *and* Postgres reachable). systemd does not consider the container started until it is healthy, so a start that never becomes ready now **fails** rather than silently serving a broken container. 2. **A failed new image rolls back automatically**, via `podman auto-update`'s rollback on healthcheck failure. ADR 0019 as written implies a deploy is one-directional: pull, restart, done. There is also a post-deploy probe and a Grafana deploy annotation, which are observability rather than mechanism and probably do not need ADR text. ## Why it needs recording rather than just a runbook entry Three consequences a reader of ADR 0019 would currently get wrong: - **A "successful" `podman auto-update` run may mean a rollback happened.** It exits 0 either way, so the published change can silently not be live. That is surprising enough to belong in the decision record, not only in a runbook procedure. - **A rollback reverts the image but not the database migration.** So the guarantee is narrower than "reverts the deploy", and anyone reasoning about safety from the ADR alone would over-trust it. - **Web startup is now coupled to Postgres availability.** A restart during a Postgres outage fails and retries. That is a deliberate trade — fail closed rather than serve a broken portal — but it is a real change to the deploy's failure modes and it should be a recorded decision rather than an emergent property of a healthcheck. ## Suggested An amendment to ADR 0019 rather than a new ADR: the decision (pull-based deploy, CI never touches production) has not changed, only its mechanism and its failure modes. If the amendment convention here prefers a superseding ADR, follow that instead. Note when writing it: as of filing, the automatic rollback is **verified on podman 5.8.3 but not yet rehearsed on production's 5.4.2**. The capability is documented as present in 5.4.2's Quadlet man page, but the end-to-end rehearsal is outstanding and the infra runbook tracks it in a results table. The ADR should not describe the rollback as proven until that row is filled in.
supernaut lade till detta till projektet Bitborg Docs 2026-08-01 15:03:35 +00:00
Logga in för att delta i denna konversation.
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-docs#62
Ingen beskrivning angiven.