docs(runbook): the automatic rollback is rehearsed on production's podman #319

Sammanfogat
supernaut sammanfogade 1 incheckning från docs/rollback-rehearsed-5.4.2 in i main 2026-08-01 19:37:44 +00:00
Ägare

Closes the one deliberate gap #315 shipped with. That PR asked, in bold, that the rollback not be
treated as working until someone had watched it on the version production actually runs. It has now
been watched.

Rehearsed on production's exact combination

podman 5.4.2, Debian GNU/Linux 13 (trixie) — verified against the services host with
podman --version first, so the match is checked rather than assumed.

Check Result
podman auto-update verdict zz-rb.service … registry **rolled back**
Image after rollback identical to the pre-deploy good image ID
Health after rollback healthy
Exit code 0 — why alerting reads the journal, not a failed unit
Journal signature start operation timed out. → State 'stop-sigterm' timed out → Started
Second run, tag still broken rolled back again, exit 0, healthy — hazard 1 reproduced

Not on the services host — and the rule is now justified

The procedure says never the services host, but the results table asked for a
services host (gitborg-prod) row. Those cannot both be satisfied, so the table now records a
podman version rather than a machine. The version was always the real unknown: a laptop
rehearsal proves the mechanism, not the version, and 5.8.3 differs from 5.4.2.

The prohibition turns out to be substantive rather than fastidious, which the runbook did not say.
Step 3 runs a bare podman auto-update, and that cannot be scoped to a single unit. On the
services host it would also evaluate bitborg-web — the only container tracking a registry tag —
and deploy a new latest if CI had pushed one, in the middle of a test deliberately breaking
things. That reasoning is now recorded next to the rule.

Also added

  • How to rehearse on production's version without touching production: a throwaway VM from the
    same Debian13 image, the whole procedure in cloud-init, result read from
    openstack console log show — so no floating IP, no keypair and no inbound rule.
  • The zero-disk flavour trap. These flavours have zero disk, so a plain server create is
    refused with "Only volume-backed servers are allowed for flavors with zero disk". The documented
    command passes an explicit delete_on_termination=true boot volume so it cascades on delete
    rather than orphaning — relevant given #305.
  • The number an operator needs mid-incident. The failed-start window measured ~61 s against
    20 s + 30 s timeouts. bitborg-web runs 90 s + 30 s, so expect roughly two minutes before a
    rollback lands. Long enough that someone watching a broken deploy would otherwise intervene early
    and conclude the gate had failed.

Verified

pnpm mdlint 0 issues · pnpm format clean · documentation only, nothing applied and no
production change.

Noted separately

Tearing the rehearsal VM down surfaced unattached boot volumes unrelated to this change; filed on
its own rather than folded in here.

Closes the one deliberate gap #315 shipped with. That PR asked, in bold, that the rollback not be treated as working until someone had watched it on the version production actually runs. It has now been watched. ## Rehearsed on production's exact combination **podman 5.4.2, Debian GNU/Linux 13 (trixie)** — verified against the services host with `podman --version` first, so the match is checked rather than assumed. | Check | Result | | ------------------------------ | ------------------------------------------------------------------ | | `podman auto-update` verdict | `zz-rb.service … registry **rolled back**` | | Image after rollback | identical to the pre-deploy good image ID | | Health after rollback | `healthy` | | Exit code | **0** — why alerting reads the journal, not a failed unit | | Journal signature | `start operation timed out.` → `State 'stop-sigterm' timed out` → `Started` | | Second run, tag still broken | rolled back **again**, exit 0, healthy — hazard 1 reproduced | ## Not on the services host — and the rule is now justified The procedure says **never** the services host, but the results table asked for a `services host (gitborg-prod)` row. Those cannot both be satisfied, so the table now records a podman **version** rather than a machine. The version was always the real unknown: a laptop rehearsal proves the mechanism, not the version, and 5.8.3 differs from 5.4.2. The prohibition turns out to be substantive rather than fastidious, which the runbook did not say. Step 3 runs a bare `podman auto-update`, and that **cannot be scoped to a single unit**. On the services host it would also evaluate `bitborg-web` — the only container tracking a registry tag — and deploy a new `latest` if CI had pushed one, in the middle of a test deliberately breaking things. That reasoning is now recorded next to the rule. ## Also added - **How to rehearse on production's version without touching production**: a throwaway VM from the same `Debian13` image, the whole procedure in cloud-init, result read from `openstack console log show` — so no floating IP, no keypair and no inbound rule. - **The zero-disk flavour trap.** These flavours have zero disk, so a plain `server create` is refused with *"Only volume-backed servers are allowed for flavors with zero disk"*. The documented command passes an explicit `delete_on_termination=true` boot volume so it cascades on delete rather than orphaning — relevant given #305. - **The number an operator needs mid-incident.** The failed-start window measured **~61 s** against 20 s + 30 s timeouts. `bitborg-web` runs 90 s + 30 s, so expect **roughly two minutes** before a rollback lands. Long enough that someone watching a broken deploy would otherwise intervene early and conclude the gate had failed. ## Verified `pnpm mdlint` 0 issues · `pnpm format` clean · documentation only, nothing applied and no production change. ## Noted separately Tearing the rehearsal VM down surfaced unattached boot volumes unrelated to this change; filed on its own rather than folded in here.
supernaut lade till 1 incheckning 2026-08-01 19:15:35 +00:00
docs(runbook): the automatic rollback is rehearsed on production's podman (#315)
Alla kontroller lyckades
ci / ci (pull_request) Successful in 15s
fc87453863
#315 shipped the readiness gate and rollback with the mechanism rehearsed only on podman 5.8.3, and
asked that it not be treated as working until someone had watched it on the version production runs.
That has now happened: 5.4.2 on Debian 13, the services host's exact combination.

Verified on a throwaway VM, not the services host. `podman auto-update` printed `rolled back`, the
container returned to the good image ID and `healthy`, the exit code was 0, the journal carried the
documented `start operation timed out` -> `stop-sigterm` -> `Started` sequence, and a second run with
the tag still broken rolled back again.

The results table asked for a row it could never legitimately get. It had a `services host
(gitborg-prod)` row marked "not yet rehearsed", while the procedure directly above it says **never**
the services host. The table now records a podman VERSION rather than a machine, which is what was
ever actually in question — a laptop rehearsal proves the mechanism but not the version, and the
version is the part that differs.

Why the host rule is hard rather than fastidious, now written down: step 3 runs a bare
`podman auto-update`, which cannot be scoped to one unit. On the services host it would also
evaluate `gitborg-web` — the only container tracking a registry tag — and deploy a new `latest` if
CI had pushed one, in the middle of a test that is deliberately breaking things.

Also records how to rehearse on production's version without touching production: a throwaway VM
from the same `Debian13` image, whole procedure in cloud-init, result read from the console log, so
no floating IP, keypair or inbound rule is needed. Includes the zero-disk flavour trap (the server
must be volume-backed) and an explicit `delete_on_termination=true` so the boot volume cascades
rather than orphaning.

And the number an operator actually needs mid-incident: the failed-start window measured ~61 s
against 20 s + 30 s timeouts, so gitborg-web's 90 s + 30 s implies roughly two minutes before a
rollback lands. Long enough that someone watching a broken deploy would otherwise intervene and
conclude the gate had failed.
supernaut sammanfogade incheckning da5c086063 till main 2026-08-01 19:37:44 +00:00
supernaut tog bort grenen docs/rollback-rehearsed-5.4.2 2026-08-01 19:37:44 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!319
Ingen beskrivning angiven.