docs(runbook): the automatic rollback is rehearsed on production's podman #319
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!319
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "docs/rollback-rehearsed-5.4.2"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Closes the one deliberate gap #315 shipped with. That PR asked, in bold, that the rollback not be
treated as working until someone had watched it on the version production actually runs. It has now
been watched.
Rehearsed on production's exact combination
podman 5.4.2, Debian GNU/Linux 13 (trixie) — verified against the services host with
podman --versionfirst, so the match is checked rather than assumed.podman auto-updateverdictzz-rb.service … registry **rolled back**healthystart operation timed out.→State 'stop-sigterm' timed out→StartedNot on the services host — and the rule is now justified
The procedure says never the services host, but the results table asked for a
services host (gitborg-prod)row. Those cannot both be satisfied, so the table now records apodman version rather than a machine. The version was always the real unknown: a laptop
rehearsal proves the mechanism, not the version, and 5.8.3 differs from 5.4.2.
The prohibition turns out to be substantive rather than fastidious, which the runbook did not say.
Step 3 runs a bare
podman auto-update, and that cannot be scoped to a single unit. On theservices host it would also evaluate
bitborg-web— the only container tracking a registry tag —and deploy a new
latestif CI had pushed one, in the middle of a test deliberately breakingthings. That reasoning is now recorded next to the rule.
Also added
same
Debian13image, the whole procedure in cloud-init, result read fromopenstack console log show— so no floating IP, no keypair and no inbound rule.server createisrefused with "Only volume-backed servers are allowed for flavors with zero disk". The documented
command passes an explicit
delete_on_termination=trueboot volume so it cascades on deleterather than orphaning — relevant given #305.
20 s + 30 s timeouts.
bitborg-webruns 90 s + 30 s, so expect roughly two minutes before arollback lands. Long enough that someone watching a broken deploy would otherwise intervene early
and conclude the gate had failed.
Verified
pnpm mdlint0 issues ·pnpm formatclean · documentation only, nothing applied and noproduction change.
Noted separately
Tearing the rehearsal VM down surfaced unattached boot volumes unrelated to this change; filed on
its own rather than folded in here.