docs(runbook): record §8b's measured verification and the RunnerPickupSlow false alarm #410
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!410
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "docs/rename-8b-measured-verification"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Docs only. The §8b register asserted a
changed=0convergence proof without the numbers behind it. This adds them, plus three things only learnable from having run it.The verification order, with its negative control
Three inventory-side checks answerable with no host contact at all, one of which is a negative control on the OLD name:
The third is the one that gets skipped. Without it a pass proves only that something resolves, not that the rename happened — a leftover directory or an un-renamed group would still return a working host.
unreachable=0is the load-bearing columnNot
changed=0—unreachable=0is what proves the renamedhost_vars/directories still carry working connection details. Also recorded: there was no unapplied drift on either host that day, sochanged=0is attributable to this change rather than masking a backlog. Explicitly flagged as not safe to assume next tranche.Provenance, stated rather than implied
The
--checkran from the PR branch pre-merge. The register now says so, along with thegit diffagainstorigin/mainshowing the squash-merged content is byte-identical and so carries the proof over. A register that implies a post-merge check it did not run is worse than one that admits the gap.RunnerPickupSlowwas the tranche's own CI, not a regressionExpected steady state is
Watchdogalone, so this read as a regression.boot_secondsready_secondspickup_secondsNormal boot means the wait is queueing, per the rule's own triage. The sharper point:
pickup_seconds_samplesis 10, so a p90 over that window is effectively the second-worst job — one queued job pins it high until ten more flush the window. Withfor: 15mon bursty CI, a single PR's runs can hold a warning up for hours. Recorded with the two metrics to check first, and pointed at #329 for the underlying queueing.