docs(runbook): record §8b's measured verification and the RunnerPickupSlow false alarm #410

Sammanfogat
supernaut sammanfogade 1 incheckning från docs/rename-8b-measured-verification in i main 2026-08-10 17:51:38 +00:00
Ägare

Docs only. The §8b register asserted a changed=0 convergence proof without the numbers behind it. This adds them, plus three things only learnable from having run it.

The verification order, with its negative control

Three inventory-side checks answerable with no host contact at all, one of which is a negative control on the OLD name:

ansible-inventory --graph
ansible-inventory --host bitborg-prod    # ansible_host + key path present
ansible-inventory --host gitborg-prod    # MUST fail

The third is the one that gets skipped. Without it a pass proves only that something resolves, not that the rename happened — a leftover directory or an un-renamed group would still return a working host.

unreachable=0 is the load-bearing column

bitborg-monitoring : ok=66   changed=0  unreachable=0  failed=0  skipped=35
bitborg-prod       : ok=288  changed=0  unreachable=0  failed=0  skipped=116

Not changed=0 — unreachable=0 is what proves the renamed host_vars/ directories still carry working connection details. Also recorded: there was no unapplied drift on either host that day, so changed=0 is attributable to this change rather than masking a backlog. Explicitly flagged as not safe to assume next tranche.

Provenance, stated rather than implied

The --check ran from the PR branch pre-merge. The register now says so, along with the git diff against origin/main showing the squash-merged content is byte-identical and so carries the proof over. A register that implies a post-merge check it did not run is worse than one that admits the gap.

RunnerPickupSlow was the tranche's own CI, not a regression

Expected steady state is Watchdog alone, so this read as a regression.

Metric p90 Reading
boot_seconds 19.8s normal
ready_seconds 36.7s normal
pickup_seconds 100.0s above the 90s target

Normal boot means the wait is queueing, per the rule's own triage. The sharper point: pickup_seconds_samples is 10, so a p90 over that window is effectively the second-worst job — one queued job pins it high until ten more flush the window. With for: 15m on bursty CI, a single PR's runs can hold a warning up for hours. Recorded with the two metrics to check first, and pointed at #329 for the underlying queueing.

Docs only. The §8b register asserted a `changed=0` convergence proof without the numbers behind it. This adds them, plus three things only learnable from having run it. ## The verification order, with its negative control Three inventory-side checks answerable with **no host contact at all**, one of which is a negative control on the OLD name: ```bash ansible-inventory --graph ansible-inventory --host bitborg-prod # ansible_host + key path present ansible-inventory --host gitborg-prod # MUST fail ``` The third is the one that gets skipped. Without it a pass proves only that *something* resolves, not that the rename happened — a leftover directory or an un-renamed group would still return a working host. ## `unreachable=0` is the load-bearing column ```text bitborg-monitoring : ok=66 changed=0 unreachable=0 failed=0 skipped=35 bitborg-prod : ok=288 changed=0 unreachable=0 failed=0 skipped=116 ``` Not `changed=0` — `unreachable=0` is what proves the renamed `host_vars/` directories still carry working connection details. Also recorded: there was no unapplied drift on either host that day, so `changed=0` is attributable to this change rather than masking a backlog. Explicitly flagged as *not* safe to assume next tranche. ## Provenance, stated rather than implied The `--check` ran from the PR branch **pre-merge**. The register now says so, along with the `git diff` against `origin/main` showing the squash-merged content is byte-identical and so carries the proof over. A register that implies a post-merge check it did not run is worse than one that admits the gap. ## `RunnerPickupSlow` was the tranche's own CI, not a regression Expected steady state is `Watchdog` alone, so this read as a regression. | Metric | p90 | Reading | | --- | --- | --- | | `boot_seconds` | 19.8s | normal | | `ready_seconds` | 36.7s | normal | | `pickup_seconds` | 100.0s | above the 90s target | Normal boot means the wait is queueing, per the rule's own triage. The sharper point: `pickup_seconds_samples` is **10**, so a p90 over that window is effectively the second-worst job — **one** queued job pins it high until ten more flush the window. With `for: 15m` on bursty CI, a single PR's runs can hold a warning up for hours. Recorded with the two metrics to check first, and pointed at #329 for the underlying queueing.
supernaut lade till 1 incheckning 2026-08-10 17:43:47 +00:00
The §8b register asserted a `changed=0` convergence proof without the
numbers behind it. This adds them, plus three things that are only
learnable from having run it.

The verification order matters and the register did not spell it out:
three inventory-side checks answerable with no host contact at all, one
of which is a negative control on the OLD name. Without that control a
pass proves only that something resolves, not that the rename happened
— a leftover directory or an un-renamed group would still return a
working host.

`unreachable=0` is the load-bearing column, not `changed=0`. It is what
proves the renamed `host_vars/` directories still carry working
connection details. Also noted: there was no unapplied drift on either
host that day, so `changed=0` is attributable to this change and not
masking a backlog — which is not safe to assume next time.

The `--check` ran from the PR branch, pre-merge. That is stated
explicitly, along with the `git diff` against `origin/main` that shows
the squash-merged content is byte-identical and so carries the proof
over. A register that implies a post-merge check it did not run is worse
than one that admits the gap.

`RunnerPickupSlow` was firing against an expected steady state of
`Watchdog` alone, so it read as a regression. It was the tranche's own
CI. Boot p90 19.8s and ready p90 36.7s are both normal, so the wait was
queueing — but the sharper point is that `pickup_seconds_samples` is 10,
making p90 effectively the second-worst job. One queued job pins it high
until ten more flush the window, so with `for: 15m` a single PR's runs
can hold the warning up for hours on bursty CI. Recorded with the two
metrics to check before treating it as an infra fault, and pointed at
 #329 for the underlying queueing.
supernaut sammanfogade incheckning d39df473aa till main 2026-08-10 17:51:38 +00:00
supernaut tog bort grenen docs/rename-8b-measured-verification 2026-08-10 17:51:38 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!410
Ingen beskrivning angiven.