observability: nothing detects an alert rule that can never fire #444
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#444
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Follow-up to #442, which fixed one vacuous alert. This is about detecting the next one.
The gap
scripts/check-alert-rules.py(#439) proves the rule set renders and parses. It cannot prove arule can fire. A selector matching zero series does both cleanly, and
StrictUndefinedissatisfied whenever the Jinja variables resolve.
That is not hypothetical.
UnitMetricStaleshipped in that state and stayed there until #442. It wasfound by hand, by someone querying its matcher out of curiosity. Nothing would have found it
otherwise, and its stated job was watching the watchdog.
The signal already exists
vmalert publishes
vmalert_alerting_rules_last_evaluation_series_fetchedper rule.== 0means thatrule's selectors matched nothing on the last evaluation.
Measured on the monitoring VM, before and after #442's apply:
series_fetched == 0["UnitMetricStale"]So the metric found exactly the rule that was broken, and no others. It also confirms the blast
radius was one rule rather than a class of them.
#442's issue body was wrong about the cost of this. It said a check would need datasource access
from CI and would be its own design problem. It does not. The measurement above is a single instant
query, and the guard can be an ordinary alert rule.
Why a runtime alert beats a CI stage
A CI stage would only check at merge time. The runtime alert catches a strictly larger set, because a
rule does not have to ship broken to become broken:
bitborg_*/gitborg_*rename hazard, already recorded in ADR 0039 §7b, where two textfilewriters were half-renamed and dual-brand matchers were the only reason nothing blanked.
file=did here.None of those are visible at merge time. All of them are visible the moment the rule next evaluates.
The two layers are complementary, not alternatives. Keep both.
Proposed rule
With
alert_vacuity_exemptas a role default inroles/monitoring/defaults/main.yml.Three things were checked before proposing this, because each could have made it unworkable:
count(...) == 51), so== 0has no absence blind spot. A rulecannot hide by not reporting.
vector(1)reports 1, not 0. TheWatchdogrule has no selectors at all and was the obviousfalse-positive candidate. It is not one.
selects nothing is possible in principle (one watching for a metric that only appears during a
failure), which is why the list exists, but it should stay empty until something earns a place and
each entry should carry a comment saying why.
for: 30mrather than5m: after a vmalert restart the first evaluation has not happened yet, andthis must not page on every apply.
Known limitation, stated rather than hidden
This rule is evaluated by vmalert, so it cannot report on a rule set that is not being evaluated at
all. That case is covered separately by the cross-probe and, since #442, by
AlertingHealthMetricStale.It also detects vacuity only after the rule reaches production, which is exactly why #439's check
stays.
Definition of done
VacuousAlertRuleadded, withalert_vacuity_exemptdefaulting to an empty listand observing the alert, per the discipline the template comment now requires
scripts/check-alert-rules.pystill passes, self-test includeddocs/runbook.mdobservability section says what to do when it fires