observability: the alert that watches the watchdog matches zero series #442
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#442
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
The defect
UnitMetricStalecannot fire. Its matcher selects zero series, and has since it was written.node_exporter labels textfile metrics with the path as it sees it, and
node-exporter.container.j2mounts
{{ node_textfile_dir }}at/textfile. So the live label is/textfile/unit.prom, neverthe bare basename.
Measured against production VictoriaMetrics:
time() - node_textfile_mtime_seconds{file="unit.prom"}time() - node_textfile_mtime_seconds{file="/textfile/unit.prom"}26.5The second row is the negative control. The metric is present and fresh. Only the matcher is wrong.
This is the alert whose own comment says "so the watchdog itself is watched". It watches nothing.
Why nothing caught it
scripts/check-alert-rules.py(#439, PR #440) renders underStrictUndefinedand parses withvmalert. Both stages pass.
{file="unit.prom"}is valid PromQL against a metric that exists, andStrictUndefinedis satisfied because every Jinja variable in the line resolves.A selector matching zero series is invisible to any check that does not query a datasource. That is
the same class of vacuous pass #440 was built to stop, one layer deeper: #440 proves the file loads,
not that the rules can fire.
Second gap: nothing watches the cross-probe
bitborg-monitoring-probe.shwrites two textfile metrics on the services host:bitborg_monitoring_vm_reachablefrommonitoring_probe.prombitborg_monitoring_alerting_healthyfrommonitoring_alerting.promNeither has a freshness alert. Both are served indefinitely at their last value by the textfile
collector, exactly like
unit.prom. Ifbitborg-monitoring-probe.timerdies, both metrics freezeat
1and the dead-man's switch is silently gone.This asymmetry is worth stating plainly. The probe covers the case where vmalert is dead. vmalert
covers the case where the probe is dead, because in that scenario vmalert is alive and evaluating.
The second half is simply not implemented.
The dotted edges are the coverage that should exist and does not.
Third: a stale claim in three places
check-alert-rules.py's docstring and the matching comment in.forgejo/workflows/ci.ymlboth saythe Watchdog "only protects against vmalert death once an EXTERNAL monitor consumes it and alarms on
silence". That is accurate about the Alertmanager watchdog receiver, which does deliberately have no
integrations. It is misleading about the system, because the cross-probe bypasses Alertmanager and
reads the
ALERTSseries directly from VictoriaMetrics. The external monitor exists and works.Confirmed live:
bitborg_monitoring_alerting_healthy{host="bitborg-prod"} = 1.What the probe does not prove is the delivery leg. It shows vmalert is evaluating, not that
Alertmanager reaches ntfy.
AlertmanagerNotificationsFailingcovers send errors, but it is itselfevaluated by vmalert, and it catches errors rather than silent misrouting. That residual gap is real
and is not in scope here.
Fix
UnitMetricStalematcher tofile=~"(.*/)?unit\\.prom". Anchoring on the fullliteral path would just move the fragility to the mount point; the regex survives both a mount
change and a node_exporter that reverts to basenames.
monitoring_probe.promandmonitoring_alerting.prom, same pattern andsame
alert_textfile_stale_minutesthreshold. The probe runs every 2 min, so 15 min is generous.Definition of done
UnitMetricStalematches its series in production, proven by query, not by inspectioncheck-alert-rules.pyand.forgejo/workflows/ci.ymlscripts/check-alert-rules.pystill passes, self-test includeddocs/runbook.mdobservability section reflects that the cross-probe is the dead-man's switchApplied to
bitborg-monitoringon 2026-08-19. Evidence, from vmalert's own accounting rather thanby inspection:
series_fetched == 0["UnitMetricStale"]UnitMetricStaleseries fetchedMonitoringProbeStaleAlertingHealthMetricStaleapply-reconcile.shmatched the dry run exactly (2 predicted, 2 changed), and a second--checkreturned
changed=0, which is what proves prod matches main. OnlyWatchdogfiring, all 8 blackboxprobes at 1.
Follow-up: filed as #444
Correction. This section previously said a check for zero-series selectors would need datasource
access from CI and was its own design problem. That was wrong, and measuring the fix proved it.
vmalert already publishes
vmalert_alerting_rules_last_evaluation_series_fetchedper rule. Queriedon the monitoring VM before this fix was applied,
== 0returned exactly one rule out of 49,UnitMetricStale, and nothing else. After the apply it returns none of 51. So the guard needs no CIdatasource access at all: it can be an ordinary alert rule.
Filed as #444, which also argues the runtime check is the better of the two, because a rule does not
have to ship broken to become broken. A renamed metric or a retired textfile writer makes a
previously-good matcher vacuous, and no merge-time check can see that.