observability: the alert that watches the watchdog matches zero series #442

Stängd
öppnade 2026-08-18 20:11:39 +00:00 av supernaut · 0 kommentarer
Ägare

The defect

UnitMetricStale cannot fire. Its matcher selects zero series, and has since it was written.

# roles/monitoring/templates/alert-rules.yml.j2:694
expr: time() - node_textfile_mtime_seconds{file="unit.prom"} > {{ alert_textfile_stale_minutes }} * 60

node_exporter labels textfile metrics with the path as it sees it, and node-exporter.container.j2
mounts {{ node_textfile_dir }} at /textfile. So the live label is /textfile/unit.prom, never
the bare basename.

Measured against production VictoriaMetrics:

Query seriesFetched Result
time() - node_textfile_mtime_seconds{file="unit.prom"} 0 empty
time() - node_textfile_mtime_seconds{file="/textfile/unit.prom"} 1 26.5

The second row is the negative control. The metric is present and fresh. Only the matcher is wrong.

This is the alert whose own comment says "so the watchdog itself is watched". It watches nothing.

Why nothing caught it

scripts/check-alert-rules.py (#439, PR #440) renders under StrictUndefined and parses with
vmalert. Both stages pass. {file="unit.prom"} is valid PromQL against a metric that exists, and
StrictUndefined is satisfied because every Jinja variable in the line resolves.

A selector matching zero series is invisible to any check that does not query a datasource. That is
the same class of vacuous pass #440 was built to stop, one layer deeper: #440 proves the file loads,
not that the rules can fire.

Second gap: nothing watches the cross-probe

bitborg-monitoring-probe.sh writes two textfile metrics on the services host:

  • bitborg_monitoring_vm_reachable from monitoring_probe.prom
  • bitborg_monitoring_alerting_healthy from monitoring_alerting.prom

Neither has a freshness alert. Both are served indefinitely at their last value by the textfile
collector, exactly like unit.prom. If bitborg-monitoring-probe.timer dies, both metrics freeze
at 1 and the dead-man's switch is silently gone.

This asymmetry is worth stating plainly. The probe covers the case where vmalert is dead. vmalert
covers the case where the probe is dead, because in that scenario vmalert is alive and evaluating.
The second half is simply not implemented.

flowchart LR
  timer["bitborg-monitoring-probe.timer<br/>(services host, 2 min)"]
  vm["monitoring VM<br/>VictoriaMetrics + vmalert"]
  smtp["Sweego SMTP<br/>(external to the VM)"]
  op(["operator"])

  timer -->|"queries ALERTS{alertname=Watchdog}"| vm
  timer -->|"transition-only mail"| smtp --> op
  vm -->|"CoreUnitDown, DiskWillFill, ..."| am["Alertmanager"] --> ntfy["ntfy"] --> op

  vm -.->|"UnitMetricStale<br/>MATCHES ZERO SERIES"| timer
  vm -.->|"no freshness alert<br/>on the probe textfiles"| timer

  classDef broken stroke-dasharray: 4 3;
  class timer,vm broken;

The dotted edges are the coverage that should exist and does not.

Third: a stale claim in three places

check-alert-rules.py's docstring and the matching comment in .forgejo/workflows/ci.yml both say
the Watchdog "only protects against vmalert death once an EXTERNAL monitor consumes it and alarms on
silence". That is accurate about the Alertmanager watchdog receiver, which does deliberately have no
integrations. It is misleading about the system, because the cross-probe bypasses Alertmanager and
reads the ALERTS series directly from VictoriaMetrics. The external monitor exists and works.

Confirmed live: bitborg_monitoring_alerting_healthy{host="bitborg-prod"} = 1.

What the probe does not prove is the delivery leg. It shows vmalert is evaluating, not that
Alertmanager reaches ntfy. AlertmanagerNotificationsFailing covers send errors, but it is itself
evaluated by vmalert, and it catches errors rather than silent misrouting. That residual gap is real
and is not in scope here.

Fix

  1. Correct the UnitMetricStale matcher to file=~"(.*/)?unit\\.prom". Anchoring on the full
    literal path would just move the fragility to the mount point; the regex survives both a mount
    change and a node_exporter that reverts to basenames.
  2. Add freshness alerts for monitoring_probe.prom and monitoring_alerting.prom, same pattern and
    same alert_textfile_stale_minutes threshold. The probe runs every 2 min, so 15 min is generous.
  3. Correct the docstring and CI comment described above.

Definition of done

  • UnitMetricStale matches its series in production, proven by query, not by inspection
  • The probe's two textfiles have freshness alerts
  • The stale Watchdog claim is corrected in check-alert-rules.py and .forgejo/workflows/ci.yml
  • Each new and changed matcher is shown able to select its series BEFORE the apply
  • scripts/check-alert-rules.py still passes, self-test included
  • docs/runbook.md observability section reflects that the cross-probe is the dead-man's switch

Applied to bitborg-monitoring on 2026-08-19. Evidence, from vmalert's own accounting rather than
by inspection:

Check Before After
Rules loaded 49 51
Rules with series_fetched == 0 ["UnitMetricStale"] none
UnitMetricStale series fetched 0 1
MonitoringProbeStale did not exist 1
AlertingHealthMetricStale did not exist 1

apply-reconcile.sh matched the dry run exactly (2 predicted, 2 changed), and a second --check
returned changed=0, which is what proves prod matches main. Only Watchdog firing, all 8 blackbox
probes at 1.

Follow-up: filed as #444

Correction. This section previously said a check for zero-series selectors would need datasource
access from CI and was its own design problem. That was wrong, and measuring the fix proved it.

vmalert already publishes vmalert_alerting_rules_last_evaluation_series_fetched per rule. Queried
on the monitoring VM before this fix was applied, == 0 returned exactly one rule out of 49,
UnitMetricStale, and nothing else. After the apply it returns none of 51. So the guard needs no CI
datasource access at all: it can be an ordinary alert rule.

Filed as #444, which also argues the runtime check is the better of the two, because a rule does not
have to ship broken to become broken. A renamed metric or a retired textfile writer makes a
previously-good matcher vacuous, and no merge-time check can see that.

## The defect `UnitMetricStale` cannot fire. Its matcher selects zero series, and has since it was written. ``` # roles/monitoring/templates/alert-rules.yml.j2:694 expr: time() - node_textfile_mtime_seconds{file="unit.prom"} > {{ alert_textfile_stale_minutes }} * 60 ``` node_exporter labels textfile metrics with the path as it sees it, and `node-exporter.container.j2` mounts `{{ node_textfile_dir }}` at `/textfile`. So the live label is `/textfile/unit.prom`, never the bare basename. Measured against production VictoriaMetrics: | Query | seriesFetched | Result | | --- | --- | --- | | `time() - node_textfile_mtime_seconds{file="unit.prom"}` | **0** | empty | | `time() - node_textfile_mtime_seconds{file="/textfile/unit.prom"}` | 1 | `26.5` | The second row is the negative control. The metric is present and fresh. Only the matcher is wrong. This is the alert whose own comment says "so the watchdog itself is watched". It watches nothing. ## Why nothing caught it `scripts/check-alert-rules.py` (#439, PR #440) renders under `StrictUndefined` and parses with vmalert. Both stages pass. `{file="unit.prom"}` is valid PromQL against a metric that exists, and `StrictUndefined` is satisfied because every Jinja variable in the line resolves. A selector matching zero series is invisible to any check that does not query a datasource. That is the same class of vacuous pass #440 was built to stop, one layer deeper: #440 proves the file loads, not that the rules can fire. ## Second gap: nothing watches the cross-probe `bitborg-monitoring-probe.sh` writes two textfile metrics on the services host: - `bitborg_monitoring_vm_reachable` from `monitoring_probe.prom` - `bitborg_monitoring_alerting_healthy` from `monitoring_alerting.prom` Neither has a freshness alert. Both are served indefinitely at their last value by the textfile collector, exactly like `unit.prom`. If `bitborg-monitoring-probe.timer` dies, both metrics freeze at `1` and the dead-man's switch is silently gone. This asymmetry is worth stating plainly. The probe covers the case where vmalert is dead. vmalert covers the case where the probe is dead, because in that scenario vmalert is alive and evaluating. The second half is simply not implemented. ```mermaid flowchart LR timer["bitborg-monitoring-probe.timer<br/>(services host, 2 min)"] vm["monitoring VM<br/>VictoriaMetrics + vmalert"] smtp["Sweego SMTP<br/>(external to the VM)"] op(["operator"]) timer -->|"queries ALERTS{alertname=Watchdog}"| vm timer -->|"transition-only mail"| smtp --> op vm -->|"CoreUnitDown, DiskWillFill, ..."| am["Alertmanager"] --> ntfy["ntfy"] --> op vm -.->|"UnitMetricStale<br/>MATCHES ZERO SERIES"| timer vm -.->|"no freshness alert<br/>on the probe textfiles"| timer classDef broken stroke-dasharray: 4 3; class timer,vm broken; ``` The dotted edges are the coverage that should exist and does not. ## Third: a stale claim in three places `check-alert-rules.py`'s docstring and the matching comment in `.forgejo/workflows/ci.yml` both say the Watchdog "only protects against vmalert death once an EXTERNAL monitor consumes it and alarms on silence". That is accurate about the Alertmanager watchdog receiver, which does deliberately have no integrations. It is misleading about the system, because the cross-probe bypasses Alertmanager and reads the `ALERTS` series directly from VictoriaMetrics. The external monitor exists and works. Confirmed live: `bitborg_monitoring_alerting_healthy{host="bitborg-prod"} = 1`. What the probe does **not** prove is the delivery leg. It shows vmalert is evaluating, not that Alertmanager reaches ntfy. `AlertmanagerNotificationsFailing` covers send errors, but it is itself evaluated by vmalert, and it catches errors rather than silent misrouting. That residual gap is real and is not in scope here. ## Fix 1. Correct the `UnitMetricStale` matcher to `file=~"(.*/)?unit\\.prom"`. Anchoring on the full literal path would just move the fragility to the mount point; the regex survives both a mount change and a node_exporter that reverts to basenames. 2. Add freshness alerts for `monitoring_probe.prom` and `monitoring_alerting.prom`, same pattern and same `alert_textfile_stale_minutes` threshold. The probe runs every 2 min, so 15 min is generous. 3. Correct the docstring and CI comment described above. ## Definition of done - [x] `UnitMetricStale` matches its series in production, proven by query, not by inspection - [x] The probe's two textfiles have freshness alerts - [x] The stale Watchdog claim is corrected in `check-alert-rules.py` and `.forgejo/workflows/ci.yml` - [x] Each new and changed matcher is shown able to select its series BEFORE the apply - [x] `scripts/check-alert-rules.py` still passes, self-test included - [x] `docs/runbook.md` observability section reflects that the cross-probe is the dead-man's switch Applied to `bitborg-monitoring` on 2026-08-19. Evidence, from vmalert's own accounting rather than by inspection: | Check | Before | After | | --- | --- | --- | | Rules loaded | 49 | 51 | | Rules with `series_fetched == 0` | `["UnitMetricStale"]` | none | | `UnitMetricStale` series fetched | 0 | 1 | | `MonitoringProbeStale` | did not exist | 1 | | `AlertingHealthMetricStale` | did not exist | 1 | `apply-reconcile.sh` matched the dry run exactly (2 predicted, 2 changed), and a second `--check` returned `changed=0`, which is what proves prod matches main. Only `Watchdog` firing, all 8 blackbox probes at 1. ## Follow-up: filed as #444 **Correction.** This section previously said a check for zero-series selectors would need datasource access from CI and was its own design problem. That was wrong, and measuring the fix proved it. vmalert already publishes `vmalert_alerting_rules_last_evaluation_series_fetched` per rule. Queried on the monitoring VM before this fix was applied, `== 0` returned exactly one rule out of 49, `UnitMetricStale`, and nothing else. After the apply it returns none of 51. So the guard needs no CI datasource access at all: it can be an ordinary alert rule. Filed as #444, which also argues the runtime check is the better of the two, because a rule does not have to ship broken to become broken. A renamed metric or a retired textfile writer makes a previously-good matcher vacuous, and no merge-time check can see that.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#442
Ingen beskrivning angiven.