fix(monitoring): the alert watching the watchdog matched zero series #443

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/textfile-staleness-alerts in i main 2026-08-19 06:16:20 +00:00
Ägare

Closes #442.

What was wrong

UnitMetricStale could never fire. Its matcher selected zero series.

expr: time() - node_textfile_mtime_seconds{file="unit.prom"} > 15 * 60

node_exporter labels textfile metrics with the path as it sees them, and
node-exporter.container.j2 mounts node_textfile_dir at /textfile. The live label is
/textfile/unit.prom, never the bare basename.

This is the alert whose own comment said "so the watchdog itself is watched".

Measured against production, with a negative control:

Query seriesFetched Result
time() - node_textfile_mtime_seconds{file="unit.prom"} 0 empty
time() - node_textfile_mtime_seconds{file="/textfile/unit.prom"} 1 26.5

Second gap: nothing watched the cross-probe

bitborg-monitoring-probe.sh writes monitoring_probe.prom and monitoring_alerting.prom. Neither
had a freshness alert, and the textfile collector serves the last value indefinitely. A dead
bitborg-monitoring-probe.timer freezes both at 1 and the dead-man's switch is silently gone.

The asymmetry is the point. The probe covers vmalert dying. vmalert can cover the probe dying,
precisely because vmalert is alive in that scenario. Only the first half existed.

What changed

  • UnitMetricStale matches file=~"(.*/)?unit\\.prom". Regex rather than the literal path, which
    would only move the fragility to the mount point. The regex also survives a node_exporter that
    reverts to basenames.
  • New MonitoringProbeStale and AlertingHealthMetricStale, same pattern and threshold. Their
    descriptions distinguish the two cases: both firing means the timer is dead, only the second
    firing means the probe runs but its watchdog query fails.
  • check-alert-rules.py and the CI comment now state what the check cannot do, and drop a stale
    claim. Both said the Watchdog "only protects against vmalert death once an EXTERNAL monitor
    consumes it and alarms on silence". True of the Alertmanager receiver, which does deliberately
    have no integrations. Misleading about the system: the cross-probe bypasses Alertmanager and
    reads ALERTS{alertname="Watchdog"} straight from VictoriaMetrics. Confirmed live at
    bitborg_monitoring_alerting_healthy{host="bitborg-prod"} = 1.
  • Dropped the hardcoded "49 rules" from both docs rather than updating it to 51. The script prints
    the live count, so the literal could only drift again.
  • Corrected the UnitMetricStale description, which sent operators to gitborg-unit-metrics. The
    live unit is bitborg-unit-metrics.

Why \\. and not \.

The expr: line is a YAML plain scalar, so backslashes pass through untouched. MetricsQL then
unescapes the string literal itself. A single \. is a parse error at query time:

cannot parse string literal "\"(.*/)?unit\\.prom\"": invalid syntax

Verified after rendering, so what vmalert receives is what was tested.

Evidence

Each rule shown able to FIRE, not merely to parse, by evaluating the real rendered expression with
the threshold inverted:

Rule file matched age
UnitMetricStale /textfile/unit.prom 21s
MonitoringProbeStale /textfile/monitoring_probe.prom 133s
AlertingHealthMetricStale /textfile/monitoring_alerting.prom 139s

The last two match the 2 min timer. At the real 900s threshold none fire, which is correct.

Negative control: {file=~"(.*/)?nosuchfile\\.prom"} returns seriesFetched 0, so the regex form is
not matching everything vacuously.

./scripts/check-alert-rules.py reports 51 rules render and parse against
victoriametrics/vmalert:v1.147.0, and --self-test is still 4/4.

Note for the applier

This touches alert rules, so the apply restarts vmalert and resets every pending for: timer.

Deliberately not in scope

A third stage for check-alert-rules.py that extracts each rule's selectors, queries production,
and flags any matching zero series. That is what would have caught this. It needs datasource access
from CI and a way to express legitimately-empty selectors, so it deserves its own issue.

Closes #442. ## What was wrong `UnitMetricStale` could never fire. Its matcher selected zero series. ``` expr: time() - node_textfile_mtime_seconds{file="unit.prom"} > 15 * 60 ``` node_exporter labels textfile metrics with the path as it sees them, and `node-exporter.container.j2` mounts `node_textfile_dir` at `/textfile`. The live label is `/textfile/unit.prom`, never the bare basename. This is the alert whose own comment said "so the watchdog itself is watched". Measured against production, with a negative control: | Query | seriesFetched | Result | | --- | --- | --- | | `time() - node_textfile_mtime_seconds{file="unit.prom"}` | **0** | empty | | `time() - node_textfile_mtime_seconds{file="/textfile/unit.prom"}` | 1 | `26.5` | ## Second gap: nothing watched the cross-probe `bitborg-monitoring-probe.sh` writes `monitoring_probe.prom` and `monitoring_alerting.prom`. Neither had a freshness alert, and the textfile collector serves the last value indefinitely. A dead `bitborg-monitoring-probe.timer` freezes both at `1` and the dead-man's switch is silently gone. The asymmetry is the point. The probe covers vmalert dying. vmalert can cover the probe dying, precisely because vmalert is alive in that scenario. Only the first half existed. ## What changed - `UnitMetricStale` matches `file=~"(.*/)?unit\\.prom"`. Regex rather than the literal path, which would only move the fragility to the mount point. The regex also survives a node_exporter that reverts to basenames. - New `MonitoringProbeStale` and `AlertingHealthMetricStale`, same pattern and threshold. Their descriptions distinguish the two cases: both firing means the timer is dead, only the second firing means the probe runs but its watchdog query fails. - `check-alert-rules.py` and the CI comment now state what the check cannot do, and drop a stale claim. Both said the Watchdog "only protects against vmalert death once an EXTERNAL monitor consumes it and alarms on silence". True of the Alertmanager receiver, which does deliberately have no integrations. Misleading about the system: the cross-probe bypasses Alertmanager and reads `ALERTS{alertname="Watchdog"}` straight from VictoriaMetrics. Confirmed live at `bitborg_monitoring_alerting_healthy{host="bitborg-prod"} = 1`. - Dropped the hardcoded "49 rules" from both docs rather than updating it to 51. The script prints the live count, so the literal could only drift again. - Corrected the `UnitMetricStale` description, which sent operators to `gitborg-unit-metrics`. The live unit is `bitborg-unit-metrics`. ## Why `\\.` and not `\.` The `expr:` line is a YAML plain scalar, so backslashes pass through untouched. MetricsQL then unescapes the string literal itself. A single `\.` is a parse error at query time: ``` cannot parse string literal "\"(.*/)?unit\\.prom\"": invalid syntax ``` Verified after rendering, so what vmalert receives is what was tested. ## Evidence Each rule shown able to FIRE, not merely to parse, by evaluating the real rendered expression with the threshold inverted: | Rule | file matched | age | | --- | --- | --- | | `UnitMetricStale` | `/textfile/unit.prom` | 21s | | `MonitoringProbeStale` | `/textfile/monitoring_probe.prom` | 133s | | `AlertingHealthMetricStale` | `/textfile/monitoring_alerting.prom` | 139s | The last two match the 2 min timer. At the real 900s threshold none fire, which is correct. Negative control: `{file=~"(.*/)?nosuchfile\\.prom"}` returns seriesFetched 0, so the regex form is not matching everything vacuously. `./scripts/check-alert-rules.py` reports 51 rules render and parse against `victoriametrics/vmalert:v1.147.0`, and `--self-test` is still 4/4. ## Note for the applier This touches alert rules, so the apply restarts vmalert and resets every pending `for:` timer. ## Deliberately not in scope A third stage for `check-alert-rules.py` that extracts each rule's selectors, queries production, and flags any matching zero series. That is what would have caught this. It needs datasource access from CI and a way to express legitimately-empty selectors, so it deserves its own issue.
supernaut lade till 1 incheckning 2026-08-18 20:19:17 +00:00
fix(monitoring): the alert watching the watchdog matched zero series
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m53s
4f6e36bd88
UnitMetricStale selected `{file="unit.prom"}`. node_exporter labels textfile
metrics with the path as it sees them, and node-exporter.container.j2 mounts
node_textfile_dir at /textfile, so the live label is `/textfile/unit.prom`.
The matcher had never selected a series, so the one alert whose stated job was
watching the watchdog could not fire.

Measured against production, with a negative control:

  {file="unit.prom"}             seriesFetched 0, empty
  {file="/textfile/unit.prom"}   seriesFetched 1, 26.5

Matched by regex rather than by the literal path, which would only move the
fragility to the mount point. `\\.` is required: YAML plain scalars pass
backslashes through untouched and MetricsQL then unescapes the string literal,
so a single `\.` is a parse error at query time.

Added MonitoringProbeStale and AlertingHealthMetricStale over the cross-probe's
two textfiles. Nothing watched the probe at all. The asymmetry matters: the
probe covers vmalert dying, and vmalert can cover the probe dying precisely
because it is alive in that case. Only the first half was implemented.

All three shown able to fire against the live datasource before this landed,
by evaluating the real expressions with the threshold inverted:

  unit.prom                 21s
  monitoring_probe.prom    133s
  monitoring_alerting.prom 139s

The last two match the 2 min timer. At the real 900s threshold none fire.

check-alert-rules.py cannot catch this class of defect and now says so. A
selector matching zero series renders and parses cleanly, so both its stages
pass. Its claim that the Watchdog only protects "once an EXTERNAL monitor
consumes it" was also stale, and repeated in the CI comment: the cross-probe
reads ALERTS{alertname="Watchdog"} straight from VictoriaMetrics, bypassing
Alertmanager, so the receiver having no integrations does not weaken it.

Dropped the hardcoded rule count from both docs rather than updating 49 to 51.
The script prints the live count.

Also corrected the stale unit name in the UnitMetricStale description, which
pointed operators at gitborg-unit-metrics. The live unit is bitborg-unit-metrics.

Closes #442
supernaut sammanfogade incheckning 2658f4f263 till main 2026-08-19 06:16:20 +00:00
supernaut tog bort grenen fix/textfile-staleness-alerts 2026-08-19 06:16:20 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!443
Ingen beskrivning angiven.