monitoring: dashboard improvements — drill-down links, thresholds, alert state, and two missing dashboards #301

Öppen
öppnade 2026-08-01 09:03:03 +00:00 av supernaut · 0 kommentarer
Ägare

Findings from a full audit of all six provisioned dashboards on 2026-08-01. None is broken; these are
gaps that make on-call slower than it needs to be. Roughly in value order.

Cross-cutting

  • No dashboard links anywhere. Not one of the six has a links block, so there is no drill-down
    from overview to host / endpoints / web / Forgejo / runners. Cheapest usability win available, and it
    also dissolves the overview↔endpoints and overview↔host panel duplication without deleting anything.
  • No alert-state surface. Grafana holds zero alert rules by design — 33 live in
    alert-rules.yml.j2 plus 2 in loki-alert-rules.yml.j2. Only bitborg-runners bridges this, with a
    text panel naming the alerts behind its metrics. Replicate that on the other five, and consider an
    alertlist panel on overview.
  • Thresholds are almost entirely absent even though the alert thresholds are known constants
    (alert_disk_warn_pct: 80, alert_disk_crit_pct: 90). Every timeseries on overview, host, endpoints
    and Forgejo has zero threshold steps, so a graph never shows where the pager sits.
  • Annotations: five of six omit the annotations.list block entirely. At minimum a deploy
    annotation, so "did the apply cause this?" is answerable on the graph.
  • ~30 existing metrics appear on no dashboard, enough for two more: a backup dashboard covering
    the five-stage pipeline (gitborg_backup_* status and age for backup → storage → verify → drill →
    offsite, today one stat on overview), and an ops dashboard for
    gitborg_renovate_*, gitborg_registry_mirror_* and gitborg_token_audit_*.

Per dashboard

bitborg-overview (27 panels) — four dashboards in one. Split the VictoriaMetrics/Loki self-health
row (p14/p16/p17) into a monitoring-stack dashboard and the reconcile-trigger row (p24–p27) into a
reconciler dashboard. Bug: p7 is titled "Disk used % (services host)" but its query has no host
filter, so it plots both hosts. p8 Memory available bytes is the wrong framing for an at-a-glance
page — use Memory used % with 80/90 to match the alert.

bitborg-host (17 panels) — best-built of the set. Missing: node_vmstat_oom_kill despite
OOMKilled being an active alert; inode usage (node_filesystem_files_free), the classic disk-full
sibling the avail_bytes panels miss; and Load (1m) should threshold on vCPU count rather than
carrying unit short with no steps. Title says "host health" but it covers both hosts.

bitborg-endpoints (9 panels) — solid. Endpoint up / down would read far better as a
state-timeline, the pattern bitborg-runners already uses. Nothing distinguishes the
blackbox-http-auth job, so stats.gitborg.se's 401-is-up semantics are invisible to a reader. An
availability-over-window stat (avg_over_time(probe_success[30d])) is the natural SLO seed.

bitborg-forgejo (14 panels) — weakest. Pure inventory gauges: no request rate, no latency, no
error rate for git.gitborg.se, though bitborg-web already proves the pattern works off
{service_name="caddy-access"}. Add a traffic and latency row plus a {container="forgejo"} log
panel. Sixteen gitea_* metrics are unused; gitea_hooktasks and gitea_updatetasks are the
queue-depth signals worth graphing.

bitborg-web (11 panels) — Bug: p2 Container health and p3 Unit active have unit: "" and
no value mappings, so they render a bare 0/1 instead of Down/Up. p5 5xx rate has only one
threshold step, no critical tier. Nothing represents the sign-up flow, even though
SignupProvisioningFailing is a severity: critical, for: 0m alert.

bitborg-runners (12 panels) — the reference implementation. Only gap: Cinder volume quota used
uses 100 * (a > -1) / (b > 0) sentinel arithmetic that silently yields no data when the controller
reports -1; add a "metrics unavailable" mapping so the panel says so rather than going blank.

SLO note

There is no error-budget view anywhere, and the data barely exists: no HTTP request counters in
VictoriaMetrics at all
— a caddy_* sweep comes back empty, because Caddy's admin metrics are not
scraped. bitborg-web derives availability from Loki access logs, fine for a panel but too expensive
for burn-rate windows. Enabling Caddy's metrics endpoint and adding a scrape job would give
caddy_http_request_duration_seconds histograms and make a real multi-window burn-rate panel possible.

Findings from a full audit of all six provisioned dashboards on 2026-08-01. None is broken; these are gaps that make on-call slower than it needs to be. Roughly in value order. ## Cross-cutting - **No dashboard links anywhere.** Not one of the six has a `links` block, so there is no drill-down from overview to host / endpoints / web / Forgejo / runners. Cheapest usability win available, and it also dissolves the overview↔endpoints and overview↔host panel duplication without deleting anything. - **No alert-state surface.** Grafana holds zero alert rules by design — 33 live in `alert-rules.yml.j2` plus 2 in `loki-alert-rules.yml.j2`. Only `bitborg-runners` bridges this, with a text panel naming the alerts behind its metrics. Replicate that on the other five, and consider an `alertlist` panel on overview. - **Thresholds are almost entirely absent** even though the alert thresholds are known constants (`alert_disk_warn_pct: 80`, `alert_disk_crit_pct: 90`). Every timeseries on overview, host, endpoints and Forgejo has zero threshold steps, so a graph never shows where the pager sits. - **Annotations:** five of six omit the `annotations.list` block entirely. At minimum a deploy annotation, so "did the apply cause this?" is answerable on the graph. - **~30 existing metrics appear on no dashboard**, enough for two more: a **backup** dashboard covering the five-stage pipeline (`gitborg_backup_*` status and age for backup → storage → verify → drill → offsite, today one stat on overview), and an **ops** dashboard for `gitborg_renovate_*`, `gitborg_registry_mirror_*` and `gitborg_token_audit_*`. ## Per dashboard **`bitborg-overview` (27 panels)** — four dashboards in one. Split the VictoriaMetrics/Loki self-health row (p14/p16/p17) into a monitoring-stack dashboard and the reconcile-trigger row (p24–p27) into a reconciler dashboard. **Bug:** p7 is titled "Disk used % (services host)" but its query has no `host` filter, so it plots both hosts. p8 `Memory available bytes` is the wrong framing for an at-a-glance page — use `Memory used %` with 80/90 to match the alert. **`bitborg-host` (17 panels)** — best-built of the set. Missing: `node_vmstat_oom_kill` despite `OOMKilled` being an active alert; inode usage (`node_filesystem_files_free`), the classic disk-full sibling the `avail_bytes` panels miss; and `Load (1m)` should threshold on vCPU count rather than carrying unit `short` with no steps. Title says "host health" but it covers both hosts. **`bitborg-endpoints` (9 panels)** — solid. `Endpoint up / down` would read far better as a `state-timeline`, the pattern `bitborg-runners` already uses. Nothing distinguishes the `blackbox-http-auth` job, so `stats.gitborg.se`'s 401-is-up semantics are invisible to a reader. An availability-over-window stat (`avg_over_time(probe_success[30d])`) is the natural SLO seed. **`bitborg-forgejo` (14 panels)** — weakest. Pure inventory gauges: no request rate, no latency, no error rate for `git.gitborg.se`, though `bitborg-web` already proves the pattern works off `{service_name="caddy-access"}`. Add a traffic and latency row plus a `{container="forgejo"}` log panel. Sixteen `gitea_*` metrics are unused; `gitea_hooktasks` and `gitea_updatetasks` are the queue-depth signals worth graphing. **`bitborg-web` (11 panels)** — **Bug:** p2 `Container health` and p3 `Unit active` have `unit: ""` and no value mappings, so they render a bare `0`/`1` instead of Down/Up. p5 `5xx rate` has only one threshold step, no critical tier. Nothing represents the sign-up flow, even though `SignupProvisioningFailing` is a `severity: critical`, `for: 0m` alert. **`bitborg-runners` (12 panels)** — the reference implementation. Only gap: `Cinder volume quota used` uses `100 * (a > -1) / (b > 0)` sentinel arithmetic that silently yields no data when the controller reports `-1`; add a "metrics unavailable" mapping so the panel says so rather than going blank. ## SLO note There is no error-budget view anywhere, and the data barely exists: **no HTTP request counters in VictoriaMetrics at all** — a `caddy_*` sweep comes back empty, because Caddy's admin metrics are not scraped. `bitborg-web` derives availability from Loki access logs, fine for a panel but too expensive for burn-rate windows. Enabling Caddy's metrics endpoint and adding a scrape job would give `caddy_http_request_duration_seconds` histograms and make a real multi-window burn-rate panel possible.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#301
Ingen beskrivning angiven.