monitoring: dashboard improvements — drill-down links, thresholds, alert state, and two missing dashboards #301
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#301
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Findings from a full audit of all six provisioned dashboards on 2026-08-01. None is broken; these are
gaps that make on-call slower than it needs to be. Roughly in value order.
Cross-cutting
linksblock, so there is no drill-downfrom overview to host / endpoints / web / Forgejo / runners. Cheapest usability win available, and it
also dissolves the overview↔endpoints and overview↔host panel duplication without deleting anything.
alert-rules.yml.j2plus 2 inloki-alert-rules.yml.j2. Onlybitborg-runnersbridges this, with atext panel naming the alerts behind its metrics. Replicate that on the other five, and consider an
alertlistpanel on overview.(
alert_disk_warn_pct: 80,alert_disk_crit_pct: 90). Every timeseries on overview, host, endpointsand Forgejo has zero threshold steps, so a graph never shows where the pager sits.
annotations.listblock entirely. At minimum a deployannotation, so "did the apply cause this?" is answerable on the graph.
the five-stage pipeline (
gitborg_backup_*status and age for backup → storage → verify → drill →offsite, today one stat on overview), and an ops dashboard for
gitborg_renovate_*,gitborg_registry_mirror_*andgitborg_token_audit_*.Per dashboard
bitborg-overview(27 panels) — four dashboards in one. Split the VictoriaMetrics/Loki self-healthrow (p14/p16/p17) into a monitoring-stack dashboard and the reconcile-trigger row (p24–p27) into a
reconciler dashboard. Bug: p7 is titled "Disk used % (services host)" but its query has no
hostfilter, so it plots both hosts. p8
Memory available bytesis the wrong framing for an at-a-glancepage — use
Memory used %with 80/90 to match the alert.bitborg-host(17 panels) — best-built of the set. Missing:node_vmstat_oom_killdespiteOOMKilledbeing an active alert; inode usage (node_filesystem_files_free), the classic disk-fullsibling the
avail_bytespanels miss; andLoad (1m)should threshold on vCPU count rather thancarrying unit
shortwith no steps. Title says "host health" but it covers both hosts.bitborg-endpoints(9 panels) — solid.Endpoint up / downwould read far better as astate-timeline, the patternbitborg-runnersalready uses. Nothing distinguishes theblackbox-http-authjob, sostats.gitborg.se's 401-is-up semantics are invisible to a reader. Anavailability-over-window stat (
avg_over_time(probe_success[30d])) is the natural SLO seed.bitborg-forgejo(14 panels) — weakest. Pure inventory gauges: no request rate, no latency, noerror rate for
git.gitborg.se, thoughbitborg-webalready proves the pattern works off{service_name="caddy-access"}. Add a traffic and latency row plus a{container="forgejo"}logpanel. Sixteen
gitea_*metrics are unused;gitea_hooktasksandgitea_updatetasksare thequeue-depth signals worth graphing.
bitborg-web(11 panels) — Bug: p2Container healthand p3Unit activehaveunit: ""andno value mappings, so they render a bare
0/1instead of Down/Up. p55xx ratehas only onethreshold step, no critical tier. Nothing represents the sign-up flow, even though
SignupProvisioningFailingis aseverity: critical,for: 0malert.bitborg-runners(12 panels) — the reference implementation. Only gap:Cinder volume quota useduses
100 * (a > -1) / (b > 0)sentinel arithmetic that silently yields no data when the controllerreports
-1; add a "metrics unavailable" mapping so the panel says so rather than going blank.SLO note
There is no error-budget view anywhere, and the data barely exists: no HTTP request counters in
VictoriaMetrics at all — a
caddy_*sweep comes back empty, because Caddy's admin metrics are notscraped.
bitborg-webderives availability from Loki access logs, fine for a panel but too expensivefor burn-rate windows. Enabling Caddy's metrics endpoint and adding a scrape job would give
caddy_http_request_duration_secondshistograms and make a real multi-window burn-rate panel possible.