feat(monitoring): rename the gitborg_* metric prefix to bitborg_* in every producer #414
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!414
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "feat/rename-7b-metric-prefix"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
ADR 0039 §7b. Stacked on #413 — base is the §7a branch so this diff shows only the producer rename. #413 must merge first; the dual-brand consumers are the safety net that makes this window-free.
166 renames across 23 files.
An allowlist, not a prefix sweep
A blanket
gitborg_→bitborg_would have renamed two values that are not metrics and are load-bearing:gitborg_webLEGACY-PINned by the §5 decision. Renaming the declaration creates a fresh empty database and the portal comes up blank.gitborg_ephemeralcontroller.py'sMETADATA_TAG— the OpenStack metadata key deciding which VMs the runner-controller may DELETE. Renamed in code but not on live VMs, the controller stops recognising its own fleet.Also excluded:
gitborg_domain,gitborg_network_id(variable / tofu-output names) andgitborg_reconciler_image*(an Ansible var prefix). 59 renamed, 5 excluded, each classified by hand.controller.pybuilds 23 metric names from_METRIC_PREFIX = "gitborg_runner_controller", which is why enumerating full names in the repo undercounts the surface. That one assignment moved and carries all 23.Reconciler prose deliberately left on the old name
Seven
gitborg_reconciler_*metrics are the only ones this repo does not produce — they come frombitborg-auth-reconciler/src/metrics.ts, a separate repo shipped as a container image. Moving them needs a code change, a release, an image-tag bump and an apply (§7d).Their
exprs are already dual-brand and match either name, so they need nothing. Their descriptions still saygitborg_reconciler_*because that is what production emits today, with a comment saying why. An operator grepping mid-incident for a metric name nothing emits is exactly the misdirection this plan keeps warning about.Verified
vmalert:v1.147.0 -dryRunPASS; 46 alerts, set identical to pre-change, samefor:/severity:countstest_controller_logic.py— the suite CI runs, and it asserts on rendered metric output: 69/69 checks passedgitborg_webstill 2 hits in group_vars;gitborg_ephemeralstill 7 in controller.py — both checked explicitlyansible-lint0/0 across 194 files atproduction; pinnedshellcheckcleanExpected transient on apply
{__name__=~"[bg]itborg_x"}matches two series while retention holds the old one. They don't overlap in time except within one 5-minute staleness lookback at the switchover, sosum(...)panels can briefly read double. No alert aggregates over these metrics, so alerting is unaffected.for:clocks reset andincrease()under-reports for one window — the documented discontinuity you chose.ADR 0039 §7b. Stacked on §7a — the dual-brand consumers must be live first, so this must merge second. 166 renames across 23 files. Producers now emit `bitborg_*`; consumers already accept either name, so there is no window where a panel is blank or an alert selects something nothing emits. ## An allowlist, not a prefix sweep A blanket `gitborg_` -> `bitborg_` would have renamed two values that are not metrics and are load-bearing: - `gitborg_web` — the live Postgres database AND role, explicitly LEGACY-PINned by the §5 decision. Renaming the declaration creates a fresh empty database and the portal comes up blank. - `gitborg_ephemeral` — `controller.py`'s `METADATA_TAG`, the OpenStack metadata key that decides which VMs the runner-controller may DELETE. Renamed in code but not on live VMs, the controller stops recognising its own fleet. Also excluded: `gitborg_domain` and `gitborg_network_id` (variable and tofu-output names) and `gitborg_reconciler_image*` (an Ansible var prefix). 59 names/prefixes renamed, 5 excluded, each verified by hand. `controller.py` constructs 23 metric names from `_METRIC_PREFIX = "gitborg_runner_controller"`, so enumerating full names in the repo undercounts. That single assignment moved and carries all 23. ## Reconciler metrics deliberately NOT renamed in prose Seven `gitborg_reconciler_*` metrics are the only ones this repo does not produce — they come from `bitborg-auth-reconciler/src/metrics.ts`, a separate repo shipped as a container image, so moving them needs a code change, a release, an image-tag bump and an apply (§7d). Their `expr`s are already dual-brand and match either name. Their descriptions are left saying `gitborg_reconciler_*` because that is what production emits today, with a comment explaining why, so nobody tidies them into naming a metric that does not exist — an operator grepping for a nonexistent metric name mid-incident is exactly the misdirection this plan keeps warning about. ## Verified - `vmalert:v1.147.0 -dryRun` PASS, 46 alerts in a set identical to pre-change, same `for:`/`severity:` counts. - `test_controller_logic.py` (the suite CI runs, and it asserts on rendered metric output): **69/69 checks passed**. - `gitborg_web` still 2 hits in group_vars, `gitborg_ephemeral` still 7 in controller.py — both untouched, checked explicitly. - Dual-brand selectors intact: 46 in dashboards, 35 in alert-rules. - Dashboard JSON valid, `ansible-lint` 0/0 across 194 files at profile `production`, pinned `shellcheck` clean. ## Expected transient on apply `{__name__=~"[bg]itborg_x"}` matches two series while retention holds the old one. They do not overlap in time except within one 5-minute staleness lookback at the switchover, so `sum(...)` panels can briefly read double. No alert rule aggregates over these metrics, so alerting is unaffected. `for:` clocks reset and `increase()` under-reports for one window — the documented discontinuity.b5961beed7e14f63a240