feat(monitoring): dual-brand every metric-name selector before the producers move #413
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!413
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "feat/rename-7a-dualbrand-consumers"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
ADR 0039 §7a. Consumers only — no producer changes, and no behaviour change, because only
gitborg_*exists at this point.Why consumers first, as a separate change
Rename producers first and there is a window where a dashboard panel is blank (visible, annoying) and, worse, where an alert rule selects a metric name nothing emits. An alert that cannot fire is indistinguishable from a healthy system.
Dual-branding consumers first is a free safety net. It is also already this repo's pattern — §6 shipped
container=~"[bg]itborg-web"for exactly this reason.45 dashboard
exprvalues (overview / runners / web) and 35 expression lines inalert-rules.yml.j2(27 inline, 8 in block scalars) moved to{__name__=~"[bg]itborg_<name>"}. Prose in panel descriptions and alert annotations is deliberately untouched — it moves with the producers so each change stays self-consistent.Proven equivalent, not assumed
Against live VictoriaMetrics:
{__name__=~"gitborg_.+"}{__name__=~"[bg]itborg_.+"}A labelled instant query
{__name__=~"[bg]itborg_unit_active",name=~"[bg]itborg-web"}matches its old form exactly.The rules file is validated with the pinned
vmalert:v1.147.0under-dryRun, with the pre-change file as a control:Dashboard JSON is edited textually, not round-tripped —
.prettierignoreexcludes these files on purpose and ajson.dumpwould reflow every panel. Diff is 45 insertions / 45 deletions, one-for-one.A YAML trap worth knowing
expr: {__name__=~"..."} != 0is invalid YAML, not invalid PromQL: a plain scalar starting with{is parsed as a flow mapping, so it fails before PromQL is consulted. 19 inline expressions are single-quoted for this reason; those starting with(don't need it, and block scalars are immune.The control-vs-new dryRun is what caught it — the first attempt failed at
yaml: line 157while the control passed, which is the only reason I knew the bug was mine rather than my Jinja-stripping harness.Transient to expect at §7b, documented not fixed
Once producers move,
{__name__=~"[bg]itborg_x"}matches two series. They don't overlap in time except within one 5-minute staleness lookback at the switchover, sosum(...)can briefly read double. No alert rule aggregates over these metrics (checked inline and block exprs — the one aggregation,max by(account)over a timestamp, is idempotent), so this is panel-only.Not fixed here, but it should be
The monitoring role validates the Caddyfile before restarting Caddy — with a comment pointing at the caddy role's identical guard — but does not validate the alert rules before restarting vmalert. "Render alert rules" simply notifies
Restart vmalert.So a malformed rules file gets written, vmalert refuses to start, alerting dies, and the thing that would have reported it is the thing that died. The
-dryRunabove is precisely the check the role is missing. Worth its own change.ADR 0039 §8c. Two changes: the prep that makes `instance_name` metadata-only, and the decision not to use it. ## Prep `keypair_name` (default `gitborg-prod-key`) and `guest_hostname` (default `gitborg-prod`) replace `"${var.instance_name}-key"` and `hostname: ${var.instance_name}`. Both of those were ForceNew, which is why changing `instance_name` alone measured `6 to add, 14 to change, 6 to destroy` on 2026-08-04 with both VMs reporting `must be replaced`. Defaults are byte-identical to what the old expressions produced, and `terraform.tfvars` sets only `instance_name`, so this is a no-op against live state: `tofu plan -detailed-exitcode` reports **zero resource actions**. `fmt -check` and `validate` clean. ## Decision: the names stay pinned The live OpenStack object names keep their `gitborg` prefix permanently. Nothing user-facing resolves any of them — no user, no DNS record, no certificate; they appear only in the OpenStack dashboard and `tofu state`. Against that, renaming means a create-new-and-repoint on the keypair, a vault edit that has to stay in lockstep with a var edit, and two security-group names whose failure mode is silent and deferred to the next runner boot and the next restore drill. ADR 0039's own principle applies unchanged: move the label, pin the value. §8a already demonstrated it on this very keypair — address moved to `openstack_compute_keypair_v2.bitborg`, id still `gitborg-prod-key`. The prep still earns its place: pinning by decision and pinning by accident are different states, and a future rebuild should not be forced into a rename it did not ask for. The 8c object inventory in the runbook is relabelled as a record rather than a task list, and the 08-04 measurements are marked as pre-prep — post-prep numbers have deliberately not been taken, so they are an expectation, not a result. ## Found while verifying `tofu plan` is not clean, and it is not this change. It exits 2 on `Changes to Outputs` alone, which OpenTofu annotates "without changing any real infrastructure". Outputs are state and only refresh on `apply`, so the stored `monitoring_ansible_host_hint` still reads `grafana.gitborg.se` — stale since §4 moved the domain. That is the outputs analogue of the #63 image-tag gap, and it matters because §8a's documented success condition is `tofu plan` reporting "No changes". An output-only apply to clear it is noted for §10.ADR 0039 §5. Two things: the one part of §5 that moved, and an explicit decision that the rest never will. ## Moved — operator-side only `$HOME/.config/gitborg/vault_pass` → `$HOME/.config/bitborg/vault_pass` and `GITBORG_ANSIBLE_VAULT_PASSWORD_FILE` → `BITBORG_ANSIBLE_VAULT_PASSWORD_FILE`, with **both variable names accepted for one release** so no operator shell breaks mid-migration. `vault-password.sh` resolves BITBORG_* first, falls back to GITBORG_* and notes the deprecation on stderr — stderr because Ansible consumes this script's stdout verbatim as the password. `apply-reconcile.sh`'s guard accepts either. `signup-drill.sh` probes both directories, with the canonical path as the final fallback so a failure message names where the file should live rather than where it used to. The variable holds an explicit path, so either name pointing at either directory keeps working; only the defaults moved. ## The distinction that shaped the scope `~/.config/gitborg/` is two directories on two machines. The operator's is above. The services host's `~gitborg/.config/gitborg/` holds `backup.pass`, `backup.env`, `reconciler.env`, `renovate.env`, `registry-mirror.env`, `registry-retention.env`, `token-audit.env` and `backup-verify-identity.txt` — an age identity. That one is data-carrying and stays pinned; every `{{ bitborg_home }}/.config/gitborg/` in the roles is deliberate. The runbook now says so in both places it used to be ambiguous. ## Pinned by decision 11 volumes and 5 `*.volume` units, ~13 podman secrets, the Postgres database and role, the unix user at uid 2000, `/srv/gitborg-*`, the backup archive prefix and the restic `--host` tag: all pinned, permanently, with a table recording what renaming each would cost. The uid-2000 user is the case the rest follow from — everything else on the host is owned by or namespaced under an identity that is not moving, so renaming objects around a pinned owner buys inconsistency rather than consistency. ## Verified The resolution matrix, since this is the path every Ansible command depends on: - deprecated var only -> works, warns - both set -> BITBORG_* wins (proved by pointing GITBORG_* at /nonexistent, so wrong precedence fails) - neither set -> exits 1, message names BITBORG_* - path absent / empty -> distinct, correct errors End to end through `ansible.cfg`, `ansible-vault view` decrypts under either variable name. `--syntax-check` clean, `ansible-lint` 0/0 across 194 files at profile `production`, pinned `shellcheck` clean.ADR 0039 §7a. Consumers only — no producer changes, no behaviour change, because only `gitborg_*` exists at this point. ## Why consumers first, as a separate change If producers were renamed first, there would be a window where a dashboard panel is blank (visible, annoying) and, worse, where an alert rule selects a metric name nothing emits. **An alert that cannot fire is indistinguishable from a healthy system.** Dual-branding consumers first is a free safety net, and it is already this repo's pattern — §6 shipped `container=~"[bg]itborg-web"` for the same reason. 45 dashboard `expr` values across overview/runners/web, and 35 expression lines in `alert-rules.yml.j2` (27 inline, 8 inside block scalars), all moved to `{__name__=~"[bg]itborg_<name>"}`. Prose in panel descriptions and alert annotations is deliberately untouched — it is renamed with the producers so each change stays self-consistent. ## Proven equivalent, not assumed Against live VictoriaMetrics, `{__name__=~"[bg]itborg_.+"}` and `{__name__=~"gitborg_.+"}` return **identical** results: 70 metric names, 104 series. A labelled instant query (`{__name__=~"[bg]itborg_unit_active",name=~"[bg]itborg-web"}`) matches its old form exactly. The rules file is validated with the pinned `vmalert:v1.147.0` under `-dryRun`, **with the pre-change file as a control** — original PASS, new PASS, 46 alerts in an identical set, same `for:` and `severity:` counts. Dashboard JSON is edited textually, not round-tripped: `.prettierignore` excludes these files on purpose, and a `json.dump` would reflow every panel. Diff is 45 insertions / 45 deletions, one-for-one. ## A YAML trap this hit `expr: {__name__=~"..."} != 0` is invalid **YAML**, not invalid PromQL — a plain scalar starting with `{` is parsed as a flow mapping, so it fails before PromQL is consulted. 19 inline expressions are single-quoted for this reason; the ones starting with `(` do not need it, and block scalars are immune. The control-vs-new dryRun is what caught it. ## Not fixed here, but it should be The monitoring role validates the Caddyfile before restarting Caddy (with a comment pointing at the caddy role's identical guard) but does **not** validate the alert rules before restarting vmalert — "Render alert rules" simply notifies `Restart vmalert`. A malformed rules file therefore gets written and vmalert refuses to start, so alerting dies and the thing that would report it is the thing that died. The dryRun above is exactly the check the role is missing. Worth its own change.5a232faa6cb85685e38d