feat(monitoring): dual-brand every metric-name selector before the producers move #413

Sammanfogat
supernaut sammanfogade 3 incheckningar från feat/rename-7a-dualbrand-consumers in i main 2026-08-10 18:29:06 +00:00
Ägare

ADR 0039 §7a. Consumers only — no producer changes, and no behaviour change, because only gitborg_* exists at this point.

Why consumers first, as a separate change

Rename producers first and there is a window where a dashboard panel is blank (visible, annoying) and, worse, where an alert rule selects a metric name nothing emits. An alert that cannot fire is indistinguishable from a healthy system.

Dual-branding consumers first is a free safety net. It is also already this repo's pattern — §6 shipped container=~"[bg]itborg-web" for exactly this reason.

45 dashboard expr values (overview / runners / web) and 35 expression lines in alert-rules.yml.j2 (27 inline, 8 in block scalars) moved to {__name__=~"[bg]itborg_<name>"}. Prose in panel descriptions and alert annotations is deliberately untouched — it moves with the producers so each change stays self-consistent.

Proven equivalent, not assumed

Against live VictoriaMetrics:

Query Names Series
{__name__=~"gitborg_.+"} 70 104
{__name__=~"[bg]itborg_.+"} 70 104

A labelled instant query {__name__=~"[bg]itborg_unit_active",name=~"[bg]itborg-web"} matches its old form exactly.

The rules file is validated with the pinned vmalert:v1.147.0 under -dryRun, with the pre-change file as a control:

=== orig === PASS
=== new  === PASS
46 alerts, identical set; same for:/severity counts

Dashboard JSON is edited textually, not round-tripped — .prettierignore excludes these files on purpose and a json.dump would reflow every panel. Diff is 45 insertions / 45 deletions, one-for-one.

A YAML trap worth knowing

expr: {__name__=~"..."} != 0 is invalid YAML, not invalid PromQL: a plain scalar starting with { is parsed as a flow mapping, so it fails before PromQL is consulted. 19 inline expressions are single-quoted for this reason; those starting with ( don't need it, and block scalars are immune.

The control-vs-new dryRun is what caught it — the first attempt failed at yaml: line 157 while the control passed, which is the only reason I knew the bug was mine rather than my Jinja-stripping harness.

Transient to expect at §7b, documented not fixed

Once producers move, {__name__=~"[bg]itborg_x"} matches two series. They don't overlap in time except within one 5-minute staleness lookback at the switchover, so sum(...) can briefly read double. No alert rule aggregates over these metrics (checked inline and block exprs — the one aggregation, max by(account) over a timestamp, is idempotent), so this is panel-only.

Not fixed here, but it should be

The monitoring role validates the Caddyfile before restarting Caddy — with a comment pointing at the caddy role's identical guard — but does not validate the alert rules before restarting vmalert. "Render alert rules" simply notifies Restart vmalert.

So a malformed rules file gets written, vmalert refuses to start, alerting dies, and the thing that would have reported it is the thing that died. The -dryRun above is precisely the check the role is missing. Worth its own change.

ADR 0039 §7a. **Consumers only** — no producer changes, and no behaviour change, because only `gitborg_*` exists at this point. ## Why consumers first, as a separate change Rename producers first and there is a window where a dashboard panel is blank (visible, annoying) and, worse, where an alert rule selects a metric name nothing emits. **An alert that cannot fire is indistinguishable from a healthy system.** Dual-branding consumers first is a free safety net. It is also already this repo's pattern — §6 shipped `container=~"[bg]itborg-web"` for exactly this reason. 45 dashboard `expr` values (overview / runners / web) and 35 expression lines in `alert-rules.yml.j2` (27 inline, 8 in block scalars) moved to `{__name__=~"[bg]itborg_<name>"}`. Prose in panel descriptions and alert annotations is deliberately **untouched** — it moves with the producers so each change stays self-consistent. ## Proven equivalent, not assumed Against live VictoriaMetrics: | Query | Names | Series | | --- | --- | --- | | `{__name__=~"gitborg_.+"}` | 70 | 104 | | `{__name__=~"[bg]itborg_.+"}` | 70 | 104 | A labelled instant query `{__name__=~"[bg]itborg_unit_active",name=~"[bg]itborg-web"}` matches its old form exactly. The rules file is validated with the **pinned** `vmalert:v1.147.0` under `-dryRun`, **with the pre-change file as a control**: ``` === orig === PASS === new === PASS 46 alerts, identical set; same for:/severity counts ``` Dashboard JSON is edited **textually**, not round-tripped — `.prettierignore` excludes these files on purpose and a `json.dump` would reflow every panel. Diff is 45 insertions / 45 deletions, one-for-one. ## A YAML trap worth knowing `expr: {__name__=~"..."} != 0` is invalid **YAML**, not invalid PromQL: a plain scalar starting with `{` is parsed as a flow mapping, so it fails before PromQL is consulted. 19 inline expressions are single-quoted for this reason; those starting with `(` don't need it, and block scalars are immune. The control-vs-new dryRun is what caught it — the first attempt failed at `yaml: line 157` while the control passed, which is the only reason I knew the bug was mine rather than my Jinja-stripping harness. ## Transient to expect at §7b, documented not fixed Once producers move, `{__name__=~"[bg]itborg_x"}` matches two series. They don't overlap in time **except within one 5-minute staleness lookback** at the switchover, so `sum(...)` can briefly read double. No alert rule aggregates over these metrics (checked inline *and* block exprs — the one aggregation, `max by(account)` over a timestamp, is idempotent), so this is panel-only. ## Not fixed here, but it should be The monitoring role validates the Caddyfile before restarting Caddy — with a comment pointing at the caddy role's identical guard — but does **not** validate the alert rules before restarting vmalert. "Render alert rules" simply notifies `Restart vmalert`. So a malformed rules file gets written, vmalert refuses to start, alerting dies, and the thing that would have reported it is the thing that died. The `-dryRun` above is precisely the check the role is missing. Worth its own change.
supernaut lade till 3 incheckningar 2026-08-10 18:19:34 +00:00
ADR 0039 §8c. Two changes: the prep that makes `instance_name`
metadata-only, and the decision not to use it.

## Prep

`keypair_name` (default `gitborg-prod-key`) and `guest_hostname`
(default `gitborg-prod`) replace `"${var.instance_name}-key"` and
`hostname: ${var.instance_name}`. Both of those were ForceNew, which is
why changing `instance_name` alone measured `6 to add, 14 to change,
6 to destroy` on 2026-08-04 with both VMs reporting `must be replaced`.

Defaults are byte-identical to what the old expressions produced, and
`terraform.tfvars` sets only `instance_name`, so this is a no-op against
live state: `tofu plan -detailed-exitcode` reports **zero resource
actions**. `fmt -check` and `validate` clean.

## Decision: the names stay pinned

The live OpenStack object names keep their `gitborg` prefix
permanently. Nothing user-facing resolves any of them — no user, no DNS
record, no certificate; they appear only in the OpenStack dashboard and
`tofu state`. Against that, renaming means a create-new-and-repoint on
the keypair, a vault edit that has to stay in lockstep with a var edit,
and two security-group names whose failure mode is silent and deferred
to the next runner boot and the next restore drill.

ADR 0039's own principle applies unchanged: move the label, pin the
value. §8a already demonstrated it on this very keypair — address moved
to `openstack_compute_keypair_v2.bitborg`, id still `gitborg-prod-key`.

The prep still earns its place: pinning by decision and pinning by
accident are different states, and a future rebuild should not be forced
into a rename it did not ask for.

The 8c object inventory in the runbook is relabelled as a record rather
than a task list, and the 08-04 measurements are marked as pre-prep —
post-prep numbers have deliberately not been taken, so they are an
expectation, not a result.

## Found while verifying

`tofu plan` is not clean, and it is not this change. It exits 2 on
`Changes to Outputs` alone, which OpenTofu annotates "without changing
any real infrastructure". Outputs are state and only refresh on `apply`,
so the stored `monitoring_ansible_host_hint` still reads
`grafana.gitborg.se` — stale since §4 moved the domain. That is the
outputs analogue of the #63 image-tag gap, and it matters because §8a's
documented success condition is `tofu plan` reporting "No changes". An
output-only apply to clear it is noted for §10.
ADR 0039 §5. Two things: the one part of §5 that moved, and an explicit
decision that the rest never will.

## Moved — operator-side only

`$HOME/.config/gitborg/vault_pass` → `$HOME/.config/bitborg/vault_pass`
and `GITBORG_ANSIBLE_VAULT_PASSWORD_FILE` →
`BITBORG_ANSIBLE_VAULT_PASSWORD_FILE`, with **both variable names
accepted for one release** so no operator shell breaks mid-migration.

`vault-password.sh` resolves BITBORG_* first, falls back to GITBORG_*
and notes the deprecation on stderr — stderr because Ansible consumes
this script's stdout verbatim as the password. `apply-reconcile.sh`'s
guard accepts either. `signup-drill.sh` probes both directories, with
the canonical path as the final fallback so a failure message names
where the file should live rather than where it used to.

The variable holds an explicit path, so either name pointing at either
directory keeps working; only the defaults moved.

## The distinction that shaped the scope

`~/.config/gitborg/` is two directories on two machines. The operator's
is above. The services host's `~gitborg/.config/gitborg/` holds
`backup.pass`, `backup.env`, `reconciler.env`, `renovate.env`,
`registry-mirror.env`, `registry-retention.env`, `token-audit.env` and
`backup-verify-identity.txt` — an age identity. That one is
data-carrying and stays pinned; every
`{{ bitborg_home }}/.config/gitborg/` in the roles is deliberate. The
runbook now says so in both places it used to be ambiguous.

## Pinned by decision

11 volumes and 5 `*.volume` units, ~13 podman secrets, the Postgres
database and role, the unix user at uid 2000, `/srv/gitborg-*`, the
backup archive prefix and the restic `--host` tag: all pinned,
permanently, with a table recording what renaming each would cost.

The uid-2000 user is the case the rest follow from — everything else on
the host is owned by or namespaced under an identity that is not
moving, so renaming objects around a pinned owner buys inconsistency
rather than consistency.

## Verified

The resolution matrix, since this is the path every Ansible command
depends on:

- deprecated var only  -> works, warns
- both set             -> BITBORG_* wins (proved by pointing GITBORG_*
                          at /nonexistent, so wrong precedence fails)
- neither set          -> exits 1, message names BITBORG_*
- path absent / empty  -> distinct, correct errors

End to end through `ansible.cfg`, `ansible-vault view` decrypts under
either variable name. `--syntax-check` clean, `ansible-lint` 0/0 across
194 files at profile `production`, pinned `shellcheck` clean.
feat(monitoring): dual-brand every metric-name selector before the producers move
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m40s
5a232faa6c
ADR 0039 §7a. Consumers only — no producer changes, no behaviour change,
because only `gitborg_*` exists at this point.

## Why consumers first, as a separate change

If producers were renamed first, there would be a window where a
dashboard panel is blank (visible, annoying) and, worse, where an alert
rule selects a metric name nothing emits. **An alert that cannot fire is
indistinguishable from a healthy system.** Dual-branding consumers first
is a free safety net, and it is already this repo's pattern — §6 shipped
`container=~"[bg]itborg-web"` for the same reason.

45 dashboard `expr` values across overview/runners/web, and 35 expression
lines in `alert-rules.yml.j2` (27 inline, 8 inside block scalars), all
moved to `{__name__=~"[bg]itborg_<name>"}`. Prose in panel descriptions
and alert annotations is deliberately untouched — it is renamed with the
producers so each change stays self-consistent.

## Proven equivalent, not assumed

Against live VictoriaMetrics, `{__name__=~"[bg]itborg_.+"}` and
`{__name__=~"gitborg_.+"}` return **identical** results: 70 metric names,
104 series. A labelled instant query
(`{__name__=~"[bg]itborg_unit_active",name=~"[bg]itborg-web"}`) matches
its old form exactly.

The rules file is validated with the pinned `vmalert:v1.147.0` under
`-dryRun`, **with the pre-change file as a control** — original PASS, new
PASS, 46 alerts in an identical set, same `for:` and `severity:` counts.

Dashboard JSON is edited textually, not round-tripped: `.prettierignore`
excludes these files on purpose, and a `json.dump` would reflow every
panel. Diff is 45 insertions / 45 deletions, one-for-one.

## A YAML trap this hit

`expr: {__name__=~"..."} != 0` is invalid **YAML**, not invalid PromQL —
a plain scalar starting with `{` is parsed as a flow mapping, so it fails
before PromQL is consulted. 19 inline expressions are single-quoted for
this reason; the ones starting with `(` do not need it, and block scalars
are immune. The control-vs-new dryRun is what caught it.

## Not fixed here, but it should be

The monitoring role validates the Caddyfile before restarting Caddy
(with a comment pointing at the caddy role's identical guard) but does
**not** validate the alert rules before restarting vmalert — "Render
alert rules" simply notifies `Restart vmalert`. A malformed rules file
therefore gets written and vmalert refuses to start, so alerting dies
and the thing that would report it is the thing that died. The dryRun
above is exactly the check the role is missing. Worth its own change.
supernaut tvångsskickade feat/rename-7a-dualbrand-consumers från 5a232faa6c
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m40s
till b85685e38d
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m34s
2026-08-10 18:26:45 +00:00
Jämför
supernaut sammanfogade incheckning 27eaafde12 till main 2026-08-10 18:29:06 +00:00
supernaut tog bort grenen feat/rename-7a-dualbrand-consumers 2026-08-10 18:29:06 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!413
Ingen beskrivning angiven.