feat(kanidm): rename the provision account and entry-manager group #407

Sammanfogat
supernaut sammanfogade 1 incheckning från feat/rename-1b-sa-and-manager-group in i main 2026-08-10 09:03:09 +00:00
Ägare

ADR 0039 §1b, identifiers ② and ③ — the coupled half, completing the tranche. Applied and verified on prod 2026-08-10.

gitborg-web-provision → bitborg-web-provision, and idm_gitborg_ent_managers → idm_bitborg_ent_managers.

Four sites, two of which the plan did not list

Site Note
group_vars/all/vars.yml kanidm_web_entry_manager_group the variable
roles/kanidm/defaults/main.yml kanidm_web_provision_account the variable
roles/kanidm/defaults/main.yml kanidm_web_provision_groups[] duplicated literal of the manager group
scripts/signup-drill.sh MANAGER_HINT hardcoded copy, and this script is the tranche's own gate

The last two are the dangerous ones. MANAGER_HINT only feeds a remediation hint, but a stale value would make the completion gate report a broken delegation while the delegation was fine.

Zero sign-up downtime

entry_managed_by is single-valued, so moving the delegation would normally revoke the old account's access instantly. Instead both service accounts were put in the new manager group first, so the portal's existing token kept working while the new one was swapped in. The old account stays a working fallback until it is deleted.

Three stale comments corrected — all actively misleading

  • "The delegation is NOT declarative." It is, and has been since the carried patch (kanidm_provision_image_tag: v1.3.0-gitborg1) made kanidm-state.json.j2 emit entryManagedBy — whose own comment says it "retires the manual kanidm group set-entry-manager step". Three places repeated the old claim.
  • "tier_basic and ent_renovate" as the delegated pair. Verified live: four groups carried the delegation (tier_participant, tier_trial, forgejo_users, ent_renovate), and tier_basic is not among them — it is the group ADR 0029 moved, whose stale duplicate broke sign-up for 13 days. Acting on that comment would have left trial sign-ups and forgejo_users broken.
  • The Loki alert annotation named the old group, so an operator debugging a sign-up outage would have checked the wrong entry_managed_by. That one is user-facing during an incident.

What genuinely is not declarative is now stated precisely: the manager group and the service account are never created by Ansible — the group appears nowhere in kanidm_groups, and kanidm-provision's serde silently discards the service-accounts block (upstream PR #29). A rebuild must create both by hand first.

Verification

Baseline before: changed=0 on both hosts, and a passing signup:drill as a known-good reference.

Apply: changed=3, failed=0 — state file, provision-token podman secret, container restart. Provision entitlement groups reported no change, confirming the manual Kanidm prep had already reached the declared state.

  • signup:drill PASSES end to end (person → tier_participant → tier_trial → forgejo_users → reset intent → cleanup 200)
  • The drill reads the token from the vault, so it proves the token's privileges but not that the container uses it. Closed separately: podman secret rewritten 08:19:46.23Z, container started 08:19:56.53Z — 10s later — and reports running healthy.
  • Portal / 200, /account 302; all health-check gates passed

Vault re-committed encrypted; the diff contains no non-hex lines.

Follow-up

The rename plan (bitborg-internal #92) repeats the "not declarative" claim, which I took from these comments before verifying it. Correcting that separately.

ADR 0039 §1b, identifiers ② and ③ — the coupled half, completing the tranche. Applied and verified on prod 2026-08-10. `gitborg-web-provision` → `bitborg-web-provision`, and `idm_gitborg_ent_managers` → `idm_bitborg_ent_managers`. ## Four sites, two of which the plan did not list | Site | Note | | --- | --- | | `group_vars/all/vars.yml` `kanidm_web_entry_manager_group` | the variable | | `roles/kanidm/defaults/main.yml` `kanidm_web_provision_account` | the variable | | `roles/kanidm/defaults/main.yml` `kanidm_web_provision_groups[]` | **duplicated literal** of the manager group | | `scripts/signup-drill.sh` `MANAGER_HINT` | **hardcoded copy**, and this script is the tranche's own gate | The last two are the dangerous ones. `MANAGER_HINT` only feeds a remediation hint, but a stale value would make the completion gate report a broken delegation while the delegation was fine. ## Zero sign-up downtime `entry_managed_by` is single-valued, so moving the delegation would normally revoke the old account's access instantly. Instead **both** service accounts were put in the new manager group first, so the portal's existing token kept working while the new one was swapped in. The old account stays a working fallback until it is deleted. ## Three stale comments corrected — all actively misleading - **"The delegation is NOT declarative."** It *is*, and has been since the carried patch (`kanidm_provision_image_tag: v1.3.0-gitborg1`) made `kanidm-state.json.j2` emit `entryManagedBy` — whose own comment says it "retires the manual `kanidm group set-entry-manager` step". Three places repeated the old claim. - **"`tier_basic` and `ent_renovate`"** as the delegated pair. Verified live: **four** groups carried the delegation (`tier_participant`, `tier_trial`, `forgejo_users`, `ent_renovate`), and `tier_basic` is not among them — it is the group ADR 0029 moved, whose stale duplicate broke sign-up for 13 days. Acting on that comment would have left trial sign-ups and `forgejo_users` broken. - The **Loki alert annotation** named the old group, so an operator debugging a sign-up outage would have checked the wrong `entry_managed_by`. That one is user-facing during an incident. What genuinely is *not* declarative is now stated precisely: the manager group and the service account are never created by Ansible — the group appears nowhere in `kanidm_groups`, and kanidm-provision's serde silently discards the `service-accounts` block (upstream PR #29). A rebuild must create both by hand first. ## Verification Baseline before: `changed=0` on both hosts, and a passing `signup:drill` as a known-good reference. Apply: `changed=3`, `failed=0` — state file, provision-token podman secret, container restart. `Provision entitlement groups` reported **no** change, confirming the manual Kanidm prep had already reached the declared state. - **`signup:drill` PASSES** end to end (person → `tier_participant` → `tier_trial` → `forgejo_users` → reset intent → cleanup 200) - The drill reads the token from the vault, so it proves the token's privileges but not that the container uses it. Closed separately: podman secret rewritten `08:19:46.23Z`, container started `08:19:56.53Z` — 10s later — and reports `running healthy`. - Portal `/` 200, `/account` 302; all `health-check` gates passed Vault re-committed encrypted; the diff contains no non-hex lines. ## Follow-up The rename plan (bitborg-internal #92) repeats the "not declarative" claim, which I took from these comments before verifying it. Correcting that separately.
supernaut lade till 1 incheckning 2026-08-10 08:22:00 +00:00
feat(kanidm): rename the provision account and entry-manager group
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m41s
550806f381
supernaut sammanfogade incheckning d823025aab till main 2026-08-10 09:03:09 +00:00
supernaut tog bort grenen feat/rename-1b-sa-and-manager-group 2026-08-10 09:03:09 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!407
Ingen beskrivning angiven.