fix(kanidm): gate, alert and drill the sign-up delegation that broke for 13 days #262
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!262
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/signup-delegation-gate"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Sign-up was broken in production from 2026-07-17 to 2026-07-30. Every new-user registration
failed and nothing noticed — it surfaced only because someone tried to register and said so.
Root cause
The portal adds each sign-up to a Kanidm tier group. Member-write on those groups is delegated to
idm_bitborg_ent_managersviaentry-managed-by, and that delegation is not declarative —kanidm-provision (pinned) has no ACP support, so it is a manual
kanidm group set-entry-managerstep.ADR 0029 moved sign-up from
tier_basictotier_participant; ADR 0036 addedtier_trial. Thedelegation followed neither. So the portal was denied member-write on the only two groups it writes:
tier_basicent_renovatetier_participanttier_trialtier_participant.memberwasnullwith three persons in the system — zero successful adds, ever.Sign-up created the person, failed the group add, and rolled back, so there is no orphaned data.
Two things that made this hard to see
group add -> 404, which readsas "missing group" and sent diagnosis in the wrong direction.
GET /v1/group/<g>returns HTTP 200 with anullbody when the caller cannot manage the entry.A status-code-only probe reads as healthy — a mistake made during this very diagnosis, which is
why the health gate below asserts on the body.
Mitigations, in the order they would have caught it
1. Apply-time health gate (
health-checkrole, taggedalways, so even a scoped apply runs it).Asserts the portal's own token can manage every group it writes — derived from the same
web_kanidm_*_groupvars the container is handed, so a tier-group change can never again outrun thedelegation. On failure it prints the exact
set-entry-managercommand per group. Read-only; needs noidm_admin.2.
SignupProvisioningFailing— critical Loki alert on[signup] Kanidmin the web container,for: 0m. There is no acceptable rate of sign-up failure and no self-healing path — the visitor justsees an error — so one occurrence pages. Previously nothing watched this.
3.
scripts/signup-drill.sh(pnpm signup:drill) — proves the chain for real, the waybackup-drill(ADR 0027) does for restores: creates a synthetic person with the portal's owntoken, adds it to the tier, trial and
forgejo_usersgroups, mints a reset intent, deletes it.A capability check can pass while a write still fails; only a write proves a write. Reads the group
names from
group_vars, so it cannot drill a different group than production uses.4. Removed
kanidm_web_default_tier_group: "tier_basic"— referenced nowhere, documenting thepre-ADR-0029 truth. A stale duplicate of a value that must match a manual delegation is precisely the
shape of this bug; there is now one source of truth.
Why not simply make the delegation declarative? kanidm-provision is pinned and has no ACP support, so
it genuinely cannot be. Given that, the next best thing is to make its absence impossible to ship
unnoticed — hence gate + alert + drill rather than a fourth place to restate the group name.
Why unit tests and local e2e could not have caught this
Neither can: the failure is a production permission/config mismatch, not logic, and there is no Kanidm
in the local harness. That is exactly why the drill exists.
Runbook
New "Sign-up is broken (
group add -> 404)" section covering both traps and the fix.Note for reviewers
The health gate's shell body contains no pipe character anywhere, deliberately: ansible-lint's
risky-shell-pipematches that character textually, so a Jinjajoinfilter, acasealternationand a shell or-list all trip it — and
/bin/shhas nopipefailto satisfy it with. Hence the listarriving via the environment,
; trueinstead of an or-list, and the flag-variable style.--syntax-checkpasses;ansible-lintclean at theproductionprofile across all three touchedroles;
shellcheckclean on the drill.Still required in production
The delegation itself must be restored as
idm_admin— this PR detects and drills it, it does notgrant it:
Then
pnpm signup:drillshould pass, and the health gate will confirm it on every subsequent apply.Sign-up was broken in production from 2026-07-17 to 2026-07-30. EVERY new-user registration failed and nothing noticed; it surfaced only because someone tried to register and said so. Root cause: the portal adds each sign-up to a Kanidm tier group, and member-write on those groups is delegated to idm_gitborg_ent_managers via `entry-managed-by`. That delegation is NOT declarative — kanidm-provision (pinned) has no ACP support, so it is a manual `kanidm group set-entry-manager` step. ADR 0029 moved sign-up from tier_basic to tier_participant and ADR 0036 added tier_trial; the delegation followed neither. The portal was therefore denied member-write on the only two groups it actually writes. Two things made it hard to see: - Kanidm returns 404 for a denied write, not 403, so the app logged `group add -> 404`, which reads as "missing group". - `GET /v1/group/<g>` returns HTTP 200 with a NULL BODY in that state. A status-code-only probe reads as healthy — which is exactly the mistake made while diagnosing this. tier_participant.member being null (with three persons in the system) is what gave it away. Mitigations, in the order they would have caught it: 1. Health gate (health-check role, tags: always — so even a scoped apply runs it). Asserts the portal's own token can MANAGE every group it writes, derived from the same web_kanidm_*_group vars the container is handed, so a tier-group change can never again outrun the delegation. Asserts on the body, not the status. On failure it prints the exact set-entry-manager command per group. Read-only and needs no idm_admin. 2. SignupProvisioningFailing — critical Loki alert on `[signup] Kanidm` in the web container, `for: 0m`. There is no acceptable rate of sign-up failure and no self-healing path, so one occurrence pages. Previously NOTHING watched this. 3. scripts/signup-drill.sh — proves the chain for real, the way backup-drill does for restores: creates a synthetic person with the PORTAL's token, adds it to the tier, trial and forgejo_users groups, mints a reset intent, deletes it. A capability check can pass while a write still fails; only a write proves a write. Reads the group names from group_vars so it can never drill a different group than production uses. `pnpm signup:drill`. 4. Removed kanidm_web_default_tier_group ("tier_basic") — referenced nowhere and documenting the pre-ADR-0029 truth. A stale duplicate of a value that must match a manual delegation is precisely the shape of this bug; there is now one source of truth. Runbook gains a "Sign-up is broken" section covering the 404-means-denied trap, the 200-with-null-body trap, and the fix. Note on the health gate's shell body: it contains no pipe character anywhere, because ansible-lint's risky-shell-pipe matches that character textually — a Jinja join filter, a `case` alternation and a shell or-list all trip it, and /bin/sh has no pipefail to satisfy it with.CI's "Shellcheck scripts" step failed while `shellcheck scripts/*.sh` passed locally. Not a flake and not a config difference — a VERSION difference: shellcheck 0.11.0 (local, Homebrew) exit 0 shellcheck 0.10.0 (CI, Debian apt) exit 1 shellcheck 0.9.0 exit 1 `scripts/bake-runner-image.sh` installs shellcheck from apt, so CI runs whatever Debian ships. 0.10 and earlier emit SC2015 on [ -n "$TIER_GROUP" ] && [ -n "$TRIAL_GROUP" ] || fail "..." ("Note that A && B || C is not if-then-else. C may run when A is true.") 0.11 dropped that check, so the newer local binary was silently more permissive than CI — the same class of trap as the renovate-config-validator resolving 37.x earlier today and rejecting config that is valid on the deployed 43. Rewritten as an explicit `if`, which is clearer anyway and passes on 0.9, 0.10 and 0.11. Reproduced and verified against all three via `podman run koalaman/shellcheck:<v>`. Note the semantics were never actually wrong here — `fail` should run whenever either variable is empty, which is what the original did. SC2015 is a readability/foot-gun warning, not a correctness finding in this case.The first version of the sign-up health gate deliberately excluded forgejo_users, with the comment "its delegation has never been in question and adding it would widen this gate beyond the regression it guards". That assumption was wrong, and scripts/signup-drill.sh disproved it within minutes of the delegation being restored: [signup-drill] step 2: added to tier_participant [signup-drill] step 3: added to tier_trial [signup-drill] FAIL: forgejo_users add -> HTTP 404 forgejo_users was undelegated as well. So restoring only the tier groups would have left sign-up broken one step later, in a quieter way: the person is created and lands in their tier, then fails to join forgejo_users — which grants the OIDC scopes for the Forgejo client. Without it a user can authenticate to Kanidm but every Forgejo login is refused (available_scopes: {}) and the account is never JIT-created. This is the case for the drill existing, made concretely. The capability gate as written would have gone green while sign-up stayed broken, because I had reasoned about which groups were at risk instead of enumerating which groups are written. Only performing the actual writes found it. So the list is now mechanical rather than judged: every group gitborg-web writes belongs in it. Grep `_attr/member` in gitborg-web/src/lib/kanidm.ts — if a group appears there and not in health_check_kanidm_portal_groups, the gate has a hole. That rule is recorded in both the defaults and the runbook. Production still needs the third delegation: kanidm group set-entry-manager forgejo_users idm_gitborg_ent_managers --name idm_admin