fix(kanidm): gate, alert and drill the sign-up delegation that broke for 13 days #262

Sammanfogat
supernaut sammanfogade 5 incheckningar från fix/signup-delegation-gate in i main 2026-07-30 18:36:27 +00:00
Ägare

Sign-up was broken in production from 2026-07-17 to 2026-07-30. Every new-user registration
failed and nothing noticed — it surfaced only because someone tried to register and said so.

Root cause

The portal adds each sign-up to a Kanidm tier group. Member-write on those groups is delegated to
idm_bitborg_ent_managers via entry-managed-by, and that delegation is not declarative —
kanidm-provision (pinned) has no ACP support, so it is a manual kanidm group set-entry-manager step.

ADR 0029 moved sign-up from tier_basic to tier_participant; ADR 0036 added tier_trial. The
delegation followed neither.
So the portal was denied member-write on the only two groups it writes:

group delegated portal writes it
tier_basic ✅ no — superseded by ADR 0029
ent_renovate ✅ yes
tier_participant ❌ every non-trial sign-up
tier_trial ❌ every trial sign-up

tier_participant.member was null with three persons in the system — zero successful adds, ever.
Sign-up created the person, failed the group add, and rolled back, so there is no orphaned data.

Two things that made this hard to see

  • Kanidm returns 404 for a denied write, not 403. The app logged group add -> 404, which reads
    as "missing group" and sent diagnosis in the wrong direction.
  • GET /v1/group/<g> returns HTTP 200 with a null body when the caller cannot manage the entry.
    A status-code-only probe reads as healthy — a mistake made during this very diagnosis, which is
    why the health gate below asserts on the body.

Mitigations, in the order they would have caught it

1. Apply-time health gate (health-check role, tagged always, so even a scoped apply runs it).
Asserts the portal's own token can manage every group it writes — derived from the same
web_kanidm_*_group vars the container is handed, so a tier-group change can never again outrun the
delegation. On failure it prints the exact set-entry-manager command per group. Read-only; needs no
idm_admin.

2. SignupProvisioningFailing — critical Loki alert on [signup] Kanidm in the web container,
for: 0m. There is no acceptable rate of sign-up failure and no self-healing path — the visitor just
sees an error — so one occurrence pages. Previously nothing watched this.

3. scripts/signup-drill.sh (pnpm signup:drill) — proves the chain for real, the way
backup-drill (ADR 0027) does for restores: creates a synthetic person with the portal's own
token
, adds it to the tier, trial and forgejo_users groups, mints a reset intent, deletes it.
A capability check can pass while a write still fails; only a write proves a write. Reads the group
names from group_vars, so it cannot drill a different group than production uses.

4. Removed kanidm_web_default_tier_group: "tier_basic" — referenced nowhere, documenting the
pre-ADR-0029 truth. A stale duplicate of a value that must match a manual delegation is precisely the
shape of this bug; there is now one source of truth.

Why not simply make the delegation declarative? kanidm-provision is pinned and has no ACP support, so
it genuinely cannot be. Given that, the next best thing is to make its absence impossible to ship
unnoticed
— hence gate + alert + drill rather than a fourth place to restate the group name.

Why unit tests and local e2e could not have caught this

Neither can: the failure is a production permission/config mismatch, not logic, and there is no Kanidm
in the local harness. That is exactly why the drill exists.

Runbook

New "Sign-up is broken (group add -> 404)" section covering both traps and the fix.

Note for reviewers

The health gate's shell body contains no pipe character anywhere, deliberately: ansible-lint's
risky-shell-pipe matches that character textually, so a Jinja join filter, a case alternation
and a shell or-list all trip it — and /bin/sh has no pipefail to satisfy it with. Hence the list
arriving via the environment, ; true instead of an or-list, and the flag-variable style.

--syntax-check passes; ansible-lint clean at the production profile across all three touched
roles; shellcheck clean on the drill.

Still required in production

The delegation itself must be restored as idm_admin — this PR detects and drills it, it does not
grant it:

kanidm group set-entry-manager tier_participant idm_bitborg_ent_managers --name idm_admin
kanidm group set-entry-manager tier_trial       idm_bitborg_ent_managers --name idm_admin

Then pnpm signup:drill should pass, and the health gate will confirm it on every subsequent apply.

Sign-up was broken in production **from 2026-07-17 to 2026-07-30**. Every new-user registration failed and nothing noticed — it surfaced only because someone tried to register and said so. ## Root cause The portal adds each sign-up to a Kanidm tier group. Member-write on those groups is delegated to `idm_bitborg_ent_managers` via `entry-managed-by`, and that delegation is **not declarative** — kanidm-provision (pinned) has no ACP support, so it is a manual `kanidm group set-entry-manager` step. ADR 0029 moved sign-up from `tier_basic` to `tier_participant`; ADR 0036 added `tier_trial`. **The delegation followed neither.** So the portal was denied member-write on the only two groups it writes: | group | delegated | portal writes it | | --- | --- | --- | | `tier_basic` | ✅ | no — superseded by ADR 0029 | | `ent_renovate` | ✅ | yes | | **`tier_participant`** | ❌ | **every non-trial sign-up** | | **`tier_trial`** | ❌ | **every trial sign-up** | `tier_participant.member` was `null` with three persons in the system — zero successful adds, ever. Sign-up created the person, failed the group add, and rolled back, so there is no orphaned data. ## Two things that made this hard to see - **Kanidm returns 404 for a denied write, not 403.** The app logged `group add -> 404`, which reads as "missing group" and sent diagnosis in the wrong direction. - **`GET /v1/group/<g>` returns HTTP 200 with a `null` body** when the caller cannot manage the entry. A status-code-only probe reads as healthy — a mistake made *during this very diagnosis*, which is why the health gate below asserts on the **body**. ## Mitigations, in the order they would have caught it **1. Apply-time health gate** (`health-check` role, tagged `always`, so even a scoped apply runs it). Asserts the portal's own token can *manage* every group it writes — derived from the same `web_kanidm_*_group` vars the container is handed, so a tier-group change can never again outrun the delegation. On failure it prints the exact `set-entry-manager` command per group. Read-only; needs no `idm_admin`. **2. `SignupProvisioningFailing`** — critical Loki alert on `[signup] Kanidm` in the web container, `for: 0m`. There is no acceptable rate of sign-up failure and no self-healing path — the visitor just sees an error — so one occurrence pages. Previously **nothing** watched this. **3. `scripts/signup-drill.sh`** (`pnpm signup:drill`) — proves the chain for real, the way `backup-drill` (ADR 0027) does for restores: creates a synthetic person **with the portal's own token**, adds it to the tier, trial and `forgejo_users` groups, mints a reset intent, deletes it. A capability check can pass while a write still fails; only a write proves a write. Reads the group names from `group_vars`, so it cannot drill a different group than production uses. **4. Removed `kanidm_web_default_tier_group: "tier_basic"`** — referenced nowhere, documenting the pre-ADR-0029 truth. A stale duplicate of a value that must match a manual delegation is precisely the shape of this bug; there is now one source of truth. Why not simply make the delegation declarative? kanidm-provision is pinned and has no ACP support, so it genuinely cannot be. Given that, the next best thing is to make its absence **impossible to ship unnoticed** — hence gate + alert + drill rather than a fourth place to restate the group name. ## Why unit tests and local e2e could not have caught this Neither can: the failure is a production permission/config mismatch, not logic, and there is no Kanidm in the local harness. That is exactly why the drill exists. ## Runbook New **"Sign-up is broken (`group add -> 404`)"** section covering both traps and the fix. ## Note for reviewers The health gate's shell body contains **no pipe character anywhere**, deliberately: ansible-lint's `risky-shell-pipe` matches that character textually, so a Jinja `join` filter, a `case` alternation and a shell or-list all trip it — and `/bin/sh` has no `pipefail` to satisfy it with. Hence the list arriving via the environment, `; true` instead of an or-list, and the flag-variable style. `--syntax-check` passes; `ansible-lint` clean at the `production` profile across all three touched roles; `shellcheck` clean on the drill. ## Still required in production The delegation itself must be restored as `idm_admin` — this PR detects and drills it, it does not grant it: ``` kanidm group set-entry-manager tier_participant idm_bitborg_ent_managers --name idm_admin kanidm group set-entry-manager tier_trial idm_bitborg_ent_managers --name idm_admin ``` Then `pnpm signup:drill` should pass, and the health gate will confirm it on every subsequent apply.
supernaut lade till 1 incheckning 2026-07-30 18:15:57 +00:00
fix(kanidm): gate, alert and drill the sign-up delegation that broke for 13 days
En del kontroller misslyckades
ci / ci (pull_request) Failing after 42s
6c70f4db67
Sign-up was broken in production from 2026-07-17 to 2026-07-30. EVERY new-user
registration failed and nothing noticed; it surfaced only because someone tried
to register and said so.

Root cause: the portal adds each sign-up to a Kanidm tier group, and member-write
on those groups is delegated to idm_gitborg_ent_managers via `entry-managed-by`.
That delegation is NOT declarative — kanidm-provision (pinned) has no ACP support,
so it is a manual `kanidm group set-entry-manager` step. ADR 0029 moved sign-up
from tier_basic to tier_participant and ADR 0036 added tier_trial; the delegation
followed neither. The portal was therefore denied member-write on the only two
groups it actually writes.

Two things made it hard to see:

- Kanidm returns 404 for a denied write, not 403, so the app logged
  `group add -> 404`, which reads as "missing group".
- `GET /v1/group/<g>` returns HTTP 200 with a NULL BODY in that state. A
  status-code-only probe reads as healthy — which is exactly the mistake made
  while diagnosing this. tier_participant.member being null (with three persons
  in the system) is what gave it away.

Mitigations, in the order they would have caught it:

1. Health gate (health-check role, tags: always — so even a scoped apply runs it).
   Asserts the portal's own token can MANAGE every group it writes, derived from
   the same web_kanidm_*_group vars the container is handed, so a tier-group change
   can never again outrun the delegation. Asserts on the body, not the status. On
   failure it prints the exact set-entry-manager command per group. Read-only and
   needs no idm_admin.

2. SignupProvisioningFailing — critical Loki alert on `[signup] Kanidm` in the web
   container, `for: 0m`. There is no acceptable rate of sign-up failure and no
   self-healing path, so one occurrence pages. Previously NOTHING watched this.

3. scripts/signup-drill.sh — proves the chain for real, the way backup-drill does
   for restores: creates a synthetic person with the PORTAL's token, adds it to the
   tier, trial and forgejo_users groups, mints a reset intent, deletes it. A
   capability check can pass while a write still fails; only a write proves a write.
   Reads the group names from group_vars so it can never drill a different group
   than production uses. `pnpm signup:drill`.

4. Removed kanidm_web_default_tier_group ("tier_basic") — referenced nowhere and
   documenting the pre-ADR-0029 truth. A stale duplicate of a value that must match
   a manual delegation is precisely the shape of this bug; there is now one source
   of truth.

Runbook gains a "Sign-up is broken" section covering the 404-means-denied trap,
the 200-with-null-body trap, and the fix.

Note on the health gate's shell body: it contains no pipe character anywhere,
because ansible-lint's risky-shell-pipe matches that character textually — a Jinja
join filter, a `case` alternation and a shell or-list all trip it, and /bin/sh has
no pipefail to satisfy it with.
supernaut lade till 1 incheckning 2026-07-30 18:19:35 +00:00
build: add shellcheck npm script
En del kontroller misslyckades
ci / ci (pull_request) Failing after 48s
e2742fdbe3
supernaut lade till 1 incheckning 2026-07-30 18:23:32 +00:00
fix(scripts): avoid SC2015 in signup-drill — CI's shellcheck is older than local
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m36s
8d4e921b88
CI's "Shellcheck scripts" step failed while `shellcheck scripts/*.sh` passed
locally. Not a flake and not a config difference — a VERSION difference:

  shellcheck 0.11.0 (local, Homebrew)   exit 0
  shellcheck 0.10.0 (CI, Debian apt)    exit 1
  shellcheck 0.9.0                      exit 1

`scripts/bake-runner-image.sh` installs shellcheck from apt, so CI runs whatever
Debian ships. 0.10 and earlier emit SC2015 on

  [ -n "$TIER_GROUP" ] && [ -n "$TRIAL_GROUP" ] || fail "..."

("Note that A && B || C is not if-then-else. C may run when A is true.") 0.11
dropped that check, so the newer local binary was silently more permissive than
CI — the same class of trap as the renovate-config-validator resolving 37.x
earlier today and rejecting config that is valid on the deployed 43.

Rewritten as an explicit `if`, which is clearer anyway and passes on 0.9, 0.10 and
0.11. Reproduced and verified against all three via
`podman run koalaman/shellcheck:<v>`.

Note the semantics were never actually wrong here — `fail` should run whenever
either variable is empty, which is what the original did. SC2015 is a
readability/foot-gun warning, not a correctness finding in this case.
supernaut lade till 1 incheckning 2026-07-30 18:29:15 +00:00
build(ci): pin shellcheck so local and CI cannot silently disagree
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m28s
23a9346564
Today CI's shellcheck step failed on code that linted clean locally. The cause was
a version difference, not config: bake-runner-image.sh installs the linter from
apt, so CI ran Debian's 0.10.x while Homebrew had 0.11.0. Versions <= 0.10 report
SC2015 on `A && B || C`; 0.11 dropped that check. So the newer local binary was
SILENTLY MORE PERMISSIVE than CI — the worst direction for the skew to point,
because "it passes locally" was true and useless.

This is the second version-skew bite of the day: renovate-config-validator
resolved 37.x via npx and rejected config that is valid on the deployed 43. The
general lesson is that a linter's VERSION is part of its configuration, and
neither repo stated it anywhere.

scripts/shellcheck.sh states it in one place:

- Preferred: a pinned koalaman/shellcheck container — byte-identical on any machine
  with podman or docker. Probes /opt/podman/bin/podman explicitly, since Podman
  Desktop on macOS is not on a non-interactive PATH.
- Fallback: the local binary, with a loud warning naming BOTH versions and stating
  that a clean run does not imply CI will pass. The fallback keeps a machine
  without a container runtime working, but never pretends to be authoritative —
  making the skew visible is the entire point.
- SHELLCHECK_NO_CONTAINER=1 forces the local path: an escape hatch for a faster
  loop or a runner without nested podman, and what makes the fallback branch
  testable at all (the absolute-path probe otherwise always succeeds locally).

Both `pnpm shellcheck` and the CI step now call it, so there is one code path.
Pinned to 0.10.0 — what CI runs today — so local matches CI now rather than
matching some future ideal; the header notes that bumping it means bumping apt in
bake-runner-image.sh too, or accepting the container path in CI.

Verified: clean on 0.9.0, 0.10.0 and 0.11.0 via the container; the fallback branch
correctly reports local=0.11.0 vs pinned=0.10.0.

One trap worth recording, hit while writing this: a comment line beginning with
the tool's own name is parsed as a DIRECTIVE, so the header's prose tripped
SC1072/SC1073 until reworded. Same shape as ansible-lint's risky-shell-pipe
matching a comment that merely mentioned a pipe character.
supernaut lade till 1 incheckning 2026-07-30 18:31:42 +00:00
fix(kanidm): the gate must cover forgejo_users too — the drill caught what it missed
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m28s
71485b0d16
The first version of the sign-up health gate deliberately excluded forgejo_users,
with the comment "its delegation has never been in question and adding it would
widen this gate beyond the regression it guards".

That assumption was wrong, and scripts/signup-drill.sh disproved it within minutes
of the delegation being restored:

  [signup-drill] step 2: added to tier_participant
  [signup-drill] step 3: added to tier_trial
  [signup-drill] FAIL: forgejo_users add -> HTTP 404

forgejo_users was undelegated as well. So restoring only the tier groups would have
left sign-up broken one step later, in a quieter way: the person is created and
lands in their tier, then fails to join forgejo_users — which grants the OIDC
scopes for the Forgejo client. Without it a user can authenticate to Kanidm but
every Forgejo login is refused (available_scopes: {}) and the account is never
JIT-created.

This is the case for the drill existing, made concretely. The capability gate as
written would have gone green while sign-up stayed broken, because I had reasoned
about which groups were at risk instead of enumerating which groups are written.
Only performing the actual writes found it.

So the list is now mechanical rather than judged: every group gitborg-web writes
belongs in it. Grep `_attr/member` in gitborg-web/src/lib/kanidm.ts — if a group
appears there and not in health_check_kanidm_portal_groups, the gate has a hole.
That rule is recorded in both the defaults and the runbook.

Production still needs the third delegation:
  kanidm group set-entry-manager forgejo_users idm_gitborg_ent_managers --name idm_admin
supernaut sammanfogade incheckning 6c1b5f08a6 till main 2026-07-30 18:36:27 +00:00
supernaut tog bort grenen fix/signup-delegation-gate 2026-07-30 18:36:27 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!262
Ingen beskrivning angiven.