kanidm: decide between portal-driven and Kanidm-native account recovery #338

Öppen
öppnade 2026-08-02 12:30:49 +00:00 av supernaut · 4 kommentarer
Ägare

Why

We now have two candidate self-service recovery flows for the same job, and only one is a decision.

What we run today: the portal exposes an anonymous "resend my setup link" endpoint. The caller
supplies a username; the portal mints a Kanidm credential-reset intent and mails the link to the
address already on file. Captcha- and rate-limit-gated, non-enumerable by design.

What the pinned Kanidm (1.10.4) also offers: a built-in account-recovery flow, off by default:

kanidm system domain set-allow-account-recovery true

Per the upstream book, when enabled "users will be able to follow the account recovery link from the
login page and have a credential reset link sent to their email. Users must prove knowledge of one
of their account email addresses to proceed", behind a browser proof-of-work challenge. It keys on
email; ours keys on username. It requires outbound mail from Kanidm, which needs the
separate kanidm-mail-sender component — not deployed here.

Check the live state first. The sign-in page currently renders a "Recover Account" link, so the
domain setting appears to be enabled in production — but it is not declared anywhere in the Ansible
role. Establish whether that is deliberate, and bring it under configuration management either way,
before deciding anything else here.

Leaving this undecided means we keep maintaining auth-adjacent security code that upstream also
maintains, while a user who has forgotten their username still has no route back in.

Options

  1. Keep portal-only. One branded mail channel on our chosen EU provider; no new container.
    Cost: we own the security-sensitive flow, and username-only recovery strands anyone who forgets
    it.
  2. Enable Kanidm's. Upstream owns the flow; recovery keys on email, which users actually
    remember. Cost: a new kanidm-mail-sender container, SMTP relay credentials, a second sending
    identity — and mail then originates from Kanidm rather than the portal, which must be checked
    against ADR 0011 (external services / sovereignty) rather than assumed acceptable.
  3. Both, keyed differently — most forgiving for users, largest surface.

Done when

The live state of the domain setting is known and declared in the role, an option is chosen and
recorded (ADR or a note in the runbook's Kanidm section), and the sovereignty question in option 2
is answered explicitly either way.

Part of gitborg/gitborg-docs#69.

## Why We now have two candidate self-service recovery flows for the same job, and only one is a decision. **What we run today:** the portal exposes an anonymous "resend my setup link" endpoint. The caller supplies a **username**; the portal mints a Kanidm credential-reset intent and mails the link to the address already on file. Captcha- and rate-limit-gated, non-enumerable by design. **What the pinned Kanidm (1.10.4) also offers:** a built-in account-recovery flow, off by default: kanidm system domain set-allow-account-recovery true Per the upstream book, when enabled "users will be able to follow the account recovery link from the login page and have a credential reset link sent to their email. Users must prove knowledge of one of their account email addresses to proceed", behind a browser proof-of-work challenge. It keys on **email**; ours keys on **username**. It requires outbound mail from Kanidm, which needs the separate `kanidm-mail-sender` component — not deployed here. **Check the live state first.** The sign-in page currently renders a "Recover Account" link, so the domain setting appears to be enabled in production — but it is not declared anywhere in the Ansible role. Establish whether that is deliberate, and bring it under configuration management either way, before deciding anything else here. Leaving this undecided means we keep maintaining auth-adjacent security code that upstream also maintains, while a user who has forgotten their *username* still has no route back in. ## Options 1. **Keep portal-only.** One branded mail channel on our chosen EU provider; no new container. Cost: we own the security-sensitive flow, and username-only recovery strands anyone who forgets it. 2. **Enable Kanidm's.** Upstream owns the flow; recovery keys on email, which users actually remember. Cost: a new `kanidm-mail-sender` container, SMTP relay credentials, a second sending identity — and mail then originates from Kanidm rather than the portal, which must be checked against ADR 0011 (external services / sovereignty) rather than assumed acceptable. 3. **Both**, keyed differently — most forgiving for users, largest surface. ## Done when The live state of the domain setting is known and declared in the role, an option is chosen and recorded (ADR or a note in the runbook's Kanidm section), and the sovereignty question in option 2 is answered explicitly either way. Part of gitborg/gitborg-docs#69.
Upphovsperson
Ägare

Verified against production on 2026-08-02, without credentials:

  • The flow is live. GET https://auth.gitborg.se/ui/recover returns 200 with
    <title>Account Recovery</title>. It is a real route, not a fallback — an unknown path such as
    /ui/nonexistent-xyz returns 404. The sign-in page links to it as <a href="/ui/recover">Recover Account</a>.
  • It is undeclared. Grepping ansible/roles/kanidm/ for allow.account.recovery /
    account_recovery returns nothing. The only related string anywhere in this repo is a comment in
    roles/kanidm/templates/override.css.j2.
  • kanidm-mail-sender is not deployed, and upstream's flow ends by emailing a credential-reset
    link. So the flow may terminate with nothing sent.
  • Requests to /ui/recover are not logged. Kanidm logs /ui/reset requests (verified: probe
    requests appear in the log store within a minute), but there are zero /ui/recover lines over
    14 days, including probes that returned 200. So an attempt that silently fails leaves no
    server-side trace at all — no log line, no metric.

That last point matters for how this is triaged: log inspection cannot tell us whether anyone has
tried to recover an account, or whether it worked. The only available evidence is whether mail
actually arrives for a test account.

Suggested order of work:

  1. Read the live value (kanidm system domain show) and establish whether enabling it was
    deliberate.
  2. Trigger the flow once with a test identity and observe whether mail arrives. This is the only
    test that distinguishes "functional but undocumented" from "linked dead end".
  3. If no mail arrives, disable it as an immediate mitigation — a linked, unmonitored dead end on the
    sign-in page is worse than no link — then decide the portal-vs-native question below.
  4. Declare the resulting state in the role either way, the way account policy is already asserted in
    roles/kanidm/defaults/main.yml.
Verified against production on 2026-08-02, without credentials: - **The flow is live.** `GET https://auth.gitborg.se/ui/recover` returns **200** with `<title>Account Recovery</title>`. It is a real route, not a fallback — an unknown path such as `/ui/nonexistent-xyz` returns 404. The sign-in page links to it as `<a href="/ui/recover">Recover Account</a>`. - **It is undeclared.** Grepping `ansible/roles/kanidm/` for `allow.account.recovery` / `account_recovery` returns nothing. The only related string anywhere in this repo is a comment in `roles/kanidm/templates/override.css.j2`. - **`kanidm-mail-sender` is not deployed**, and upstream's flow ends by emailing a credential-reset link. So the flow may terminate with nothing sent. - **Requests to `/ui/recover` are not logged.** Kanidm logs `/ui/reset` requests (verified: probe requests appear in the log store within a minute), but there are **zero** `/ui/recover` lines over 14 days, including probes that returned 200. So an attempt that silently fails leaves no server-side trace at all — no log line, no metric. That last point matters for how this is triaged: log inspection cannot tell us whether anyone has tried to recover an account, or whether it worked. The only available evidence is whether mail actually arrives for a test account. Suggested order of work: 1. Read the live value (`kanidm system domain show`) and establish whether enabling it was deliberate. 2. Trigger the flow once with a test identity and observe whether mail arrives. This is the only test that distinguishes "functional but undocumented" from "linked dead end". 3. If no mail arrives, disable it as an immediate mitigation — a linked, unmonitored dead end on the sign-in page is worse than no link — then decide the portal-vs-native question below. 4. Declare the resulting state in the role either way, the way account policy is already asserted in `roles/kanidm/defaults/main.yml`.
Upphovsperson
Ägare

Follow-up: reading the domain configuration is itself a problem, and it is probably the cause.

kanidm system domain show as idm_admin returns EmptyResponse. Domain operations are a
system_admins action while idm_admin is in idm_admins, and Kanidm filters entries the caller
cannot read rather than returning an authorisation error — so an empty response is the expected
symptom of insufficient read access, not a broken command.

Getting a system_admins session is expensive here. The vault holds
vault_kanidm_idm_admin_password and three service tokens, and no admin credential at all;
roles/kanidm/tasks/main.yml:62 states plainly that admin credentials are recovered out-of-band.
So inspecting domain configuration requires kanidmd recover-account admin, which resets that
account's credential
.

Two consequences worth folding into this issue's scope:

  1. This is very likely how the drift happened. Someone recovered admin once, set
    set-allow-account-recovery, and there was nowhere for the setting to be recorded. Nothing would
    ever surface it, because checking costs a credential reset.
  2. The fix should not stop at declaring this one setting. Domain-level configuration should be
    asserted by the role the way account policy already is in
    roles/kanidm/defaults/main.yml:216-241, so that drift is both visible and cheap to detect.
    Otherwise the next undeclared domain setting is equally invisible and equally expensive to audit.

Note also that the flag's current value did not require a login to establish: the sign-in page
renders the recovery anchor only when it is enabled, and that anchor is present. Confirming it via
domain show is therefore optional — but auditing the rest of the domain entry is not currently
possible without the credential reset described above.

Follow-up: reading the domain configuration is itself a problem, and it is probably the cause. `kanidm system domain show` as `idm_admin` returns `EmptyResponse`. Domain operations are a `system_admins` action while `idm_admin` is in `idm_admins`, and Kanidm filters entries the caller cannot read rather than returning an authorisation error — so an empty response is the expected symptom of insufficient read access, not a broken command. Getting a `system_admins` session is expensive here. The vault holds `vault_kanidm_idm_admin_password` and three service tokens, and no `admin` credential at all; `roles/kanidm/tasks/main.yml:62` states plainly that admin credentials are recovered out-of-band. So inspecting domain configuration requires `kanidmd recover-account admin`, which **resets that account's credential**. Two consequences worth folding into this issue's scope: 1. **This is very likely how the drift happened.** Someone recovered `admin` once, set `set-allow-account-recovery`, and there was nowhere for the setting to be recorded. Nothing would ever surface it, because checking costs a credential reset. 2. **The fix should not stop at declaring this one setting.** Domain-level configuration should be asserted by the role the way account policy already is in `roles/kanidm/defaults/main.yml:216-241`, so that drift is both visible and cheap to detect. Otherwise the next undeclared domain setting is equally invisible and equally expensive to audit. Note also that the flag's current value did not require a login to establish: the sign-in page renders the recovery anchor only when it is enabled, and that anchor is present. Confirming it via `domain show` is therefore optional — but auditing the *rest* of the domain entry is not currently possible without the credential reset described above.
Upphovsperson
Ägare

Read succeeded as admin (a system_admins session; the credential exists out-of-band, so no
credential reset was needed after all). The flag is confirmed:

domain_allow_account_recovery: true
domain_display_name: bitborg auth
domain_name: auth.gitborg.se
image: 01f64f15c1709de77f3dbd7afc5421641d6eed1512ed20386412710562740738
version: 14

The scope is wider than this one setting: the domain entry is entirely unmanaged.
roles/kanidm/templates/kanidm-state.json.j2 declares only persons, groups and systems
(OAuth2 clients). It has no domain-level section, and the string bitborg auth does not appear
anywhere in this repository. So all three operator-settable fields above are live, hand-set and
recorded nowhere:

Live value Declared Controls
domain_allow_account_recovery: true no the "Recover Account" link on the sign-in page
domain_display_name: bitborg auth no the branding shown on every sign-in page
image: 01f64f15… no the site logo

The latter two are user-visible branding on the identity provider. They live in the database, so a
restore from backup carries them — but a rebuild from scratch would silently produce an unbranded
sign-in page, and nothing in the health gate would notice. The concealment gate added for ADR 0038
checks the CSS rules, not these.

Revised scope for this issue:

  • Declare the domain entry in the provisioning state alongside persons/groups/systems — display
    name, image and the recovery flag — so all three are asserted rather than remembered. Confirm
    first whether the provisioning tool supports a domain-level section; if it does not, an
    idempotent task using the admin credential is the fallback.
  • Extend the health gate to assert the sign-in page still carries the expected display name and
    image, the same way the concealment rules are gated.
  • Then take the portal-vs-native recovery decision below, with the flag under management either
    way.

Still outstanding and unanswerable from configuration alone: whether the recovery flow actually
sends mail, given kanidm-mail-sender is not deployed. /ui/recover produces no log lines, so a
test with a real address is the only way to find out.

Read succeeded as `admin` (a `system_admins` session; the credential exists out-of-band, so no credential reset was needed after all). The flag is confirmed: ```text domain_allow_account_recovery: true domain_display_name: bitborg auth domain_name: auth.gitborg.se image: 01f64f15c1709de77f3dbd7afc5421641d6eed1512ed20386412710562740738 version: 14 ``` **The scope is wider than this one setting: the domain entry is entirely unmanaged.** `roles/kanidm/templates/kanidm-state.json.j2` declares only `persons`, `groups` and `systems` (OAuth2 clients). It has no domain-level section, and the string `bitborg auth` does not appear anywhere in this repository. So all three operator-settable fields above are live, hand-set and recorded nowhere: | Live value | Declared | Controls | | --- | --- | --- | | `domain_allow_account_recovery: true` | no | the "Recover Account" link on the sign-in page | | `domain_display_name: bitborg auth` | no | the branding shown on every sign-in page | | `image: 01f64f15…` | no | the site logo | The latter two are user-visible branding on the identity provider. They live in the database, so a restore from backup carries them — but a rebuild from scratch would silently produce an unbranded sign-in page, and nothing in the health gate would notice. The concealment gate added for ADR 0038 checks the CSS rules, not these. Revised scope for this issue: - [ ] Declare the domain entry in the provisioning state alongside persons/groups/systems — display name, image and the recovery flag — so all three are asserted rather than remembered. Confirm first whether the provisioning tool supports a domain-level section; if it does not, an idempotent task using the `admin` credential is the fallback. - [ ] Extend the health gate to assert the sign-in page still carries the expected display name and image, the same way the concealment rules are gated. - [ ] Then take the portal-vs-native recovery decision below, with the flag under management either way. Still outstanding and unanswerable from configuration alone: whether the recovery flow actually sends mail, given `kanidm-mail-sender` is not deployed. `/ui/recover` produces no log lines, so a test with a real address is the only way to find out.
Upphovsperson
Ägare

Resolved for now: the flag is disabled in production (2026-08-02).

The mail test settled it. A recovery attempt for a test identity produced no mail, and the
supporting evidence is consistent: no kanidm-mail-sender container runs on the host, and there is
no SMTP, relay or mail configuration anywhere in the role's templates — server.toml.j2 mentions
email only in a comment about the WebAuthn/OAuth2 security context. So the flow could not have sent
mail for any address. It was a linked dead end on the sign-in page, and because the endpoint emits
no log lines, nothing would ever have surfaced it.

kanidm system domain set-allow-account-recovery false has been run. Verified after the change:

  • the /ui/recover anchor is gone from the sign-in page;
  • domain branding is unaffected (bitborg auth still renders);
  • but GET /ui/recover still returns 200 with <title>Account Recovery</title>.

That last point is a residual worth closing rather than assuming: disabling the flag unadvertises
the page, it does not unroute it, so a bookmarked or cached URL still reaches it. Whether the POST
handler now refuses submissions with the flag off has not been tested — upstream probably gates
it on the same domain_info check that gates the anchor, but that is an inference. "Link hidden,
endpoint still accepts and silently drops" would be worse than the original state, since there would
be neither an affordance nor a signal.

Remaining scope for this issue is unchanged and now better motivated:

  • Declare the domain entry in provisioning — display name, image, and this flag — so the value
    is asserted rather than remembered, and so re-enabling is a reviewed change rather than a
    recalled command.
  • Confirm the POST handler refuses while the flag is off.
  • Extend the health gate to assert the recovery anchor stays absent and the branding stays
    present.
  • Take the portal-vs-native recovery decision. Adopting the native flow now has a hard
    prerequisite: kanidm-mail-sender must be deployed first, with the ADR 0011 sovereignty
    question answered, or the same dead end returns.

No user-facing capability was lost by disabling it: the portal's username-keyed "resend setup link"
remains the working recovery path.

**Resolved for now: the flag is disabled in production (2026-08-02).** The mail test settled it. A recovery attempt for a test identity produced no mail, and the supporting evidence is consistent: no `kanidm-mail-sender` container runs on the host, and there is no SMTP, relay or mail configuration anywhere in the role's templates — `server.toml.j2` mentions email only in a comment about the WebAuthn/OAuth2 security context. So the flow could not have sent mail for any address. It was a linked dead end on the sign-in page, and because the endpoint emits no log lines, nothing would ever have surfaced it. `kanidm system domain set-allow-account-recovery false` has been run. Verified after the change: - the `/ui/recover` anchor is **gone** from the sign-in page; - domain branding is unaffected (`bitborg auth` still renders); - **but `GET /ui/recover` still returns 200** with `<title>Account Recovery</title>`. That last point is a residual worth closing rather than assuming: disabling the flag unadvertises the page, it does not unroute it, so a bookmarked or cached URL still reaches it. Whether the POST handler now refuses submissions with the flag off has **not** been tested — upstream probably gates it on the same `domain_info` check that gates the anchor, but that is an inference. "Link hidden, endpoint still accepts and silently drops" would be worse than the original state, since there would be neither an affordance nor a signal. Remaining scope for this issue is unchanged and now better motivated: - [ ] Declare the domain entry in provisioning — display name, image, and this flag — so the value is asserted rather than remembered, and so re-enabling is a reviewed change rather than a recalled command. - [ ] Confirm the POST handler refuses while the flag is off. - [ ] Extend the health gate to assert the recovery anchor stays absent and the branding stays present. - [ ] Take the portal-vs-native recovery decision. Adopting the native flow now has a hard prerequisite: `kanidm-mail-sender` must be deployed first, with the ADR 0011 sovereignty question answered, or the same dead end returns. No user-facing capability was lost by disabling it: the portal's username-keyed "resend setup link" remains the working recovery path.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#338
Ingen beskrivning angiven.