kanidm: decide between portal-driven and Kanidm-native account recovery #338
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#338
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Why
We now have two candidate self-service recovery flows for the same job, and only one is a decision.
What we run today: the portal exposes an anonymous "resend my setup link" endpoint. The caller
supplies a username; the portal mints a Kanidm credential-reset intent and mails the link to the
address already on file. Captcha- and rate-limit-gated, non-enumerable by design.
What the pinned Kanidm (1.10.4) also offers: a built-in account-recovery flow, off by default:
Per the upstream book, when enabled "users will be able to follow the account recovery link from the
login page and have a credential reset link sent to their email. Users must prove knowledge of one
of their account email addresses to proceed", behind a browser proof-of-work challenge. It keys on
email; ours keys on username. It requires outbound mail from Kanidm, which needs the
separate
kanidm-mail-sendercomponent — not deployed here.Check the live state first. The sign-in page currently renders a "Recover Account" link, so the
domain setting appears to be enabled in production — but it is not declared anywhere in the Ansible
role. Establish whether that is deliberate, and bring it under configuration management either way,
before deciding anything else here.
Leaving this undecided means we keep maintaining auth-adjacent security code that upstream also
maintains, while a user who has forgotten their username still has no route back in.
Options
Cost: we own the security-sensitive flow, and username-only recovery strands anyone who forgets
it.
remember. Cost: a new
kanidm-mail-sendercontainer, SMTP relay credentials, a second sendingidentity — and mail then originates from Kanidm rather than the portal, which must be checked
against ADR 0011 (external services / sovereignty) rather than assumed acceptable.
Done when
The live state of the domain setting is known and declared in the role, an option is chosen and
recorded (ADR or a note in the runbook's Kanidm section), and the sovereignty question in option 2
is answered explicitly either way.
Part of gitborg/gitborg-docs#69.
supernaut refererade till detta ärende från bitborg/bitborg-docs2026-08-02 12:32:47 +00:00
Verified against production on 2026-08-02, without credentials:
GET https://auth.gitborg.se/ui/recoverreturns 200 with<title>Account Recovery</title>. It is a real route, not a fallback — an unknown path such as/ui/nonexistent-xyzreturns 404. The sign-in page links to it as<a href="/ui/recover">Recover Account</a>.ansible/roles/kanidm/forallow.account.recovery/account_recoveryreturns nothing. The only related string anywhere in this repo is a comment inroles/kanidm/templates/override.css.j2.kanidm-mail-senderis not deployed, and upstream's flow ends by emailing a credential-resetlink. So the flow may terminate with nothing sent.
/ui/recoverare not logged. Kanidm logs/ui/resetrequests (verified: proberequests appear in the log store within a minute), but there are zero
/ui/recoverlines over14 days, including probes that returned 200. So an attempt that silently fails leaves no
server-side trace at all — no log line, no metric.
That last point matters for how this is triaged: log inspection cannot tell us whether anyone has
tried to recover an account, or whether it worked. The only available evidence is whether mail
actually arrives for a test account.
Suggested order of work:
kanidm system domain show) and establish whether enabling it wasdeliberate.
test that distinguishes "functional but undocumented" from "linked dead end".
sign-in page is worse than no link — then decide the portal-vs-native question below.
roles/kanidm/defaults/main.yml.Follow-up: reading the domain configuration is itself a problem, and it is probably the cause.
kanidm system domain showasidm_adminreturnsEmptyResponse. Domain operations are asystem_adminsaction whileidm_adminis inidm_admins, and Kanidm filters entries the callercannot read rather than returning an authorisation error — so an empty response is the expected
symptom of insufficient read access, not a broken command.
Getting a
system_adminssession is expensive here. The vault holdsvault_kanidm_idm_admin_passwordand three service tokens, and noadmincredential at all;roles/kanidm/tasks/main.yml:62states plainly that admin credentials are recovered out-of-band.So inspecting domain configuration requires
kanidmd recover-account admin, which resets thataccount's credential.
Two consequences worth folding into this issue's scope:
adminonce, setset-allow-account-recovery, and there was nowhere for the setting to be recorded. Nothing wouldever surface it, because checking costs a credential reset.
asserted by the role the way account policy already is in
roles/kanidm/defaults/main.yml:216-241, so that drift is both visible and cheap to detect.Otherwise the next undeclared domain setting is equally invisible and equally expensive to audit.
Note also that the flag's current value did not require a login to establish: the sign-in page
renders the recovery anchor only when it is enabled, and that anchor is present. Confirming it via
domain showis therefore optional — but auditing the rest of the domain entry is not currentlypossible without the credential reset described above.
Read succeeded as
admin(asystem_adminssession; the credential exists out-of-band, so nocredential reset was needed after all). The flag is confirmed:
The scope is wider than this one setting: the domain entry is entirely unmanaged.
roles/kanidm/templates/kanidm-state.json.j2declares onlypersons,groupsandsystems(OAuth2 clients). It has no domain-level section, and the string
bitborg authdoes not appearanywhere in this repository. So all three operator-settable fields above are live, hand-set and
recorded nowhere:
domain_allow_account_recovery: truedomain_display_name: bitborg authimage: 01f64f15…The latter two are user-visible branding on the identity provider. They live in the database, so a
restore from backup carries them — but a rebuild from scratch would silently produce an unbranded
sign-in page, and nothing in the health gate would notice. The concealment gate added for ADR 0038
checks the CSS rules, not these.
Revised scope for this issue:
name, image and the recovery flag — so all three are asserted rather than remembered. Confirm
first whether the provisioning tool supports a domain-level section; if it does not, an
idempotent task using the
admincredential is the fallback.image, the same way the concealment rules are gated.
way.
Still outstanding and unanswerable from configuration alone: whether the recovery flow actually
sends mail, given
kanidm-mail-senderis not deployed./ui/recoverproduces no log lines, so atest with a real address is the only way to find out.
Resolved for now: the flag is disabled in production (2026-08-02).
The mail test settled it. A recovery attempt for a test identity produced no mail, and the
supporting evidence is consistent: no
kanidm-mail-sendercontainer runs on the host, and there isno SMTP, relay or mail configuration anywhere in the role's templates —
server.toml.j2mentionsemail only in a comment about the WebAuthn/OAuth2 security context. So the flow could not have sent
mail for any address. It was a linked dead end on the sign-in page, and because the endpoint emits
no log lines, nothing would ever have surfaced it.
kanidm system domain set-allow-account-recovery falsehas been run. Verified after the change:/ui/recoveranchor is gone from the sign-in page;bitborg authstill renders);GET /ui/recoverstill returns 200 with<title>Account Recovery</title>.That last point is a residual worth closing rather than assuming: disabling the flag unadvertises
the page, it does not unroute it, so a bookmarked or cached URL still reaches it. Whether the POST
handler now refuses submissions with the flag off has not been tested — upstream probably gates
it on the same
domain_infocheck that gates the anchor, but that is an inference. "Link hidden,endpoint still accepts and silently drops" would be worse than the original state, since there would
be neither an affordance nor a signal.
Remaining scope for this issue is unchanged and now better motivated:
is asserted rather than remembered, and so re-enabling is a reviewed change rather than a
recalled command.
present.
prerequisite:
kanidm-mail-sendermust be deployed first, with the ADR 0011 sovereigntyquestion answered, or the same dead end returns.
No user-facing capability was lost by disabling it: the portal's username-keyed "resend setup link"
remains the working recovery path.