kanidm: 1.10.x leaves support around 2026-09-01, upgrade to 1.11.1 #424

Stängd
öppnade 2026-08-16 12:09:57 +00:00 av supernaut · 3 kommentarer
Ägare

Deadline

Production runs Kanidm 1.10.4 (ansible/roles/kanidm/defaults/main.yml:12). Current stable is
1.11.1, released 2026-08-14.

Kanidm supports each stable release for 4 months from its release date, on a quarterly cadence
(1 Feb, 1 May, 1 Aug, 1 Nov). Verified at
https://kanidm.github.io/kanidm/stable/support.html: "Stable releases will be supported for 4
months after their release date. This allows a 1 month support overlap between N and N+1 versions."

1.10.0 shipped 2026-05-01. Support ends around 2026-09-01. That is roughly two weeks out.

Why this cannot be deferred

Three constraints compound, and all three are enforced in code, not just documented.

  1. No minor skipping. "Upgrades are supported from 1 release (minor version) before the current
    stable release." The server raises MG0008SkipUpgradeAttempted rather than migrating.
  2. No downgrades. MG0010DowngradeNotAllowed. The project states this is an architectural
    conclusion, not an oversight.
  3. Restore is locked to the same minor series. The backup envelope stamps the series, and a
    mismatched restore fails with DB0001MismatchedRestoreVersion. A 1.10 backup restores onto any
    1.10.x and never onto 1.11.

1.12.0 is due 2026-11-01. If we are still on 1.10.x then, the direct path is gone. It becomes
1.10 to 1.11 to 1.12, with a stop-the-world step at each hop on a single-node deployment.

Plan

Take the latest patch of the current series first. 1.10.3's release notes fix "an incorrect
internal query that can prevent upgrades from 1.9 to 1.10", which is the pattern.

  • Read the GitHub release bodies for 1.10.5, 1.11.0 and 1.11.1. Do not rely on
    RELEASE_NOTES.md: it carries only X.Y.0 entries and has drifted from the release bodies.
  • Run kanidmd domain upgrade-check and clear anything it flags.
  • Back up at the current version. Keep the current image tag pinned and recoverable, because
    rollback is "start the previous container", with restore-from-backup as the fallback.
  • 1.10.4 to 1.10.5.
  • 1.10.5 to 1.11.1.
  • Verify after each hop: OIDC login to Forgejo, OIDC login to the portal, and the reconciler's
    next tick projecting unchanged.

Known 1.11.0 breaking change

Passwords are capped at 128 UTF-8 characters / 512 bytes. Unlikely to matter under SSO-only, but
worth confirming no service account exceeds it.

Check alongside

  • The override.css selectors and the concealment health gate against 1.11 markup.
  • kanidm-provision compatibility. Its patch set targets Kanidm 1.7 and 1.8, and it has had no
    commits since 2025-11-22. See the separate issue for replacing it.

Standing consequence

This is not a one-off. The 4-month window plus the no-skip rule means four forced upgrades per
year
, each a stop-the-world event. That cost is inherent to running Kanidm and should be planned
for rather than rediscovered each quarter.

## Deadline Production runs Kanidm **1.10.4** (`ansible/roles/kanidm/defaults/main.yml:12`). Current stable is **1.11.1**, released 2026-08-14. Kanidm supports each stable release for **4 months** from its release date, on a quarterly cadence (1 Feb, 1 May, 1 Aug, 1 Nov). Verified at <https://kanidm.github.io/kanidm/stable/support.html>: "Stable releases will be supported for 4 months after their release date. This allows a 1 month support overlap between N and N+1 versions." 1.10.0 shipped 2026-05-01. **Support ends around 2026-09-01.** That is roughly two weeks out. ## Why this cannot be deferred Three constraints compound, and all three are enforced in code, not just documented. 1. **No minor skipping.** "Upgrades are supported from 1 release (minor version) before the current stable release." The server raises `MG0008SkipUpgradeAttempted` rather than migrating. 2. **No downgrades.** `MG0010DowngradeNotAllowed`. The project states this is an architectural conclusion, not an oversight. 3. **Restore is locked to the same minor series.** The backup envelope stamps the series, and a mismatched restore fails with `DB0001MismatchedRestoreVersion`. A 1.10 backup restores onto any 1.10.x and never onto 1.11. 1.12.0 is due 2026-11-01. If we are still on 1.10.x then, the direct path is gone. It becomes 1.10 to 1.11 to 1.12, with a stop-the-world step at each hop on a single-node deployment. ## Plan Take the latest patch of the current series first. 1.10.3's release notes fix "an incorrect internal query that can prevent upgrades from 1.9 to 1.10", which is the pattern. - [ ] Read the **GitHub release bodies** for 1.10.5, 1.11.0 and 1.11.1. Do not rely on `RELEASE_NOTES.md`: it carries only `X.Y.0` entries and has drifted from the release bodies. - [ ] Run `kanidmd domain upgrade-check` and clear anything it flags. - [ ] Back up at the current version. Keep the current image tag pinned and recoverable, because rollback is "start the previous container", with restore-from-backup as the fallback. - [ ] 1.10.4 to 1.10.5. - [ ] 1.10.5 to 1.11.1. - [ ] Verify after each hop: OIDC login to Forgejo, OIDC login to the portal, and the reconciler's next tick projecting unchanged. ## Known 1.11.0 breaking change Passwords are capped at 128 UTF-8 characters / 512 bytes. Unlikely to matter under SSO-only, but worth confirming no service account exceeds it. ## Check alongside - The `override.css` selectors and the concealment health gate against 1.11 markup. - `kanidm-provision` compatibility. Its patch set targets Kanidm 1.7 and 1.8, and it has had no commits since 2025-11-22. See the separate issue for replacing it. ## Standing consequence This is not a one-off. The 4-month window plus the no-skip rule means **four forced upgrades per year**, each a stop-the-world event. That cost is inherent to running Kanidm and should be planned for rather than rediscovered each quarter.
Upphovsperson
Ägare

Correction: one hop, not two. PR #425 stages it.

The plan above says 1.10.4 to 1.10.5 to 1.11.1. Having read the actual release bodies, the
intermediate hop is unnecessary.

The upgrade policy supports one minor behind current stable, so 1.10 to 1.11 is a direct
supported step
. The reason to take the latest patch of the current series first is the 1.10.3
precedent, where a patch fixed "an incorrect internal query that can prevent upgrades from 1.9 to
1.10". Nothing of that kind exists in 1.10.5.

Read from the GitHub release bodies, not RELEASE_NOTES.md:

  • 1.10.5 contains exactly one change: the High-severity SCIM filter parsing depth fix.
  • 1.11.0 lists that same fix among its own highlights.
  • Both were published 2026-08-02, one minute apart.

1.10.5 therefore has nothing 1.11.1 lacks, and the extra hop is one more database migration and one
more restart for no gain.

What the upgrade actually delivers

Severity Change
Security High SCIM filter parsing depth, stack exhaustion DoS (via 1.11.0)
Security Moderate LDAP BER messages length-checked too late, OOM/DoS
Fix Incorrect URN in OAuth2 AccessTokenResponses
Fix Wrong HTTP status codes on OAuth2 responses
Change Wider character set allowed in OAuth2 scopes
Internal Schema moved from the database into memory

The OAuth2 entries matter because both the git host and the portal authenticate through that
surface. They are fixes, so the expected effect is neutral-to-better, but exercise both logins
after the apply.

Revised checklist

  • kanidmd domain upgrade-check, clear anything it flags.
  • Back up at the current version.
  • Confirm no service account holds a secret over 512 bytes (1.11.0 caps passwords at 128 UTF-8
    characters / 512 bytes).
  • Merge #425 and apply.
  • Confirm the new verify task passes. That task is now the proof the upgrade took.
  • OIDC sign-in to the git host, and to the portal.
  • Next reconcile tick projects unchanged.
  • override.css selectors and the concealment health gate still hold against 1.11 markup.

The 1.10.4 to 1.10.5 step is struck.

Rollback, corrected

The plan above says to keep the current image tag recoverable. Worth stating why, because it is
counter-intuitive: the rollback is to start the previous tag, not to restore the backup. A 1.10
backup cannot be loaded by 1.11 (DB0001MismatchedRestoreVersion). The backup protects against
data loss, not against a bad upgrade.

Now covered by #425

The role had no check that the running container is on the pinned image, which the restart handler
had flagged as an outstanding follow-up since 2026-08-04. #425 adds a pull before the handler flush
and a verify after start, matching forgejo, postgres, caddy and web.

Without it, a tag bump that silently no-ops would leave the operator believing a migration ran when
it did not, and planning the next hop from a version the server is not on.

Risk to watch

kanidm-provision is pinned at v1.3.0, built from source with a carried entryManagedBy patch.
Upstream has had no commits since 2025-11-22, and its own patch set targets Kanidm 1.7 and 1.8.
Compatibility with a 1.11 server is unverified. If entitlement groups stop converging after the
upgrade, look there first.

## Correction: one hop, not two. PR #425 stages it. The plan above says 1.10.4 to 1.10.5 to 1.11.1. Having read the actual release bodies, the intermediate hop is unnecessary. The upgrade policy supports one minor behind current stable, so **1.10 to 1.11 is a direct supported step**. The reason to take the latest patch of the current series first is the 1.10.3 precedent, where a patch fixed "an incorrect internal query that can prevent upgrades from 1.9 to 1.10". Nothing of that kind exists in 1.10.5. Read from the GitHub release bodies, not `RELEASE_NOTES.md`: - **1.10.5** contains exactly one change: the High-severity SCIM filter parsing depth fix. - **1.11.0** lists that same fix among its own highlights. - Both were published **2026-08-02, one minute apart**. 1.10.5 therefore has nothing 1.11.1 lacks, and the extra hop is one more database migration and one more restart for no gain. ## What the upgrade actually delivers | Severity | Change | | --- | --- | | Security High | SCIM filter parsing depth, stack exhaustion DoS (via 1.11.0) | | Security Moderate | LDAP BER messages length-checked too late, OOM/DoS | | Fix | Incorrect URN in OAuth2 `AccessTokenResponses` | | Fix | Wrong HTTP status codes on OAuth2 responses | | Change | Wider character set allowed in OAuth2 scopes | | Internal | Schema moved from the database into memory | The OAuth2 entries matter because both the git host and the portal authenticate through that surface. They are fixes, so the expected effect is neutral-to-better, but exercise both logins after the apply. ## Revised checklist - [ ] `kanidmd domain upgrade-check`, clear anything it flags. - [ ] Back up at the current version. - [ ] Confirm no service account holds a secret over 512 bytes (1.11.0 caps passwords at 128 UTF-8 characters / 512 bytes). - [ ] Merge #425 and apply. - [ ] Confirm the new verify task passes. That task is now the proof the upgrade took. - [ ] OIDC sign-in to the git host, and to the portal. - [ ] Next reconcile tick projects unchanged. - [ ] `override.css` selectors and the concealment health gate still hold against 1.11 markup. The 1.10.4 to 1.10.5 step is struck. ## Rollback, corrected The plan above says to keep the current image tag recoverable. Worth stating why, because it is counter-intuitive: **the rollback is to start the previous tag, not to restore the backup.** A 1.10 backup cannot be loaded by 1.11 (`DB0001MismatchedRestoreVersion`). The backup protects against data loss, not against a bad upgrade. ## Now covered by #425 The role had no check that the running container is on the pinned image, which the restart handler had flagged as an outstanding follow-up since 2026-08-04. #425 adds a pull before the handler flush and a verify after start, matching forgejo, postgres, caddy and web. Without it, a tag bump that silently no-ops would leave the operator believing a migration ran when it did not, and planning the next hop from a version the server is not on. ## Risk to watch `kanidm-provision` is pinned at `v1.3.0`, built from source with a carried `entryManagedBy` patch. Upstream has had no commits since 2025-11-22, and its own patch set targets Kanidm 1.7 and 1.8. Compatibility with a 1.11 server is unverified. If entitlement groups stop converging after the upgrade, look there first.
Upphovsperson
Ägare

Applied to production 2026-08-16. Kanidm is on 1.11.1.

One hop, as corrected above. Applied together with the merged backup-tooling bumps and some
pre-existing rename drift, via ./scripts/apply-reconcile.sh.

Result

PLAY RECAP  bitborg-prod : ok=351  changed=7  unreachable=0  failed=0

── reconcile: --check vs apply, by task name ──
  predicted changed : 7
  actually changed  : 7
  ✓ apply matched the dry run exactly.

A second --check afterwards returned changed=0, which is the only thing that proves prod
matches main.

Verified

Check Result
Running image docker.io/kanidm/server:1.11.1 (was 1.10.4)
kanidm unit active
New pinned-image verify task ok
Health gates 14 ran, no failures
git.bitborg.se/api/healthz 200
auth.bitborg.se/status 200
www.bitborg.se/ 200
bitborg_reconciler_last_run_status 0

Pre-flight: backup 16.6 h old and restore-verified 15.6 h old, both taken at 1.10.x so they were
restore-compatible with the version being replaced.

The flagged risk did not materialise

kanidm-provision is pinned at v1.3.0 with a patch set targeting Kanidm 1.7/1.8, and upstream has
had no commits since 2025-11-22. It is invisible to --check because it is a command task, so it
was the one thing the dry run could not clear. It ran and returned ok against the 1.11 server.

That risk is retired for 1.11. It is not retired for 1.12, and the tool is still dormant upstream.

Two things still open

1. The ADR 0038 concealment gate was overridden for this run.

The gate correctly blocked the apply: Kanidm: running 1.11.1, selectors verified against 1.10.4.
There is a chicken-and-egg, since 1.11's markup cannot be inspected until 1.11 is running.

It was overridden at runtime with -e health_check_concealment=false, not by committing a
disabled gate or by pre-bumping the verified tag, so nothing in the repository asserts a
verification that has not happened.

The six surfaces now need checking by hand against the running 1.11.1, per the gate's own list:

  • /user/settings → no "Full Name" field
  • /user/settings/account → no "Set as primary", no password section, no delete
  • auth.bitborg.se/ui/profile → no username / display name inputs
  • auth.bitborg.se/ui/login → the "No account yet?" note is present, signed out
  • www.bitborg.se signed out → "Create account" in the navbar
  • auth.bitborg.se/ui/reset?token=<t> → the passkey NAME input is visible (the over-match
    direction, which is what broke passkey enrolment on 2026-08-04)

Then bump health_check_kanidm_concealment_verified_tag to 1.11.1 and re-apply with the gate on.
Until that happens the estate is running unverified concealment selectors.

2. A local .netrc breaks Ansible uri tasks.

bad follower token 'region' (~/.netrc, line 4) makes every uri task delegated to localhost fail
with Status code was -1, which killed the first dry run at
roles/forgejo/tasks/system-webhook.yml. Worked around with NETRC= pointed at an empty file. It
is an operator-workstation problem, not a repository one, but it will hit anyone whose .netrc
carries a non-netrc token.

## Applied to production 2026-08-16. Kanidm is on 1.11.1. One hop, as corrected above. Applied together with the merged backup-tooling bumps and some pre-existing rename drift, via `./scripts/apply-reconcile.sh`. ### Result ``` PLAY RECAP bitborg-prod : ok=351 changed=7 unreachable=0 failed=0 ── reconcile: --check vs apply, by task name ── predicted changed : 7 actually changed : 7 ✓ apply matched the dry run exactly. ``` A second `--check` afterwards returned **`changed=0`**, which is the only thing that proves prod matches main. ### Verified | Check | Result | | --- | --- | | Running image | `docker.io/kanidm/server:1.11.1` (was 1.10.4) | | `kanidm` unit | active | | New pinned-image verify task | `ok` | | Health gates | 14 ran, no failures | | `git.bitborg.se/api/healthz` | 200 | | `auth.bitborg.se/status` | 200 | | `www.bitborg.se/` | 200 | | `bitborg_reconciler_last_run_status` | 0 | Pre-flight: backup 16.6 h old and restore-verified 15.6 h old, both taken at 1.10.x so they were restore-compatible with the version being replaced. ### The flagged risk did not materialise `kanidm-provision` is pinned at `v1.3.0` with a patch set targeting Kanidm 1.7/1.8, and upstream has had no commits since 2025-11-22. It is invisible to `--check` because it is a `command` task, so it was the one thing the dry run could not clear. It ran and returned **`ok`** against the 1.11 server. That risk is retired for 1.11. It is not retired for 1.12, and the tool is still dormant upstream. ### Two things still open **1. The ADR 0038 concealment gate was overridden for this run.** The gate correctly blocked the apply: `Kanidm: running 1.11.1, selectors verified against 1.10.4`. There is a chicken-and-egg, since 1.11's markup cannot be inspected until 1.11 is running. It was overridden at runtime with `-e health_check_concealment=false`, **not** by committing a disabled gate or by pre-bumping the verified tag, so nothing in the repository asserts a verification that has not happened. The six surfaces now need checking by hand against the running 1.11.1, per the gate's own list: - `/user/settings` → no "Full Name" field - `/user/settings/account` → no "Set as primary", no password section, no delete - `auth.bitborg.se/ui/profile` → no username / display name inputs - `auth.bitborg.se/ui/login` → the "No account yet?" note is present, signed out - `www.bitborg.se` signed out → "Create account" in the navbar - `auth.bitborg.se/ui/reset?token=<t>` → the passkey NAME input is **visible** (the over-match direction, which is what broke passkey enrolment on 2026-08-04) Then bump `health_check_kanidm_concealment_verified_tag` to `1.11.1` and re-apply with the gate on. Until that happens the estate is running unverified concealment selectors. **2. A local `.netrc` breaks Ansible `uri` tasks.** `bad follower token 'region' (~/.netrc, line 4)` makes every `uri` task delegated to localhost fail with `Status code was -1`, which killed the first dry run at `roles/forgejo/tasks/system-webhook.yml`. Worked around with `NETRC=` pointed at an empty file. It is an operator-workstation problem, not a repository one, but it will hit anyone whose `.netrc` carries a non-netrc token.
Upphovsperson
Ägare

Closing: the upgrade is complete and both flagged items are resolved

Kanidm has run 1.11.1 in production since 2026-08-16, and the two things the apply comment left
open are now closed. Recording the evidence here so the state is not carried in anyone's head.

1. The ADR 0038 concealment gate is back on

Closed by #436 (e675d44, 2026-08-17). All six surfaces were checked by hand, in both directions
the gate cares about, including the over-match direction that broke passkey enrolment on 2026-08-04.
health_check_kanidm_concealment_verified_tag is now 1.11.1, matching kanidm_image_tag, and a
full --check with the gate on and no override returned:

TASK [health-check : Health gate — identity concealment ...]
ok: [bitborg-prod] => "Concealment selectors verified against the deployed Forgejo + Kanidm tags."

The comment above said "until that happens the estate is running unverified concealment selectors".
That is no longer true, and this note supersedes it.

2. The ~/.netrc breakage no longer affects Ansible

Closed by #434 (4270f87, 2026-08-17): every ansible.builtin.uri task now carries
use_netrc: false. It remains an operator-workstation problem for other tools, but it can no longer
kill a play.

The revised checklist, honestly

Not ticking the list above item by item, because three of its entries were never recorded and
ticking them would assert a verification nobody can point at.

Item State
Back up at the current version Evidenced. Backup 16.6 h old, restore-verified 15.6 h old, both at 1.10.x
Merge #425 and apply Evidenced. changed=7, matching the dry run exactly; second --check returned changed=0
Pinned-image verify task passes Evidenced. ok, and the running image reads 1.11.1
Next reconcile tick projects unchanged Evidenced. bitborg_reconciler_last_run_status 0
override.css + concealment gate hold on 1.11 Evidenced. #436
kanidmd domain upgrade-check Not recorded. A pre-upgrade guard, now moot: the migration ran and the server is up on 1.11.1
No service account secret over 512 bytes Not recorded. Empirically fine (OIDC-dependent services are all healthy), but never audited
OIDC sign-in to the git host and the portal Partially evidenced. git.bitborg.se/api/healthz, auth.bitborg.se/status and www.bitborg.se/ all returned 200, which is not the same as an interactive sign-in. Note that no probe covers /user/login at all, which is #384

Closing on the substance: the deadline this issue existed for (1.10.x leaving support around
2026-09-01) is retired, the server is on a supported release, and the gate that guards the next one
is armed. The two genuinely unaudited items above are small and non-blocking; the sign-in coverage
gap is tracked separately as #384.

Carried forward, not closed here

  • kanidm-provision returned ok against the 1.11 server, so that risk is retired for 1.11
    only
    . It is still pinned at v1.3.0 with a patch set targeting Kanidm 1.7/1.8 and no upstream
    commits since 2025-11-22.
  • 1.12.0 is due 2026-11-01, and the 4-month support window plus the no-minor-skip rule makes that
    upgrade mandatory rather than optional. This recurs four times a year.
  • The mirror entry for Kanidm was left at 1.10.4 by this upgrade and nothing could notice. Filed
    as #441.
## Closing: the upgrade is complete and both flagged items are resolved Kanidm has run **1.11.1** in production since 2026-08-16, and the two things the apply comment left open are now closed. Recording the evidence here so the state is not carried in anyone's head. ### 1. The ADR 0038 concealment gate is back on Closed by #436 (`e675d44`, 2026-08-17). All six surfaces were checked by hand, in both directions the gate cares about, including the over-match direction that broke passkey enrolment on 2026-08-04. `health_check_kanidm_concealment_verified_tag` is now `1.11.1`, matching `kanidm_image_tag`, and a full `--check` with the gate **on** and no override returned: ```text TASK [health-check : Health gate — identity concealment ...] ok: [bitborg-prod] => "Concealment selectors verified against the deployed Forgejo + Kanidm tags." ``` The comment above said "until that happens the estate is running unverified concealment selectors". That is no longer true, and this note supersedes it. ### 2. The `~/.netrc` breakage no longer affects Ansible Closed by #434 (`4270f87`, 2026-08-17): every `ansible.builtin.uri` task now carries `use_netrc: false`. It remains an operator-workstation problem for other tools, but it can no longer kill a play. ### The revised checklist, honestly Not ticking the list above item by item, because three of its entries were never recorded and ticking them would assert a verification nobody can point at. | Item | State | | --- | --- | | Back up at the current version | **Evidenced.** Backup 16.6 h old, restore-verified 15.6 h old, both at 1.10.x | | Merge #425 and apply | **Evidenced.** `changed=7`, matching the dry run exactly; second `--check` returned `changed=0` | | Pinned-image verify task passes | **Evidenced.** `ok`, and the running image reads `1.11.1` | | Next reconcile tick projects unchanged | **Evidenced.** `bitborg_reconciler_last_run_status 0` | | `override.css` + concealment gate hold on 1.11 | **Evidenced.** #436 | | `kanidmd domain upgrade-check` | **Not recorded.** A pre-upgrade guard, now moot: the migration ran and the server is up on 1.11.1 | | No service account secret over 512 bytes | **Not recorded.** Empirically fine (OIDC-dependent services are all healthy), but never audited | | OIDC sign-in to the git host and the portal | **Partially evidenced.** `git.bitborg.se/api/healthz`, `auth.bitborg.se/status` and `www.bitborg.se/` all returned 200, which is not the same as an interactive sign-in. Note that no probe covers `/user/login` at all, which is #384 | Closing on the substance: the deadline this issue existed for (1.10.x leaving support around 2026-09-01) is retired, the server is on a supported release, and the gate that guards the next one is armed. The two genuinely unaudited items above are small and non-blocking; the sign-in coverage gap is tracked separately as #384. ### Carried forward, not closed here - `kanidm-provision` returned `ok` against the 1.11 server, so that risk is retired **for 1.11 only**. It is still pinned at `v1.3.0` with a patch set targeting Kanidm 1.7/1.8 and no upstream commits since 2025-11-22. - 1.12.0 is due 2026-11-01, and the 4-month support window plus the no-minor-skip rule makes that upgrade mandatory rather than optional. This recurs four times a year. - The mirror entry for Kanidm was left at 1.10.4 by this upgrade and nothing could notice. Filed as #441.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#424
Ingen beskrivning angiven.