feat(kanidm): upgrade to 1.11.1, and verify the running image matches the pin #425

Sammanfogat
supernaut sammanfogade 1 incheckning från feat/kanidm-1-11-1 in i main 2026-08-16 14:44:55 +00:00
Ägare

Prepares #424. Staged for review, not applied.

Why

The 1.10 series leaves support around 2026-09-01. Kanidm supports each stable release for 4
months on a quarterly cadence, and 1.10.0 shipped 2026-05-01. Current stable is 1.11.1.

One hop, not two

#424's plan said 1.10.4 to 1.10.5 to 1.11.1. The intermediate hop is unnecessary.

The upgrade policy supports one minor behind current stable, so 1.10 to 1.11 is a direct supported
step. The reason to take the latest patch of the current series first is the 1.10.3 precedent,
where a patch fixed "an incorrect internal query that can prevent upgrades from 1.9 to 1.10". That
does not apply here.

Read from the GitHub release bodies rather than RELEASE_NOTES.md, which carries only X.Y.0
entries and has drifted:

  • 1.10.5 contains exactly one change: the High-severity SCIM filter parsing depth fix.
  • 1.11.0 lists that same fix in its own highlights.
  • Both were published 2026-08-02, one minute apart.

So 1.10.5 has nothing 1.11.1 lacks. I will correct #424.

What 1.11.1 brings over 1.10.4

Severity Change
Security High SCIM filter parsing depth, stack exhaustion DoS (via 1.11.0)
Security Moderate LDAP BER messages length-checked too late, OOM/DoS
Fix Incorrect URN in OAuth2 AccessTokenResponses
Fix Wrong HTTP status codes on OAuth2 responses
Change Wider character set allowed in OAuth2 scopes
Internal Schema moved from the database into memory

The OAuth2 items are worth noting because both the git host and the portal authenticate through
that surface. They are fixes, so the expected effect is neutral-to-better, but both logins should
be exercised after the apply.

One behaviour change: passwords are capped at 128 UTF-8 characters / 512 bytes. Sign-in is
SSO-only so this is low risk, but a service account with a longer secret would be affected.

Second change: the missing image check

The restart handler has carried this follow-up since the 2026-08-04 incident, in its own words:

unlike forgejo/postgres/caddy/web, this role has no "running container is on the pinned image"
check to catch it (follow-up)

Now done, matching the existing pattern:

  • Pull before the flush, so the image is on the host before the restart handler fires. This also
    moves failure earlier and makes it louder: a bad tag or unreachable registry now fails with Kanidm
    still up, rather than failing the restart and leaving the identity provider down.
  • Verify after start, failing the play if the running container is not on the pinned tag.
    Skipped under --check, where a pending bump legitimately differs.

That check earns its place here more than elsewhere. Kanidm migrates its database on start, refuses
downgrades, and restores a backup only into its own minor series. A silent no-op would mean
believing a migration ran when it did not, and planning the next hop from a version the server is
not on.

Rollback

Recorded in the defaults alongside the pin, because it is counter-intuitive: start the previous
image tag
, do not restore the backup. A 1.10 backup cannot be loaded by 1.11
(DB0001MismatchedRestoreVersion). Keep 1.10.4 recoverable until the new version has run.

Before applying

  • kanidmd domain upgrade-check and clear anything it flags.
  • Back up at the current version.
  • Confirm no service account holds a secret over 512 bytes.

After applying

  • The new verify task passes, which is now the proof the upgrade actually took.
  • OIDC sign-in to the git host.
  • OIDC sign-in to the portal.
  • The next reconcile tick projects unchanged.
  • override.css selectors and the concealment health gate still hold against 1.11 markup.

Known risk, not addressed here

kanidm-provision is pinned at v1.3.0 and built from source with a carried entryManagedBy
patch. Upstream has had no commits since 2025-11-22 and its own patch set targets Kanidm 1.7 and
1.8. Compatibility with a 1.11 server is unverified. If provisioning breaks after the upgrade, that
is the first place to look, and the entitlement groups it manages would stop converging.

Verification

  • pnpm ansible:check passes.
  • ansible-lint ansible/roles/kanidm/ reports the same 12 pre-existing findings before and after
    this change, so nothing was regressed. None are in the added tasks.
Prepares #424. Staged for review, not applied. ## Why The 1.10 series leaves support around **2026-09-01**. Kanidm supports each stable release for 4 months on a quarterly cadence, and 1.10.0 shipped 2026-05-01. Current stable is 1.11.1. ## One hop, not two **#424's plan said 1.10.4 to 1.10.5 to 1.11.1. The intermediate hop is unnecessary.** The upgrade policy supports one minor behind current stable, so 1.10 to 1.11 is a direct supported step. The reason to take the latest patch of the current series first is the 1.10.3 precedent, where a patch fixed "an incorrect internal query that can prevent upgrades from 1.9 to 1.10". That does not apply here. Read from the GitHub release bodies rather than `RELEASE_NOTES.md`, which carries only `X.Y.0` entries and has drifted: - **1.10.5** contains exactly one change: the High-severity SCIM filter parsing depth fix. - **1.11.0** lists that same fix in its own highlights. - Both were published **2026-08-02, one minute apart**. So 1.10.5 has nothing 1.11.1 lacks. I will correct #424. ## What 1.11.1 brings over 1.10.4 | Severity | Change | | --- | --- | | Security High | SCIM filter parsing depth, stack exhaustion DoS (via 1.11.0) | | Security Moderate | LDAP BER messages length-checked too late, OOM/DoS | | Fix | Incorrect URN in OAuth2 `AccessTokenResponses` | | Fix | Wrong HTTP status codes on OAuth2 responses | | Change | Wider character set allowed in OAuth2 scopes | | Internal | Schema moved from the database into memory | The OAuth2 items are worth noting because both the git host and the portal authenticate through that surface. They are fixes, so the expected effect is neutral-to-better, but both logins should be exercised after the apply. **One behaviour change:** passwords are capped at 128 UTF-8 characters / 512 bytes. Sign-in is SSO-only so this is low risk, but a service account with a longer secret would be affected. ## Second change: the missing image check The restart handler has carried this follow-up since the 2026-08-04 incident, in its own words: > unlike forgejo/postgres/caddy/web, this role has no "running container is on the pinned image" > check to catch it (follow-up) Now done, matching the existing pattern: - **Pull before the flush**, so the image is on the host before the restart handler fires. This also moves failure earlier and makes it louder: a bad tag or unreachable registry now fails with Kanidm still up, rather than failing the restart and leaving the identity provider down. - **Verify after start**, failing the play if the running container is not on the pinned tag. Skipped under `--check`, where a pending bump legitimately differs. That check earns its place here more than elsewhere. Kanidm migrates its database on start, refuses downgrades, and restores a backup only into its own minor series. A silent no-op would mean believing a migration ran when it did not, and planning the next hop from a version the server is not on. ## Rollback Recorded in the defaults alongside the pin, because it is counter-intuitive: **start the previous image tag**, do not restore the backup. A 1.10 backup cannot be loaded by 1.11 (`DB0001MismatchedRestoreVersion`). Keep `1.10.4` recoverable until the new version has run. ## Before applying - [ ] `kanidmd domain upgrade-check` and clear anything it flags. - [ ] Back up at the current version. - [ ] Confirm no service account holds a secret over 512 bytes. ## After applying - [ ] The new verify task passes, which is now the proof the upgrade actually took. - [ ] OIDC sign-in to the git host. - [ ] OIDC sign-in to the portal. - [ ] The next reconcile tick projects unchanged. - [ ] `override.css` selectors and the concealment health gate still hold against 1.11 markup. ## Known risk, not addressed here `kanidm-provision` is pinned at `v1.3.0` and built from source with a carried `entryManagedBy` patch. Upstream has had no commits since 2025-11-22 and its own patch set targets Kanidm 1.7 and 1.8. Compatibility with a 1.11 server is unverified. If provisioning breaks after the upgrade, that is the first place to look, and the entitlement groups it manages would stop converging. ## Verification - `pnpm ansible:check` passes. - `ansible-lint ansible/roles/kanidm/` reports the same 12 pre-existing findings before and after this change, so nothing was regressed. None are in the added tasks.
supernaut lade till 1 incheckning 2026-08-16 14:38:20 +00:00
feat(kanidm): upgrade to 1.11.1, and verify the running image matches the pin
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m37s
f5600d3120
The 1.10 series leaves support around 2026-09-01. Kanidm supports each stable
release for 4 months on a quarterly cadence, and 1.10.0 shipped 2026-05-01.

One hop, not two. The upgrade policy allows one minor behind current stable, so
1.10 to 1.11 is a supported direct step. An intermediate 1.10.5 buys nothing:
its only content is the High-severity SCIM filter parsing fix, and that same fix
is listed in 1.11.0's own highlights. Both were published on 2026-08-02, one
minute apart.

What 1.11.1 brings over 1.10.4:

  * Security High: SCIM filter parsing depth, stack exhaustion DoS (via 1.11.0)
  * Security Moderate: LDAP BER messages length-checked too late, OOM/DoS
  * OAuth2 fixes: incorrect URN in AccessTokenResponses, wrong HTTP status
    codes on responses, wider character set allowed in scopes
  * schema moved from the database into memory

One behaviour change to be aware of: passwords are now capped at 128 UTF-8
characters, 512 bytes. Sign-in is SSO-only so this is low risk, but a service
account with a longer secret would be affected.

Also closes the follow-up the restart handler has been carrying since the
2026-08-04 incident. The role now pulls the pinned image before the handler
flush and verifies the running container is on it afterwards, matching what
forgejo, postgres, caddy and web already do.

That check earns its place here more than elsewhere. Kanidm migrates its
database on start, refuses downgrades, and restores a backup only into its own
minor series. A tag bump that silently no-ops does not just leave drift: it
means the operator believes a migration ran when it did not, and plans the next
hop from a version the server is not on.

The rollback path is recorded in the defaults for the same reason. It is "start
the previous image tag", not "restore the backup", because a 1.10 backup cannot
be loaded by 1.11.
supernaut sammanfogade incheckning 0ee69fd53b till main 2026-08-16 14:44:55 +00:00
supernaut tog bort grenen feat/kanidm-1-11-1 2026-08-16 14:44:55 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!425
Ingen beskrivning angiven.