fix(health-check): the sign-up gate failed the apply on the monitoring host #263

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/signup-gate-scope-and-visibility in i main 2026-07-30 19:05:11 +00:00
Ägare

The sign-up health gate added in #262 broke ansible-playbook site.yml. gitborg-prod passed;
gitborg-monitoring failed with:

fatal: [gitborg-monitoring]: FAILED! => {"censored": "the output has been hidden due to the fact
that 'no_log: true' was specified for this result"}

Three defects, all mine, all in the mitigation rather than the thing it guards.

1. Wrong host

The gate reaches Kanidm through the local Caddy (--resolve <domain>:443:127.0.0.1 — the
no-NAT-hairpin trick the other gates use), which only exists on the services host. But health-check
also runs in the monitoring play, where 127.0.0.1 is not Caddy. So every group read came back empty,
the gate concluded "denied", and it failed a play with nothing to do with sign-up.

health_check_kanidm_portal_groups now defaults to [] and is populated only by the bitborg play, so
each play opts in explicitly.

2. The failure was undebuggable

no_log: true is genuinely required — the provisioning token is in the task environment — but it
censors the result. So a gate whose entire purpose is a loud, actionable message reported nothing:
not the failing groups, not the remediation command.

The shell task no longer fails. A separate fail: task without no_log does, printing the captured
stdout/stderr. The registered variable always held that output; no_log only suppressed the automatic
echo. What gets printed is group names and a set-entry-manager hint — never the token, which existed
only in the previous task's environment.

3. It could not be dry-run — which is why 1 and 2 reached production

The task was gated on not ansible_check_mode, so it was skipped in every --check run. The
dry-run I ran before applying could not possibly have revealed either defect.

It is a read-only GET, so it now runs under --check via check_mode: false.

This is the exact trap #257 fixed elsewhere in this repo — a task skipped under --check making the
dry-run gate meaningless — reintroduced two commits later, in the role whose job is gating.

Verification

before after
--check on gitborg-prod skipped ok
--check on gitborg-monitoring skipped skipping (opted out)
real apply on gitborg-monitoring fatal n/a — skipped
failed= both hosts 1 0

Also confirmed after the failed apply

The apply reached the gate only after the monitoring role had done its work, so #262's Loki alert
did land. Three things checked, because the failed run left questions:

  • kanidm-provision does not clobber entry_managed_by. It reported changed on the same run,
    which was the obvious worry given it manages groups declaratively. All four groups remain delegated.
  • pnpm signup:drill still passes end to end — person create, all three group adds, reset intent.
  • The predicted web restart on prod did not happen, confirming it was a check-mode artefact of
    the podman container inspect that registers web_running_image being skipped under --check.

Unrelated finding, not fixed here

base : Enable and start fail2ban reported changed during the apply and still reports
changed in a dry-run afterwards, on gitborg-monitoring only. That suggests fail2ban is not staying
started there. I could not verify — the monitoring VM does not ship logs to Loki and I have no shell on
it. Worth a systemctl status fail2ban on that host.

The sign-up health gate added in #262 **broke `ansible-playbook site.yml`**. gitborg-prod passed; gitborg-monitoring failed with: ``` fatal: [gitborg-monitoring]: FAILED! => {"censored": "the output has been hidden due to the fact that 'no_log: true' was specified for this result"} ``` Three defects, all mine, all in the mitigation rather than the thing it guards. ## 1. Wrong host The gate reaches Kanidm through the **local Caddy** (`--resolve <domain>:443:127.0.0.1` — the no-NAT-hairpin trick the other gates use), which only exists on the services host. But `health-check` also runs in the monitoring play, where `127.0.0.1` is not Caddy. So every group read came back empty, the gate concluded "denied", and it failed a play with nothing to do with sign-up. `health_check_kanidm_portal_groups` now defaults to `[]` and is populated only by the bitborg play, so each play opts in explicitly. ## 2. The failure was undebuggable `no_log: true` is genuinely required — the provisioning token is in the task environment — but it censors the **result**. So a gate whose entire purpose is a loud, actionable message reported *nothing*: not the failing groups, not the remediation command. The shell task no longer fails. A separate `fail:` task **without** `no_log` does, printing the captured stdout/stderr. The registered variable always held that output; `no_log` only suppressed the automatic echo. What gets printed is group names and a `set-entry-manager` hint — never the token, which existed only in the previous task's environment. ## 3. It could not be dry-run — which is why 1 and 2 reached production The task was gated on `not ansible_check_mode`, so it was **skipped in every `--check` run**. The dry-run I ran before applying could not possibly have revealed either defect. It is a read-only GET, so it now runs under `--check` via `check_mode: false`. This is the exact trap #257 fixed elsewhere in this repo — a task skipped under `--check` making the dry-run gate meaningless — reintroduced two commits later, in the role whose job is gating. ## Verification | | before | after | | --- | --- | --- | | `--check` on gitborg-prod | *skipped* | **`ok`** | | `--check` on gitborg-monitoring | *skipped* | **`skipping`** (opted out) | | real apply on gitborg-monitoring | **fatal** | n/a — skipped | | `failed=` both hosts | 1 | **0** | ## Also confirmed after the failed apply The apply reached the gate only *after* the monitoring role had done its work, so #262's Loki alert did land. Three things checked, because the failed run left questions: - **`kanidm-provision` does not clobber `entry_managed_by`.** It reported `changed` on the same run, which was the obvious worry given it manages groups declaratively. All four groups remain delegated. - **`pnpm signup:drill` still passes end to end** — person create, all three group adds, reset intent. - The predicted `web` restart on prod did **not** happen, confirming it was a check-mode artefact of the `podman container inspect` that registers `web_running_image` being skipped under `--check`. ## Unrelated finding, not fixed here `base : Enable and start fail2ban` reported `changed` **during** the apply and **still** reports `changed` in a dry-run afterwards, on gitborg-monitoring only. That suggests fail2ban is not *staying* started there. I could not verify — the monitoring VM does not ship logs to Loki and I have no shell on it. Worth a `systemctl status fail2ban` on that host.
supernaut lade till 1 incheckning 2026-07-30 18:59:24 +00:00
fix(health-check): the sign-up gate failed the apply on the monitoring host
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m28s
c5586cbf5c
The gate added in #262 broke `ansible-playbook site.yml`: gitborg-prod passed, but
gitborg-monitoring reported

  fatal: [gitborg-monitoring]: FAILED! => {"censored": "the output has been hidden
  due to the fact that 'no_log: true' was specified for this result"}

Three defects, all mine, all in the mitigation rather than the thing it guards.

1. WRONG HOST. The gate reaches Kanidm through the LOCAL Caddy
   (`--resolve <domain>:443:127.0.0.1`, the no-NAT-hairpin trick the other gates
   use), which only exists on the services host. The health-check role also runs in
   the monitoring play, where 127.0.0.1 is not Caddy — so every group read returned
   empty, the gate concluded "denied", and it failed a play that has nothing to do
   with sign-up. health_check_kanidm_portal_groups now defaults to [] and is
   populated only by the gitborg play, so each play opts in explicitly.

2. UNDEBUGGABLE FAILURE. `no_log: true` is required (the provisioning token is in
   the task environment) but it censors the RESULT, so a gate whose entire purpose
   is a loud, actionable message reported nothing at all — not the failing groups,
   not the remediation command. The shell task no longer fails; a separate
   `fail:` task without no_log does, printing the captured stdout/stderr. The
   registered variable always held that output; no_log only suppressed the echo.
   The printed text is group names and a `set-entry-manager` hint — never the token.

3. NOT DRY-RUNNABLE — and this is why the other two reached production. The task
   was gated on `not ansible_check_mode`, so it was SKIPPED in every `--check` run;
   the dry-run could not have revealed either defect. It is a read-only GET, so it
   now runs under --check via `check_mode: false`. Exactly the trap #257 fixed
   elsewhere, reintroduced two commits later.

Verified: `--check` now reports `ok: [gitborg-prod]` (previously skipped) and
`skipping: [gitborg-monitoring]` (previously fatal), failed=0 on both hosts.

Also confirmed, after the failed apply: kanidm-provision does NOT clobber
entry_managed_by (it reported changed on the same run, which was the obvious
worry), all four groups remain delegated, and `pnpm signup:drill` still passes end
to end.
supernaut sammanfogade incheckning 1221e6b542 till main 2026-07-30 19:05:11 +00:00
supernaut tog bort grenen fix/signup-gate-scope-and-visibility 2026-07-30 19:05:12 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!263
Ingen beskrivning angiven.