fix(health-check): the sign-up gate failed the apply on the monitoring host #263
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!263
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/signup-gate-scope-and-visibility"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
The sign-up health gate added in #262 broke
ansible-playbook site.yml. gitborg-prod passed;gitborg-monitoring failed with:
Three defects, all mine, all in the mitigation rather than the thing it guards.
1. Wrong host
The gate reaches Kanidm through the local Caddy (
--resolve <domain>:443:127.0.0.1— theno-NAT-hairpin trick the other gates use), which only exists on the services host. But
health-checkalso runs in the monitoring play, where
127.0.0.1is not Caddy. So every group read came back empty,the gate concluded "denied", and it failed a play with nothing to do with sign-up.
health_check_kanidm_portal_groupsnow defaults to[]and is populated only by the bitborg play, soeach play opts in explicitly.
2. The failure was undebuggable
no_log: trueis genuinely required — the provisioning token is in the task environment — but itcensors the result. So a gate whose entire purpose is a loud, actionable message reported nothing:
not the failing groups, not the remediation command.
The shell task no longer fails. A separate
fail:task withoutno_logdoes, printing the capturedstdout/stderr. The registered variable always held that output;
no_logonly suppressed the automaticecho. What gets printed is group names and a
set-entry-managerhint — never the token, which existedonly in the previous task's environment.
3. It could not be dry-run — which is why 1 and 2 reached production
The task was gated on
not ansible_check_mode, so it was skipped in every--checkrun. Thedry-run I ran before applying could not possibly have revealed either defect.
It is a read-only GET, so it now runs under
--checkviacheck_mode: false.This is the exact trap #257 fixed elsewhere in this repo — a task skipped under
--checkmaking thedry-run gate meaningless — reintroduced two commits later, in the role whose job is gating.
Verification
--checkon gitborg-prodok--checkon gitborg-monitoringskipping(opted out)failed=both hostsAlso confirmed after the failed apply
The apply reached the gate only after the monitoring role had done its work, so #262's Loki alert
did land. Three things checked, because the failed run left questions:
kanidm-provisiondoes not clobberentry_managed_by. It reportedchangedon the same run,which was the obvious worry given it manages groups declaratively. All four groups remain delegated.
pnpm signup:drillstill passes end to end — person create, all three group adds, reset intent.webrestart on prod did not happen, confirming it was a check-mode artefact ofthe
podman container inspectthat registersweb_running_imagebeing skipped under--check.Unrelated finding, not fixed here
base : Enable and start fail2banreportedchangedduring the apply and still reportschangedin a dry-run afterwards, on gitborg-monitoring only. That suggests fail2ban is not stayingstarted there. I could not verify — the monitoring VM does not ship logs to Loki and I have no shell on
it. Worth a
systemctl status fail2banon that host.