fix(base): fail2ban was DEAD on the monitoring host; stop zoekt churn #264
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!264
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/fail2ban-monitoring-and-zoekt-churn"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Two findings from reviewing the last apply's
changedtasks. Both were invisible to--check, whichis why #258's churn sweep missed them — and one of them is a real security gap, not churn.
1. fail2ban has not been running on gitborg-monitoring
Confirmed on the host:
base/defaultsenables thecaddy-authjail on every host with apollingbackend readingfail2ban_caddy_log_path(~gitborg/caddy/logs/access.log). The two hosts log completely differently:gitborg-prodoutput file /var/log/caddy/access.logVolume={{ caddy_log_dir }}:/var/log/caddy✅gitborg-monitoringoutput stdout(journal)So the path is never created, and fail2ban refuses to start with an unresolvable polling logpath.
The failure mode is the point: a jail added as defence-in-depth took the whole daemon down, so the
[sshd]jail — the protection that actually matters on a box with no Caddy file log — silently did notexist either. The monitoring VM has had no SSH brute-force protection.
The only symptom was
Enable and start fail2banreportingchangedon three consecutive applies:exactly the kind of line that reads as noise.
Duration: 107msshows it died immediately each time anapply started it.
Disabled for the
monitoringgroup rather than ported to the systemd backend: grafana/stats have theirown auth, and a journal filter that silently matched nothing would be worse than no jail.
[sshd]nowapplies, which is the whole objective.
2. The zoekt index removal reported changed on every apply
rm -rfexits 0 whether or not anything was there, so this meant changed on every apply, forever.The deletion is idempotent in effect; the reporting never was.
changed_when: rc == 0is always wrong on a command designed to succeed —rm -rf,mkdir -p,touch. The information lives in a probe, not the exit code. (Same root confusion as arm -fused asa test assertion, which passes on a file that was never there.)
Now probes first and reports
changedonly onREMOVED.Deliberately not fixed here
kanidm : Provision entitlement groupsalso reportschangedevery run — itschanged_whenis astring match on tool output. Diagnosing it requires seeing that output, but the task is
no_log:becausethe idm_admin password is in its environment, and kanidm-provision provisions OAuth2 clients, so its
stdout may carry a client secret. Printing that to console and CI logs to fix cosmetic churn is a bad
trade. Left as documented known-churn; the proper fix is a dry-run flag on the tool, if it has one.
Verification
--syntax-checkpasses;ansible-lintclean at theproductionprofile.The post-apply check is precise and needs no shell access: a follow-up dry-run should report
changed=0on gitborg-monitoring (fail2ban finally staying up) and no zoekt change on prod.