fix(base): fail2ban was DEAD on the monitoring host; stop zoekt churn #264

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/fail2ban-monitoring-and-zoekt-churn in i main 2026-07-30 19:36:41 +00:00
Ägare

Two findings from reviewing the last apply's changed tasks. Both were invisible to --check, which
is why #258's churn sweep missed them — and one of them is a real security gap, not churn.

1. fail2ban has not been running on gitborg-monitoring

Confirmed on the host:

× fail2ban.service - Fail2Ban Service
     Active: failed (Result: exit-code) since Thu 2026-07-30 19:14:20 UTC
   Duration: 107ms
    Process: ExecStart=/usr/bin/fail2ban-server -xf start (code=exited, status=255/EXCEPTION)

fail2ban-server: ERROR  Failed during configuration: Have not found any log file for caddy-auth jail
fail2ban-server: ERROR  Async configuration of server failed

base/defaults enables the caddy-auth jail on every host with a polling backend reading
fail2ban_caddy_log_path (~gitborg/caddy/logs/access.log). The two hosts log completely differently:

Caddy access log log bind mount
gitborg-prod output file /var/log/caddy/access.log Volume={{ caddy_log_dir }}:/var/log/caddy ✅
gitborg-monitoring output stdout (journal) none ❌

So the path is never created, and fail2ban refuses to start with an unresolvable polling logpath.

The failure mode is the point: a jail added as defence-in-depth took the whole daemon down, so the
[sshd] jail — the protection that actually matters on a box with no Caddy file log — silently did not
exist either. The monitoring VM has had no SSH brute-force protection.

The only symptom was Enable and start fail2ban reporting changed on three consecutive applies:
exactly the kind of line that reads as noise. Duration: 107ms shows it died immediately each time an
apply started it.

Disabled for the monitoring group rather than ported to the systemd backend: grafana/stats have their
own auth, and a journal filter that silently matched nothing would be worse than no jail. [sshd] now
applies, which is the whole objective.

2. The zoekt index removal reported changed on every apply

cmd: podman exec forgejo rm -rf .../indexers
changed_when: _forgejo_index_rm.rc == 0

rm -rf exits 0 whether or not anything was there, so this meant changed on every apply, forever.
The deletion is idempotent in effect; the reporting never was.

changed_when: rc == 0 is always wrong on a command designed to succeed — rm -rf, mkdir -p,
touch. The information lives in a probe, not the exit code. (Same root confusion as a rm -f used as
a test assertion, which passes on a file that was never there.)

Now probes first and reports changed only on REMOVED.

Deliberately not fixed here

kanidm : Provision entitlement groups also reports changed every run — its changed_when is a
string match on tool output. Diagnosing it requires seeing that output, but the task is no_log: because
the idm_admin password is in its environment, and kanidm-provision provisions OAuth2 clients, so its
stdout may carry a client secret. Printing that to console and CI logs to fix cosmetic churn is a bad
trade. Left as documented known-churn; the proper fix is a dry-run flag on the tool, if it has one.

Verification

--syntax-check passes; ansible-lint clean at the production profile.

The post-apply check is precise and needs no shell access: a follow-up dry-run should report
changed=0 on gitborg-monitoring (fail2ban finally staying up) and no zoekt change on prod.

Two findings from reviewing the last apply's `changed` tasks. Both were invisible to `--check`, which is why #258's churn sweep missed them — and one of them is a real security gap, not churn. ## 1. fail2ban has not been running on gitborg-monitoring Confirmed on the host: ``` × fail2ban.service - Fail2Ban Service Active: failed (Result: exit-code) since Thu 2026-07-30 19:14:20 UTC Duration: 107ms Process: ExecStart=/usr/bin/fail2ban-server -xf start (code=exited, status=255/EXCEPTION) fail2ban-server: ERROR Failed during configuration: Have not found any log file for caddy-auth jail fail2ban-server: ERROR Async configuration of server failed ``` `base/defaults` enables the `caddy-auth` jail on **every** host with a `polling` backend reading `fail2ban_caddy_log_path` (`~gitborg/caddy/logs/access.log`). The two hosts log completely differently: | | Caddy access log | log bind mount | | --- | --- | --- | | `gitborg-prod` | `output file /var/log/caddy/access.log` | `Volume={{ caddy_log_dir }}:/var/log/caddy` ✅ | | `gitborg-monitoring` | **`output stdout`** (journal) | **none** ❌ | So the path is never created, and fail2ban refuses to start with an unresolvable polling logpath. **The failure mode is the point:** a jail added as *defence-in-depth* took the whole daemon down, so the `[sshd]` jail — the protection that actually matters on a box with no Caddy file log — silently did not exist either. **The monitoring VM has had no SSH brute-force protection.** The only symptom was `Enable and start fail2ban` reporting `changed` on three consecutive applies: exactly the kind of line that reads as noise. `Duration: 107ms` shows it died immediately each time an apply started it. Disabled for the `monitoring` group rather than ported to the systemd backend: grafana/stats have their own auth, and a journal filter that silently matched nothing would be worse than no jail. `[sshd]` now applies, which is the whole objective. ## 2. The zoekt index removal reported changed on every apply ```yaml cmd: podman exec forgejo rm -rf .../indexers changed_when: _forgejo_index_rm.rc == 0 ``` `rm -rf` exits 0 whether or not anything was there, so this meant *changed on every apply, forever*. The deletion is idempotent in **effect**; the **reporting** never was. `changed_when: rc == 0` is always wrong on a command designed to succeed — `rm -rf`, `mkdir -p`, `touch`. The information lives in a probe, not the exit code. (Same root confusion as a `rm -f` used as a test assertion, which passes on a file that was never there.) Now probes first and reports `changed` only on `REMOVED`. ## Deliberately not fixed here `kanidm : Provision entitlement groups` also reports `changed` every run — its `changed_when` is a string match on tool output. Diagnosing it requires seeing that output, but the task is `no_log:` because the idm_admin password is in its environment, **and** kanidm-provision provisions OAuth2 clients, so its stdout may carry a client secret. Printing that to console and CI logs to fix cosmetic churn is a bad trade. Left as documented known-churn; the proper fix is a dry-run flag on the tool, if it has one. ## Verification `--syntax-check` passes; `ansible-lint` clean at the `production` profile. The post-apply check is precise and needs no shell access: a follow-up dry-run should report **`changed=0` on gitborg-monitoring** (fail2ban finally staying up) and no zoekt change on prod.
supernaut lade till 1 incheckning 2026-07-30 19:34:38 +00:00
fix(base): fail2ban was DEAD on the monitoring host; stop zoekt churn
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m32s
1d6ea5ee27
Two findings from the post-apply review, both invisible to `--check`, which is why
#258's churn sweep missed them.

1. fail2ban has not been running on gitborg-monitoring

Confirmed on the host:

  × fail2ban.service - Fail2Ban Service
       Active: failed (Result: exit-code) since Thu 2026-07-30 19:14:20 UTC
      Process: ExecStart=/usr/bin/fail2ban-server -xf start (code=exited, status=255/EXCEPTION)
    fail2ban-server: ERROR Failed during configuration: Have not found any log file
                           for caddy-auth jail
    fail2ban-server: ERROR Async configuration of server failed

base/defaults enables the caddy-auth jail on every host with a POLLING backend
reading fail2ban_caddy_log_path ({{ gitborg_home }}/caddy/logs/access.log). That
file only exists on the services host, whose Caddy writes `output file
/var/log/caddy/access.log` through a bind mount. THIS host's Caddy logs `output
stdout` into the journal and mounts no log directory at all, so the path is never
created and fail2ban refuses to start.

The failure mode is the notable part: a jail intended as defence-in-depth took the
whole daemon down, so the [sshd] jail — the protection that actually matters on a
box with no Caddy file log — silently did not exist either. The only symptom was
`Enable and start fail2ban` reporting changed on three consecutive applies, which
is exactly the kind of line that gets read as noise. `Duration: 107ms` shows it
died immediately each time the apply started it.

Disabled for the monitoring group rather than ported to the systemd backend:
grafana/stats have their own auth, and a journal filter that silently matched
nothing would be worse than no jail. [sshd] now applies, which is the point.

2. The zoekt index removal reported changed on every apply

  cmd: podman exec forgejo rm -rf .../indexers
  changed_when: _forgejo_index_rm.rc == 0

`rm -rf` exits 0 whether or not anything was there, so this meant "changed on
every single apply, forever" — the deletion is idempotent in effect but the
REPORTING never was. `changed_when: rc == 0` is always wrong on a command designed
to succeed (rm -rf, mkdir -p, touch): the information lives in a probe, not the
exit code.

Now probes first and reports changed only on REMOVED.

Deliberately NOT fixed here: `kanidm : Provision entitlement groups` also reports
changed every run (changed_when is a string match on tool output). Diagnosing it
needs that output, but the task is no_log because the idm_admin password is in its
environment AND kanidm-provision provisions OAuth2 clients, so stdout may carry a
client secret. Printing it to console and CI logs to fix cosmetic churn is a bad
trade; left as a documented known-churn item.
supernaut sammanfogade incheckning 9f68282735 till main 2026-07-30 19:36:41 +00:00
supernaut tog bort grenen fix/fail2ban-monitoring-and-zoekt-churn 2026-07-30 19:36:41 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!264
Ingen beskrivning angiven.