fail2ban caddy-auth: flip 401 deny-list to a credential-endpoint allow-list #196
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#196
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Problem
The
caddy-authfail2ban jail bans on any repeated HTTP 401, with a growingignoreregexexclusion list for endpoints that answer 401 by protocol (not by wrong credential). This deny-list design has now broken CI twice:/v2/handshake 401s banned the runner mid-push./api/actions/runner.v1.RunnerService/*401s banned the shared runner egress IP → runners couldn't declare → boot-loop → Action run 133 stuck 'waiting' for hours.Every new endpoint that returns 401-by-protocol is a latent CI outage waiting to happen. Whack-a-mole exclusions don't scale.
Proposal
Invert the filter: instead of ban all 401s except an exclusion list, only count 401s on genuine credential-auth endpoints (git-over-HTTPS Basic auth paths,
/user/login,/api/v1token auth). Protocol handshakes (/v2/,/api/actions/, and any future ones) then never trip the jail without needing per-endpoint exclusions.This is security-sensitive: the allow-list must enumerate the real brute-forceable surfaces carefully so we don't weaken brute-force protection. Verify with
fail2ban-regexagainst a live access-log sample (both directions: real 401s still matched, protocol 401s still ignored) before deploying.Also
RunnerQueueStalledfired only as warning and lagged the stall onset by ~3h — review its severity/for/routing so a runner boot-loop pages promptly instead of running silently.Follow-up complete: the RunnerQueueStalled severity/lag review (noted at close) shipped in PR #206 — added
RunnerQueueStalledCritical(queued_jobs > 0for: 30m, severity critical → pages), applied to the monitoring host + verified. Data confirmed the existing warning tier fired correctly ~15 min into the sustained stall; the only gap was that warning doesn't page, now covered by the 30-min critical escalation. No separate issue was tracked for this small follow-up.