fail2ban caddy-auth: flip 401 deny-list to a credential-endpoint allow-list #196

Stängd
öppnade 2026-07-21 22:07:14 +00:00 av supernaut · 1 kommentar
Ägare

Problem

The caddy-auth fail2ban jail bans on any repeated HTTP 401, with a growing ignoreregex exclusion list for endpoints that answer 401 by protocol (not by wrong credential). This deny-list design has now broken CI twice:

  • #127/#175 → #179: the container-registry /v2/ handshake 401s banned the runner mid-push.
  • #195: the Forgejo Actions runner gRPC protocol /api/actions/runner.v1.RunnerService/* 401s banned the shared runner egress IP → runners couldn't declare → boot-loop → Action run 133 stuck 'waiting' for hours.

Every new endpoint that returns 401-by-protocol is a latent CI outage waiting to happen. Whack-a-mole exclusions don't scale.

Proposal

Invert the filter: instead of ban all 401s except an exclusion list, only count 401s on genuine credential-auth endpoints (git-over-HTTPS Basic auth paths, /user/login, /api/v1 token auth). Protocol handshakes (/v2/, /api/actions/, and any future ones) then never trip the jail without needing per-endpoint exclusions.

This is security-sensitive: the allow-list must enumerate the real brute-forceable surfaces carefully so we don't weaken brute-force protection. Verify with fail2ban-regex against a live access-log sample (both directions: real 401s still matched, protocol 401s still ignored) before deploying.

Also

  • RunnerQueueStalled fired only as warning and lagged the stall onset by ~3h — review its severity/for/routing so a runner boot-loop pages promptly instead of running silently.
  • Consider a pre-promotion smoke test of the runner image + a controller-side guard that flags 'runners boot but never declare (agent_labels null / never online)' as a distinct alert from plain queue depth.
## Problem The `caddy-auth` fail2ban jail bans on **any** repeated HTTP 401, with a growing `ignoreregex` exclusion list for endpoints that answer 401 *by protocol* (not by wrong credential). This deny-list design has now broken CI **twice**: - **#127/#175 → #179**: the container-registry `/v2/` handshake 401s banned the runner mid-push. - **#195**: the Forgejo Actions runner gRPC protocol `/api/actions/runner.v1.RunnerService/*` 401s banned the shared runner egress IP → runners couldn't declare → boot-loop → Action run 133 stuck 'waiting' for hours. Every new endpoint that returns 401-by-protocol is a latent CI outage waiting to happen. Whack-a-mole exclusions don't scale. ## Proposal Invert the filter: instead of *ban all 401s except an exclusion list*, **only count 401s on genuine credential-auth endpoints** (git-over-HTTPS Basic auth paths, `/user/login`, `/api/v1` token auth). Protocol handshakes (`/v2/`, `/api/actions/`, and any future ones) then never trip the jail without needing per-endpoint exclusions. This is security-sensitive: the allow-list must enumerate the real brute-forceable surfaces carefully so we don't *weaken* brute-force protection. Verify with `fail2ban-regex` against a live access-log sample (both directions: real 401s still matched, protocol 401s still ignored) before deploying. ## Also - `RunnerQueueStalled` fired only as **warning** and lagged the stall onset by ~3h — review its severity/`for`/routing so a runner boot-loop pages promptly instead of running silently. - Consider a pre-promotion smoke test of the runner image + a controller-side guard that flags 'runners boot but never declare (agent_labels null / never online)' as a distinct alert from plain queue depth.
Upphovsperson
Ägare

Follow-up complete: the RunnerQueueStalled severity/lag review (noted at close) shipped in PR #206 — added RunnerQueueStalledCritical (queued_jobs > 0 for: 30m, severity critical → pages), applied to the monitoring host + verified. Data confirmed the existing warning tier fired correctly ~15 min into the sustained stall; the only gap was that warning doesn't page, now covered by the 30-min critical escalation. No separate issue was tracked for this small follow-up.

Follow-up complete: the RunnerQueueStalled severity/lag review (noted at close) shipped in **PR #206** — added `RunnerQueueStalledCritical` (`queued_jobs > 0` `for: 30m`, severity **critical** → pages), applied to the monitoring host + verified. Data confirmed the existing warning tier fired correctly ~15 min into the sustained stall; the only gap was that warning doesn't page, now covered by the 30-min critical escalation. No separate issue was tracked for this small follow-up.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#196
Ingen beskrivning angiven.