fix(fail2ban): exclude /api/actions runner-protocol 401s from caddy-auth #195

Sammanfogat
supernaut sammanfogade 2 incheckningar från fix/caddy-auth-ignore-actions-401 in i main 2026-07-21 22:15:45 +00:00
Ägare

Root cause (Action run 133 stuck 'waiting' for hours)

The Forgejo Actions runner gRPC protocol (/api/actions/runner.v1.RunnerService/*, e.g. UpdateTask/UpdateLog) answers 401 by protocol during the ephemeral runner's token/handshake lifecycle — not because a credential is wrong.

Our ephemeral runners share one SNAT egress IP (no floating IP), so those 401s pile up fast. The caddy-auth fail2ban jail counted them and banned the shared egress IP at nftables. Every subsequent runner was then dropped before it could declare itself — its registration row stayed agent_labels = null, last_online = epoch (never online) — so queued jobs (runs_on: [ci]) never matched a runner, the job sat waiting, and the runner-controller boot-looped (boot → can't declare → one-job exits → poweroff → reap → repeat), each cycle refreshing the ban.

This is the same class as the /v2/ registry-handshake regression (#127/#175, fixed in #179) — one endpoint further along.

Fix

Extend the caddy-auth ignoreregex alternation to also exclude /api/actions/ 401s.

Verification (fail2ban-regex against the live access log)

filter matched (would-ban)
old (/v2/ only) 10 — all UpdateTask/UpdateLog (= maxretry → ban)
new (+/api/actions/) 0

Real credential brute force (/user/login, /api/v1 tokens, git-HTTPS Basic auth) still hits failregex and is jailed.

Applied to prod + unbanned the egress IP: runner declared [ci] + came online within seconds, run 133 moved waiting → running, queue drained to 0.

Follow-up

Filing a separate issue for the systemic fix: flip the filter from a deny-list (ban all 401s minus a growing exclusion list) to an allow-list of genuine credential endpoints, so the next protocol-401 endpoint can't re-break CI. Plus verify the RunnerQueueStalled alert fires (this ran silently).

## Root cause (Action run 133 stuck 'waiting' for hours) The Forgejo Actions runner gRPC protocol (`/api/actions/runner.v1.RunnerService/*`, e.g. `UpdateTask`/`UpdateLog`) answers **401 by protocol** during the ephemeral runner's token/handshake lifecycle — not because a credential is wrong. Our ephemeral runners share **one SNAT egress IP** (no floating IP), so those 401s pile up fast. The `caddy-auth` fail2ban jail counted them and **banned the shared egress IP at nftables**. Every subsequent runner was then dropped before it could declare itself — its registration row stayed `agent_labels = null`, `last_online = epoch` (never online) — so queued jobs (`runs_on: [ci]`) never matched a runner, the job sat `waiting`, and the runner-controller boot-looped (boot → can't declare → `one-job` exits → `poweroff` → reap → repeat), each cycle refreshing the ban. This is the **same class** as the `/v2/` registry-handshake regression (#127/#175, fixed in #179) — one endpoint further along. ## Fix Extend the `caddy-auth` `ignoreregex` alternation to also exclude `/api/actions/` 401s. ## Verification (fail2ban-regex against the live access log) | filter | matched (would-ban) | |---|---| | old (`/v2/` only) | **10** — all `UpdateTask`/`UpdateLog` (= maxretry → ban) | | new (`+/api/actions/`) | **0** | Real credential brute force (`/user/login`, `/api/v1` tokens, git-HTTPS Basic auth) still hits `failregex` and is jailed. Applied to prod + unbanned the egress IP: runner declared `[ci]` + came online within seconds, run 133 moved `waiting → running`, queue drained to 0. ## Follow-up Filing a separate issue for the systemic fix: flip the filter from a deny-list (ban all 401s minus a growing exclusion list) to an allow-list of genuine credential endpoints, so the next protocol-401 endpoint can't re-break CI. Plus verify the `RunnerQueueStalled` alert fires (this ran silently).
supernaut lade till 1 incheckning 2026-07-21 22:06:13 +00:00
fix(fail2ban): exclude /api/actions runner-protocol 401s from caddy-auth
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m25s
83c7375b02
The Forgejo Actions runner gRPC protocol (/api/actions/runner.v1.RunnerService/*,
e.g. UpdateTask/UpdateLog) answers 401 by protocol during the ephemeral runner's
token/handshake lifecycle — not a wrong credential. Our ephemeral runners share one
SNAT egress IP (no floating IP), so those 401s piled up and the caddy-auth jail
banned the shared egress at nftables. Every runner was then dropped before it could
declare itself (agent_labels stayed null, never online), so queued jobs never got
claimed and the runner boot-looped for hours (Action run 133 stuck 'waiting').

Same class as the /v2/ registry-handshake regression (#127/#175, fixed in #179),
one endpoint further along. Extend the ignoreregex alternation to also exclude
/api/actions/ 401s. Verified with fail2ban-regex against the live access log:
old filter matched 10 (=maxretry) UpdateTask/UpdateLog 401s (would-ban); new
filter matches 0. Real credential brute force (/user/login, /api/v1, git HTTPS
Basic auth) is unaffected.
supernaut lade till 1 incheckning 2026-07-21 22:08:15 +00:00
docs(runbook): diagnose stuck-waiting jobs / runner boot-loop from fail2ban egress ban
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m26s
d55eb45d2c
The action_runner.agent_labels=null + last_online=epoch signal identifies a runner
blocked before it can declare itself; usual cause is fail2ban banning the runners'
shared SNAT egress IP. Cross-refs #195 (/api/actions fix) and #196 (allow-list).
supernaut sammanfogade incheckning 61b7c5a06c till main 2026-07-21 22:15:45 +00:00
supernaut tog bort grenen fix/caddy-auth-ignore-actions-401 2026-07-21 22:15:45 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!195
Ingen beskrivning angiven.