fix(caddy): tune the ADR 0032 rate limits from measured 429s #309
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!309
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/caddy-ratelimit-tuning"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Closes #302, closes #304.
⚠ Apply this first — it fixes a live bug
Over 30 days of Caddy access logs there were 59 HTTP 429s in total, and not one was abuse. All three
clusters were first-party traffic being throttled:
/api/actions/runner.v1.RunnerService/{UpdateTask,UpdateLog}GET /signup127.0.0.1/api/v1/*The runner cluster is the urgent one: CI agents are being rate-limited while reporting task status and
logs.
private_rangescannot save them, because the runners reach the edge from a public SNATaddress — the same shared-egress shape as the fail2ban ban in #195. They need a path exemption, which
this adds.
So nothing was tightened
The ADR asks for tuning from observed 429s. The observation is that the limits are not too loose. Peak
per-minute per client IP on each zone's own match set:
forgejo_general293 (runner egress, againsta 300 limit) and 112 (a human browsing);
forgejo_api108 against a 120 limit;auth19 forone interactive login.
Changes:
/api/actions/*fromforgejo_general. The live bug.forgejo_api120 → 300. One legitimate client was already at 90% of the limit, and that figure isa floor: these are 1-minute counts sampled every 10 minutes.
authstays at 300. The ADR's literal ~20/min would 429 a real interactive login with one requestof margin, on the only sign-in path there is.
The signup zone was wrong in both directions (#304)
Rescoped to
POST /api/signup,/api/captcha/redeemand/api/resend-setup-link, raised 10 → 30.GET /signup, so page views spent the budget meant for sign-ups —which is what produced the 9 human 429s.
pathmatcher is an exact match, so/signup/(30 hits) and/en/signup(5 hits) sat outside the zone entirely. A trailing slash bypassed the limitaltogether. That was not in the issue and is the more serious of the two.
/api/resend-setup-linksends mail and had no edge limit at all. Now covered.not remote_ip private_ranges.bitborg-web's in-process limiter is not fine eitherReported rather than changed — it is a different repo, filed separately. It keys on
clientIp()with noIPv6 prefixing, so it has the same shared-egress problem, per full IPv6 address. The consequence that
matters here: at 10 events the edge zone was tighter than the app limiter it exists to back up, so
the inner ring never engaged at all. At 30 the layering is the right way up.
New
auth_credentialszone — built, shipped DISABLEDThe issue's actual ask was to establish which paths are credential submission. They are
POST /v1/authand
POST /v1/reauth— exactly 3 POSTs per login attempt, with everything else on that host a GET. Thezone is written at 30/min (≈10 attempts/min) and defaults to
false, because enabling a new limit onthe instance's only sign-in path on announcement day trades a theoretical attack for a real lockout. The
runbook says how and when to flip it.
Verified
--syntax-checkclean;ansible-linton the caddy role unchanged frommain. The new nested matchersets were run through the same parser stock Caddy uses for named matchers and then
caddy validate→"Valid configuration", with the file rendered both with
auth_credentialson and off.Behaviourally probed against a locally-run Caddy: the three POSTs are inside the signup zone while
GET /signup,/signup/,/en/signupandPOST /api/captcha/challengeare not; the web UI is insideforgejo_generalwhile/api/actions/*,/api/v1/*,/v2/*, git transport and LFS batch are not;POST /v1/authand/v1/reauthare insideauth_credentialswhileGET /v1/auth,/pkg/style.cssand/v1/oauth2/*are not; and probing from 127.0.0.1 falls out of zone, confirming the exemption works.Applying
--tags caddy. Config only, no manual prerequisite. The diff is the rendered Caddyfile, so Caddyrestarts (~1 s) — the role runs
caddy validatein a throwaway container with the module-bearing imagefirst, so a bad file fails the play there rather than on the live edge.
Verify after: a CI run completes with no 429s on
/api/actions/*, andsum by (request_uri) (count_over_time({job="caddy-access"} | json | status="429" [1h]))returns nothing.The final threshold table, the reproduction LogQL and the traps are in
docs/runbook.md → Edge rate limiting (ADR 0032). ADR 0032 itself lives inbitborg-docs; thethreshold table should be copied there.
ADR 0032 requires the thresholds to be tuned from observed 429s and the result recorded. The Phase 1/2 values were starting points; these are the outcome of that tuning, measured from Loki `{job="caddy-access"}` over the 30 days to 2026-08-01 14:00Z. The edge has served 59 HTTP 429s in total, ever, and not one was abuse: 44 CI runners' shared SNAT egress → /api/actions/runner.v1.RunnerService/ {UpdateTask,UpdateLog}, all inside one minute (2026-08-01 11:56Z) 9 one human → GET /signup (2026-07-31 23:01Z, reloading the page) 6 127.0.0.1 → /api/v1/* — the reconciler, 00:35Z–00:40Z, the six-minute incident the private_ranges exemption already fixed So nothing is tightened here. Two limits are loosened, two match sets are corrected, and the auth figure is confirmed with numbers instead of reasoning. forgejo_general: exempt /api/actions/*. This is a live bug, not a tuning nit. `/api/actions/*` is not `/api/v1/*`, so the Actions runner protocol fell into the general zone by default, and the ephemeral runners share ONE SNAT egress — measured peak 293 requests/min from that egress against a 300/min zone, then 44 429s against a running CI job. Note what makes this different from the reconciler: the runners arrive from a PUBLIC address (runner-controller bakes the public URL into their cloud-init deliberately), so `not remote_ip private_ranges` cannot save them. It needs a path exemption, exactly like git and /v2/. Runner auth failures stay fail2ban's job on that path. Same shape as the shared-SNAT ban of #195, one layer up. forgejo_api: 120 → 300/min. Measured peak from a single legitimate public API client is 108/min — 90% of the old limit, and a floor rather than a peak (1-min counts sampled every 10 min). A limit one legitimate client routinely brushes is mistuned. signup: rescoped to the state-changing POSTs (POST /api/signup, /api/captcha/redeem, /api/resend-setup-link) and raised 10 → 30/min. Two measured problems with the old set. It counted page views: one genuine sign-up cost four events — GET /signup, both Cap captcha calls, the submit — so at 10/min a visitor who mistyped and resubmitted burned ~7 of 10 alone, and people merely reading the page spent the same budget as conversions. Under shared IPv4 (corporate NAT, CGNAT) that budget is shared across the whole egress, which is exactly the traffic an announcement produces, and a 429 mid-captcha reads as a broken site at the moment of conversion. Second, `path` is an EXACT match, so `/signup/` (30 hits) and `/en/signup` (5 hits) sat outside the zone entirely — a trailing slash bypassed the limit. Counting POSTs fixes both and closes the hole. /api/resend-setup-link joins the set: it sends mail, making it the most abusable endpoint on the host, and it had no edge limit at all. That also un-inverts the layering. gitborg-web's in-process limiter allows 5 POST /api/signup, 30 captcha calls and 3 resends per minute per IP (src/lib/rate-limit.ts) — and it is keyed on the client IP with no IPv6 prefixing, so it shares the shared-egress caveat. At 10 events the edge zone was TIGHTER than the app limiter it exists to back up, so the inner ring never engaged. The edge is now the looser, outer ring. auth: stays 300/min host-wide, now with a measurement behind it. One interactive login costs 19 requests in one minute from a single IP (Kanidm serves its SPA assets and several XHRs from the same host), so the ADR's literal ~20/min would have 429'd a real user mid-sign-in with one request of margin — on the only sign-in path there is. That host has never served a 429. auth_credentials: the tighter zone the ADR wanted, built and shipped DISABLED. The issue noted this "needs someone to establish which paths those actually are"; established — POST /v1/auth (and /v1/reauth), costing exactly 3 POSTs per login attempt in every sample across 30 days, with everything else on the host a GET. 30/min is ~10 attempts/min: a 10x tighter brake on credential stuffing than the host-wide zone, unable to touch the assets a login needs. Default false because enabling a new zone on the instance's only sign-in path on announcement day would trade a theoretical attack for a real lockout; the runbook says how and when to flip it. Every zone keeps `not remote_ip private_ranges` — including signup, which previously lacked it. Recorded in docs/runbook.md → Edge rate limiting (ADR 0032): the final thresholds, the LogQL to reproduce the measurement, and the traps in the order they were learned. Verified: - `ansible-playbook site.yml --syntax-check` clean; `ansible-lint` on the caddy role unchanged from `main` (2 pre-existing findings). - The three new/changed matcher sets run through the SAME nested-matcher-set parser stock Caddy uses for named matchers: `caddy validate` → "Valid configuration". - Behaviourally probed against a locally-run Caddy: POST /api/signup, /api/captcha/redeem and /api/resend-setup-link are in the signup zone while GET /signup, /signup/, /en/signup and POST /api/captcha/challenge are not; the Forgejo web UI is in forgejo_general while /api/actions/*, /api/v1/*, /v2/*, git transport and LFS batch are not; POST /v1/auth and /v1/reauth are in auth_credentials while GET /v1/auth, /pkg/style.css and /v1/oauth2/* are not. Probing the same matcher from 127.0.0.1 falls out of zone, which is the private_ranges exemption doing its job. - The full rendered Caddyfile adapts to JSON under stock `caddy validate` with the module directives off (the only part stock Caddy cannot parse), braces balance, and it renders correctly both with auth_credentials off and on. A true validate with the module happens on the host: the caddy role runs `caddy validate` in a throwaway container using the module-bearing image before restarting the live one.