fix(caddy): tune the ADR 0032 rate limits from measured 429s #309

Sammanfogat
supernaut sammanfogade 1 incheckning från fix/caddy-ratelimit-tuning in i main 2026-08-01 16:58:05 +00:00
Ägare

Closes #302, closes #304.

⚠ Apply this first — it fixes a live bug

Over 30 days of Caddy access logs there were 59 HTTP 429s in total, and not one was abuse. All three
clusters were first-party traffic being throttled:

Count Client Host Path Verdict
44 runner shared SNAT git /api/actions/runner.v1.RunnerService/{UpdateTask,UpdateLog} live bug, all inside one minute (2026-08-01 11:56Z)
9 one human www GET /signup bug — the zone counted page views
6 127.0.0.1 git /api/v1/* the reconciler; the known six-minute incident

The runner cluster is the urgent one: CI agents are being rate-limited while reporting task status and
logs.
private_ranges cannot save them, because the runners reach the edge from a public SNAT
address — the same shared-egress shape as the fail2ban ban in #195. They need a path exemption, which
this adds.

So nothing was tightened

The ADR asks for tuning from observed 429s. The observation is that the limits are not too loose. Peak
per-minute per client IP on each zone's own match set: forgejo_general 293 (runner egress, against
a 300 limit) and 112 (a human browsing); forgejo_api 108 against a 120 limit; auth 19 for
one interactive login.

Changes:

  • Exempt /api/actions/* from forgejo_general. The live bug.
  • forgejo_api 120 → 300. One legitimate client was already at 90% of the limit, and that figure is
    a floor: these are 1-minute counts sampled every 10 minutes.
  • auth stays at 300. The ADR's literal ~20/min would 429 a real interactive login with one request
    of margin, on the only sign-in path there is.

The signup zone was wrong in both directions (#304)

Rescoped to POST /api/signup, /api/captcha/redeem and /api/resend-setup-link, raised 10 → 30.

  • Too tight: the zone counted GET /signup, so page views spent the budget meant for sign-ups —
    which is what produced the 9 human 429s.
  • Too loose, and worse: Caddy's path matcher is an exact match, so /signup/ (30 hits) and
    /en/signup (5 hits) sat outside the zone entirely. A trailing slash bypassed the limit
    altogether. That was not in the issue and is the more serious of the two.
  • /api/resend-setup-link sends mail and had no edge limit at all. Now covered.
  • Every zone, signup included, now carries not remote_ip private_ranges.

bitborg-web's in-process limiter is not fine either

Reported rather than changed — it is a different repo, filed separately. It keys on clientIp() with no
IPv6 prefixing, so it has the same shared-egress problem, per full IPv6 address. The consequence that
matters here: at 10 events the edge zone was tighter than the app limiter it exists to back up, so
the inner ring never engaged at all. At 30 the layering is the right way up.

New auth_credentials zone — built, shipped DISABLED

The issue's actual ask was to establish which paths are credential submission. They are POST /v1/auth
and POST /v1/reauth — exactly 3 POSTs per login attempt, with everything else on that host a GET. The
zone is written at 30/min (≈10 attempts/min) and defaults to false, because enabling a new limit on
the instance's only sign-in path on announcement day trades a theoretical attack for a real lockout. The
runbook says how and when to flip it.

Verified

--syntax-check clean; ansible-lint on the caddy role unchanged from main. The new nested matcher
sets were run through the same parser stock Caddy uses for named matchers and then caddy validate →
"Valid configuration", with the file rendered both with auth_credentials on and off.

Behaviourally probed against a locally-run Caddy: the three POSTs are inside the signup zone while
GET /signup, /signup/, /en/signup and POST /api/captcha/challenge are not; the web UI is inside
forgejo_general while /api/actions/*, /api/v1/*, /v2/*, git transport and LFS batch are not;
POST /v1/auth and /v1/reauth are inside auth_credentials while GET /v1/auth, /pkg/style.css and
/v1/oauth2/* are not; and probing from 127.0.0.1 falls out of zone, confirming the exemption works.

Applying

--tags caddy. Config only, no manual prerequisite. The diff is the rendered Caddyfile, so Caddy
restarts
(~1 s) — the role runs caddy validate in a throwaway container with the module-bearing image
first, so a bad file fails the play there rather than on the live edge.

Verify after: a CI run completes with no 429s on /api/actions/*, and
sum by (request_uri) (count_over_time({job="caddy-access"} | json | status="429" [1h])) returns nothing.

The final threshold table, the reproduction LogQL and the traps are in
docs/runbook.md → Edge rate limiting (ADR 0032). ADR 0032 itself lives in bitborg-docs; the
threshold table should be copied there.

Closes #302, closes #304. ## ⚠ Apply this first — it fixes a live bug Over 30 days of Caddy access logs there were **59 HTTP 429s in total, and not one was abuse.** All three clusters were first-party traffic being throttled: | Count | Client | Host | Path | Verdict | | --- | --- | --- | --- | --- | | 44 | runner shared SNAT | git | `/api/actions/runner.v1.RunnerService/{UpdateTask,UpdateLog}` | **live bug**, all inside one minute (2026-08-01 11:56Z) | | 9 | one human | www | `GET /signup` | **bug** — the zone counted page views | | 6 | `127.0.0.1` | git | `/api/v1/*` | the reconciler; the known six-minute incident | The runner cluster is the urgent one: **CI agents are being rate-limited while reporting task status and logs.** `private_ranges` cannot save them, because the runners reach the edge from a **public** SNAT address — the same shared-egress shape as the fail2ban ban in #195. They need a path exemption, which this adds. ## So nothing was tightened The ADR asks for tuning from observed 429s. The observation is that the limits are not too loose. Peak per-minute per client IP on each zone's own match set: `forgejo_general` **293** (runner egress, against a 300 limit) and **112** (a human browsing); `forgejo_api` **108** against a 120 limit; `auth` **19** for one interactive login. Changes: - **Exempt `/api/actions/*` from `forgejo_general`.** The live bug. - **`forgejo_api` 120 → 300.** One legitimate client was already at 90% of the limit, and that figure is a floor: these are 1-minute counts sampled every 10 minutes. - **`auth` stays at 300.** The ADR's literal ~20/min would 429 a real interactive login with one request of margin, on the only sign-in path there is. ## The signup zone was wrong in both directions (#304) Rescoped to `POST /api/signup`, `/api/captcha/redeem` and `/api/resend-setup-link`, raised 10 → 30. - **Too tight:** the zone counted `GET /signup`, so page views spent the budget meant for sign-ups — which is what produced the 9 human 429s. - **Too loose, and worse:** Caddy's `path` matcher is an **exact** match, so `/signup/` (30 hits) and `/en/signup` (5 hits) sat **outside the zone entirely**. A trailing slash bypassed the limit altogether. That was not in the issue and is the more serious of the two. - `/api/resend-setup-link` sends mail and had **no** edge limit at all. Now covered. - Every zone, signup included, now carries `not remote_ip private_ranges`. ## `bitborg-web`'s in-process limiter is not fine either Reported rather than changed — it is a different repo, filed separately. It keys on `clientIp()` with no IPv6 prefixing, so it has the same shared-egress problem, per full IPv6 address. The consequence that matters here: at 10 events **the edge zone was tighter than the app limiter it exists to back up**, so the inner ring never engaged at all. At 30 the layering is the right way up. ## New `auth_credentials` zone — built, shipped DISABLED The issue's actual ask was to establish which paths are credential submission. They are `POST /v1/auth` and `POST /v1/reauth` — exactly 3 POSTs per login attempt, with everything else on that host a GET. The zone is written at 30/min (≈10 attempts/min) and defaults to **`false`**, because enabling a new limit on the instance's only sign-in path on announcement day trades a theoretical attack for a real lockout. The runbook says how and when to flip it. ## Verified `--syntax-check` clean; `ansible-lint` on the caddy role unchanged from `main`. The new nested matcher sets were run through the same parser stock Caddy uses for named matchers and then `caddy validate` → **"Valid configuration"**, with the file rendered both with `auth_credentials` on and off. Behaviourally probed against a locally-run Caddy: the three POSTs are inside the signup zone while `GET /signup`, `/signup/`, `/en/signup` and `POST /api/captcha/challenge` are not; the web UI is inside `forgejo_general` while `/api/actions/*`, `/api/v1/*`, `/v2/*`, git transport and LFS batch are not; `POST /v1/auth` and `/v1/reauth` are inside `auth_credentials` while `GET /v1/auth`, `/pkg/style.css` and `/v1/oauth2/*` are not; and probing from 127.0.0.1 falls out of zone, confirming the exemption works. ## Applying `--tags caddy`. Config only, no manual prerequisite. The diff is the rendered Caddyfile, so **Caddy restarts** (~1 s) — the role runs `caddy validate` in a throwaway container with the module-bearing image first, so a bad file fails the play there rather than on the live edge. Verify after: a CI run completes with no 429s on `/api/actions/*`, and `sum by (request_uri) (count_over_time({job="caddy-access"} | json | status="429" [1h]))` returns nothing. The final threshold table, the reproduction LogQL and the traps are in `docs/runbook.md → Edge rate limiting (ADR 0032)`. ADR 0032 itself lives in `bitborg-docs`; the threshold table should be copied there.
supernaut lade till 1 incheckning 2026-08-01 14:34:55 +00:00
fix(caddy): tune the ADR 0032 rate limits from measured 429s (#302)
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m29s
0d65ef9107
ADR 0032 requires the thresholds to be tuned from observed 429s and the result
recorded. The Phase 1/2 values were starting points; these are the outcome of
that tuning, measured from Loki `{job="caddy-access"}` over the 30 days to
2026-08-01 14:00Z.

The edge has served 59 HTTP 429s in total, ever, and not one was abuse:

  44  CI runners' shared SNAT egress → /api/actions/runner.v1.RunnerService/
      {UpdateTask,UpdateLog}, all inside one minute (2026-08-01 11:56Z)
   9  one human → GET /signup (2026-07-31 23:01Z, reloading the page)
   6  127.0.0.1 → /api/v1/* — the reconciler, 00:35Z–00:40Z, the six-minute
      incident the private_ranges exemption already fixed

So nothing is tightened here. Two limits are loosened, two match sets are
corrected, and the auth figure is confirmed with numbers instead of reasoning.

forgejo_general: exempt /api/actions/*. This is a live bug, not a tuning nit.
`/api/actions/*` is not `/api/v1/*`, so the Actions runner protocol fell into
the general zone by default, and the ephemeral runners share ONE SNAT egress —
measured peak 293 requests/min from that egress against a 300/min zone, then 44
429s against a running CI job. Note what makes this different from the
reconciler: the runners arrive from a PUBLIC address (runner-controller bakes
the public URL into their cloud-init deliberately), so `not remote_ip
private_ranges` cannot save them. It needs a path exemption, exactly like git
and /v2/. Runner auth failures stay fail2ban's job on that path. Same shape as
the shared-SNAT ban of #195, one layer up.

forgejo_api: 120 → 300/min. Measured peak from a single legitimate public API
client is 108/min — 90% of the old limit, and a floor rather than a peak (1-min
counts sampled every 10 min). A limit one legitimate client routinely brushes is
mistuned.

signup: rescoped to the state-changing POSTs (POST /api/signup,
/api/captcha/redeem, /api/resend-setup-link) and raised 10 → 30/min. Two
measured problems with the old set. It counted page views: one genuine sign-up
cost four events — GET /signup, both Cap captcha calls, the submit — so at
10/min a visitor who mistyped and resubmitted burned ~7 of 10 alone, and people
merely reading the page spent the same budget as conversions. Under shared IPv4
(corporate NAT, CGNAT) that budget is shared across the whole egress, which is
exactly the traffic an announcement produces, and a 429 mid-captcha reads as a
broken site at the moment of conversion. Second, `path` is an EXACT match, so
`/signup/` (30 hits) and `/en/signup` (5 hits) sat outside the zone entirely — a
trailing slash bypassed the limit. Counting POSTs fixes both and closes the
hole. /api/resend-setup-link joins the set: it sends mail, making it the most
abusable endpoint on the host, and it had no edge limit at all.

That also un-inverts the layering. gitborg-web's in-process limiter allows 5
POST /api/signup, 30 captcha calls and 3 resends per minute per IP
(src/lib/rate-limit.ts) — and it is keyed on the client IP with no IPv6
prefixing, so it shares the shared-egress caveat. At 10 events the edge zone was
TIGHTER than the app limiter it exists to back up, so the inner ring never
engaged. The edge is now the looser, outer ring.

auth: stays 300/min host-wide, now with a measurement behind it. One interactive
login costs 19 requests in one minute from a single IP (Kanidm serves its SPA
assets and several XHRs from the same host), so the ADR's literal ~20/min would
have 429'd a real user mid-sign-in with one request of margin — on the only
sign-in path there is. That host has never served a 429.

auth_credentials: the tighter zone the ADR wanted, built and shipped DISABLED.
The issue noted this "needs someone to establish which paths those actually
are"; established — POST /v1/auth (and /v1/reauth), costing exactly 3 POSTs per
login attempt in every sample across 30 days, with everything else on the host a
GET. 30/min is ~10 attempts/min: a 10x tighter brake on credential stuffing than
the host-wide zone, unable to touch the assets a login needs. Default false
because enabling a new zone on the instance's only sign-in path on announcement
day would trade a theoretical attack for a real lockout; the runbook says how and
when to flip it.

Every zone keeps `not remote_ip private_ranges` — including signup, which
previously lacked it. Recorded in docs/runbook.md → Edge rate limiting (ADR
0032): the final thresholds, the LogQL to reproduce the measurement, and the
traps in the order they were learned.

Verified:
- `ansible-playbook site.yml --syntax-check` clean; `ansible-lint` on the caddy
  role unchanged from `main` (2 pre-existing findings).
- The three new/changed matcher sets run through the SAME nested-matcher-set
  parser stock Caddy uses for named matchers: `caddy validate` → "Valid
  configuration".
- Behaviourally probed against a locally-run Caddy: POST /api/signup,
  /api/captcha/redeem and /api/resend-setup-link are in the signup zone while
  GET /signup, /signup/, /en/signup and POST /api/captcha/challenge are not;
  the Forgejo web UI is in forgejo_general while /api/actions/*, /api/v1/*,
  /v2/*, git transport and LFS batch are not; POST /v1/auth and /v1/reauth are
  in auth_credentials while GET /v1/auth, /pkg/style.css and /v1/oauth2/* are
  not. Probing the same matcher from 127.0.0.1 falls out of zone, which is the
  private_ranges exemption doing its job.
- The full rendered Caddyfile adapts to JSON under stock `caddy validate` with
  the module directives off (the only part stock Caddy cannot parse), braces
  balance, and it renders correctly both with auth_credentials off and on. A
  true validate with the module happens on the host: the caddy role runs
  `caddy validate` in a throwaway container using the module-bearing image
  before restarting the live one.
supernaut sammanfogade incheckning f93afb8b4c till main 2026-08-01 16:58:05 +00:00
supernaut tog bort grenen fix/caddy-ratelimit-tuning 2026-08-01 16:58:06 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!309
Ingen beskrivning angiven.