feat(monitoring): probe the sign-in route, not just the root #392

Sammanfogat
supernaut sammanfogade 1 incheckning från feat/signin-route-probe in i main 2026-08-06 06:54:51 +00:00
Ägare

Closes #384.

⚠️ Not applied — deliberately left for a waking operator

Built and gated but not merged or applied. The acceptance criterion requires driving a probe red
to prove it can fail, and alerting goes to email and ntfy — doing that unattended overnight would
page someone awake. Everything that could be proven read-only has been; see below.

Shape: two probes, one hop each, redirects OFF

blackbox-signin-redirect   /user/login                      302 + Location must match
blackbox-signin-handoff    /user/oauth2/<source name>       307

Two single-hop probes rather than one follow_redirects probe of /user/login:

  • Following the chain ends inside Kanidm (/ui/oauth2), and an anonymous request there makes
    Kanidm log an Invalid identity: NotAuthenticated ERROR span. That is the documented reason
    monitoring_probe_targets probes kanidm at /status rather than root — worth ~2900 false error
    lines a day. Following would reintroduce exactly that.
  • The broken state WAS a 302. Only the second hop 500'd, so a reachability probe or a lenient
    2xx/3xx module passes straight through the failure this exists to catch.

Hop 1 asserts the Location header, not just the status. This closes a gap in the issue's own
suggestion: a 302 to the wrong slug is still a 302, and Caddy-vs-Forgejo disagreement about the
source name is precisely the failure mode. Status alone would not see it.

Evidence gathered read-only, against production

GET /user/login                        -> 302  Location: https://git.bitborg.se/user/oauth2/Bitborg%20Auth
GET /user/oauth2/Bitborg%20Auth        -> 307  -> https://auth.bitborg.se/ui/oauth2?...
GET /user/oauth2/Nonexistent%20Auth    -> 500     <-- the failure signal, confirmed to discriminate

The third line is most of the "prove it can fail" criterion, obtained without applying anything or
turning a production probe red
: the 500 that hid during the rename window is reproducible on
demand, and hop 2 pins valid_status_codes: [302]/[307] so it cannot pass on a 500.

What still needs an apply to confirm: that the probe itself goes red and that SignInRouteBroken
routes. Point forgejo_oidc_source_name at a wrong name, apply, confirm red, put it back.

Also measured, and it shaped the design: Caddy emits an absolute Location with no trailing
?
on an empty query. Both the host prefix and query suffix are optional groups in the pattern —
matching what is served without pinning so tightly that a user arriving at
/user/login?redirect_to=… reads as an outage.

forgejo_oidc_source_name moves to group_vars/all

It has to. The probe runs from the monitoring host, where the forgejo role never runs, so a role
default is simply undefined there. Verified by differential before moving it:

ansible gitborg-monitoring -m debug
  forgejo_oidc_source_name       = UNDEFINED     <- role-scoped
  bitborg_domain                 = git.bitborg.se
  forgejo_enable_internal_signin = False          <- group-scoped, fine

Deliberately not mirrored back as a role default "so the role still renders standalone" — that is
the pattern already removed from alert_email_from for cause, and here a stale fallback would be a
probe that passes against the wrong source name, which is worse than no probe. local/ loads neither
group_vars nor role defaults and sets forgejo_enable_internal_signin: true, so it never evaluates
this.

Verified inert: site.yml --check --limit gitborg-prod → changed=0. The move changes where the
value is declared, not what anything renders.

Alerting

SignInRouteBroken is its own alert rather than EndpointDown, because the two need different first
moves. This failure looks like nothing is wrong — root probe green, Forgejo healthy, every existing
session still working, only signed-out users affected and they cannot tell you. A generic "Endpoint
down" would send the operator to check Forgejo's health, which is fine. The annotation names the
actual suspect and says which job narrows it to which side.

The sign-in jobs are excluded from EndpointDown (no double page) and from
CertificateExpiringSoon — they share a certificate with the existing root target, so including them
would fire three identical cert warnings for one cert.

⚠️ Open question, not resolved here

GET /user/oauth2/<source> begins an OAuth flow — the 307 carries a fresh state and
code_challenge, so Forgejo mints session state per probe, and blackbox keeps no cookies. I set
scrape_interval: 60s on these jobs rather than the global 30s to halve the churn (alerting waits 5m
regardless, so nothing is detected meaningfully later), but I have not measured the actual growth
of Forgejo's session store.
Confirm it is bounded before trusting this long-term; if it is not, hop
2 needs rethinking. Flagged in the template too.

Gates

  • --syntax-check clean
  • --check --limit gitborg-monitoring → 6 changed, 0 failed: renders scrape config, blackbox config
    and alert rules, restarts victoria-metrics, blackbox-exporter, vmalert. Monitoring host only.
  • --check --limit gitborg-prod → changed=0
Closes #384. ## ⚠️ Not applied — deliberately left for a waking operator Built and gated but **not merged or applied**. The acceptance criterion requires driving a probe red to prove it can fail, and alerting goes to email and ntfy — doing that unattended overnight would page someone awake. Everything that could be proven read-only has been; see below. ## Shape: two probes, one hop each, redirects OFF ```text blackbox-signin-redirect /user/login 302 + Location must match blackbox-signin-handoff /user/oauth2/<source name> 307 ``` Two single-hop probes rather than one `follow_redirects` probe of `/user/login`: - **Following the chain ends inside Kanidm** (`/ui/oauth2`), and an anonymous request there makes Kanidm log an `Invalid identity: NotAuthenticated` ERROR span. That is the documented reason `monitoring_probe_targets` probes kanidm at `/status` rather than root — worth ~2900 false error lines a day. Following would reintroduce exactly that. - **The broken state WAS a 302.** Only the second hop 500'd, so a reachability probe or a lenient 2xx/3xx module passes straight through the failure this exists to catch. **Hop 1 asserts the `Location` header, not just the status.** This closes a gap in the issue's own suggestion: a `302` to the *wrong* slug is still a `302`, and Caddy-vs-Forgejo disagreement about the source name is precisely the failure mode. Status alone would not see it. ## Evidence gathered read-only, against production ```text GET /user/login -> 302 Location: https://git.bitborg.se/user/oauth2/Bitborg%20Auth GET /user/oauth2/Bitborg%20Auth -> 307 -> https://auth.bitborg.se/ui/oauth2?... GET /user/oauth2/Nonexistent%20Auth -> 500 <-- the failure signal, confirmed to discriminate ``` The third line is most of the "prove it can fail" criterion, obtained **without applying anything or turning a production probe red**: the 500 that hid during the rename window is reproducible on demand, and hop 2 pins `valid_status_codes: [302]`/`[307]` so it cannot pass on a 500. What still needs an apply to confirm: that the probe *itself* goes red and that `SignInRouteBroken` routes. Point `forgejo_oidc_source_name` at a wrong name, apply, confirm red, put it back. Also measured, and it shaped the design: Caddy emits an **absolute** Location with **no trailing `?`** on an empty query. Both the host prefix and query suffix are optional groups in the pattern — matching what is served without pinning so tightly that a user arriving at `/user/login?redirect_to=…` reads as an outage. ## `forgejo_oidc_source_name` moves to `group_vars/all` It has to. The probe runs from the **monitoring host**, where the forgejo role never runs, so a role default is simply undefined there. Verified by differential *before* moving it: ```text ansible gitborg-monitoring -m debug forgejo_oidc_source_name = UNDEFINED <- role-scoped bitborg_domain = git.bitborg.se forgejo_enable_internal_signin = False <- group-scoped, fine ``` Deliberately **not** mirrored back as a role default "so the role still renders standalone" — that is the pattern already removed from `alert_email_from` for cause, and here a stale fallback would be a probe that passes against the wrong source name, which is worse than no probe. `local/` loads neither group_vars nor role defaults and sets `forgejo_enable_internal_signin: true`, so it never evaluates this. **Verified inert:** `site.yml --check --limit gitborg-prod` → `changed=0`. The move changes where the value is declared, not what anything renders. ## Alerting `SignInRouteBroken` is its own alert rather than `EndpointDown`, because the two need different first moves. This failure *looks like nothing is wrong* — root probe green, Forgejo healthy, every existing session still working, only signed-out users affected and they cannot tell you. A generic "Endpoint down" would send the operator to check Forgejo's health, which is fine. The annotation names the actual suspect and says which job narrows it to which side. The sign-in jobs are excluded from `EndpointDown` (no double page) and from `CertificateExpiringSoon` — they share a certificate with the existing root target, so including them would fire three identical cert warnings for one cert. ## ⚠️ Open question, not resolved here `GET /user/oauth2/<source>` **begins an OAuth flow** — the 307 carries a fresh `state` and `code_challenge`, so Forgejo mints session state per probe, and blackbox keeps no cookies. I set `scrape_interval: 60s` on these jobs rather than the global 30s to halve the churn (alerting waits 5m regardless, so nothing is detected meaningfully later), **but I have not measured the actual growth of Forgejo's session store.** Confirm it is bounded before trusting this long-term; if it is not, hop 2 needs rethinking. Flagged in the template too. ## Gates - `--syntax-check` clean - `--check --limit gitborg-monitoring` → 6 changed, 0 failed: renders scrape config, blackbox config and alert rules, restarts victoria-metrics, blackbox-exporter, vmalert. **Monitoring host only.** - `--check --limit gitborg-prod` → `changed=0`
supernaut lade till 1 incheckning 2026-08-05 22:04:15 +00:00
feat(monitoring): probe the sign-in route, not just the root
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m33s
e2565e0ee5
The 2026-08-05 rename window served 500 on /user/login while all six
probes read 1, because they probe the root — which is healthy whether or
not anyone can sign in. It was found by hand.

Adds two blackbox probes asserting one hop each, with redirects OFF:

  blackbox-signin-redirect  /user/login              302 + Location match
  blackbox-signin-handoff   /user/oauth2/<source>    307

Two hops rather than one following probe, for two reasons. Following the
chain ends inside Kanidm's /ui/oauth2, and an anonymous request there
logs an `Invalid identity: NotAuthenticated` ERROR span — the same reason
kanidm is probed at /status rather than root. And the broken state WAS a
302: only the second hop failed, so a reachability or lenient 2xx/3xx
probe would have passed straight through it.

Hop 1 asserts the Location header, not just the status. A 302 to the
WRONG slug is still a 302, which is exactly the Caddy/Forgejo
desynchronisation this is meant to catch.

Moves forgejo_oidc_source_name from roles/forgejo/defaults to
group_vars/all. It has to be group-scoped: the probe runs from the
monitoring host, where the forgejo role never runs, so a role default is
undefined there — verified by differential before moving it. Deliberately
not mirrored back as a fallback, following the reasoning already recorded
beside alert_email_from: a stale fallback would be a probe that passes
against the wrong name. Verified inert — prod is changed=0 after the move.

SignInRouteBroken is its own alert rather than EndpointDown because the
two need different first moves; this failure looks like nothing is wrong,
since the root is green and existing sessions keep working. The sign-in
jobs are excluded from EndpointDown and from CertificateExpiringSoon, the
latter because they share a certificate with the root target and would
otherwise fire three warnings for one cert.

Refs #384.
supernaut sammanfogade incheckning 04c03538d8 till main 2026-08-06 06:54:51 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!392
Ingen beskrivning angiven.