monitoring: no probe covers /user/login — a 500 on the sign-in route is invisible #384
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#384
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
The gap
monitoring_probe_targetsprobeshttps://git.bitborg.se— the root, which serves200regardless of whether anyone can actually sign in. Nothing probes the login entry point.
This is not hypothetical. During the
Gitborg Auth→Bitborg Authauth-source rename (#383),production spent a window in this state:
The canonical sign-in URL was returning a 500 to every visitor. All six probes reported
1theentire time and the only firing alert was
Watchdog, because the root was healthy and no probelooked at
/user/login. It was found by hand, not by monitoring.The same blind spot applies to Grafana (
/api/healthis probed, the OAuth button is not) and toanything else where "the service is up" and "a user can get in" are different questions.
Why the login route is especially exposed
/user/loginis not served by Forgejo in the normal case — Caddy rewrites it:So the route depends on
forgejo_oidc_source_nameagreeing with the live Forgejo auth source name.Those live in two different systems, which means any future rename, restore, or partial apply can
desynchronise them — and the failure is a 500 on the primary way in, with every existing session
still working fine, so nobody notices until a logged-out user complains.
Suggested shape
Add a probe for the login route that asserts the redirect chain, not just reachability:
https://git.bitborg.se/user/loginshould302to/user/oauth2/<source name>, which should307tohttps://auth.bitborg.se/ui/oauth2?....2xx/3xx-tolerant module is not enough on its own — the broken state was a302. Thefailing hop was the second one. Either probe
/user/oauth2/<source name>directly (expect307,and specifically not
5xx), or use a blackbox module withfail_if_body_matches_regexp/an explicit
valid_status_codeson the final hop withfollow_redirects: true.Worth deciding whether this is a
probe_successtarget or a dedicated alert, since a login-routefailure is arguably higher severity than a generic
EndpointDown.Acceptance
it goes red, then point it back. A probe that has never failed is not yet evidence of anything.
Implemented in #392 — open, not applied. The acceptance criterion here requires driving a probe
red, and alerting goes to email and ntfy, so the deliberate-failure step is left for a waking
operator rather than run overnight.
Most of "prove it can fail" is already satisfied read-only, though, against production:
The 500 that hid during the rename window discriminates cleanly, and the probes pin exact status
codes (
[302]/[307]), so neither can pass on it. What still needs an apply is that the probegoes red and that
SignInRouteBrokenroutes.Three decisions worth recording against this issue:
follow_redirectsprobe of/user/login.Following the chain terminates inside Kanidm's
/ui/oauth2, and an anonymous request there logs anInvalid identity: NotAuthenticatedERROR span; that is the documented reason kanidm is probed at/statusrather than root, and following would reintroduce it.Locationheader, which this issue's suggested shape does not. A302tothe wrong slug is still a
302— and Caddy disagreeing with the live source name is exactly thefailure described above, so a status-only check would miss it.
SignInRouteBrokenalert, answering the "probe_success target or dedicated alert?"question here: dedicated. The generic wording would actively mislead, because everything else is
green during this failure.
One open question I could not settle:
/user/oauth2/<source>begins an OAuth flow, so Forgejo mintssession state per probe (the 307 carries a fresh
stateandcode_challenge). Interval set to 60sto halve the churn, but the session store's actual growth is unmeasured — worth confirming it is
bounded before trusting hop 2 long-term.
Done — applied and verified
PR #392 merged and applied 2026-08-06. Two probes live on the monitoring host:
Apply made exactly the 6 predicted changed tasks (3 renders + 3 restarts), monitoring host only;
gitborg-prodstayedchanged=0, confirming theforgejo_oidc_source_namemove togroup_vars/allis inert on the services host. Both targets register
health: up, both reportprobe_success=1inVictoriaMetrics,
SignInRouteBrokenis loaded in vmalert with a clean log, and onlyWatchdogfires."Prove it can fail" — done at the probe level, with matched controls
Each failure case is paired with its own positive control, run directly against the blackbox modules
so no config was touched and no alert was fired:
probe_success500307301302?redirect_to=/explore)302The last row matters as much as the failures. Caddy appends
?{query}, so a signed-out user arrivingat
/user/login?redirect_to=…producesLocation: …/user/oauth2/Bitborg%20Auth?redirect_to=%2Fexplore.The pattern carries an optional query group specifically for this; without it, every such visitor
would have made the probe red — a permanent false outage, and precisely the kind of probe-that-cries-
wolf this issue warns against.
What was NOT tested, and why — read this before trusting it further
Two honest gaps:
The
Locationassertion has no isolated negative. The hop-1 failure above was rejected onstatus (
301≠302), not by the header regex —probe_failed_due_to_regexstayed0. Provingthe regex specifically needs a target returning exactly
302with a wrongLocation, which doesnot exist on this estate without a config change. The assertion is known to be evaluated (it
reports
probe_failed_due_to_regex 0rather than absent) and known to pass on both real shapes,but a live differential against a mismatching Location was not obtained.
The end-to-end alert path was deliberately not exercised. Decided rather than skipped: firing
it means a real critical page to
admin@bitborg.seand ntfy, and what it would add isprobe_success=0 → vmalert → Alertmanager → email/ntfy, whichEndpointDownalready exercises onthe same wiring, plus a rule whose expression is
EndpointDown's with a job matcher swapped in.The substantive risk this issue names — a probe that has never failed and is therefore not
evidence — is disproven above.
If either gap ever matters, the recipe is: set
forgejo_oidc_source_nameto a wrong value and apply--limit gitborg-monitoring— that moves only the probe target, leaving Caddy and Forgejo aloneso real sign-in keeps working. Forgetting the
--limitbreaks actual logins.Also worth knowing
forgejo_oidc_source_namemoved fromroles/forgejo/defaults/togroup_vars/all/, becausethe probe runs on the monitoring host where the forgejo role never runs, so a role default is
undefined there. Caught by an
ansible -m debugdifferential before anything depended on it. It isdeliberately not mirrored back as a fallback — same reasoning as
alert_email_from.EndpointDown(no double page) and fromCertificateExpiringSoon(they share a cert with the root target and would otherwise fire three warnings for one cert).
⚠️ Open, not resolved here
GET /user/oauth2/<source>begins an OAuth flow — the 307 carries a freshstateandcode_challenge, so Forgejo mints session state on every probe and blackbox keeps no cookies.Interval is 60s rather than the global 30s to halve that churn, but the session store's actual
growth is unmeasured. Worth confirming it is bounded; if it is not, hop 2 needs rethinking.