feat(monitoring): probe the sign-in route, not just the root #392
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!392
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "feat/signin-route-probe"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Closes #384.
⚠️ Not applied — deliberately left for a waking operator
Built and gated but not merged or applied. The acceptance criterion requires driving a probe red
to prove it can fail, and alerting goes to email and ntfy — doing that unattended overnight would
page someone awake. Everything that could be proven read-only has been; see below.
Shape: two probes, one hop each, redirects OFF
Two single-hop probes rather than one
follow_redirectsprobe of/user/login:/ui/oauth2), and an anonymous request there makesKanidm log an
Invalid identity: NotAuthenticatedERROR span. That is the documented reasonmonitoring_probe_targetsprobes kanidm at/statusrather than root — worth ~2900 false errorlines a day. Following would reintroduce exactly that.
2xx/3xx module passes straight through the failure this exists to catch.
Hop 1 asserts the
Locationheader, not just the status. This closes a gap in the issue's ownsuggestion: a
302to the wrong slug is still a302, and Caddy-vs-Forgejo disagreement about thesource name is precisely the failure mode. Status alone would not see it.
Evidence gathered read-only, against production
The third line is most of the "prove it can fail" criterion, obtained without applying anything or
turning a production probe red: the 500 that hid during the rename window is reproducible on
demand, and hop 2 pins
valid_status_codes: [302]/[307]so it cannot pass on a 500.What still needs an apply to confirm: that the probe itself goes red and that
SignInRouteBrokenroutes. Point
forgejo_oidc_source_nameat a wrong name, apply, confirm red, put it back.Also measured, and it shaped the design: Caddy emits an absolute Location with no trailing
?on an empty query. Both the host prefix and query suffix are optional groups in the pattern —matching what is served without pinning so tightly that a user arriving at
/user/login?redirect_to=…reads as an outage.forgejo_oidc_source_namemoves togroup_vars/allIt has to. The probe runs from the monitoring host, where the forgejo role never runs, so a role
default is simply undefined there. Verified by differential before moving it:
Deliberately not mirrored back as a role default "so the role still renders standalone" — that is
the pattern already removed from
alert_email_fromfor cause, and here a stale fallback would be aprobe that passes against the wrong source name, which is worse than no probe.
local/loads neithergroup_vars nor role defaults and sets
forgejo_enable_internal_signin: true, so it never evaluatesthis.
Verified inert:
site.yml --check --limit gitborg-prod→changed=0. The move changes where thevalue is declared, not what anything renders.
Alerting
SignInRouteBrokenis its own alert rather thanEndpointDown, because the two need different firstmoves. This failure looks like nothing is wrong — root probe green, Forgejo healthy, every existing
session still working, only signed-out users affected and they cannot tell you. A generic "Endpoint
down" would send the operator to check Forgejo's health, which is fine. The annotation names the
actual suspect and says which job narrows it to which side.
The sign-in jobs are excluded from
EndpointDown(no double page) and fromCertificateExpiringSoon— they share a certificate with the existing root target, so including themwould fire three identical cert warnings for one cert.
⚠️ Open question, not resolved here
GET /user/oauth2/<source>begins an OAuth flow — the 307 carries a freshstateandcode_challenge, so Forgejo mints session state per probe, and blackbox keeps no cookies. I setscrape_interval: 60son these jobs rather than the global 30s to halve the churn (alerting waits 5mregardless, so nothing is detected meaningfully later), but I have not measured the actual growth
of Forgejo's session store. Confirm it is bounded before trusting this long-term; if it is not, hop
2 needs rethinking. Flagged in the template too.
Gates
--syntax-checkclean--check --limit gitborg-monitoring→ 6 changed, 0 failed: renders scrape config, blackbox configand alert rules, restarts victoria-metrics, blackbox-exporter, vmalert. Monitoring host only.
--check --limit gitborg-prod→changed=0