mail: nothing in automation detects a provider-side sending-domain or credential failure #389
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#389
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Raised by #381 and left open when it closed.
senderAllowed()in bitborg-web closes the infra/portal disagreement it was built for: infrasupplies
EMAIL_ALLOWED_SENDER_DOMAINS, the portal checks its hardcodedEMAIL_FROMagainst it, anda divergence becomes a logged refusal instead of silently-rejected mail.
It has no view of the provider at all. On 2026-08-05 infra had declared
email.bitborg.seallowed, so the guard passed — and the provider refused every message. The guard cannot catch that
class of failure by construction, and it is the class that actually caused the outage.
Today the only defence is step 0, a runbook instruction to send a two-domain differential by hand
before moving. It worked, but it is a step a human must remember at exactly the moment they are
confident the move is safe — which is when it gets skipped. It was skipped once already.
What would actually help
A periodic check that the configured sending domain is accepted by the provider right now, so a
dead or de-authorised credential surfaces before the next send that matters rather than during it.
That is a different signal from "the key is deployed" (already covered by comparing sha256
fingerprints) and from "mail was sent" (only observable when something happens to send).
Design notes, mostly about what does not work:
2xxalone is not evidence./channelsanswers200completely unauthenticated. Any probemust pair the real request with a deliberately invalid credential and require the two to differ.
401is not diagnosable — Sweego returns an identical401for a dead key and for a valid keywith an unauthorised sender, and no reachable route separates them. So the check can prove
"something is wrong with our ability to send as this domain" but cannot say which. That is
still worth alerting on; it just must not claim a cause.
up). Worth considering whether the existing cross-probe's transition-only pattern applies — alert
on a state change, not every run.
Acceptance
before trusting a pass, by pointing it at a deliberately wrong sender and confirming it goes red,
then putting it back.
that is not observable from here.