email: a wrong-account API key broke all portal mail silently — no reason logged, nothing verifies the sender #115
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-web#115
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
A misconfigured Sweego API key broke every portal email for roughly 25 minutes and the system
reported almost nothing useful. Sign-ups completed, accounts were provisioned, and no
credential-reset link was ever delivered. The only evidence was one line:
Diagnosis was possible only because the operator happened to know the provider account was wrong.
Nothing in the code, CI, or the infra health gates could have told us.
What happened
The API key in use was valid, but issued from a different Sweego account — one with no verified
mail.gitborg.sesending domain. So it authenticated (hence 422, not 401) and Sweego thenrejected the sender. Meanwhile Forgejo's SMTP path used credentials from the correct account and
delivered fine with
dkim=pass, so the outage looked partial and contradictory. Fixed operationallyin bitborg-infra#273; this issue is about the fact that we could not see it.
Gap 1 — the failure reason is discarded
sendEmailinsrc/lib/email.tskeeps the status and throws the body away:Sweego returns a JSON body explaining a 422. We never read it, so a precise, self-describing error
was reduced to a bare integer.
Proposal: on a non-2xx, read the body (size-capped, e.g. first 512 bytes) and log it alongside
the status. Do not put it in the returned value if that risks reaching a user-facing surface —
this is operator telemetry. Take care not to log the recipient address any more than we already do,
and never the
Api-Keyheader.Gap 2 — nothing verifies the key's account can send as us
EMAIL_FROMis a hardcoded constant; the API key arrives as a podman secret from infra. Any validkey for any account satisfies every check we have —
pnpm check, CI, the e2e suite, thehealth-checkrole, and the Ansible apply itself all pass. The two values must agree and nothingasserts that they do.
This is the same defect class as #113 (the From address duplicated across two repos with nothing
keeping them in sync), and arguably the sharper half: #113 is drift risk, this is an
authorisation pairing that fails silently in production.
Options:
lists
EMAIL_FROM's domain as verified. Fail loudly — or at least log an unmissable warning —rather than discovering it at the first sign-up. Needs a Sweego endpoint that exposes verified
domains; worth confirming one exists.
health-checkrole to send to a sink address andassert 2xx. Catches the real path end to end, at the cost of a mail per deploy.
three — it still requires a real user to hit the broken path first — but it is the cheapest and
composes with either of the others.
Option 1 plus 3 is probably the right pairing: refuse to start quietly misconfigured, and alert if
sends start failing later.
Gap 3 — best-effort masks systematic failure
Portal email is deliberately best-effort so a mail hiccup cannot fail a sign-up that already
succeeded. That is the right call and should stay. But it means a totally broken sender is
indistinguishable from a transient one, and the affected user is left with an account and no way to
set a password.
Worth considering: surface something to the user when the setup-link send fails outright — a "we
could not send your email, contact support" state, or a self-service resend — so the failure is not
invisible on their side too.
Done when