email: a wrong-account API key broke all portal mail silently — no reason logged, nothing verifies the sender #115

Stängd
öppnade 2026-07-31 01:35:45 +00:00 av supernaut · 0 kommentarer
Ägare

A misconfigured Sweego API key broke every portal email for roughly 25 minutes and the system
reported almost nothing useful. Sign-ups completed, accounts were provisioned, and no
credential-reset link was ever delivered. The only evidence was one line:

[signup] setup-link email failed (status 422)

Diagnosis was possible only because the operator happened to know the provider account was wrong.
Nothing in the code, CI, or the infra health gates could have told us.

What happened

The API key in use was valid, but issued from a different Sweego account — one with no verified
mail.gitborg.se sending domain. So it authenticated (hence 422, not 401) and Sweego then
rejected the sender. Meanwhile Forgejo's SMTP path used credentials from the correct account and
delivered fine with dkim=pass, so the outage looked partial and contradictory. Fixed operationally
in bitborg-infra#273; this issue is about the fact that we could not see it.

Gap 1 — the failure reason is discarded

sendEmail in src/lib/email.ts keeps the status and throws the body away:

if (res.ok) {
  return { ok: true };
}
return { ok: false, reason: "error", status: res.status };

Sweego returns a JSON body explaining a 422. We never read it, so a precise, self-describing error
was reduced to a bare integer.

Proposal: on a non-2xx, read the body (size-capped, e.g. first 512 bytes) and log it alongside
the status. Do not put it in the returned value if that risks reaching a user-facing surface —
this is operator telemetry. Take care not to log the recipient address any more than we already do,
and never the Api-Key header.

Gap 2 — nothing verifies the key's account can send as us

EMAIL_FROM is a hardcoded constant; the API key arrives as a podman secret from infra. Any valid
key for any account satisfies every check we have
— pnpm check, CI, the e2e suite, the
health-check role, and the Ansible apply itself all pass. The two values must agree and nothing
asserts that they do.

This is the same defect class as #113 (the From address duplicated across two repos with nothing
keeping them in sync), and arguably the sharper half: #113 is drift risk, this is an
authorisation pairing that fails silently in production.

Options:

  1. Startup assertion. On boot, call a cheap Sweego read endpoint and verify the key's account
    lists EMAIL_FROM's domain as verified. Fail loudly — or at least log an unmissable warning —
    rather than discovering it at the first sign-up. Needs a Sweego endpoint that exposes verified
    domains; worth confirming one exists.
  2. Post-deploy smoke check. Extend the infra health-check role to send to a sink address and
    assert 2xx. Catches the real path end to end, at the cost of a mail per deploy.
  3. Alert on the failure instead. Emit a metric on send failure and alert on it. Weakest of the
    three — it still requires a real user to hit the broken path first — but it is the cheapest and
    composes with either of the others.

Option 1 plus 3 is probably the right pairing: refuse to start quietly misconfigured, and alert if
sends start failing later.

Gap 3 — best-effort masks systematic failure

Portal email is deliberately best-effort so a mail hiccup cannot fail a sign-up that already
succeeded. That is the right call and should stay. But it means a totally broken sender is
indistinguishable from a transient one, and the affected user is left with an account and no way to
set a password.

Worth considering: surface something to the user when the setup-link send fails outright — a "we
could not send your email, contact support" state, or a self-service resend — so the failure is not
invisible on their side too.

Done when

  • A non-2xx from Sweego logs the provider's explanation, not just a status code.
  • A wrong-account or wrong-domain credential is caught before a real user hits it.
  • A sustained send failure is alertable rather than silent.
A misconfigured Sweego API key broke every portal email for roughly 25 minutes and the system reported almost nothing useful. Sign-ups completed, accounts were provisioned, and no credential-reset link was ever delivered. The only evidence was one line: ``` [signup] setup-link email failed (status 422) ``` Diagnosis was possible only because the operator happened to know the provider account was wrong. Nothing in the code, CI, or the infra health gates could have told us. ## What happened The API key in use was valid, but issued from a **different Sweego account** — one with no verified `mail.gitborg.se` sending domain. So it authenticated (hence **422**, not 401) and Sweego then rejected the sender. Meanwhile Forgejo's SMTP path used credentials from the correct account and delivered fine with `dkim=pass`, so the outage looked partial and contradictory. Fixed operationally in bitborg-infra#273; this issue is about the fact that we could not see it. ## Gap 1 — the failure reason is discarded `sendEmail` in `src/lib/email.ts` keeps the status and throws the body away: ```ts if (res.ok) { return { ok: true }; } return { ok: false, reason: "error", status: res.status }; ``` Sweego returns a JSON body explaining a 422. We never read it, so a precise, self-describing error was reduced to a bare integer. **Proposal:** on a non-2xx, read the body (size-capped, e.g. first 512 bytes) and log it alongside the status. Do **not** put it in the returned value if that risks reaching a user-facing surface — this is operator telemetry. Take care not to log the recipient address any more than we already do, and never the `Api-Key` header. ## Gap 2 — nothing verifies the key's account can send as us `EMAIL_FROM` is a hardcoded constant; the API key arrives as a podman secret from infra. **Any valid key for any account satisfies every check we have** — `pnpm check`, CI, the e2e suite, the `health-check` role, and the Ansible apply itself all pass. The two values must agree and nothing asserts that they do. This is the same defect class as #113 (the From address duplicated across two repos with nothing keeping them in sync), and arguably the sharper half: #113 is drift risk, this is an authorisation pairing that fails silently in production. **Options:** 1. **Startup assertion.** On boot, call a cheap Sweego read endpoint and verify the key's account lists `EMAIL_FROM`'s domain as verified. Fail loudly — or at least log an unmissable warning — rather than discovering it at the first sign-up. Needs a Sweego endpoint that exposes verified domains; worth confirming one exists. 2. **Post-deploy smoke check.** Extend the infra `health-check` role to send to a sink address and assert 2xx. Catches the real path end to end, at the cost of a mail per deploy. 3. **Alert on the failure instead.** Emit a metric on send failure and alert on it. Weakest of the three — it still requires a real user to hit the broken path first — but it is the cheapest and composes with either of the others. Option 1 plus 3 is probably the right pairing: refuse to start quietly misconfigured, and alert if sends start failing later. ## Gap 3 — best-effort masks systematic failure Portal email is deliberately best-effort so a mail hiccup cannot fail a sign-up that already succeeded. That is the right call and should stay. But it means a **totally** broken sender is indistinguishable from a transient one, and the affected user is left with an account and no way to set a password. Worth considering: surface something to the user when the setup-link send fails outright — a "we could not send your email, contact support" state, or a self-service resend — so the failure is not invisible on their side too. ## Done when - A non-2xx from Sweego logs the provider's explanation, not just a status code. - A wrong-account or wrong-domain credential is caught before a real user hits it. - A sustained send failure is alertable rather than silent.
supernaut lade till detta till projektet Bitborg Web 2026-07-31 01:36:00 +00:00
Logga in för att delta i denna konversation.
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-web#115
Ingen beskrivning angiven.