mail: nothing in automation detects a provider-side sending-domain or credential failure #389

Öppen
öppnade 2026-08-05 21:02:28 +00:00 av supernaut · 0 kommentarer
Ägare

Raised by #381 and left open when it closed.

senderAllowed() in bitborg-web closes the infra/portal disagreement it was built for: infra
supplies EMAIL_ALLOWED_SENDER_DOMAINS, the portal checks its hardcoded EMAIL_FROM against it, and
a divergence becomes a logged refusal instead of silently-rejected mail.

It has no view of the provider at all. On 2026-08-05 infra had declared email.bitborg.se
allowed, so the guard passed — and the provider refused every message. The guard cannot catch that
class of failure by construction, and it is the class that actually caused the outage.

Today the only defence is step 0, a runbook instruction to send a two-domain differential by hand
before moving. It worked, but it is a step a human must remember at exactly the moment they are
confident the move is safe — which is when it gets skipped. It was skipped once already.

What would actually help

A periodic check that the configured sending domain is accepted by the provider right now, so a
dead or de-authorised credential surfaces before the next send that matters rather than during it.
That is a different signal from "the key is deployed" (already covered by comparing sha256
fingerprints) and from "mail was sent" (only observable when something happens to send).

Design notes, mostly about what does not work:

  • A 2xx alone is not evidence. /channels answers 200 completely unauthenticated. Any probe
    must pair the real request with a deliberately invalid credential and require the two to differ.
  • 401 is not diagnosable — Sweego returns an identical 401 for a dead key and for a valid key
    with an unauthorised sender, and no reachable route separates them. So the check can prove
    "something is wrong with our ability to send as this domain" but cannot say which. That is
    still worth alerting on; it just must not claim a cause.
  • A check that sends real mail on a timer has a cost (deliverability reputation, a mailbox filling
    up). Worth considering whether the existing cross-probe's transition-only pattern applies — alert
    on a state change, not every run.

Acceptance

  • The check fails loudly when pointed at a domain the provider does not accept — prove it can fail
    before trusting a pass
    , by pointing it at a deliberately wrong sender and confirming it goes red,
    then putting it back.
  • It does not page on a single transient provider blip.
  • Whatever it asserts, it does not claim to distinguish a dead key from an unauthorised sender, since
    that is not observable from here.
Raised by #381 and left open when it closed. `senderAllowed()` in bitborg-web closes the **infra/portal** disagreement it was built for: infra supplies `EMAIL_ALLOWED_SENDER_DOMAINS`, the portal checks its hardcoded `EMAIL_FROM` against it, and a divergence becomes a logged refusal instead of silently-rejected mail. It has **no view of the provider at all**. On 2026-08-05 infra had declared `email.bitborg.se` allowed, so the guard passed — and the provider refused every message. The guard cannot catch that class of failure by construction, and it is the class that actually caused the outage. Today the only defence is **step 0**, a runbook instruction to send a two-domain differential by hand before moving. It worked, but it is a step a human must remember at exactly the moment they are confident the move is safe — which is when it gets skipped. It was skipped once already. ## What would actually help A periodic check that the configured sending domain is **accepted by the provider right now**, so a dead or de-authorised credential surfaces before the next send that matters rather than during it. That is a different signal from "the key is deployed" (already covered by comparing sha256 fingerprints) and from "mail was sent" (only observable when something happens to send). Design notes, mostly about what does not work: - **A `2xx` alone is not evidence.** `/channels` answers `200` completely unauthenticated. Any probe must pair the real request with a deliberately invalid credential and require the two to *differ*. - **`401` is not diagnosable** — Sweego returns an identical `401` for a dead key and for a valid key with an unauthorised sender, and no reachable route separates them. So the check can prove "something is wrong with our ability to send as this domain" but **cannot** say which. That is still worth alerting on; it just must not claim a cause. - A check that sends real mail on a timer has a cost (deliverability reputation, a mailbox filling up). Worth considering whether the existing cross-probe's transition-only pattern applies — alert on a state change, not every run. ## Acceptance - The check fails loudly when pointed at a domain the provider does not accept — **prove it can fail before trusting a pass**, by pointing it at a deliberately wrong sender and confirming it goes red, then putting it back. - It does not page on a single transient provider blip. - Whatever it asserts, it does not claim to distinguish a dead key from an unauthorised sender, since that is not observable from here.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#389
Ingen beskrivning angiven.