mail: verify email.bitborg.se as a Sweego sending domain before re-landing the move #381

Stängd
öppnade 2026-08-05 14:19:58 +00:00 av supernaut · 3 kommentarer
Ägare

The move of the outbound sending domain to email.bitborg.se is blocked at the provider. Attempted 2026-08-05: steps 1 and 2 landed, portal mail broke immediately, step 2 was reverted and step 3 was never applied.

The finding

Sweego does not authorise email.bitborg.se. Measured from inside the running bitborg-web container using the portal's own SWEEGO_API_KEY, so the key was not a variable — same request shape, only the From domain differing:

no-reply@mail.gitborg.se   -> HTTP 200, accepted (transaction id returned)
no-reply@email.bitborg.se  -> HTTP 401 {"detail":"Unauthorized"}

Published DNS is not provider authorisation. The pre-flight that passed had checked sweego1._domainkey.email.bitborg.se (resolves, same key as the old domain) and _dmarc.email.bitborg.se (v=DMARC1; p=none;) — the same shape as the working old domain. None of that makes the domain a verified sending identity on the account.

Two things made it hide: Sweego answers 401 rather than 422 for an unauthorised sender, so it reads as a credential fault rather than a domain one; and a single-domain probe returning a bare 401 is equally consistent with a wrong key or a malformed request. Only the two-domain differential isolates the cause.

To unblock

  1. Verify email.bitborg.se as a sending domain in the Sweego dashboard. ⚠️ The assumption that "the credentials are unchanged" may itself be wrong — the API key can be bound to a sending identity, so it may need re-scoping or re-minting (which would mean a vault edit after all).
  2. Re-run the two-domain send check (recorded as step 0 beside mail_allowed_sender_domains in ansible/group_vars/all/vars.yml) and require 2xx on the new domain.
  3. Re-land the portal-side change (bitborg-web #196, reverted by #197).
  4. Apply step 3: mail_sending_domain = email.bitborg.se, allowlist narrowed back to one entry.

Current state

  • Step 1 is applied and inert. EMAIL_ALLOWED_SENDER_DOMAINS=mail.gitborg.se,email.bitborg.se permits a domain nothing sends as. Harmless while no consumer uses it, but it is a widened allowlist outside a move window — so either finish the move or drop the entry; do not leave it ambiguous.
  • Mail works: everything sends as mail.gitborg.se, and both hosts converge at changed=0 failed=0.
  • No user impact from the incident — Loki shows no sign-up or mail-send activity during the roughly 20-minute window.

Step 3's blast radius, from the dry-run

Six changed tasks across both hosts: renders app.ini (restarts Forgejo), the web Quadlet unit (restarts the portal), the monitoring-VM cross-probe script, and Alertmanager's config (restarts Alertmanager).

Note for the guard

senderAllowed() compares the sender against what infra declares, and infra had declared the new domain allowed — so the guard passed and the provider refused. It closes the infra/portal disagreement it was built for and has no view of the provider. Worth considering whether the provider check belongs in automation rather than a runbook step.

The move of the outbound sending domain to `email.bitborg.se` is **blocked at the provider**. Attempted 2026-08-05: steps 1 and 2 landed, portal mail broke immediately, step 2 was reverted and step 3 was never applied. ## The finding **Sweego does not authorise `email.bitborg.se`.** Measured from inside the running `bitborg-web` container using the portal's own `SWEEGO_API_KEY`, so the key was not a variable — same request shape, only the From domain differing: ```text no-reply@mail.gitborg.se -> HTTP 200, accepted (transaction id returned) no-reply@email.bitborg.se -> HTTP 401 {"detail":"Unauthorized"} ``` **Published DNS is not provider authorisation.** The pre-flight that passed had checked `sweego1._domainkey.email.bitborg.se` (resolves, same key as the old domain) and `_dmarc.email.bitborg.se` (`v=DMARC1; p=none;`) — the same shape as the working old domain. None of that makes the domain a verified sending identity on the account. Two things made it hide: Sweego answers `401` rather than `422` for an unauthorised sender, so it reads as a credential fault rather than a domain one; and a single-domain probe returning a bare `401` is equally consistent with a wrong key or a malformed request. Only the two-domain differential isolates the cause. ## To unblock 1. Verify `email.bitborg.se` as a sending domain in the Sweego dashboard. ⚠️ The assumption that "the credentials are unchanged" may itself be wrong — the API key can be bound to a sending identity, so it may need re-scoping or re-minting (which would mean a vault edit after all). 2. Re-run the two-domain send check (recorded as **step 0** beside `mail_allowed_sender_domains` in `ansible/group_vars/all/vars.yml`) and require **2xx on the new domain**. 3. Re-land the portal-side change (bitborg-web #196, reverted by #197). 4. Apply step 3: `mail_sending_domain = email.bitborg.se`, allowlist narrowed back to one entry. ## Current state - **Step 1 is applied and inert.** `EMAIL_ALLOWED_SENDER_DOMAINS=mail.gitborg.se,email.bitborg.se` permits a domain nothing sends as. Harmless while no consumer uses it, but it is a widened allowlist outside a move window — so either finish the move or drop the entry; do not leave it ambiguous. - Mail works: everything sends as `mail.gitborg.se`, and both hosts converge at `changed=0 failed=0`. - No user impact from the incident — Loki shows no sign-up or mail-send activity during the roughly 20-minute window. ## Step 3's blast radius, from the dry-run Six changed tasks across both hosts: renders `app.ini` (**restarts Forgejo**), the web Quadlet unit (**restarts the portal**), the monitoring-VM cross-probe script, and Alertmanager's config (**restarts Alertmanager**). ## Note for the guard `senderAllowed()` compares the sender against what infra declares, and infra had declared the new domain allowed — so the guard passed and the provider refused. It closes the infra/portal disagreement it was built for and has no view of the provider. Worth considering whether the provider check belongs in automation rather than a runbook step.
Upphovsperson
Ägare

Tightened diagnosis, 2026-08-05 — the credentials are fine; the domain is not authorised

Re-tested after the suggestion that existing credentials should be sufficient. They are — and that turns out to be compatible with the block, not a contradiction of it.

Authorisation is per-DOMAIN, not per-address

The discriminating test is a local part that has never been used via the API, on the known-good domain:

alerts@mail.gitborg.se     -> HTTP 200   (brand-new local part, sends immediately)
no-reply@email.bitborg.se  -> HTTP 401 {"detail":"Unauthorized"}

So the key authorises a domain, and any address under an authorised domain works. mail.gitborg.se is authorised; email.bitborg.se is not.

The alternative explanations are ruled out

  • Not ordering or rate limiting — sending the new domain first still returns 401, and the old domain returns 200 mid-sequence.
  • Not flapping — the new domain returns 401 on repeat within the same run.
  • Not the payload — the requests are byte-identical apart from the From address (the first probe had also varied subject and body; this one does not).
  • Not the key — the same key, same process, same request returns 200 for the old domain.

DNS is correct and the domain is known to Sweego

Only two domains carry Sweego DKIM CNAMEs, each with its own account token:

email.bitborg.se   ae4a3bca-f912-4a96-91c2-7d25bc433cbb-dkim.sweego.co.
mail.gitborg.se    d65c0acd-fab8-4809-9428-087273ed6750-dkim.sweego.co.
bitborg.se / mail.bitborg.se / gitborg.se / email.gitborg.se   <none>

A distinct token means email.bitborg.se has its own registration — it is not unknown to Sweego. So the gap is its state, not its existence.

Supporting detail: the 401 comes from the application layer, not the gateway. Routes that do not exist return {"error_msg":"404 Route Not Found"} from the edge, whereas /send returns {"detail":"Unauthorized"} — the same response shape as authenticated routes. Sweego accepted the key and then refused the sender.

What is left to do, and why it cannot be done from here

Complete the domain's validation in the Sweego dashboard (or confirm it lives in the same project/workspace as the API key). There is no API surface to inspect or change this: /domains, /senders, /sending-domains, /identities, /account, /me, /v1/domains, /domain, /sending_domains, /transac/domains all 404.

Then re-run the two-domain check (step 0, beside mail_allowed_sender_domains), require 2xx on the new domain, re-land bitborg-web #196, and apply step 3.

No vault edit is expected — the key itself is demonstrably working.

## Tightened diagnosis, 2026-08-05 — the credentials are fine; the domain is not authorised Re-tested after the suggestion that existing credentials should be sufficient. **They are** — and that turns out to be compatible with the block, not a contradiction of it. ### Authorisation is per-DOMAIN, not per-address The discriminating test is a local part that has never been used via the API, on the known-good domain: ```text alerts@mail.gitborg.se -> HTTP 200 (brand-new local part, sends immediately) no-reply@email.bitborg.se -> HTTP 401 {"detail":"Unauthorized"} ``` So the key authorises a *domain*, and any address under an authorised domain works. `mail.gitborg.se` is authorised; `email.bitborg.se` is not. ### The alternative explanations are ruled out - **Not ordering or rate limiting** — sending the new domain **first** still returns 401, and the old domain returns 200 mid-sequence. - **Not flapping** — the new domain returns 401 on repeat within the same run. - **Not the payload** — the requests are byte-identical apart from the From address (the first probe had also varied subject and body; this one does not). - **Not the key** — the same key, same process, same request returns 200 for the old domain. ### DNS is correct and the domain *is* known to Sweego Only two domains carry Sweego DKIM CNAMEs, each with its own account token: ```text email.bitborg.se ae4a3bca-f912-4a96-91c2-7d25bc433cbb-dkim.sweego.co. mail.gitborg.se d65c0acd-fab8-4809-9428-087273ed6750-dkim.sweego.co. bitborg.se / mail.bitborg.se / gitborg.se / email.gitborg.se <none> ``` A distinct token means `email.bitborg.se` has its **own registration** — it is not unknown to Sweego. So the gap is its **state**, not its existence. Supporting detail: the 401 comes from the **application** layer, not the gateway. Routes that do not exist return `{"error_msg":"404 Route Not Found"}` from the edge, whereas `/send` returns `{"detail":"Unauthorized"}` — the same response shape as authenticated routes. Sweego accepted the key and then refused the sender. ### What is left to do, and why it cannot be done from here Complete the domain's validation in the Sweego dashboard (or confirm it lives in the same project/workspace as the API key). There is **no API surface** to inspect or change this: `/domains`, `/senders`, `/sending-domains`, `/identities`, `/account`, `/me`, `/v1/domains`, `/domain`, `/sending_domains`, `/transac/domains` all 404. Then re-run the two-domain check (step 0, beside `mail_allowed_sender_domains`), require **2xx on the new domain**, re-land bitborg-web #196, and apply step 3. No vault edit is expected — the key itself is demonstrably working.
Upphovsperson
Ägare

Unblocked as of 2026-08-05 evening — step 0 now passes on both paths.

The blocker was never email.bitborg.se's validation state. It was an invalidated API key:
/send began answering 401 for every sender, including one that had returned 200 that morning,
while the key value was unchanged, correctly deployed (vault and in-container sha256 prefixes
identical), and the dashboard showed both domains allowed for API and SMTP. Re-minting the API key
and SMTP credentials fixed it immediately (#385).

Step 0 results, both paths:

path mail.gitborg.se email.bitborg.se
HTTP API (portal) 200 200
SMTP, services host (Forgejo, cross-probe) rc=0 rc=0
SMTP, monitoring host (Alertmanager) rc=0 rc=0

SMTP was tested separately and deliberately (#386) because step 0 as written only covered the HTTP
API — the portal's path — while step 3 repoints forgejo_mailer_from, alert_email_from and
monitoring_probe_mail_from, all of which send over SMTP with a different credential and its own
authorisation list.

Both were run with negative controls, since uniform passes are not evidence: an unauthorised sender
domain is rejected (rc=8 / 401) and a wrong password is rejected (rc=67), proving the relay
enforces the sender and that AUTH is exercised.

Remaining: steps 2 and 3. Step 1 (plural allowlist) is applied and stays. Step 2 is re-landing
bitborg-web #196, which #197 reverted. Step 3 moves mail_sending_domain — measured blast radius is
six changed tasks across both hosts, restarting Forgejo, the portal and Alertmanager. Keep 2 and 3
close together: step 1 is currently a widened allowlist outside a move window.

**Unblocked as of 2026-08-05 evening — step 0 now passes on both paths.** The blocker was never `email.bitborg.se`'s validation state. It was an **invalidated API key**: `/send` began answering 401 for _every_ sender, including one that had returned 200 that morning, while the key value was unchanged, correctly deployed (vault and in-container sha256 prefixes identical), and the dashboard showed both domains allowed for API and SMTP. Re-minting the API key and SMTP credentials fixed it immediately (#385). Step 0 results, both paths: | path | `mail.gitborg.se` | `email.bitborg.se` | | ------------------------------------------ | ----------------- | ------------------ | | HTTP API (portal) | 200 | **200** | | SMTP, services host (Forgejo, cross-probe) | `rc=0` | **`rc=0`** | | SMTP, monitoring host (Alertmanager) | `rc=0` | **`rc=0`** | SMTP was tested separately and deliberately (#386) because step 0 as written only covered the HTTP API — the portal's path — while step 3 repoints `forgejo_mailer_from`, `alert_email_from` and `monitoring_probe_mail_from`, all of which send over SMTP with a different credential and its own authorisation list. Both were run with negative controls, since uniform passes are not evidence: an unauthorised sender domain is rejected (`rc=8` / 401) and a wrong password is rejected (`rc=67`), proving the relay enforces the sender and that AUTH is exercised. **Remaining: steps 2 and 3.** Step 1 (plural allowlist) is applied and stays. Step 2 is re-landing bitborg-web #196, which #197 reverted. Step 3 moves `mail_sending_domain` — measured blast radius is six changed tasks across both hosts, restarting Forgejo, the portal and Alertmanager. Keep 2 and 3 close together: step 1 is currently a widened allowlist outside a move window.
Upphovsperson
Ägare

Done — the move is applied and verified

Steps 2 and 3 both landed (bitborg-web #198, this repo #387) and prod is converged.

The premise in the description above was wrong, and is worth correcting

This issue was filed as "Sweego does not authorise email.bitborg.se". It did. The blocker was an
invalidated API key — the same key that had returned 200 for the old domain hours earlier began
answering 401 for every sender while its value stayed unchanged and correctly deployed (vault
and in-container sha256 fingerprints matched). Fresh credentials fixed both domains at once.

No dashboard authorisation work was ever needed. The 401-not-422 reasoning in the description is
still sound as a diagnostic observation, but it pointed at the wrong cause: a bare 401 is equally
consistent with a dead key, and that is what it was.

Evidence, with controls

Step 0, re-run from inside the container immediately before landing step 2 — three cases, one
request shape:

OLD domain + real key   -> 200 (transaction id)   positive control
NEW domain + real key   -> 200 (transaction id)   the gate
NEW domain + bogus key  -> 401                    negative control

The negative control is the point: it proves the credential is actually evaluated on /send, so the
200s are not vacuous. (/channels answers 200 completely unauthenticated, which is exactly the
trap that makes an uncontrolled 200 worthless here.)

Step 2, deployed and verified on the artefact rather than on "the container started":

ImageID   ca75d81d9f69 -> 14af5f7c3ee5
sender baked into the running code:  no-reply@email.bitborg.se

Step 3 apply — seven changed tasks against six predicted, reconciled by name:

host task
prod forgejo : Render app.ini
prod forgejo : Restart Forgejo (handler)
prod web : Install the bitborg-web container Quadlet unit
prod web : Enable and start the web container ← not in --check
prod monitoring-agent : Install the monitoring-VM cross-probe script
monitoring monitoring : Render Alertmanager config
monitoring monitoring : Restart alertmanager (handler)

The extra task is the portal restart, invisible to --check — another instance of the known
command/podman_secret blind spot, and the reason the delta was reconciled by name rather than count.

Post-apply verification, each with a control or a positive assertion

portal EMAIL_ALLOWED_SENDER_DOMAINS  = email.bitborg.se          (narrowed; window closed)
portal sendEmail() real send         = {"ok":true}
  negative control, foreign allowlist -> refused, naming no-reply@email.bitborg.se
app.ini FROM                         = Bitborg <no-reply@email.bitborg.se>
cross-probe MAIL_FROM                = alerts@email.bitborg.se
SMTP send as the new From            = curl rc=0
  negative control, wrong password    -> rc=67 Login denied
Alertmanager loaded config           = smtp_from: alerts@email.bitborg.se  (from /api/v2/status)
Alertmanager /-/healthy              = OK, restarted clean, no errors in log
six blackbox probes                  = 1
firing alerts                        = Watchdog only
site.yml --check                     = changed=0 failed=0 unreachable=0, both hosts

Mail delivery was confirmed received, not merely accepted with a 2xx.

Follow-ups, not done here

  • The description's closing question stands and is a real gap: senderAllowed() has no view of the
    provider, so nothing in automation would catch an unauthorised-or-dead sending identity. Step 0 is
    a runbook step a human must remember. Worth a separate issue.
  • One note found while verifying: the cross-probe's email header is still From: gitborg monitoring <…> — the address moved, the display name did not. Filed separately.
  • The first --check after merge aborted on an SSH transport drop to prod (Data could not be sent to remote host) mid-play; the re-run was clean and matched the branch dry-run exactly. Second
    connection hiccup to that host this session, so noting it rather than dismissing it.
## Done — the move is applied and verified Steps 2 and 3 both landed (bitborg-web #198, this repo #387) and prod is converged. ### The premise in the description above was wrong, and is worth correcting This issue was filed as *"Sweego does not authorise `email.bitborg.se`"*. It did. The blocker was an **invalidated API key** — the same key that had returned `200` for the old domain hours earlier began answering `401` for **every** sender while its value stayed unchanged and correctly deployed (vault and in-container sha256 fingerprints matched). Fresh credentials fixed both domains at once. No dashboard authorisation work was ever needed. The `401`-not-`422` reasoning in the description is still sound as a *diagnostic* observation, but it pointed at the wrong cause: a bare `401` is equally consistent with a dead key, and that is what it was. ### Evidence, with controls Step 0, re-run from inside the container immediately before landing step 2 — three cases, one request shape: ```text OLD domain + real key -> 200 (transaction id) positive control NEW domain + real key -> 200 (transaction id) the gate NEW domain + bogus key -> 401 negative control ``` The negative control is the point: it proves the credential is actually evaluated on `/send`, so the `200`s are not vacuous. (`/channels` answers `200` completely unauthenticated, which is exactly the trap that makes an uncontrolled `200` worthless here.) Step 2, deployed and verified on the artefact rather than on "the container started": ```text ImageID ca75d81d9f69 -> 14af5f7c3ee5 sender baked into the running code: no-reply@email.bitborg.se ``` Step 3 apply — **seven** changed tasks against six predicted, reconciled **by name**: | host | task | | ---------- | --------------------------------------------------------------- | | prod | `forgejo : Render app.ini` | | prod | `forgejo : Restart Forgejo` (handler) | | prod | `web : Install the bitborg-web container Quadlet unit` | | prod | `web : Enable and start the web container` ← **not in `--check`** | | prod | `monitoring-agent : Install the monitoring-VM cross-probe script` | | monitoring | `monitoring : Render Alertmanager config` | | monitoring | `monitoring : Restart alertmanager` (handler) | The extra task is the portal restart, invisible to `--check` — another instance of the known `command`/`podman_secret` blind spot, and the reason the delta was reconciled by name rather than count. ### Post-apply verification, each with a control or a positive assertion ```text portal EMAIL_ALLOWED_SENDER_DOMAINS = email.bitborg.se (narrowed; window closed) portal sendEmail() real send = {"ok":true} negative control, foreign allowlist -> refused, naming no-reply@email.bitborg.se app.ini FROM = Bitborg <no-reply@email.bitborg.se> cross-probe MAIL_FROM = alerts@email.bitborg.se SMTP send as the new From = curl rc=0 negative control, wrong password -> rc=67 Login denied Alertmanager loaded config = smtp_from: alerts@email.bitborg.se (from /api/v2/status) Alertmanager /-/healthy = OK, restarted clean, no errors in log six blackbox probes = 1 firing alerts = Watchdog only site.yml --check = changed=0 failed=0 unreachable=0, both hosts ``` Mail delivery was confirmed **received**, not merely accepted with a `2xx`. ### Follow-ups, not done here - The description's closing question stands and is a real gap: `senderAllowed()` has no view of the provider, so nothing in automation would catch an unauthorised-or-dead sending identity. Step 0 is a runbook step a human must remember. Worth a separate issue. - One note found while verifying: the cross-probe's email header is still `From: gitborg monitoring <…>` — the address moved, the display name did not. Filed separately. - The first `--check` after merge aborted on an SSH transport drop to prod (`Data could not be sent to remote host`) mid-play; the re-run was clean and matched the branch dry-run exactly. Second connection hiccup to that host this session, so noting it rather than dismissing it.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#381
Ingen beskrivning angiven.