mail: verify email.bitborg.se as a Sweego sending domain before re-landing the move #381
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#381
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
The move of the outbound sending domain to
email.bitborg.seis blocked at the provider. Attempted 2026-08-05: steps 1 and 2 landed, portal mail broke immediately, step 2 was reverted and step 3 was never applied.The finding
Sweego does not authorise
email.bitborg.se. Measured from inside the runningbitborg-webcontainer using the portal's ownSWEEGO_API_KEY, so the key was not a variable — same request shape, only the From domain differing:Published DNS is not provider authorisation. The pre-flight that passed had checked
sweego1._domainkey.email.bitborg.se(resolves, same key as the old domain) and_dmarc.email.bitborg.se(v=DMARC1; p=none;) — the same shape as the working old domain. None of that makes the domain a verified sending identity on the account.Two things made it hide: Sweego answers
401rather than422for an unauthorised sender, so it reads as a credential fault rather than a domain one; and a single-domain probe returning a bare401is equally consistent with a wrong key or a malformed request. Only the two-domain differential isolates the cause.To unblock
email.bitborg.seas a sending domain in the Sweego dashboard. ⚠️ The assumption that "the credentials are unchanged" may itself be wrong — the API key can be bound to a sending identity, so it may need re-scoping or re-minting (which would mean a vault edit after all).mail_allowed_sender_domainsinansible/group_vars/all/vars.yml) and require 2xx on the new domain.mail_sending_domain = email.bitborg.se, allowlist narrowed back to one entry.Current state
EMAIL_ALLOWED_SENDER_DOMAINS=mail.gitborg.se,email.bitborg.sepermits a domain nothing sends as. Harmless while no consumer uses it, but it is a widened allowlist outside a move window — so either finish the move or drop the entry; do not leave it ambiguous.mail.gitborg.se, and both hosts converge atchanged=0 failed=0.Step 3's blast radius, from the dry-run
Six changed tasks across both hosts: renders
app.ini(restarts Forgejo), the web Quadlet unit (restarts the portal), the monitoring-VM cross-probe script, and Alertmanager's config (restarts Alertmanager).Note for the guard
senderAllowed()compares the sender against what infra declares, and infra had declared the new domain allowed — so the guard passed and the provider refused. It closes the infra/portal disagreement it was built for and has no view of the provider. Worth considering whether the provider check belongs in automation rather than a runbook step.Tightened diagnosis, 2026-08-05 — the credentials are fine; the domain is not authorised
Re-tested after the suggestion that existing credentials should be sufficient. They are — and that turns out to be compatible with the block, not a contradiction of it.
Authorisation is per-DOMAIN, not per-address
The discriminating test is a local part that has never been used via the API, on the known-good domain:
So the key authorises a domain, and any address under an authorised domain works.
mail.gitborg.seis authorised;email.bitborg.seis not.The alternative explanations are ruled out
DNS is correct and the domain is known to Sweego
Only two domains carry Sweego DKIM CNAMEs, each with its own account token:
A distinct token means
email.bitborg.sehas its own registration — it is not unknown to Sweego. So the gap is its state, not its existence.Supporting detail: the 401 comes from the application layer, not the gateway. Routes that do not exist return
{"error_msg":"404 Route Not Found"}from the edge, whereas/sendreturns{"detail":"Unauthorized"}— the same response shape as authenticated routes. Sweego accepted the key and then refused the sender.What is left to do, and why it cannot be done from here
Complete the domain's validation in the Sweego dashboard (or confirm it lives in the same project/workspace as the API key). There is no API surface to inspect or change this:
/domains,/senders,/sending-domains,/identities,/account,/me,/v1/domains,/domain,/sending_domains,/transac/domainsall 404.Then re-run the two-domain check (step 0, beside
mail_allowed_sender_domains), require 2xx on the new domain, re-land bitborg-web #196, and apply step 3.No vault edit is expected — the key itself is demonstrably working.
Unblocked as of 2026-08-05 evening — step 0 now passes on both paths.
The blocker was never
email.bitborg.se's validation state. It was an invalidated API key:/sendbegan answering 401 for every sender, including one that had returned 200 that morning,while the key value was unchanged, correctly deployed (vault and in-container sha256 prefixes
identical), and the dashboard showed both domains allowed for API and SMTP. Re-minting the API key
and SMTP credentials fixed it immediately (#385).
Step 0 results, both paths:
mail.gitborg.seemail.bitborg.serc=0rc=0rc=0rc=0SMTP was tested separately and deliberately (#386) because step 0 as written only covered the HTTP
API — the portal's path — while step 3 repoints
forgejo_mailer_from,alert_email_fromandmonitoring_probe_mail_from, all of which send over SMTP with a different credential and its ownauthorisation list.
Both were run with negative controls, since uniform passes are not evidence: an unauthorised sender
domain is rejected (
rc=8/ 401) and a wrong password is rejected (rc=67), proving the relayenforces the sender and that AUTH is exercised.
Remaining: steps 2 and 3. Step 1 (plural allowlist) is applied and stays. Step 2 is re-landing
bitborg-web #196, which #197 reverted. Step 3 moves
mail_sending_domain— measured blast radius issix changed tasks across both hosts, restarting Forgejo, the portal and Alertmanager. Keep 2 and 3
close together: step 1 is currently a widened allowlist outside a move window.
Done — the move is applied and verified
Steps 2 and 3 both landed (bitborg-web #198, this repo #387) and prod is converged.
The premise in the description above was wrong, and is worth correcting
This issue was filed as "Sweego does not authorise
email.bitborg.se". It did. The blocker was aninvalidated API key — the same key that had returned
200for the old domain hours earlier begananswering
401for every sender while its value stayed unchanged and correctly deployed (vaultand in-container sha256 fingerprints matched). Fresh credentials fixed both domains at once.
No dashboard authorisation work was ever needed. The
401-not-422reasoning in the description isstill sound as a diagnostic observation, but it pointed at the wrong cause: a bare
401is equallyconsistent with a dead key, and that is what it was.
Evidence, with controls
Step 0, re-run from inside the container immediately before landing step 2 — three cases, one
request shape:
The negative control is the point: it proves the credential is actually evaluated on
/send, so the200s are not vacuous. (/channelsanswers200completely unauthenticated, which is exactly thetrap that makes an uncontrolled
200worthless here.)Step 2, deployed and verified on the artefact rather than on "the container started":
Step 3 apply — seven changed tasks against six predicted, reconciled by name:
forgejo : Render app.iniforgejo : Restart Forgejo(handler)web : Install the bitborg-web container Quadlet unitweb : Enable and start the web container← not in--checkmonitoring-agent : Install the monitoring-VM cross-probe scriptmonitoring : Render Alertmanager configmonitoring : Restart alertmanager(handler)The extra task is the portal restart, invisible to
--check— another instance of the knowncommand/podman_secretblind spot, and the reason the delta was reconciled by name rather than count.Post-apply verification, each with a control or a positive assertion
Mail delivery was confirmed received, not merely accepted with a
2xx.Follow-ups, not done here
senderAllowed()has no view of theprovider, so nothing in automation would catch an unauthorised-or-dead sending identity. Step 0 is
a runbook step a human must remember. Worth a separate issue.
From: gitborg monitoring <…>— the address moved, the display name did not. Filed separately.--checkafter merge aborted on an SSH transport drop to prod (Data could not be sent to remote host) mid-play; the re-run was clean and matched the branch dry-run exactly. Secondconnection hiccup to that host this session, so noting it rather than dismissing it.