fix(ansible): stop uri tasks reading the operator's ~/.netrc #434
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!434
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/uri-tasks-ignore-netrc"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Every
ansible.builtin.uritask in the tree read the operator's~/.netrc, because that is the default. None of them need it.Four carry an explicit
Authorizationheader (a Forgejo admin token or a Kanidm bearer token) and two hit unauthenticated health endpoints. So netrc can only do one of two things here: send an unrelated credential to one of our own hosts, or fail the task outright.It did the second. A single malformed line in a developer's
~/.netrc(aregiontoken, which is not netrc syntax) made every one of these fail:That reads as an outage. It killed a production dry-run at
roles/forgejo/tasks/system-webhook.yml, and killed the credential-reset onboarding flow atroles/kanidm/tasks/reset-links.yml, on a completely healthy host. Diagnosing it costs real time because the message names neither Ansible nor the host.Sets
use_netrc: falseon all six, with the reasoning recorded at each site.Verification
A full
--checkagainst production, with noNETRCworkaround in the environment:Zero netrc references in the log, and the previously-failing task returns
ok. Before this change the same command gavefailed=1.The
changed=2is unrelated pending drift: the merged Alloy bump (#433) is not yet applied.Every ansible.builtin.uri task in the tree read the operator's ~/.netrc, because that is the default. None of them need it. Four carry an explicit Authorization header (a Forgejo admin token or a Kanidm bearer token) and two hit unauthenticated health endpoints. So netrc can only do one of two things here: send an unrelated credential to one of our own hosts, or fail the task outright. It did the second. A single malformed line in a developer's ~/.netrc — an `region` token, which is not netrc syntax — made every one of these fail with Status code was -1 and not [200]: An unknown error occurred: bad follower token 'region' (~/.netrc, line 4) That reads as an outage. It killed a production dry-run at roles/forgejo/tasks/system-webhook.yml, and killed the credential-reset onboarding flow at roles/kanidm/tasks/reset-links.yml, on a host that was completely healthy. Sets use_netrc: false on all six, with the reasoning recorded at each site. Verified: a full --check against production now completes with failed=0 and zero netrc references in the log, with no NETRC workaround in the environment. The task that previously failed returns ok.