fix(backup-drill): log restic errors in the hot leg and create the work dir 0700 #524
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!524
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "fix/158-drill-hot-restore"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
What
The first production drill (2026-10-02 17:19 UTC) passed the daily leg but failed the hot leg at
restic restore(exit 1, under a second). restic's stderr was not captured, so the cause is not visible. A manual restore of the same snapshot outside the drill works (201.6 MiB, 7 s).This PR makes the next run tell us why:
0700right before the restore (the umask made it 0755).Ruled out by reading the code: a missing work dir (it is created before
podman run -v). Still open: something that differs under the systemd unit.Verification
--syntax-check, ansible-lint 0, rendered orchestrator passespy_compile. Local test with a fakepodmanthat exits 1: the error and mode are logged and the dir is 0700.Apply
Tag
backup-drill: only the orchestrator script changes. Then run the drill by hand and read the logged restic error.Refs #158
497102de3c2b82a2bfe2