backups: abort incomplete s3 multipart uploads (lifecycle rule) #181
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!181
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "feat/160-s3-abort-mpu-lifecycle"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Closes #160.
Problem
The off-site backup uploads each (multi-GB) age-encrypted archive to two S3-compatible
destinations (Glesys primary, Hetzner secondary) as S3 multipart uploads. When a run is
aborted or fails mid-upload (OOM, network drop, reboot), the object store keeps the
already-uploaded parts — invisible in a normal bucket listing but still billed — until
something aborts them. Cost accrues silently, run after run.
Mechanism (which path I took, and why)
The whole backup stack talks to the buckets exclusively via rclone in a podman container, and
bitborg-backup.shalready wires per-destinationRCLONE_CONFIG_<NAME>_*credentials. rclonehas no command to set a server-side bucket lifecycle rule, and this project deliberately avoids
the
aws-cliimage (principle 1 — see thebackup_rclone_imagecomment; picking rclone overamazon/aws-cli was a deliberate sovereignty choice). Pulling a lifecycle-capable client (aws-cli /
an unofficial s3cmd image) would add a principle-1 or supply-chain cost for a marginal gain.
So the applied mechanism reuses the existing, vetted tooling: the backup script now sweeps
each destination with
which aborts any incomplete multipart upload older than the window. This runs after each upload
attempt and regardless of its outcome (a failed upload is precisely what leaves parts behind),
per destination, every run — so it is inherently idempotent. A sweep failure is cost hygiene, not
data integrity, so it is logged non-fatally and never fails the backup nor flips
gitborg_backup_offsite_last_status(which drives BackupOffsiteFailed).I also documented the equivalent server-side lifecycle policy (AbortIncompleteMultipartUpload /
DaysAfterInitiation) and a one-time provider command in the runbook, for operators who additionally
want the object store to enforce it server-side (persists even if the host is gone). The two are
complementary — running both is harmless. Both Glesys (Ceph RGW) and Hetzner support the policy.
Day threshold
New tunable
backup_s3_abort_incomplete_mpu_days(default 7) in the backup role defaults.7 days is a safe grace window: far longer than a real archive upload (minutes), so a legitimately
in-flight upload from a concurrent run is never killed, yet orphaned parts can't linger (and bill)
for more than a week.
Buckets
Both off-site destinations in
backup_s3_destinations: glesys (primary, Falkenberg) andhetzner (secondary, hel1). The sweep loops the same per-destination
upload_tofunction, so itautomatically covers any future destination too.
Changes
ansible/roles/backup/defaults/main.yml— newbackup_s3_abort_incomplete_mpu_days: 7with arationale comment.
ansible/roles/backup/templates/bitborg-backup.sh.j2— per-destinationrclone backend cleanupin the off-site phase.
docs/runbook.md— new "Incomplete multipart uploads (cost hygiene)" subsection with theserver-side lifecycle policy + one-time command.
Verification
pnpm ansible:check(syntax-check) — clean.ansible-lint roles/backup—Passed: 0 failure(s), 0 warning(s)(production profile).bash -non the result — syntax OK.prettier -c+markdownlint-cli2clean (lefthook pre-commit passed).Not applied to prod — deploy with
--tags backup(the backup role) at the maintainer'sdiscretion. No new credentials or images are required; it reuses the existing rclone container and
vaulted per-destination S3 keys.