backups: abort incomplete s3 multipart uploads (lifecycle rule) #181

Sammanfogat
supernaut sammanfogade 1 incheckning från feat/160-s3-abort-mpu-lifecycle in i main 2026-07-21 11:21:00 +00:00
Ägare

Closes #160.

Problem

The off-site backup uploads each (multi-GB) age-encrypted archive to two S3-compatible
destinations (Glesys primary, Hetzner secondary) as S3 multipart uploads. When a run is
aborted or fails mid-upload (OOM, network drop, reboot), the object store keeps the
already-uploaded parts — invisible in a normal bucket listing but still billed — until
something aborts them. Cost accrues silently, run after run.

Mechanism (which path I took, and why)

The whole backup stack talks to the buckets exclusively via rclone in a podman container, and
bitborg-backup.sh already wires per-destination RCLONE_CONFIG_<NAME>_* credentials. rclone
has no command to set a server-side bucket lifecycle rule, and this project deliberately avoids
the aws-cli image (principle 1 — see the backup_rclone_image comment; picking rclone over
amazon/aws-cli was a deliberate sovereignty choice). Pulling a lifecycle-capable client (aws-cli /
an unofficial s3cmd image) would add a principle-1 or supply-chain cost for a marginal gain.

So the applied mechanism reuses the existing, vetted tooling: the backup script now sweeps
each destination with

rclone backend cleanup <dest>:<bucket> -o max-age=<days>d

which aborts any incomplete multipart upload older than the window. This runs after each upload
attempt and regardless of its outcome
(a failed upload is precisely what leaves parts behind),
per destination, every run — so it is inherently idempotent. A sweep failure is cost hygiene, not
data integrity, so it is logged non-fatally and never fails the backup nor flips
gitborg_backup_offsite_last_status (which drives BackupOffsiteFailed).

I also documented the equivalent server-side lifecycle policy (AbortIncompleteMultipartUpload /
DaysAfterInitiation) and a one-time provider command in the runbook, for operators who additionally
want the object store to enforce it server-side (persists even if the host is gone). The two are
complementary — running both is harmless. Both Glesys (Ceph RGW) and Hetzner support the policy.

Day threshold

New tunable backup_s3_abort_incomplete_mpu_days (default 7) in the backup role defaults.
7 days is a safe grace window: far longer than a real archive upload (minutes), so a legitimately
in-flight upload from a concurrent run is never killed, yet orphaned parts can't linger (and bill)
for more than a week.

Buckets

Both off-site destinations in backup_s3_destinations: glesys (primary, Falkenberg) and
hetzner (secondary, hel1). The sweep loops the same per-destination upload_to function, so it
automatically covers any future destination too.

Changes

  • ansible/roles/backup/defaults/main.yml — new backup_s3_abort_incomplete_mpu_days: 7 with a
    rationale comment.
  • ansible/roles/backup/templates/bitborg-backup.sh.j2 — per-destination rclone backend cleanup
    in the off-site phase.
  • docs/runbook.md — new "Incomplete multipart uploads (cost hygiene)" subsection with the
    server-side lifecycle policy + one-time command.

Verification

  • pnpm ansible:check (syntax-check) — clean.
  • ansible-lint roles/backup — Passed: 0 failure(s), 0 warning(s) (production profile).
  • Rendered the Jinja template and ran bash -n on the result — syntax OK.
  • Validated the documented lifecycle XML parses.
  • prettier -c + markdownlint-cli2 clean (lefthook pre-commit passed).

Not applied to prod — deploy with --tags backup (the backup role) at the maintainer's
discretion. No new credentials or images are required; it reuses the existing rclone container and
vaulted per-destination S3 keys.

Closes #160. ## Problem The off-site backup uploads each (multi-GB) age-encrypted archive to two S3-compatible destinations (Glesys primary, Hetzner secondary) as **S3 multipart uploads**. When a run is aborted or fails mid-upload (OOM, network drop, reboot), the object store keeps the already-uploaded **parts** — invisible in a normal bucket listing but **still billed** — until something aborts them. Cost accrues silently, run after run. ## Mechanism (which path I took, and why) The whole backup stack talks to the buckets exclusively via **rclone in a podman container**, and `bitborg-backup.sh` already wires per-destination `RCLONE_CONFIG_<NAME>_*` credentials. rclone has **no** command to set a server-side bucket lifecycle rule, and this project deliberately avoids the `aws-cli` image (principle 1 — see the `backup_rclone_image` comment; picking rclone over amazon/aws-cli was a deliberate sovereignty choice). Pulling a lifecycle-capable client (aws-cli / an unofficial s3cmd image) would add a principle-1 or supply-chain cost for a marginal gain. So the **applied** mechanism reuses the existing, vetted tooling: the backup script now sweeps **each** destination with ``` rclone backend cleanup <dest>:<bucket> -o max-age=<days>d ``` which aborts any incomplete multipart upload older than the window. This runs **after each upload attempt and regardless of its outcome** (a failed upload is precisely what leaves parts behind), per destination, every run — so it is inherently idempotent. A sweep failure is cost hygiene, not data integrity, so it is logged non-fatally and never fails the backup nor flips `gitborg_backup_offsite_last_status` (which drives BackupOffsiteFailed). I also **documented the equivalent server-side lifecycle policy** (AbortIncompleteMultipartUpload / DaysAfterInitiation) and a one-time provider command in the runbook, for operators who additionally want the object store to enforce it server-side (persists even if the host is gone). The two are complementary — running both is harmless. Both Glesys (Ceph RGW) and Hetzner support the policy. ## Day threshold New tunable **`backup_s3_abort_incomplete_mpu_days`** (default **7**) in the backup role defaults. 7 days is a safe grace window: far longer than a real archive upload (minutes), so a legitimately in-flight upload from a concurrent run is never killed, yet orphaned parts can't linger (and bill) for more than a week. ## Buckets Both off-site destinations in `backup_s3_destinations`: **glesys** (primary, Falkenberg) and **hetzner** (secondary, hel1). The sweep loops the same per-destination `upload_to` function, so it automatically covers any future destination too. ## Changes - `ansible/roles/backup/defaults/main.yml` — new `backup_s3_abort_incomplete_mpu_days: 7` with a rationale comment. - `ansible/roles/backup/templates/bitborg-backup.sh.j2` — per-destination `rclone backend cleanup` in the off-site phase. - `docs/runbook.md` — new "Incomplete multipart uploads (cost hygiene)" subsection with the server-side lifecycle policy + one-time command. ## Verification - `pnpm ansible:check` (syntax-check) — clean. - `ansible-lint roles/backup` — `Passed: 0 failure(s), 0 warning(s)` (production profile). - Rendered the Jinja template and ran `bash -n` on the result — syntax OK. - Validated the documented lifecycle XML parses. - `prettier -c` + `markdownlint-cli2` clean (lefthook pre-commit passed). **Not applied to prod** — deploy with `--tags backup` (the backup role) at the maintainer's discretion. No new credentials or images are required; it reuses the existing rclone container and vaulted per-destination S3 keys.
supernaut lade till 1 incheckning 2026-07-21 11:13:02 +00:00
feat(backup): abort incomplete s3 multipart uploads (#160)
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m34s
f561f60df3
Aborted/failed off-site uploads leave incomplete S3 multipart uploads
behind: the object store keeps (and bills for) the already-uploaded
parts, invisible in a normal listing, until something aborts them.

The whole backup stack talks to the buckets via rclone in a podman
container (principle 1 rejects the aws-cli image), and rclone cannot set
a server-side bucket lifecycle rule. So the backup script now sweeps each
destination each run with `rclone backend cleanup <dest>:<bucket>
-o max-age=<days>d`, aborting incomplete multipart uploads older than the
new tunable backup_s3_abort_incomplete_mpu_days (default 7 — a grace
window far longer than a real upload, so an in-flight upload is never
killed). It runs after each upload attempt and regardless of outcome (a
failed upload is exactly what leaves parts behind); a sweep failure is
non-fatal and never flips the per-destination offsite status. Applies to
both off-site buckets (glesys primary, hetzner secondary).

The equivalent server-side lifecycle policy
(AbortIncompleteMultipartUpload/DaysAfterInitiation) plus a one-time
provider command is documented in the runbook for operators who prefer
to enforce it in the console; the two are complementary.

Closes #160
supernaut sammanfogade incheckning 8f57c0da6a till main 2026-07-21 11:21:00 +00:00
supernaut tog bort grenen feat/160-s3-abort-mpu-lifecycle 2026-07-21 11:21:01 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!181
Ingen beskrivning angiven.