renovate: six updates stuck in awaiting_schedule for 22-32 days, never becoming PRs #487

Stängd
öppnade 2026-09-20 00:03:27 +00:00 av supernaut · 1 kommentar
Ägare

Six Dependency Dashboard entries have sat in awaiting_schedule for 22 to 32 days without ever becoming a PR. RenovateUpdateHeldTooLong is firing for all six. The alert is right; these are stuck, not waiting.

The entries

Measured from bitborg_renovate_update_first_seen_timestamp_seconds, 2026-09-20:

repo branch first seen days held
bitborg-infra renovate/lock-file-maintenance 2026-08-18 32.4
bitborg-infra renovate/toolchain 2026-08-18 32.4
bitborg-infra renovate/victoriametrics 2026-08-20 31.0
bitborg-web renovate/lock-file-maintenance 2026-08-18 32.4
bitborg-web renovate/satteri-0.x 2026-08-19 32.0
bitborg-web renovate/zod-4.x 2026-08-29 22.0

Three of them carry the exporter's rollout timestamp (PR #437, 2026-08-18), so they have been held since the guard's very first scan and have never once cleared. They may have been stuck longer than that; 32 days is a floor, not a measurement of when it started.

Why this is not a legitimate wait

The alert's threshold is derived, not guessed. minimumReleaseAge is 3 days and the schedule window is before 6am on monday, at most 7 days, so 10 days is the smallest value that cannot false-fire (alert-rules.yml.j2:418-428, monitoring/defaults/main.yml:192, threshold introduced in #437). Every entry above is 2 to 3 times that ceiling.

Why it is these six specifically

The same branch names are not stuck elsewhere. Live metric values for renovate/toolchain:

bitborg-infra            1787059754   (32 days, stuck)
bitborg-docs             1789432477   (5 days)
bitborg-internal         1789432477   (5 days)
bitborg-payment          1789518935   (4 days)
bitborg-reconcile-trigger 1789518935  (4 days)
bitborg-auth-reconciler  1789518935   (4 days)

Same shape for renovate/lock-file-maintenance: stuck on bitborg-infra and bitborg-web, recent everywhere else. A recent timestamp means the clock cleared, so the update became a PR and a fresh one was detected. The mechanism works; it is failing for these repo/branch pairs only.

Supporting evidence that the pipeline is otherwise healthy:

  • bitborg_renovate_last_run_status = 0 every night for 35 days.
  • bitborg_renovate_prs_created shows clean weekly bursts exactly 7 days apart, so the Monday window fires and PRs do get created.
  • bitborg_renovate_dashboard_unknown_sections = 0, so no renamed heading is being misparsed into the wrong bucket.
  • Last exporter scan well inside the 30h staleness threshold, so RenovateDashboardScanStale is not masking anything.

What this is NOT

Not the docker-timestamp trap from #430 / #332. That fix is in place and working: default.json:42-45 sets minimumReleaseAgeBehaviour: "timestamp-optional" scoped to matchDatasources: ["docker"], and the logs carry the expected WARN: Some upgrade(s) did not have a releaseTimestamp, but as we're running with minimumReleaseAgeBehaviour=timestamp-optional, proceeding. That trap also parks updates in Pending Status Checks, a different section, and all six here are npm-datasource groups the docker scope does not cover.

Not yet determined

The Renovate-internal reason these six never reach branch creation. Across 29 days of {container="bitborg-renovate"} logs there is no mention of them at all, not even a WARN, while sibling branches in the same nightly runs log normally. Silence rather than an error is itself the finding.

Next step is the one the alert's own annotation prescribes and #430 records how to do without touching production: run Renovate at debug with platform=local and dry-run=full, and read the per-update decision for these six groups.

Six Dependency Dashboard entries have sat in `awaiting_schedule` for 22 to 32 days without ever becoming a PR. `RenovateUpdateHeldTooLong` is firing for all six. The alert is right; these are stuck, not waiting. ## The entries Measured from `bitborg_renovate_update_first_seen_timestamp_seconds`, 2026-09-20: | repo | branch | first seen | days held | | --- | --- | --- | --- | | bitborg-infra | renovate/lock-file-maintenance | 2026-08-18 | 32.4 | | bitborg-infra | renovate/toolchain | 2026-08-18 | 32.4 | | bitborg-infra | renovate/victoriametrics | 2026-08-20 | 31.0 | | bitborg-web | renovate/lock-file-maintenance | 2026-08-18 | 32.4 | | bitborg-web | renovate/satteri-0.x | 2026-08-19 | 32.0 | | bitborg-web | renovate/zod-4.x | 2026-08-29 | 22.0 | Three of them carry the exporter's rollout timestamp (PR #437, 2026-08-18), so they have been held since the guard's very first scan and have never once cleared. **They may have been stuck longer than that; 32 days is a floor, not a measurement of when it started.** ## Why this is not a legitimate wait The alert's threshold is derived, not guessed. `minimumReleaseAge` is 3 days and the schedule window is `before 6am on monday`, at most 7 days, so 10 days is the smallest value that cannot false-fire (`alert-rules.yml.j2:418-428`, `monitoring/defaults/main.yml:192`, threshold introduced in #437). Every entry above is 2 to 3 times that ceiling. ## Why it is these six specifically The same branch names are **not** stuck elsewhere. Live metric values for `renovate/toolchain`: ``` bitborg-infra 1787059754 (32 days, stuck) bitborg-docs 1789432477 (5 days) bitborg-internal 1789432477 (5 days) bitborg-payment 1789518935 (4 days) bitborg-reconcile-trigger 1789518935 (4 days) bitborg-auth-reconciler 1789518935 (4 days) ``` Same shape for `renovate/lock-file-maintenance`: stuck on `bitborg-infra` and `bitborg-web`, recent everywhere else. A recent timestamp means the clock cleared, so the update became a PR and a fresh one was detected. The mechanism works; it is failing for these repo/branch pairs only. Supporting evidence that the pipeline is otherwise healthy: - `bitborg_renovate_last_run_status = 0` every night for 35 days. - `bitborg_renovate_prs_created` shows clean weekly bursts exactly 7 days apart, so the Monday window fires and PRs do get created. - `bitborg_renovate_dashboard_unknown_sections = 0`, so no renamed heading is being misparsed into the wrong bucket. - Last exporter scan well inside the 30h staleness threshold, so `RenovateDashboardScanStale` is not masking anything. ## What this is NOT Not the docker-timestamp trap from #430 / #332. That fix is in place and working: `default.json:42-45` sets `minimumReleaseAgeBehaviour: "timestamp-optional"` scoped to `matchDatasources: ["docker"]`, and the logs carry the expected `WARN: Some upgrade(s) did not have a releaseTimestamp, but as we're running with minimumReleaseAgeBehaviour=timestamp-optional, proceeding`. That trap also parks updates in **Pending Status Checks**, a different section, and all six here are npm-datasource groups the docker scope does not cover. ## Not yet determined The Renovate-internal reason these six never reach branch creation. Across 29 days of `{container="bitborg-renovate"}` logs there is no mention of them at all, not even a WARN, while sibling branches in the same nightly runs log normally. Silence rather than an error is itself the finding. Next step is the one the alert's own annotation prescribes and #430 records how to do without touching production: run Renovate at debug with `platform=local` and `dry-run=full`, and read the per-update decision for these six groups.
Upphovsperson
Ägare

All six stuck updates became PRs and merged on 2026-09-21: bitborg-infra #491 (toolchain), #492 (victoriametrics), #493 (lock file maintenance) and bitborg-web #249 (satteri), #250 (zod), #251 (lock file maintenance). The Renovate-internal reason they were held was never determined.

Closing as resolved. If RenovateUpdateHeldTooLong fires again for npm groups, reopen and run the local repro described above (platform=local, dry-run=full, debug).

All six stuck updates became PRs and merged on 2026-09-21: bitborg-infra #491 (toolchain), #492 (victoriametrics), #493 (lock file maintenance) and bitborg-web #249 (satteri), #250 (zod), #251 (lock file maintenance). The Renovate-internal reason they were held was never determined. Closing as resolved. If `RenovateUpdateHeldTooLong` fires again for npm groups, reopen and run the local repro described above (`platform=local`, `dry-run=full`, debug).
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#487
Ingen beskrivning angiven.