renovate: forgejo 16.0.2 and postgres 17.11 stuck in Pending Status Checks with no branch #430

Stängd
öppnade 2026-08-16 17:58:25 +00:00 av supernaut · 5 kommentarer
Ägare

Two updates have sat in the Dependency Dashboard's Pending Status Checks section without
producing a PR:

  • codeberg.org/forgejo/forgejo 16.0.1-rootless → 16.0.2-rootless
  • docker.io/library/postgres 17.10-trixie → 17.11-trixie

Forgejo is the git server itself, so this is the one dependency we most need to hear about when a
CVE lands. It is the same class of failure ADR 0023 and check-renovate-annotations.py exist to
prevent, arriving by a different route.

What was established

Renovate is healthy. It runs daily and completes cleanly against this repo. From Loki
({container="bitborg-renovate"}), three consecutive nightly runs:

INFO: Dependency extraction complete (repository=bitborg/bitborg-infra, baseBranch=main)

Version 43.288.0. No errors, no rate-limit messages, no skip messages, no repository problems.

The branches do not exist. The dashboard entries carry approvePr-branch= actions, which
implies Renovate has created a branch and is waiting before opening the PR. But
git branch -r shows no renovate/* branch on the remote at all.

The quarantine is satisfied, so minimumReleaseAge is not the cause. Both tags resolve on the
registry, and 16.0.2 shipped 2026-07-30, seventeen days ago against a three-day
minimumReleaseAge:

16.0.1-rootless -> HTTP 200
16.0.2-rootless -> HTTP 200

The concurrency limits are not the cause. prConcurrentLimit: 3 and prHourlyLimit: 2 are set
in the shared preset, but there are zero open Renovate PRs on this repo, and a limited update would
be listed under "Rate Limited" rather than "Pending Status Checks".

The dashboard has not been rewritten since 2026-08-14T00:34:49Z, despite successful runs on the
15th and 16th. That may be benign, since Renovate only rewrites the issue when its content changes,
but it is worth confirming rather than assuming.

Why this could not be finished

renovate_log_level is info (roles/renovate/defaults/main.yml). Renovate does not log
per-update branch or PR decisions at that level, so nothing in Loki explains the deferral. A search
across 14 days for 16.0.2, forgejo-16, minimumReleaseAge, releaseTimestamp,
internalChecks and pending returns zero lines.

Next step

Run once at debug and read the decision:

  1. Set renovate_log_level: "debug" and apply the renovate role.
  2. Trigger a run (systemctl --user start bitborg-renovate.service) rather than waiting for 00:30.
  3. Query Loki for the two depNames and look for the branch/PR decision.
  4. Set the level back to info. Debug is verbose and this is a shared log pipeline.

The likely candidates to confirm or rule out, in order:

  • internalChecksFilter. Its default is strict, and "Pending Status Checks" covers updates
    held by Renovate's own internal checks, not only external CI. If an internal check is unsatisfied
    for a reason other than release age, Renovate defers branch creation entirely, which would match
    the observation of a listed update with no branch.
  • A branch push that failed silently. Would explain the dashboard believing a branch exists.
  • Stale dashboard state, where the section is simply not being refreshed.

A faster but less informative alternative: tick the checkbox on the dashboard entry to force
creation. That may unstick it, but it tells us nothing about why, and this is likely to recur.

Acceptance

  • The cause is identified and recorded here.
  • forgejo 16.0.2 and postgres 17.11 either open as PRs or are documented as deliberately held.
  • If the cause is systemic, a check catches it. A dependency that silently stops being watched is
    exactly what scripts/check-renovate-annotations.py was written for, and that script cannot see
    this failure mode because the annotation here is present and correct.
Two updates have sat in the Dependency Dashboard's **Pending Status Checks** section without producing a PR: - `codeberg.org/forgejo/forgejo` 16.0.1-rootless → 16.0.2-rootless - `docker.io/library/postgres` 17.10-trixie → 17.11-trixie Forgejo is the git server itself, so this is the one dependency we most need to hear about when a CVE lands. It is the same class of failure ADR 0023 and `check-renovate-annotations.py` exist to prevent, arriving by a different route. ## What was established **Renovate is healthy.** It runs daily and completes cleanly against this repo. From Loki (`{container="bitborg-renovate"}`), three consecutive nightly runs: ``` INFO: Dependency extraction complete (repository=bitborg/bitborg-infra, baseBranch=main) ``` Version 43.288.0. No errors, no rate-limit messages, no skip messages, no repository problems. **The branches do not exist.** The dashboard entries carry `approvePr-branch=` actions, which implies Renovate has created a branch and is waiting before opening the PR. But `git branch -r` shows no `renovate/*` branch on the remote at all. **The quarantine is satisfied, so `minimumReleaseAge` is not the cause.** Both tags resolve on the registry, and 16.0.2 shipped 2026-07-30, seventeen days ago against a three-day `minimumReleaseAge`: ``` 16.0.1-rootless -> HTTP 200 16.0.2-rootless -> HTTP 200 ``` **The concurrency limits are not the cause.** `prConcurrentLimit: 3` and `prHourlyLimit: 2` are set in the shared preset, but there are zero open Renovate PRs on this repo, and a limited update would be listed under "Rate Limited" rather than "Pending Status Checks". **The dashboard has not been rewritten since 2026-08-14T00:34:49Z**, despite successful runs on the 15th and 16th. That may be benign, since Renovate only rewrites the issue when its content changes, but it is worth confirming rather than assuming. ## Why this could not be finished `renovate_log_level` is `info` (`roles/renovate/defaults/main.yml`). Renovate does not log per-update branch or PR decisions at that level, so nothing in Loki explains the deferral. A search across 14 days for `16.0.2`, `forgejo-16`, `minimumReleaseAge`, `releaseTimestamp`, `internalChecks` and `pending` returns **zero** lines. ## Next step Run once at debug and read the decision: 1. Set `renovate_log_level: "debug"` and apply the renovate role. 2. Trigger a run (`systemctl --user start bitborg-renovate.service`) rather than waiting for 00:30. 3. Query Loki for the two depNames and look for the branch/PR decision. 4. Set the level back to `info`. Debug is verbose and this is a shared log pipeline. The likely candidates to confirm or rule out, in order: - **`internalChecksFilter`.** Its default is `strict`, and "Pending Status Checks" covers updates held by Renovate's own internal checks, not only external CI. If an internal check is unsatisfied for a reason other than release age, Renovate defers branch creation entirely, which would match the observation of a listed update with no branch. - **A branch push that failed silently.** Would explain the dashboard believing a branch exists. - **Stale dashboard state**, where the section is simply not being refreshed. A faster but less informative alternative: tick the checkbox on the dashboard entry to force creation. That may unstick it, but it tells us nothing about why, and this is likely to recur. ## Acceptance - The cause is identified and recorded here. - `forgejo` 16.0.2 and `postgres` 17.11 either open as PRs or are documented as deliberately held. - If the cause is systemic, a check catches it. A dependency that silently stops being watched is exactly what `scripts/check-renovate-annotations.py` was written for, and that script cannot see this failure mode because the annotation here is present and correct.
Upphovsperson
Ägare

Root cause

minimumReleaseAgeBehaviour defaults to timestamp-required (Renovate's schema: enum
timestamp-required / timestamp-optional, default timestamp-required). The docker datasource
often returns no releaseTimestamp
, and a check that needs a timestamp can never pass without one.
So the update is deferred on every run and the branch is never created. Renovate says it itself:

DEBUG: Marking 1 release(s) as pending, as they do not have a releaseTimestamp
and we're running with minimumReleaseAgeBehaviour=timestamp-required
       "depName": "codeberg.org/forgejo/forgejo"
       "versions": ["16.0.2-rootless"]
       "check": "minimumReleaseAge"

This is not a misconfiguration. Nothing in the shared preset or this repo sets the option; the
upstream default fails closed, and the docker datasource cannot satisfy it.

Three corrections to the issue as filed

It is seven updates, not two. The whole "Pending Status Checks" section is affected:

Update Registry release age
codeberg.org/forgejo/forgejo 16.0.2-rootless 18.2 d
docker.io/grafana/loki 3.7.6 12.2 d
docker.io/victoriametrics/victoria-metrics v1.149.0 17.4 d
docker.io/victoriametrics/vmagent v1.149.0 17.4 d
docker.io/victoriametrics/vmalert v1.149.0 17.4 d
docker.io/library/postgres 17.11-trixie 4.1 d
quay.io/prometheus/alertmanager v0.34.0 1.2 d

docker.io/grafana/alloy opened a PR (#433) in the same run from the same defaults file as the
held vmagent. That rules out the manager, the file, the datasource and the update type, and is the
observation that cracked this.

The dashboard is not stale. #15 was rewritten 2026-08-17T00:32:46Z, after this issue was filed.
"Stale dashboard state" is ruled out.

"The quarantine is satisfied" was not evidence of anything. Renovate never compares ages when the
timestamp is absent, so the release dates in the issue (correct as facts) could not have told us
whether the update would ever be offered. alertmanager at 1.2 d is held for the same reason as
forgejo at 18.2 d.

Why it reads as a status-check problem

  • With prCreation at its immediate default, "Pending Status Checks" can only mean an unmet
    internal check. No external CI is involved despite the section name.
  • The approvePr-branch=<branch> action reads as though a branch exists. None is created. Note that
    git branch -r only reads local remote-tracking refs; git ls-remote is the check that actually
    asks the server.
  • The internal-checks gate runs before PR-limit accounting, so a held update never reaches
    prHourlyLimit and never shows under "Rate-Limited". The two rate-limited entries and the seven
    pending ones are two separate gates, which is why the concurrency limits looked innocent (they
    are).

No debug apply needed

The plan in this issue was to set renovate_log_level: debug, apply, trigger, read, set it back.
That is not necessary. Renovate's decision reproduces with no production access, no token and no
Ansible apply, on 43.288.0 (the version production runs):

# repo copy in a podman volume; local> presets inlined; schedule overridden to "at any time"
podman run --rm -v bitborg-rn:/repo -w /repo \
  -e RENOVATE_PLATFORM=local -e RENOVATE_DRY_RUN=full -e LOG_LEVEL=debug \
  docker.io/renovate/renovate:43

It named exactly the seven dashboard entries and nothing else.

Fix

bitborg/bitborg-docs#91 adds one packageRule to the shared preset, scoped to
matchDatasources: ["docker"], setting minimumReleaseAgeBehaviour: "timestamp-optional".
Datasources that do supply timestamps keep failing closed. Verified: holds go 8 → 0 and all seven
produce branch intents.

The shared preset is resolved from main at run time, so merging is the whole deployment. No
apply, no restart.

Trade-off, stated plainly: for images whose timestamp never arrives, the 3-day quarantine stops
applying. It was not protecting those images before, it was permanently blocking them, which is the
worse failure for the git server and the database. Image-tag bumps are never automerged here, so a
human still reviews each one, and Renovate logs a WARN at info naming the releases it proceeded on,
so a skipped quarantine stays visible without raising the log level.

On the acceptance criterion "if the cause is systemic, a check catches it"

Two notes for whoever writes that check.

renovate-config-validator will not catch this class. It validates option names but not enum
values: minimumReleaseAgeBehaviour: "totally-bogus-value" passes with
Config validated successfully, while a typo'd option name is rejected.

The observable failure is "an entry sits in Pending Status Checks indefinitely". A check that ages
the dashboard's entries and fails past a threshold would catch this and anything else that stalls
the same way. scripts/check-renovate-annotations.py cannot see it, as the issue says, because the
annotation here is present and correct.

Also worth a separate look: renovate_image is pinned to a floating major
(.../bitborg/renovate:43), so an upstream default can change without any Renovate PR. #203 opened
forgejo 16.0.1 on 2026-07-22, so this did work before. I have not confirmed that a version bump
introduced the behaviour, so that is a hypothesis, not a finding.

#332 is the same root cause and is answered by this: the release-age quarantine does not release
docker-datasource updates.

## Root cause `minimumReleaseAgeBehaviour` defaults to **`timestamp-required`** (Renovate's schema: enum `timestamp-required` / `timestamp-optional`, default `timestamp-required`). The **docker datasource often returns no `releaseTimestamp`**, and a check that needs a timestamp can never pass without one. So the update is deferred on every run and the branch is never created. Renovate says it itself: ``` DEBUG: Marking 1 release(s) as pending, as they do not have a releaseTimestamp and we're running with minimumReleaseAgeBehaviour=timestamp-required "depName": "codeberg.org/forgejo/forgejo" "versions": ["16.0.2-rootless"] "check": "minimumReleaseAge" ``` This is not a misconfiguration. Nothing in the shared preset or this repo sets the option; the upstream default fails closed, and the docker datasource cannot satisfy it. ## Three corrections to the issue as filed **It is seven updates, not two.** The whole "Pending Status Checks" section is affected: | Update | Registry release age | | --- | --- | | `codeberg.org/forgejo/forgejo` 16.0.2-rootless | 18.2 d | | `docker.io/grafana/loki` 3.7.6 | 12.2 d | | `docker.io/victoriametrics/victoria-metrics` v1.149.0 | 17.4 d | | `docker.io/victoriametrics/vmagent` v1.149.0 | 17.4 d | | `docker.io/victoriametrics/vmalert` v1.149.0 | 17.4 d | | `docker.io/library/postgres` 17.11-trixie | 4.1 d | | `quay.io/prometheus/alertmanager` v0.34.0 | 1.2 d | `docker.io/grafana/alloy` opened a PR (#433) in the same run from the **same defaults file** as the held `vmagent`. That rules out the manager, the file, the datasource and the update type, and is the observation that cracked this. **The dashboard is not stale.** #15 was rewritten `2026-08-17T00:32:46Z`, after this issue was filed. "Stale dashboard state" is ruled out. **"The quarantine is satisfied" was not evidence of anything.** Renovate never compares ages when the timestamp is absent, so the release dates in the issue (correct as facts) could not have told us whether the update would ever be offered. `alertmanager` at 1.2 d is held for the same reason as `forgejo` at 18.2 d. ## Why it reads as a status-check problem - With `prCreation` at its `immediate` default, "Pending Status Checks" can **only** mean an unmet *internal* check. No external CI is involved despite the section name. - The `approvePr-branch=<branch>` action reads as though a branch exists. None is created. Note that `git branch -r` only reads local remote-tracking refs; `git ls-remote` is the check that actually asks the server. - The internal-checks gate runs **before** PR-limit accounting, so a held update never reaches `prHourlyLimit` and never shows under "Rate-Limited". The two rate-limited entries and the seven pending ones are two separate gates, which is why the concurrency limits looked innocent (they are). ## No debug apply needed The plan in this issue was to set `renovate_log_level: debug`, apply, trigger, read, set it back. That is not necessary. Renovate's decision reproduces with no production access, no token and no Ansible apply, on **43.288.0** (the version production runs): ```bash # repo copy in a podman volume; local> presets inlined; schedule overridden to "at any time" podman run --rm -v bitborg-rn:/repo -w /repo \ -e RENOVATE_PLATFORM=local -e RENOVATE_DRY_RUN=full -e LOG_LEVEL=debug \ docker.io/renovate/renovate:43 ``` It named exactly the seven dashboard entries and nothing else. ## Fix bitborg/bitborg-docs#91 adds one `packageRule` to the shared preset, scoped to `matchDatasources: ["docker"]`, setting `minimumReleaseAgeBehaviour: "timestamp-optional"`. Datasources that do supply timestamps keep failing closed. Verified: holds go 8 → 0 and all seven produce branch intents. The shared preset is resolved from `main` at run time, so **merging is the whole deployment**. No apply, no restart. Trade-off, stated plainly: for images whose timestamp never arrives, the 3-day quarantine stops applying. It was not protecting those images before, it was permanently blocking them, which is the worse failure for the git server and the database. Image-tag bumps are never automerged here, so a human still reviews each one, and Renovate logs a WARN at `info` naming the releases it proceeded on, so a skipped quarantine stays visible without raising the log level. ## On the acceptance criterion "if the cause is systemic, a check catches it" Two notes for whoever writes that check. `renovate-config-validator` will not catch this class. It validates option **names** but **not** enum **values**: `minimumReleaseAgeBehaviour: "totally-bogus-value"` passes with `Config validated successfully`, while a typo'd option name is rejected. The observable failure is "an entry sits in Pending Status Checks indefinitely". A check that ages the dashboard's entries and fails past a threshold would catch this and anything else that stalls the same way. `scripts/check-renovate-annotations.py` cannot see it, as the issue says, because the annotation here is present and correct. Also worth a separate look: `renovate_image` is pinned to a **floating major** (`.../bitborg/renovate:43`), so an upstream default can change without any Renovate PR. #203 opened `forgejo` 16.0.1 on 2026-07-22, so this did work before. I have not confirmed that a version bump introduced the behaviour, so that is a hypothesis, not a finding. #332 is the same root cause and is answered by this: the release-age quarantine does **not** release docker-datasource updates.
Upphovsperson
Ägare

Verified on production

bitborg/bitborg-docs#91 merged 09:04Z. One Renovate run triggered by hand at 09:19Z
(systemctl --user start bitborg-renovate.service, the same command the nightly timer runs).

Dashboard #15, before and after that run:

Section Before After
Pending Status Checks 7 1
Awaiting Schedule 2 8
PRs created — 0

The six held image updates moved to "Awaiting Schedule": forgejo 16.0.2, loki 3.7.6,
victoria-metrics, vmagent, vmalert (all now v1.150.0) and alertmanager v0.34.0. No PRs were
created, which is correct: the shared preset's global schedule is before 6am on monday, so they
open on Monday 2026-08-24. Run metrics: last_run_status 0, prs_created 0,
repos_processed 8.

The move from "Pending Status Checks" to "Awaiting Schedule" is the proof, and it is a better one
than PRs appearing would have been. It shows the internal release-age check now passes while the
weekly window is what holds the PRs. Renovate confirms the reason in the journal, at info, without
debug logging:

WARN: Some release(s) did not have a releaseTimestamp, but as we're running with
minimumReleaseAgeBehaviour=timestamp-optional, proceeding.
(repository=bitborg/bitborg-infra)

The one still held is the quarantine working, not the bug

postgres 17.11-trixie stays in "Pending Status Checks", and that is the desired outcome rather
than an incomplete fix. timestamp-optional only changes behaviour when the timestamp is absent,
so a remaining hold means Renovate now has a real timestamp and is applying the 3-day quarantine to
it properly. The tag was re-pushed 2026-08-16T01:08:18Z, which was 2.34 days before the run. It
clears at 2026-08-19T01:08Z and then joins the others in "Awaiting Schedule".

So the change did not blanket-disable the supply-chain quarantine. It removed a permanent block
while leaving the quarantine in force wherever it can actually be evaluated, which was the intent of
scoping the rule to matchDatasources: ["docker"].

A correction, to this issue and to my own earlier comment

I argued above that the dashboard was not stale because #15's updated_at had moved. That reasoning
was wrong, in the same way the original issue's was. Any cross-reference bumps an issue's
updated_at
— the 05:08:07Z bump I cited was my own previous comment referencing "#15", not a
Renovate run. Neither "it has not been updated since the 14th" nor "it was updated today" tells you
anything about when Renovate last rewrote it.

The trustworthy signal is bitborg_renovate_last_run_timestamp_seconds, which showed the real
sequence: last run 00:32:23Z, merge 09:04:03Z, triggered run 09:19:05Z. Worth using that rather than
updated_at next time.

Remaining on this issue

Only the third acceptance criterion: a check that catches this class. Two notes for it, one of which
is a trap.

renovate-config-validator cannot catch it. It validates option names but not enum
values — minimumReleaseAgeBehaviour: "totally-bogus-value" passes with
Config validated successfully, while a typo'd option name is rejected. Verified both directions.

The observable failure is "an entry sits in a holding section longer than that section can explain".
A check that ages the dashboard's entries and fails past a threshold would catch this and anything
else that stalls the same way. It has to read the run timestamp rather than the issue's updated_at,
per the correction above.

## Verified on production bitborg/bitborg-docs#91 merged 09:04Z. One Renovate run triggered by hand at 09:19Z (`systemctl --user start bitborg-renovate.service`, the same command the nightly timer runs). Dashboard #15, before and after that run: | Section | Before | After | | --- | --- | --- | | Pending Status Checks | 7 | **1** | | Awaiting Schedule | 2 | **8** | | PRs created | — | **0** | The six held image updates moved to "Awaiting Schedule": `forgejo` 16.0.2, `loki` 3.7.6, `victoria-metrics`, `vmagent`, `vmalert` (all now v1.150.0) and `alertmanager` v0.34.0. No PRs were created, which is correct: the shared preset's global `schedule` is `before 6am on monday`, so they open on **Monday 2026-08-24**. Run metrics: `last_run_status 0`, `prs_created 0`, `repos_processed 8`. The move from "Pending Status Checks" to "Awaiting Schedule" is the proof, and it is a better one than PRs appearing would have been. It shows the internal release-age check now passes while the weekly window is what holds the PRs. Renovate confirms the reason in the journal, at `info`, without debug logging: ``` WARN: Some release(s) did not have a releaseTimestamp, but as we're running with minimumReleaseAgeBehaviour=timestamp-optional, proceeding. (repository=bitborg/bitborg-infra) ``` ## The one still held is the quarantine working, not the bug `postgres` 17.11-trixie stays in "Pending Status Checks", and that is the desired outcome rather than an incomplete fix. `timestamp-optional` only changes behaviour when the timestamp is **absent**, so a remaining hold means Renovate now has a real timestamp and is applying the 3-day quarantine to it properly. The tag was re-pushed `2026-08-16T01:08:18Z`, which was 2.34 days before the run. It clears at **2026-08-19T01:08Z** and then joins the others in "Awaiting Schedule". So the change did not blanket-disable the supply-chain quarantine. It removed a permanent block while leaving the quarantine in force wherever it can actually be evaluated, which was the intent of scoping the rule to `matchDatasources: ["docker"]`. ## A correction, to this issue and to my own earlier comment I argued above that the dashboard was not stale because #15's `updated_at` had moved. That reasoning was wrong, in the same way the original issue's was. **Any cross-reference bumps an issue's `updated_at`** — the 05:08:07Z bump I cited was my own previous comment referencing "#15", not a Renovate run. Neither "it has not been updated since the 14th" nor "it was updated today" tells you anything about when Renovate last rewrote it. The trustworthy signal is `bitborg_renovate_last_run_timestamp_seconds`, which showed the real sequence: last run 00:32:23Z, merge 09:04:03Z, triggered run 09:19:05Z. Worth using that rather than `updated_at` next time. ## Remaining on this issue Only the third acceptance criterion: a check that catches this class. Two notes for it, one of which is a trap. `renovate-config-validator` cannot catch it. It validates option **names** but **not** enum **values** — `minimumReleaseAgeBehaviour: "totally-bogus-value"` passes with `Config validated successfully`, while a typo'd option name is rejected. Verified both directions. The observable failure is "an entry sits in a holding section longer than that section can explain". A check that ages the dashboard's entries and fails past a threshold would catch this and anything else that stalls the same way. It has to read the run timestamp rather than the issue's `updated_at`, per the correction above.
Upphovsperson
Ägare

The check is in production and verified

#437 and #438 are applied. All three acceptance criteria are now met.

Applies, each through scripts/apply-reconcile.sh, reconciled by task name:

Apply Predicted Actual Reconcile
#437 bitborg-prod --tags renovate 3 3 matched exactly
#437 bitborg-monitoring --tags monitoring 2 2 matched exactly
#438 bitborg-prod --tags renovate 1 1 matched exactly

A second --check returned changed=0 on both hosts after each, which is the only thing that proves
the hosts match main.

The rules are not merely rendered, they are loaded and evaluating:

vmalert_alerting_rules_last_evaluation_samples{alertname=~"Renovate.*"}
  RenovateUpdateHeldTooLong        group=bitborg-observability
  RenovateDashboardGuardDegraded   group=bitborg-observability
  RenovateDashboardScanStale       group=bitborg-observability
vmalert_config_last_reload_successful 1
vmalert_config_last_reload_errors_total 0

The exporter was run by hand rather than waiting for the nightly, and its metrics are queryable in
VictoriaMetrics with the labels the alerts select on: 30 held updates across all 8 repositories
Renovate autodiscovers. Only Watchdog is firing.

The guard found its own blind spot on the first run

Worth recording, because it is the argument for the cheapest part of the design. That first manual
run reported:

renovate-dashboard: 30 held across 8 repo(s), oldest 0.0d, unknown sections 1, status 0

unknown sections 1. A dashboard carried a heading the parser did not know — PR Closed (Blocked),
which Renovate renders with a recreate-branch action for updates whose PR a human closed to reject
them. Renovate never recreates those unaided, so they are terminal, not waiting; ageing them would
have alerted for ever about a decision already made. #438 maps it as not-held. The same signal
confirms the fix: the count is now 0 after a re-run and a scrape.

That gap would not have been caught by any of the tests, because the tests assert my model of the
dashboard and the section was outside it. The counter questioned the model instead. It is four lines
of code and one alert.

One expectation to keep calibrated

The clocks start at first observation, so this alert will not retroactively flag the current
stall.
State is fail-open by design: losing or lacking it resets the ages. Six of the seven held
updates become PRs on Monday 2026-08-24 when the weekly schedule window opens, so nothing is missed
in practice, but do not wait for an alert about the backlog that prompted this issue.

Left open deliberately

Keeping this issue open until Monday, when the six updates should actually open as PRs. The mechanism
is proven — they moved from "Pending Status Checks" to "Awaiting Schedule" and the release-age check
now passes — but a Forgejo security update reaching a reviewable PR is the outcome that matters, and
that has not happened yet. Closing before observing it would be trusting a green check that has not
been seen to fire.

Follow-up not folded into either PR: nothing in this repo validates alert rules. promtool check rules was run by hand against the rendered template (49 rules, and a deliberately unbalanced paren
was confirmed to fail it). Wiring it into CI needs Jinja rendering with resolved role vars, so it
deserves its own issue.

## The check is in production and verified #437 and #438 are applied. All three acceptance criteria are now met. Applies, each through `scripts/apply-reconcile.sh`, reconciled by task name: | Apply | Predicted | Actual | Reconcile | | --- | --- | --- | --- | | #437 `bitborg-prod` `--tags renovate` | 3 | 3 | matched exactly | | #437 `bitborg-monitoring` `--tags monitoring` | 2 | 2 | matched exactly | | #438 `bitborg-prod` `--tags renovate` | 1 | 1 | matched exactly | A second `--check` returned `changed=0` on both hosts after each, which is the only thing that proves the hosts match `main`. The rules are not merely rendered, they are loaded and evaluating: ``` vmalert_alerting_rules_last_evaluation_samples{alertname=~"Renovate.*"} RenovateUpdateHeldTooLong group=bitborg-observability RenovateDashboardGuardDegraded group=bitborg-observability RenovateDashboardScanStale group=bitborg-observability vmalert_config_last_reload_successful 1 vmalert_config_last_reload_errors_total 0 ``` The exporter was run by hand rather than waiting for the nightly, and its metrics are queryable in VictoriaMetrics with the labels the alerts select on: 30 held updates across all 8 repositories Renovate autodiscovers. Only `Watchdog` is firing. ## The guard found its own blind spot on the first run Worth recording, because it is the argument for the cheapest part of the design. That first manual run reported: ``` renovate-dashboard: 30 held across 8 repo(s), oldest 0.0d, unknown sections 1, status 0 ``` `unknown sections 1`. A dashboard carried a heading the parser did not know — `PR Closed (Blocked)`, which Renovate renders with a `recreate-branch` action for updates whose PR a human closed to reject them. Renovate never recreates those unaided, so they are terminal, not waiting; ageing them would have alerted for ever about a decision already made. #438 maps it as not-held. The same signal confirms the fix: the count is now `0` after a re-run and a scrape. That gap would not have been caught by any of the tests, because the tests assert my model of the dashboard and the section was outside it. The counter questioned the model instead. It is four lines of code and one alert. ## One expectation to keep calibrated The clocks start at first observation, so **this alert will not retroactively flag the current stall.** State is fail-open by design: losing or lacking it resets the ages. Six of the seven held updates become PRs on Monday 2026-08-24 when the weekly schedule window opens, so nothing is missed in practice, but do not wait for an alert about the backlog that prompted this issue. ## Left open deliberately Keeping this issue open until Monday, when the six updates should actually open as PRs. The mechanism is proven — they moved from "Pending Status Checks" to "Awaiting Schedule" and the release-age check now passes — but a Forgejo security update reaching a reviewable PR is the outcome that matters, and that has not happened yet. Closing before observing it would be trusting a green check that has not been seen to fire. Follow-up not folded into either PR: **nothing in this repo validates alert rules.** `promtool check rules` was run by hand against the rendered template (49 rules, and a deliberately unbalanced paren was confirmed to fail it). Wiring it into CI needs Jinja rendering with resolved role vars, so it deserves its own issue.
Upphovsperson
Ägare

The fix is confirmed working. Do not close this yet, but the mechanism is no longer in doubt.

Checked 2026-08-19, after last night's run (Renovate healthy: bitborg_renovate_last_run_status = 0,
last run 6.1 h ago).

The dashboard now carries the WARN the fix predicts:

⚠️ WARN: Some release(s) did not have a releaseTimestamp, but as we're running with
   minimumReleaseAgeBehaviour=timestamp-optional, proceeding.

And forgejo 16.0.2 has moved out of "Pending Status Checks" into "Awaiting Schedule". That
transition is the evidence: the update is no longer deferred by an internal check that could never
pass, it is simply waiting for the Monday window.

Two corrections to what was expected:

  • It is eight updates awaiting schedule, not six: forgejo 16.0.2, loki 3.7.6, victoria-metrics /
    vmagent / vmalert 1.150.0, alertmanager 0.34.0, pnpm 11.22.0, and lock-file maintenance.
  • postgres 17.11 is still in "Pending Status Checks", not awaiting schedule. That is consistent
    rather than a new fault: it is the one image whose releaseTimestamp did arrive, so it is under the
    genuine three-day quarantine, which had not elapsed at 00:32 today. Expect it to move on tonight's
    run. If it has not by 2026-08-20, that is a new problem rather than this one.

Also new, and working as designed: postgres 18 now sits under "Pending Approval", held by the
major gate.

Filed separately: #446

This issue's acceptance says the outcome that matters is "a Forgejo security update reaching a
reviewable PR". Watching 16.0.2 open on Monday proves the timestamp-required bug here is fixed. It
does not prove security updates arrive promptly, because a Monday-scheduled update opens on Monday
either way. One observation cannot separate the two.

#446 records why: vulnerabilityAlerts sets schedule: ["at any time"], but it only reaches updates
Renovate itself classifies, and there is no vulnerability feed for a docker image tag. So a git-server
security patch waits up to six days.

Closing this on Monday is correct. It should not be read as closing that.

Prepared ahead of Monday

Three of the eight are the VictoriaMetrics trio, which two role-defaults comments require to move in
lockstep and which nothing grouped. Ungrouped they arrive as three PRs. PR #445 groups them into one,
reproduced with platform=local and a control run.

Still to do by hand on Monday: the ADR 0038 concealment gate. The 16.0.2 bump must carry a
health_check_forgejo_concealment_verified_tag bump in the same change, after checking the surfaces.
Never pre-bump the tag, and never commit a disabled gate.

## The fix is confirmed working. Do not close this yet, but the mechanism is no longer in doubt. Checked 2026-08-19, after last night's run (Renovate healthy: `bitborg_renovate_last_run_status = 0`, last run 6.1 h ago). **The dashboard now carries the WARN the fix predicts:** ``` ⚠️ WARN: Some release(s) did not have a releaseTimestamp, but as we're running with minimumReleaseAgeBehaviour=timestamp-optional, proceeding. ``` **And `forgejo 16.0.2` has moved out of "Pending Status Checks" into "Awaiting Schedule".** That transition is the evidence: the update is no longer deferred by an internal check that could never pass, it is simply waiting for the Monday window. Two corrections to what was expected: - It is **eight** updates awaiting schedule, not six: forgejo 16.0.2, loki 3.7.6, victoria-metrics / vmagent / vmalert 1.150.0, alertmanager 0.34.0, pnpm 11.22.0, and lock-file maintenance. - **postgres 17.11 is still in "Pending Status Checks"**, not awaiting schedule. That is consistent rather than a new fault: it is the one image whose `releaseTimestamp` did arrive, so it is under the genuine three-day quarantine, which had not elapsed at 00:32 today. Expect it to move on tonight's run. If it has not by 2026-08-20, that is a new problem rather than this one. Also new, and working as designed: **postgres 18** now sits under "Pending Approval", held by the major gate. ## Filed separately: #446 This issue's acceptance says the outcome that matters is "a Forgejo **security** update reaching a reviewable PR". Watching 16.0.2 open on Monday proves the `timestamp-required` bug here is fixed. It does **not** prove security updates arrive promptly, because a Monday-scheduled update opens on Monday either way. One observation cannot separate the two. #446 records why: `vulnerabilityAlerts` sets `schedule: ["at any time"]`, but it only reaches updates Renovate itself classifies, and there is no vulnerability feed for a docker image tag. So a git-server security patch waits up to six days. Closing this on Monday is correct. It should not be read as closing that. ## Prepared ahead of Monday Three of the eight are the VictoriaMetrics trio, which two role-defaults comments require to move in lockstep and which nothing grouped. Ungrouped they arrive as three PRs. PR #445 groups them into one, reproduced with `platform=local` and a control run. Still to do by hand on Monday: the ADR 0038 concealment gate. The 16.0.2 bump must carry a `health_check_forgejo_concealment_verified_tag` bump in the same change, after checking the surfaces. Never pre-bump the tag, and never commit a disabled gate.
Upphovsperson
Ägare

Closing as done. The timestamp-required stall was fixed by bitborg/bitborg-docs#91 (minimumReleaseAgeBehaviour: timestamp-optional for the docker datasource) and confirmed on 2026-08-19. Forgejo updates have since arrived as normal PRs, through 16.0.5.

As the last comment says, this does not prove security updates arrive promptly. A docker-tag security patch still waits for the weekly schedule. That gap is tracked in #446 and #487.

Closing as done. The `timestamp-required` stall was fixed by bitborg/bitborg-docs#91 (`minimumReleaseAgeBehaviour: timestamp-optional` for the docker datasource) and confirmed on 2026-08-19. Forgejo updates have since arrived as normal PRs, through 16.0.5. As the last comment says, this does not prove security updates arrive promptly. A docker-tag security patch still waits for the weekly schedule. That gap is tracked in #446 and #487.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#430
Ingen beskrivning angiven.