renovate: forgejo 16.0.2 and postgres 17.11 stuck in Pending Status Checks with no branch #430
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra#430
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "%!s()"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
Two updates have sat in the Dependency Dashboard's Pending Status Checks section without
producing a PR:
codeberg.org/forgejo/forgejo16.0.1-rootless → 16.0.2-rootlessdocker.io/library/postgres17.10-trixie → 17.11-trixieForgejo is the git server itself, so this is the one dependency we most need to hear about when a
CVE lands. It is the same class of failure ADR 0023 and
check-renovate-annotations.pyexist toprevent, arriving by a different route.
What was established
Renovate is healthy. It runs daily and completes cleanly against this repo. From Loki
(
{container="bitborg-renovate"}), three consecutive nightly runs:Version 43.288.0. No errors, no rate-limit messages, no skip messages, no repository problems.
The branches do not exist. The dashboard entries carry
approvePr-branch=actions, whichimplies Renovate has created a branch and is waiting before opening the PR. But
git branch -rshows norenovate/*branch on the remote at all.The quarantine is satisfied, so
minimumReleaseAgeis not the cause. Both tags resolve on theregistry, and 16.0.2 shipped 2026-07-30, seventeen days ago against a three-day
minimumReleaseAge:The concurrency limits are not the cause.
prConcurrentLimit: 3andprHourlyLimit: 2are setin the shared preset, but there are zero open Renovate PRs on this repo, and a limited update would
be listed under "Rate Limited" rather than "Pending Status Checks".
The dashboard has not been rewritten since 2026-08-14T00:34:49Z, despite successful runs on the
15th and 16th. That may be benign, since Renovate only rewrites the issue when its content changes,
but it is worth confirming rather than assuming.
Why this could not be finished
renovate_log_levelisinfo(roles/renovate/defaults/main.yml). Renovate does not logper-update branch or PR decisions at that level, so nothing in Loki explains the deferral. A search
across 14 days for
16.0.2,forgejo-16,minimumReleaseAge,releaseTimestamp,internalChecksandpendingreturns zero lines.Next step
Run once at debug and read the decision:
renovate_log_level: "debug"and apply the renovate role.systemctl --user start bitborg-renovate.service) rather than waiting for 00:30.info. Debug is verbose and this is a shared log pipeline.The likely candidates to confirm or rule out, in order:
internalChecksFilter. Its default isstrict, and "Pending Status Checks" covers updatesheld by Renovate's own internal checks, not only external CI. If an internal check is unsatisfied
for a reason other than release age, Renovate defers branch creation entirely, which would match
the observation of a listed update with no branch.
A faster but less informative alternative: tick the checkbox on the dashboard entry to force
creation. That may unstick it, but it tells us nothing about why, and this is likely to recur.
Acceptance
forgejo16.0.2 andpostgres17.11 either open as PRs or are documented as deliberately held.exactly what
scripts/check-renovate-annotations.pywas written for, and that script cannot seethis failure mode because the annotation here is present and correct.
Root cause
minimumReleaseAgeBehaviourdefaults totimestamp-required(Renovate's schema: enumtimestamp-required/timestamp-optional, defaulttimestamp-required). The docker datasourceoften returns no
releaseTimestamp, and a check that needs a timestamp can never pass without one.So the update is deferred on every run and the branch is never created. Renovate says it itself:
This is not a misconfiguration. Nothing in the shared preset or this repo sets the option; the
upstream default fails closed, and the docker datasource cannot satisfy it.
Three corrections to the issue as filed
It is seven updates, not two. The whole "Pending Status Checks" section is affected:
codeberg.org/forgejo/forgejo16.0.2-rootlessdocker.io/grafana/loki3.7.6docker.io/victoriametrics/victoria-metricsv1.149.0docker.io/victoriametrics/vmagentv1.149.0docker.io/victoriametrics/vmalertv1.149.0docker.io/library/postgres17.11-trixiequay.io/prometheus/alertmanagerv0.34.0docker.io/grafana/alloyopened a PR (#433) in the same run from the same defaults file as theheld
vmagent. That rules out the manager, the file, the datasource and the update type, and is theobservation that cracked this.
The dashboard is not stale. #15 was rewritten
2026-08-17T00:32:46Z, after this issue was filed."Stale dashboard state" is ruled out.
"The quarantine is satisfied" was not evidence of anything. Renovate never compares ages when the
timestamp is absent, so the release dates in the issue (correct as facts) could not have told us
whether the update would ever be offered.
alertmanagerat 1.2 d is held for the same reason asforgejoat 18.2 d.Why it reads as a status-check problem
prCreationat itsimmediatedefault, "Pending Status Checks" can only mean an unmetinternal check. No external CI is involved despite the section name.
approvePr-branch=<branch>action reads as though a branch exists. None is created. Note thatgit branch -ronly reads local remote-tracking refs;git ls-remoteis the check that actuallyasks the server.
prHourlyLimitand never shows under "Rate-Limited". The two rate-limited entries and the sevenpending ones are two separate gates, which is why the concurrency limits looked innocent (they
are).
No debug apply needed
The plan in this issue was to set
renovate_log_level: debug, apply, trigger, read, set it back.That is not necessary. Renovate's decision reproduces with no production access, no token and no
Ansible apply, on 43.288.0 (the version production runs):
It named exactly the seven dashboard entries and nothing else.
Fix
bitborg/bitborg-docs#91 adds one
packageRuleto the shared preset, scoped tomatchDatasources: ["docker"], settingminimumReleaseAgeBehaviour: "timestamp-optional".Datasources that do supply timestamps keep failing closed. Verified: holds go 8 → 0 and all seven
produce branch intents.
The shared preset is resolved from
mainat run time, so merging is the whole deployment. Noapply, no restart.
Trade-off, stated plainly: for images whose timestamp never arrives, the 3-day quarantine stops
applying. It was not protecting those images before, it was permanently blocking them, which is the
worse failure for the git server and the database. Image-tag bumps are never automerged here, so a
human still reviews each one, and Renovate logs a WARN at
infonaming the releases it proceeded on,so a skipped quarantine stays visible without raising the log level.
On the acceptance criterion "if the cause is systemic, a check catches it"
Two notes for whoever writes that check.
renovate-config-validatorwill not catch this class. It validates option names but not enumvalues:
minimumReleaseAgeBehaviour: "totally-bogus-value"passes withConfig validated successfully, while a typo'd option name is rejected.The observable failure is "an entry sits in Pending Status Checks indefinitely". A check that ages
the dashboard's entries and fails past a threshold would catch this and anything else that stalls
the same way.
scripts/check-renovate-annotations.pycannot see it, as the issue says, because theannotation here is present and correct.
Also worth a separate look:
renovate_imageis pinned to a floating major(
.../bitborg/renovate:43), so an upstream default can change without any Renovate PR. #203 openedforgejo16.0.1 on 2026-07-22, so this did work before. I have not confirmed that a version bumpintroduced the behaviour, so that is a hypothesis, not a finding.
#332 is the same root cause and is answered by this: the release-age quarantine does not release
docker-datasource updates.
Verified on production
bitborg/bitborg-docs#91 merged 09:04Z. One Renovate run triggered by hand at 09:19Z
(
systemctl --user start bitborg-renovate.service, the same command the nightly timer runs).Dashboard #15, before and after that run:
The six held image updates moved to "Awaiting Schedule":
forgejo16.0.2,loki3.7.6,victoria-metrics,vmagent,vmalert(all now v1.150.0) andalertmanagerv0.34.0. No PRs werecreated, which is correct: the shared preset's global
scheduleisbefore 6am on monday, so theyopen on Monday 2026-08-24. Run metrics:
last_run_status 0,prs_created 0,repos_processed 8.The move from "Pending Status Checks" to "Awaiting Schedule" is the proof, and it is a better one
than PRs appearing would have been. It shows the internal release-age check now passes while the
weekly window is what holds the PRs. Renovate confirms the reason in the journal, at
info, withoutdebug logging:
The one still held is the quarantine working, not the bug
postgres17.11-trixie stays in "Pending Status Checks", and that is the desired outcome ratherthan an incomplete fix.
timestamp-optionalonly changes behaviour when the timestamp is absent,so a remaining hold means Renovate now has a real timestamp and is applying the 3-day quarantine to
it properly. The tag was re-pushed
2026-08-16T01:08:18Z, which was 2.34 days before the run. Itclears at 2026-08-19T01:08Z and then joins the others in "Awaiting Schedule".
So the change did not blanket-disable the supply-chain quarantine. It removed a permanent block
while leaving the quarantine in force wherever it can actually be evaluated, which was the intent of
scoping the rule to
matchDatasources: ["docker"].A correction, to this issue and to my own earlier comment
I argued above that the dashboard was not stale because #15's
updated_athad moved. That reasoningwas wrong, in the same way the original issue's was. Any cross-reference bumps an issue's
updated_at— the 05:08:07Z bump I cited was my own previous comment referencing "#15", not aRenovate run. Neither "it has not been updated since the 14th" nor "it was updated today" tells you
anything about when Renovate last rewrote it.
The trustworthy signal is
bitborg_renovate_last_run_timestamp_seconds, which showed the realsequence: last run 00:32:23Z, merge 09:04:03Z, triggered run 09:19:05Z. Worth using that rather than
updated_atnext time.Remaining on this issue
Only the third acceptance criterion: a check that catches this class. Two notes for it, one of which
is a trap.
renovate-config-validatorcannot catch it. It validates option names but not enumvalues —
minimumReleaseAgeBehaviour: "totally-bogus-value"passes withConfig validated successfully, while a typo'd option name is rejected. Verified both directions.The observable failure is "an entry sits in a holding section longer than that section can explain".
A check that ages the dashboard's entries and fails past a threshold would catch this and anything
else that stalls the same way. It has to read the run timestamp rather than the issue's
updated_at,per the correction above.
The check is in production and verified
#437 and #438 are applied. All three acceptance criteria are now met.
Applies, each through
scripts/apply-reconcile.sh, reconciled by task name:bitborg-prod--tags renovatebitborg-monitoring--tags monitoringbitborg-prod--tags renovateA second
--checkreturnedchanged=0on both hosts after each, which is the only thing that provesthe hosts match
main.The rules are not merely rendered, they are loaded and evaluating:
The exporter was run by hand rather than waiting for the nightly, and its metrics are queryable in
VictoriaMetrics with the labels the alerts select on: 30 held updates across all 8 repositories
Renovate autodiscovers. Only
Watchdogis firing.The guard found its own blind spot on the first run
Worth recording, because it is the argument for the cheapest part of the design. That first manual
run reported:
unknown sections 1. A dashboard carried a heading the parser did not know —PR Closed (Blocked),which Renovate renders with a
recreate-branchaction for updates whose PR a human closed to rejectthem. Renovate never recreates those unaided, so they are terminal, not waiting; ageing them would
have alerted for ever about a decision already made. #438 maps it as not-held. The same signal
confirms the fix: the count is now
0after a re-run and a scrape.That gap would not have been caught by any of the tests, because the tests assert my model of the
dashboard and the section was outside it. The counter questioned the model instead. It is four lines
of code and one alert.
One expectation to keep calibrated
The clocks start at first observation, so this alert will not retroactively flag the current
stall. State is fail-open by design: losing or lacking it resets the ages. Six of the seven held
updates become PRs on Monday 2026-08-24 when the weekly schedule window opens, so nothing is missed
in practice, but do not wait for an alert about the backlog that prompted this issue.
Left open deliberately
Keeping this issue open until Monday, when the six updates should actually open as PRs. The mechanism
is proven — they moved from "Pending Status Checks" to "Awaiting Schedule" and the release-age check
now passes — but a Forgejo security update reaching a reviewable PR is the outcome that matters, and
that has not happened yet. Closing before observing it would be trusting a green check that has not
been seen to fire.
Follow-up not folded into either PR: nothing in this repo validates alert rules.
promtool check ruleswas run by hand against the rendered template (49 rules, and a deliberately unbalanced parenwas confirmed to fail it). Wiring it into CI needs Jinja rendering with resolved role vars, so it
deserves its own issue.
The fix is confirmed working. Do not close this yet, but the mechanism is no longer in doubt.
Checked 2026-08-19, after last night's run (Renovate healthy:
bitborg_renovate_last_run_status = 0,last run 6.1 h ago).
The dashboard now carries the WARN the fix predicts:
And
forgejo 16.0.2has moved out of "Pending Status Checks" into "Awaiting Schedule". Thattransition is the evidence: the update is no longer deferred by an internal check that could never
pass, it is simply waiting for the Monday window.
Two corrections to what was expected:
vmagent / vmalert 1.150.0, alertmanager 0.34.0, pnpm 11.22.0, and lock-file maintenance.
rather than a new fault: it is the one image whose
releaseTimestampdid arrive, so it is under thegenuine three-day quarantine, which had not elapsed at 00:32 today. Expect it to move on tonight's
run. If it has not by 2026-08-20, that is a new problem rather than this one.
Also new, and working as designed: postgres 18 now sits under "Pending Approval", held by the
major gate.
Filed separately: #446
This issue's acceptance says the outcome that matters is "a Forgejo security update reaching a
reviewable PR". Watching 16.0.2 open on Monday proves the
timestamp-requiredbug here is fixed. Itdoes not prove security updates arrive promptly, because a Monday-scheduled update opens on Monday
either way. One observation cannot separate the two.
#446 records why:
vulnerabilityAlertssetsschedule: ["at any time"], but it only reaches updatesRenovate itself classifies, and there is no vulnerability feed for a docker image tag. So a git-server
security patch waits up to six days.
Closing this on Monday is correct. It should not be read as closing that.
Prepared ahead of Monday
Three of the eight are the VictoriaMetrics trio, which two role-defaults comments require to move in
lockstep and which nothing grouped. Ungrouped they arrive as three PRs. PR #445 groups them into one,
reproduced with
platform=localand a control run.Still to do by hand on Monday: the ADR 0038 concealment gate. The 16.0.2 bump must carry a
health_check_forgejo_concealment_verified_tagbump in the same change, after checking the surfaces.Never pre-bump the tag, and never commit a disabled gate.
Closing as done. The
timestamp-requiredstall was fixed by bitborg/bitborg-docs#91 (minimumReleaseAgeBehaviour: timestamp-optionalfor the docker datasource) and confirmed on 2026-08-19. Forgejo updates have since arrived as normal PRs, through 16.0.5.As the last comment says, this does not prove security updates arrive promptly. A docker-tag security patch still waits for the weekly schedule. That gap is tracked in #446 and #487.