feat(runner-controller): cancel orphaned CI runs on VM reap + plan log-export/artifact-cleanup (#78) #176
Inga granskare
Etiketter
Inga etiketter
area/backups
area/ci
area/control-panel
area/identity
area/infra
area/observability
area/payments
area/security
area/storage
area/web
blocked
needs-info
needs-triage
ready-for-implementation
type
bug
type
chore
type
docs
type
epic
type
feature
type
task
wontfix
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Inget förfallodatum satt.
Beroenden
Inga beroenden satta
Referens
bitborg/bitborg-infra!176
Läser in…
Hänvisa till i nytt ärende
Ingen beskrivning angiven.
Ta bort grenen "feat/78-actions-lifecycle"
Borttagning av en gren är permanent. Även om den borttagna grenen kan fortsätta existera en kort tid innan den faktiskt tas bort, kan det INTE ångras i de flesta fall. Vill du fortsätta?
What
Forgejo v16 is live, so wire the new Actions lifecycle APIs into the runner-controller. This PR
implements the runner-controller cancel path (the self-contained, highest-value piece; ties
into the #170 zombie-VM reap) and plans the other two candidate uses rather than building them.
Refs #78 (multi-part). Refs ADR 0021.
Implemented — runner-controller orphaned-run cancellation (checkbox done)
When the controller reaps a VM (SHUTOFF/ERROR/too-old, or a #170 zombie) that was mid-job, the
job can hang in
runninguntil Forgejo'sjob_timeout(default 1h) — wasted queue time and aconfusing "Waiting/Running" UI. The controller now cancels such a stuck run.
What it does, per reconcile cycle (only when enabled):
have a live runner (
_runner_label_set).runningstate for the pool label (GET /api/v1/admin/runners/jobs)._select_runs_to_cancel(pure, unit-tested) marks a run orphaned iff it has >=1 running joband none of its running jobs is coverable by a live runner label.
(~10-20 s), so a momentary snapshot skew (a runner registering as a job dispatches) never cancels.
repo_id -> owner/repo(GET /api/v1/repositories/{id}) and issuesPOST /api/v1/repos/{owner}/{repo}/actions/runs/{run_id}/cancel.Safety / fail-safe (the design contract):
runner_controller_cancel_orphaned_jobs: false. Enable via group_vars onceverified live.
runner_controller_dry_run(dry-run only logs).never aborts the reconcile cycle or triggers mass cancellation (any read failure -> judge
nothing that cycle).
gitborg_runner_controller_orphaned_runs_cancelled.VM -> job mapping: the limitation (and why the subset is what it is)
Endpoints verified 2026-07-20 against the LIVE instance swagger
(
git.gitborg.se/swagger.v1.json, Forgejo v16) + Forgejo source:POST /api/v1/repos/{owner}/{repo}/actions/runs/{run_id}/cancel-> 204 ("pending or runningjobs of the run are cancelled"; an already-finished run is left unchanged, still 204).
GET /api/v1/admin/runners/jobs?labels=->[ActionRunJob]; source(
routers/api/v1/shared/runners.goGetActionRunJobs) confirms the query isFindTaskOptions{Status: [StatusWaiting, StatusRunning]}-> running jobs are returned.GET /api/v1/repositories/{id}->Repository{full_name}.But
ActionRunJobcarriesrun_id,repo_id,runs_on,status— and NOrunner_id. AndActionRunnerexposes no current task/job. So a specific reaped VM cannot be mapped to aspecific run with the available APIs. The only unambiguous, zero-false-positive signal is
label coverage: a
runningjob whose required label has zero live runners cannot beprogressing -> its run is safely cancellable.
Consequence (documented): if the label still has any live runner (e.g. two concurrent VMs,
one reaped mid-job while the other runs a different job), the stuck job is not cancelled and
falls back to Forgejo's
job_timeout. Given this pool (min_idle: 0, low concurrency), the commonreal case — the single runner reaped mid-job -> label has no live runner -> cancel — is handled. If
a future Forgejo version adds
runner_idonActionRunJob(or a task->runner lookup), this can betightened to exact per-VM targeting.
Config knob
runner_controller_cancel_orphaned_jobsfalseRendered into
config.yamlasforgejo.cancel_orphaned_jobs.Verification
python3 -m py_compile .../controller.py— OKpython3 .../test_controller_logic.py-> 9/9 (orphaned+confirmed->cancel; first-sightingdefer; label-covered->no-op; runners-read-None->no-op; mixed run with a coverable job->no-op;
empty
runs_on->no-op; no running jobs; two-orphaned-one-confirmed; live-runner-wrong-label->orphaned)
pnpm ansible:check(syntax-check) — OKansible-lint roles/runner-controller-> Passed, profileproductioncontainer; the feature stays off until
runner_controller_cancel_orphaned_jobs: trueis set.Planned only (NOT built in this PR)
1. CI job-log export -> Loki (ADR 0020) — mitigates
LOG_RETENTION_DAYS=30Goal: persist CI job logs beyond Forgejo's 30-day
LOG_RETENTION_DAYSby shipping finished-runlogs to the monitoring VM's Loki, so post-incident CI forensics survive log expiry.
Sketch:
recently-completed runs via
GET /api/v1/repos/{owner}/{repo}/actions/runs?status=...(+
.../runs/{run_id}/jobs), downloads each job log with the v16 job-log API(
download-job-log, Forgejo PR 12666 — verify the exact path/operationId against live swaggerbefore building) and pushes to Loki's HTTP push API (
/loki/api/v1/push).{job="forgejo_ci", repo, workflow, run_id, job, conclusion}; the log line body is theraw step output. Keep cardinality low (no per-line labels).
run_id(or per-runupdatedtimestamp) in asmall state file (mirror the
_orphan_cancel_seen/_sweep_stuckpatterns) so re-runs do notdouble-ship. Only export runs in a terminal status (success/failure/cancelled).
the CI-log retention knob — decouple it from Forgejo's
LOG_RETENTION_DAYS, which can then staylow to bound Forgejo disk.
monitoring-agent; Loki push over the private network. Best-effort: an export failure must never
affect CI.
the download-job-log endpoint verified on v16, plus care on log volume/cost.
2. Artifact retention / cleanup — beyond the blanket
ARTIFACT_RETENTION_DAYS=30Goal: keep artifacts that matter (release/tag builds) while purging cheap, high-churn PR-build
artifacts early, instead of one flat 30-day TTL for everything.
Sketch:
(
GET /api/v1/repos/{owner}/{repo}/actions/runs/{run_id}/artifacts— verify the delete-artifactendpoint/operationId against live v16 swagger; confirm whether deletion is per-artifact or via a
retention setter) and apply a policy:
refs/tags/*or the release workflow -> keep (long TTL / never auto-purge).merged/closed.
ARTIFACT_RETENTION_DAYSas the coarse backstop; the timer does the policy-aware earlypurge on top.
run, and match releases conservatively (prefer keep-on-doubt).
unknown is the exact v16 artifact-delete API surface.
Both plan items should be filed as their own follow-up issues under #78's epic and reference the
ADRs above.
Issue #78 checkboxes
90ea02f77b86668b287d