feat(runner-controller): cancel orphaned CI runs on VM reap + plan log-export/artifact-cleanup (#78) #176

Sammanfogat
supernaut sammanfogade 3 incheckningar från feat/78-actions-lifecycle in i main 2026-07-20 20:55:43 +00:00
Ägare

What

Forgejo v16 is live, so wire the new Actions lifecycle APIs into the runner-controller. This PR
implements the runner-controller cancel path (the self-contained, highest-value piece; ties
into the #170 zombie-VM reap) and plans the other two candidate uses rather than building them.

Refs #78 (multi-part). Refs ADR 0021.

Implemented — runner-controller orphaned-run cancellation (checkbox done)

When the controller reaps a VM (SHUTOFF/ERROR/too-old, or a #170 zombie) that was mid-job, the
job can hang in running until Forgejo's job_timeout (default 1h) — wasted queue time and a
confusing "Waiting/Running" UI. The controller now cancels such a stuck run.

What it does, per reconcile cycle (only when enabled):

  1. Reads live runners once (shared with the #170 zombie check) → derives the set of labels that
    have a live runner (_runner_label_set).
  2. Lists jobs in running state for the pool label (GET /api/v1/admin/runners/jobs).
  3. _select_runs_to_cancel (pure, unit-tested) marks a run orphaned iff it has >=1 running job
    and none of its running jobs is coverable by a live runner label.
  4. Two-cycle confirmation: a run is cancelled only if it looked orphaned last cycle too
    (~10-20 s), so a momentary snapshot skew (a runner registering as a job dispatches) never cancels.
  5. Resolves repo_id -> owner/repo (GET /api/v1/repositories/{id}) and issues
    POST /api/v1/repos/{owner}/{repo}/actions/runs/{run_id}/cancel.

Safety / fail-safe (the design contract):

  • Off by default — runner_controller_cancel_orphaned_jobs: false. Enable via group_vars once
    verified live.
  • Honours runner_controller_dry_run (dry-run only logs).
  • Best-effort end-to-end: every read/POST is wrapped; a Forgejo error skips the cancel step and
    never aborts the reconcile cycle or triggers mass cancellation (any read failure -> judge
    nothing that cycle).
  • New metric gitborg_runner_controller_orphaned_runs_cancelled.

VM -> job mapping: the limitation (and why the subset is what it is)

Endpoints verified 2026-07-20 against the LIVE instance swagger
(git.gitborg.se/swagger.v1.json, Forgejo v16) + Forgejo source:

  • POST /api/v1/repos/{owner}/{repo}/actions/runs/{run_id}/cancel -> 204 ("pending or running
    jobs of the run are cancelled"; an already-finished run is left unchanged, still 204).
  • GET /api/v1/admin/runners/jobs?labels= -> [ActionRunJob]; source
    (routers/api/v1/shared/runners.go GetActionRunJobs) confirms the query is
    FindTaskOptions{Status: [StatusWaiting, StatusRunning]} -> running jobs are returned.
  • GET /api/v1/repositories/{id} -> Repository{full_name}.

But ActionRunJob carries run_id, repo_id, runs_on, status — and NO runner_id. And
ActionRunner exposes no current task/job.
So a specific reaped VM cannot be mapped to a
specific run
with the available APIs. The only unambiguous, zero-false-positive signal is
label coverage: a running job whose required label has zero live runners cannot be
progressing -> its run is safely cancellable.

Consequence (documented): if the label still has any live runner (e.g. two concurrent VMs,
one reaped mid-job while the other runs a different job), the stuck job is not cancelled and
falls back to Forgejo's job_timeout. Given this pool (min_idle: 0, low concurrency), the common
real case — the single runner reaped mid-job -> label has no live runner -> cancel — is handled. If
a future Forgejo version adds runner_id on ActionRunJob (or a task->runner lookup), this can be
tightened to exact per-VM targeting.

Config knob

Var Default Meaning
runner_controller_cancel_orphaned_jobs false Master switch for the cancel path (off-by-default-safe).

Rendered into config.yaml as forgejo.cancel_orphaned_jobs.

Verification

  • python3 -m py_compile .../controller.py — OK
  • python3 .../test_controller_logic.py -> 9/9 (orphaned+confirmed->cancel; first-sighting
    defer; label-covered->no-op; runners-read-None->no-op; mixed run with a coverable job->no-op;
    empty runs_on->no-op; no running jobs; two-orphaned-one-confirmed; live-runner-wrong-label->
    orphaned)
  • pnpm ansible:check (syntax-check) — OK
  • ansible-lint roles/runner-controller -> Passed, profile production
  • Not applied to prod (per the task). Deploying rebuilds the content-hash image + restarts the
    container; the feature stays off until runner_controller_cancel_orphaned_jobs: true is set.

Planned only (NOT built in this PR)

1. CI job-log export -> Loki (ADR 0020) — mitigates LOG_RETENTION_DAYS=30

Goal: persist CI job logs beyond Forgejo's 30-day LOG_RETENTION_DAYS by shipping finished-run
logs to the monitoring VM's Loki, so post-incident CI forensics survive log expiry.

Sketch:

  • A small exporter (new systemd timer on the services host, or a mode in the controller) walks
    recently-completed runs via GET /api/v1/repos/{owner}/{repo}/actions/runs?status=...
    (+ .../runs/{run_id}/jobs), downloads each job log with the v16 job-log API
    (download-job-log, Forgejo PR 12666 — verify the exact path/operationId against live swagger
    before building) and pushes to Loki's HTTP push API (/loki/api/v1/push).
  • Labels: {job="forgejo_ci", repo, workflow, run_id, job, conclusion}; the log line body is the
    raw step output. Keep cardinality low (no per-line labels).
  • Idempotency/state: track the last-exported run_id (or per-run updated timestamp) in a
    small state file (mirror the _orphan_cancel_seen / _sweep_stuck patterns) so re-runs do not
    double-ship. Only export runs in a terminal status (success/failure/cancelled).
  • Retention: Loki's own retention (already provisioned on the monitoring VM, ADR 0020) becomes
    the CI-log retention knob — decouple it from Forgejo's LOG_RETENTION_DAYS, which can then stay
    low to bound Forgejo disk.
  • Auth/network: reuse the existing services-host -> monitoring-VM path used by the
    monitoring-agent; Loki push over the private network. Best-effort: an export failure must never
    affect CI.
  • Risk/size: medium — a new component with its own state + a Loki datasource contract; needs
    the download-job-log endpoint verified on v16, plus care on log volume/cost.

2. Artifact retention / cleanup — beyond the blanket ARTIFACT_RETENTION_DAYS=30

Goal: keep artifacts that matter (release/tag builds) while purging cheap, high-churn PR-build
artifacts early, instead of one flat 30-day TTL for everything.

Sketch:

  • Programmatic cleanup timer: enumerate runs + their artifacts
    (GET /api/v1/repos/{owner}/{repo}/actions/runs/{run_id}/artifacts — verify the delete-artifact
    endpoint/operationId against live v16 swagger; confirm whether deletion is per-artifact or via a
    retention setter) and apply a policy:
    • Runs on refs/tags/* or the release workflow -> keep (long TTL / never auto-purge).
    • PR / branch-CI runs -> purge artifacts after a short window (e.g. 3-7 days) once the PR is
      merged/closed.
  • Keep ARTIFACT_RETENTION_DAYS as the coarse backstop; the timer does the policy-aware early
    purge on top.
  • Safety: dry-run-first (log what it would delete), never delete artifacts of an in-progress
    run, and match releases conservatively (prefer keep-on-doubt).
  • Risk/size: small-to-medium — mostly an API-driven cleanup script + a policy table; main
    unknown is the exact v16 artifact-delete API surface.

Both plan items should be filed as their own follow-up issues under #78's epic and reference the
ADRs above.

Issue #78 checkboxes

  • runner-controller cancel path (implemented, off-by-default, verified)
  • CI job-log export -> Loki (planned above; not built)
  • Artifact retention/cleanup (planned above; not built)
## What Forgejo v16 is live, so wire the new Actions lifecycle APIs into the runner-controller. This PR **implements the runner-controller cancel path** (the self-contained, highest-value piece; ties into the #170 zombie-VM reap) and **plans** the other two candidate uses rather than building them. Refs #78 (multi-part). Refs ADR 0021. ### Implemented — runner-controller orphaned-run cancellation (checkbox done) When the controller reaps a VM (SHUTOFF/ERROR/too-old, or a #170 zombie) that was **mid-job**, the job can hang in `running` until Forgejo's `job_timeout` (default 1h) — wasted queue time and a confusing "Waiting/Running" UI. The controller now cancels such a stuck run. **What it does, per reconcile cycle (only when enabled):** 1. Reads live runners once (shared with the #170 zombie check) → derives the set of labels that have a live runner (`_runner_label_set`). 2. Lists jobs in `running` state for the pool label (`GET /api/v1/admin/runners/jobs`). 3. `_select_runs_to_cancel` (pure, unit-tested) marks a **run** orphaned iff it has >=1 running job and **none** of its running jobs is coverable by a live runner label. 4. **Two-cycle confirmation**: a run is cancelled only if it looked orphaned last cycle too (~10-20 s), so a momentary snapshot skew (a runner registering as a job dispatches) never cancels. 5. Resolves `repo_id -> owner/repo` (`GET /api/v1/repositories/{id}`) and issues `POST /api/v1/repos/{owner}/{repo}/actions/runs/{run_id}/cancel`. **Safety / fail-safe (the design contract):** - **Off by default** — `runner_controller_cancel_orphaned_jobs: false`. Enable via group_vars once verified live. - Honours `runner_controller_dry_run` (dry-run only logs). - Best-effort end-to-end: every read/POST is wrapped; a Forgejo error **skips** the cancel step and **never** aborts the reconcile cycle or triggers mass cancellation (any read failure -> judge nothing that cycle). - New metric `gitborg_runner_controller_orphaned_runs_cancelled`. ### VM -> job mapping: the limitation (and why the subset is what it is) Endpoints verified **2026-07-20 against the LIVE instance swagger** (`git.gitborg.se/swagger.v1.json`, Forgejo v16) + Forgejo source: - `POST /api/v1/repos/{owner}/{repo}/actions/runs/{run_id}/cancel` -> **204** ("pending or running jobs of the run are cancelled"; an already-finished run is left unchanged, still 204). - `GET /api/v1/admin/runners/jobs?labels=` -> `[ActionRunJob]`; source (`routers/api/v1/shared/runners.go` `GetActionRunJobs`) confirms the query is `FindTaskOptions{Status: [StatusWaiting, StatusRunning]}` -> **running jobs are returned**. - `GET /api/v1/repositories/{id}` -> `Repository{full_name}`. **But `ActionRunJob` carries `run_id`, `repo_id`, `runs_on`, `status` — and NO `runner_id`. And `ActionRunner` exposes no current task/job.** So a **specific reaped VM cannot be mapped to a specific run** with the available APIs. The only unambiguous, zero-false-positive signal is label coverage: a `running` job whose required label has **zero** live runners cannot be progressing -> its run is safely cancellable. **Consequence (documented):** if the label still has *any* live runner (e.g. two concurrent VMs, one reaped mid-job while the other runs a different job), the stuck job is **not** cancelled and falls back to Forgejo's `job_timeout`. Given this pool (`min_idle: 0`, low concurrency), the common real case — the single runner reaped mid-job -> label has no live runner -> cancel — is handled. If a future Forgejo version adds `runner_id` on `ActionRunJob` (or a task->runner lookup), this can be tightened to exact per-VM targeting. ### Config knob | Var | Default | Meaning | | --- | --- | --- | | `runner_controller_cancel_orphaned_jobs` | `false` | Master switch for the cancel path (off-by-default-safe). | Rendered into `config.yaml` as `forgejo.cancel_orphaned_jobs`. ### Verification - `python3 -m py_compile .../controller.py` — OK - `python3 .../test_controller_logic.py` -> **9/9** (orphaned+confirmed->cancel; first-sighting defer; label-covered->no-op; runners-read-None->no-op; mixed run with a coverable job->no-op; empty `runs_on`->no-op; no running jobs; two-orphaned-one-confirmed; live-runner-wrong-label-> orphaned) - `pnpm ansible:check` (syntax-check) — OK - `ansible-lint roles/runner-controller` -> Passed, profile `production` - **Not applied to prod** (per the task). Deploying rebuilds the content-hash image + restarts the container; the feature stays off until `runner_controller_cancel_orphaned_jobs: true` is set. --- ## Planned only (NOT built in this PR) ### 1. CI job-log export -> Loki (ADR 0020) — mitigates `LOG_RETENTION_DAYS=30` **Goal:** persist CI job logs beyond Forgejo's 30-day `LOG_RETENTION_DAYS` by shipping finished-run logs to the monitoring VM's Loki, so post-incident CI forensics survive log expiry. **Sketch:** - A small exporter (new systemd timer on the services host, or a mode in the controller) walks recently-**completed** runs via `GET /api/v1/repos/{owner}/{repo}/actions/runs?status=...` (+ `.../runs/{run_id}/jobs`), downloads each job log with the v16 job-log API (`download-job-log`, Forgejo PR 12666 — verify the exact path/operationId against live swagger before building) and pushes to Loki's HTTP push API (`/loki/api/v1/push`). - Labels: `{job="forgejo_ci", repo, workflow, run_id, job, conclusion}`; the log line body is the raw step output. Keep cardinality low (no per-line labels). - **Idempotency/state:** track the last-exported `run_id` (or per-run `updated` timestamp) in a small state file (mirror the `_orphan_cancel_seen` / `_sweep_stuck` patterns) so re-runs do not double-ship. Only export runs in a terminal status (success/failure/cancelled). - **Retention:** Loki's own retention (already provisioned on the monitoring VM, ADR 0020) becomes the CI-log retention knob — decouple it from Forgejo's `LOG_RETENTION_DAYS`, which can then stay low to bound Forgejo disk. - **Auth/network:** reuse the existing services-host -> monitoring-VM path used by the monitoring-agent; Loki push over the private network. Best-effort: an export failure must never affect CI. - **Risk/size:** medium — a new component with its own state + a Loki datasource contract; needs the download-job-log endpoint verified on v16, plus care on log volume/cost. ### 2. Artifact retention / cleanup — beyond the blanket `ARTIFACT_RETENTION_DAYS=30` **Goal:** keep artifacts that matter (release/tag builds) while purging cheap, high-churn PR-build artifacts early, instead of one flat 30-day TTL for everything. **Sketch:** - Programmatic cleanup timer: enumerate runs + their artifacts (`GET /api/v1/repos/{owner}/{repo}/actions/runs/{run_id}/artifacts` — verify the delete-artifact endpoint/operationId against live v16 swagger; confirm whether deletion is per-artifact or via a retention setter) and apply a **policy**: - Runs on `refs/tags/*` or the release workflow -> **keep** (long TTL / never auto-purge). - PR / branch-CI runs -> purge artifacts after a short window (e.g. 3-7 days) once the PR is merged/closed. - Keep `ARTIFACT_RETENTION_DAYS` as the coarse backstop; the timer does the policy-aware early purge on top. - **Safety:** dry-run-first (log what it *would* delete), never delete artifacts of an in-progress run, and match releases conservatively (prefer keep-on-doubt). - **Risk/size:** small-to-medium — mostly an API-driven cleanup script + a policy table; main unknown is the exact v16 artifact-delete API surface. Both plan items should be filed as their own follow-up issues under #78's epic and reference the ADRs above. ## Issue #78 checkboxes - [x] runner-controller cancel path (implemented, off-by-default, verified) - [ ] CI job-log export -> Loki (planned above; not built) - [ ] Artifact retention/cleanup (planned above; not built)
supernaut tvångsskickade feat/78-actions-lifecycle från 90ea02f77b
Alla kontroller lyckades
ci / ci (pull_request) Successful in 2m48s
till 86668b287d
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m24s
2026-07-20 18:08:49 +00:00
Jämför
supernaut lade till 2 incheckningar 2026-07-20 20:49:13 +00:00
- group_vars: runner_controller_cancel_orphaned_jobs=true (turn the #176 cancel
  path on in prod — reclaims a reaped VM's run instead of hanging to timeout).
- controller.py: emit gitborg_runner_controller_queued_jobs gauge (waiting jobs
  for the label; -1=unknown) so an intermittent runner-start/job-assignment
  stall is observable.
- monitoring: RunnerQueueStalled alert (queued_jobs > 0 for
  alert_runner_queue_stalled_minutes=15) — turns the hard-to-reproduce stall
  into an evidence-backed page.

py_compile + logic test 9/9 + ansible-lint + rendered-alert YAML all clean.
Merge remote-tracking branch 'origin/main' into feat/78-actions-lifecycle
Alla kontroller lyckades
ci / ci (pull_request) Successful in 1m30s
fef205139e
supernaut sammanfogade incheckning 03d86d744b till main 2026-07-20 20:55:43 +00:00
Logga in för att delta i denna konversation.
Inga granskare
Ingen milstolpe
Inget projekt
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra!176
Ingen beskrivning angiven.