feat(runner-controller): instrument CI pickup latency per stage #328

Stängd
öppnade 2026-08-02 12:30:46 +00:00 av supernaut · 0 kommentarer
Ägare

Problem

gitborg_runner_controller_* and the bitborg CI runners — health dashboard expose health
(active_vms, queued_jobs, last_loop_ok, boot_error_vms, orphan_volumes_swept,
os_volume_gb_*) but no latency at all. Nothing records how long it takes a triggered job to
start, and nothing records where that time goes. Every number in the capacity-formula issue had to
be reconstructed by hand from the Forgejo API cross-correlated against the log store.

Consequences:

  • A 60–200 s controller stall affecting 20 % of runs went unnoticed for weeks despite a dedicated
    dashboard and two alert rules.
  • Any boot-time optimisation is unverifiable — we cannot tell whether a change helped.

What is dark today

Stage Instrumented?
push → run created in Forgejo no
run created → controller notices no
create_server → VM ACTIVE no — wait=False, no ACTIVE ever logged
ACTIVE → cloud-init done → runner registered no (runner VM journals are not drained)
runner registered → job claimed no
end-to-end pickup latency no

Proposal

  1. In controller.py, track each ephemeral VM from its create_server call and log + export:
    • gitborg_runner_controller_boot_seconds — create_server → OpenStack ACTIVE
      (a cheap per-cycle status check on VMs the controller already lists, no extra API calls);
    • gitborg_runner_controller_ready_seconds — create_server → the runner appearing in
      GET /api/v1/admin/actions/runners (already read every cycle);
    • gitborg_runner_controller_pickup_seconds — queued-job first-seen → that VM claiming a job.
      Export as summary gauges (last / p50 / p90 over a rolling window) via the existing atomic
      textfile writer — no new scrape target, no new dependency.
  2. Add a "Pickup latency" row to roles/monitoring/files/bitborg-runners.json: boot seconds,
    ready seconds, end-to-end pickup, each with a p90 stat tile.
  3. Tighten RunnerQueueStalled — or add a companion — so a sub-minute systematic stall is
    visible. The current multi-minute for: window is correct for outages but blind to the
    latency class this issue exists to surface.

Acceptance

  • Boot and ready durations visible in Grafana for every ephemeral VM.
  • End-to-end pickup p50/p90 visible without touching the Forgejo API by hand.
  • Documented in docs/runbook.md alongside the existing runner section.

Part of gitborg/gitborg-docs#68.

## Problem `gitborg_runner_controller_*` and the `bitborg CI runners — health` dashboard expose **health** (`active_vms`, `queued_jobs`, `last_loop_ok`, `boot_error_vms`, `orphan_volumes_swept`, `os_volume_gb_*`) but **no latency at all**. Nothing records how long it takes a triggered job to start, and nothing records where that time goes. Every number in the capacity-formula issue had to be reconstructed by hand from the Forgejo API cross-correlated against the log store. Consequences: - A 60–200 s controller stall affecting 20 % of runs went unnoticed for weeks despite a dedicated dashboard and two alert rules. - Any boot-time optimisation is unverifiable — we cannot tell whether a change helped. ## What is dark today | Stage | Instrumented? | | --- | --- | | push → run created in Forgejo | no | | run created → controller notices | no | | `create_server` → VM ACTIVE | **no** — `wait=False`, no ACTIVE ever logged | | ACTIVE → cloud-init done → runner registered | no (runner VM journals are not drained) | | runner registered → job claimed | no | | end-to-end pickup latency | no | ## Proposal 1. In `controller.py`, track each ephemeral VM from its `create_server` call and log + export: - `gitborg_runner_controller_boot_seconds` — create_server → OpenStack `ACTIVE` (a cheap per-cycle status check on VMs the controller already lists, no extra API calls); - `gitborg_runner_controller_ready_seconds` — create_server → the runner appearing in `GET /api/v1/admin/actions/runners` (already read every cycle); - `gitborg_runner_controller_pickup_seconds` — queued-job first-seen → that VM claiming a job. Export as summary gauges (last / p50 / p90 over a rolling window) via the existing atomic textfile writer — no new scrape target, no new dependency. 2. Add a **"Pickup latency"** row to `roles/monitoring/files/bitborg-runners.json`: boot seconds, ready seconds, end-to-end pickup, each with a p90 stat tile. 3. Tighten `RunnerQueueStalled` — or add a companion — so a *sub-minute* systematic stall is visible. The current multi-minute `for:` window is correct for outages but blind to the latency class this issue exists to surface. ## Acceptance - [ ] Boot and ready durations visible in Grafana for every ephemeral VM. - [ ] End-to-end pickup p50/p90 visible without touching the Forgejo API by hand. - [ ] Documented in `docs/runbook.md` alongside the existing runner section. Part of gitborg/gitborg-docs#68.
Logga in för att delta i denna konversation.
Ingen milstolpe
Inga tilldelade
1 deltagare
Notiser
Förfallodatum
Förfallodatumet är ogiltigt eller utanför gränserna. Använd formatet "åååå-mm-dd".

Inget förfallodatum satt.

Beroenden

Inga beroenden satta

Referens
bitborg/bitborg-infra#328
Ingen beskrivning angiven.