Reliability of AI Bots Footprints in GitHub Actions CI/CD Workflows

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shah, Syed Muhammad Ashhar, Habib, Sehrish, Hussain, Muizz, Ghafoor, Maryam Abdul, Bangash, Abdul Ali
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911608266031104
author Shah, Syed Muhammad Ashhar
Habib, Sehrish
Hussain, Muizz
Ghafoor, Maryam Abdul
Bangash, Abdul Ali
author_facet Shah, Syed Muhammad Ashhar
Habib, Sehrish
Hussain, Muizz
Ghafoor, Maryam Abdul
Bangash, Abdul Ali
contents Continuous Integration and Deployment (CI/CD) workflows are central to modern software delivery, yet the reliability of agentic AI bots operating within these workflows remain underexplored. Using pull requests (PRs), commits, and repositories from the AIDev dataset, we retrieved associated CI/CD workflow runs via the GitHub Actions API and analyzed 61,837 runs from 2,355 repositories, all triggered by PRs generated by five AI bots: Claude, Devin, Cursor, Copilot, and Codex. We observed substantial agent-dependent differences in workflow reliability, with Copilot and Codex achieving the highest success rates ~93% and ~94% respectively. At the repository level, we find a negative correlation between AI agent contribution frequency and workflow success rate, suggesting that a higher frequency of Agentic PRs may hinder CI/CD workflow reliability. We defined a taxonomy of 13 categories against 3,067 agentic PRs whose associated workflows failed, and observed a trend analysis that indicates visually observable shifts from functional to non-functional PR categories over time, although these trends are not statistically significant. Our findings motivate the need for actionable guidance on integrating AI agents into CI/CD workflows and prioritizing safeguards in workflows where failures are most likely to occur.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18334
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Reliability of AI Bots Footprints in GitHub Actions CI/CD Workflows
Shah, Syed Muhammad Ashhar
Habib, Sehrish
Hussain, Muizz
Ghafoor, Maryam Abdul
Bangash, Abdul Ali
Software Engineering
D.2.7; H.2.8
Continuous Integration and Deployment (CI/CD) workflows are central to modern software delivery, yet the reliability of agentic AI bots operating within these workflows remain underexplored. Using pull requests (PRs), commits, and repositories from the AIDev dataset, we retrieved associated CI/CD workflow runs via the GitHub Actions API and analyzed 61,837 runs from 2,355 repositories, all triggered by PRs generated by five AI bots: Claude, Devin, Cursor, Copilot, and Codex. We observed substantial agent-dependent differences in workflow reliability, with Copilot and Codex achieving the highest success rates ~93% and ~94% respectively. At the repository level, we find a negative correlation between AI agent contribution frequency and workflow success rate, suggesting that a higher frequency of Agentic PRs may hinder CI/CD workflow reliability. We defined a taxonomy of 13 categories against 3,067 agentic PRs whose associated workflows failed, and observed a trend analysis that indicates visually observable shifts from functional to non-functional PR categories over time, although these trends are not statistically significant. Our findings motivate the need for actionable guidance on integrating AI agents into CI/CD workflows and prioritizing safeguards in workflows where failures are most likely to occur.
title Reliability of AI Bots Footprints in GitHub Actions CI/CD Workflows
topic Software Engineering
D.2.7; H.2.8
url https://arxiv.org/abs/2604.18334