Why Are Agentic Pull Requests Merged or Rejected? An Empirical Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peralta, Sien Reeve O., Hoshi, Fumika, Washizaki, Hironori, Ubayashi, Naoyasu, Kondo, Inase, Higo, Yoshiki, Mukai, Hiroki, Yoshida, Norihiro, Kusama, Kazuki, Tanaka, Hidetake, Fan, Youmei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911704740265984
author Peralta, Sien Reeve O.
Hoshi, Fumika
Washizaki, Hironori
Ubayashi, Naoyasu
Kondo, Inase
Higo, Yoshiki
Mukai, Hiroki
Yoshida, Norihiro
Kusama, Kazuki
Tanaka, Hidetake
Fan, Youmei
author_facet Peralta, Sien Reeve O.
Hoshi, Fumika
Washizaki, Hironori
Ubayashi, Naoyasu
Kondo, Inase
Higo, Yoshiki
Mukai, Hiroki
Yoshida, Norihiro
Kusama, Kazuki
Tanaka, Hidetake
Fan, Youmei
contents AI coding agents increasingly submit pull requests (Agentic-PRs) to open-source repositories, yet their performance is commonly assessed using merge and rejection outcomes alone. We hypothesized that these outcome labels do not reliably reflect agent capability without considering review interactions. To test this, we conducted a decision-oriented analysis of 11,048 closed Agentic Pull Requests, refined to 9,799 human-reviewed PRs, and manually inspected 717 representative cases to recover decision rationale from interaction artifacts. We found that rejection outcomes substantially overstate agent error: only 35.7% of rejected PRs reflected clear agentic failures, while 31.2% were driven by workflow constraints and 33.1% lacked observable decision rationale. Among merged PRs, 15.4% required explicit reviewer involvement through feedback or direct commits, and 5.5% showed no visible interaction trace. We further observed systematic differences across agents, with Copilot and Devin more often embedded in reviewer-mediated workflows, while Codex and Cursor PRs were typically merged with minimal interaction. These results reject the assumption that PR outcomes alone capture agent performance and demonstrate the need for interaction-aware evaluation grounded in review behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22534
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Why Are Agentic Pull Requests Merged or Rejected? An Empirical Study
Peralta, Sien Reeve O.
Hoshi, Fumika
Washizaki, Hironori
Ubayashi, Naoyasu
Kondo, Inase
Higo, Yoshiki
Mukai, Hiroki
Yoshida, Norihiro
Kusama, Kazuki
Tanaka, Hidetake
Fan, Youmei
Software Engineering
AI coding agents increasingly submit pull requests (Agentic-PRs) to open-source repositories, yet their performance is commonly assessed using merge and rejection outcomes alone. We hypothesized that these outcome labels do not reliably reflect agent capability without considering review interactions. To test this, we conducted a decision-oriented analysis of 11,048 closed Agentic Pull Requests, refined to 9,799 human-reviewed PRs, and manually inspected 717 representative cases to recover decision rationale from interaction artifacts. We found that rejection outcomes substantially overstate agent error: only 35.7% of rejected PRs reflected clear agentic failures, while 31.2% were driven by workflow constraints and 33.1% lacked observable decision rationale. Among merged PRs, 15.4% required explicit reviewer involvement through feedback or direct commits, and 5.5% showed no visible interaction trace. We further observed systematic differences across agents, with Copilot and Devin more often embedded in reviewer-mediated workflows, while Codex and Cursor PRs were typically merged with minimal interaction. These results reject the assumption that PR outcomes alone capture agent performance and demonstrate the need for interaction-aware evaluation grounded in review behavior.
title Why Are Agentic Pull Requests Merged or Rejected? An Empirical Study
topic Software Engineering
url https://arxiv.org/abs/2605.22534