SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback
Fuente:
arXiv
Saved in:
| Main Author: | Kumar, Deepak |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analyzing Message-Code Inconsistency in AI Coding Agent-Authored Pull Requests
by: Gong, Jingzhi, et al.
Published: (2026)
by: Gong, Jingzhi, et al.
Published: (2026)
Quality and Security Signals in AI-Generated Python Refactoring Pull Requests
by: Almukhtar, Mohamed, et al.
Published: (2026)
by: Almukhtar, Mohamed, et al.
Published: (2026)
More Code, Less Reuse: Investigating Code Quality and Reviewer Sentiment towards AI-generated Pull Requests
by: Huang, Haoming, et al.
Published: (2026)
by: Huang, Haoming, et al.
Published: (2026)
How AI Coding Agents Communicate: A Study of Pull Request Description Characteristics and Human Review Responses
by: Watanabe, Kan, et al.
Published: (2026)
by: Watanabe, Kan, et al.
Published: (2026)
When AI Teammates Meet Code Review: Collaboration Signals Shaping the Integration of Agent-Authored Pull Requests
by: Nachuma, Costain, et al.
Published: (2026)
by: Nachuma, Costain, et al.
Published: (2026)
Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
by: Ehsani, Ramtin, et al.
Published: (2026)
by: Ehsani, Ramtin, et al.
Published: (2026)
On Unified Prompt Tuning for Request Quality Assurance in Public Code Review
by: Chen, Xinyu, et al.
Published: (2024)
by: Chen, Xinyu, et al.
Published: (2024)
SWE Context Bench: A Benchmark for Context Learning in Coding
by: Zhu, Jiayuan, et al.
Published: (2026)
by: Zhu, Jiayuan, et al.
Published: (2026)
SWE-QA: A Dataset and Benchmark for Complex Code Understanding
by: Elkoussy, Laïla, et al.
Published: (2026)
by: Elkoussy, Laïla, et al.
Published: (2026)
Knowledge-Guided Prompt Learning for Request Quality Assurance in Public Code Review
by: Li, Lin, et al.
Published: (2024)
by: Li, Lin, et al.
Published: (2024)
Pull Requests as a Training Signal for Repo-Level Code Editing
by: Zhu, Qinglin, et al.
Published: (2026)
by: Zhu, Qinglin, et al.
Published: (2026)
Why Are AI Agent Involved Pull Requests (Fix-Related) Remain Unmerged? An Empirical Study
by: Alam, Khairul, et al.
Published: (2026)
by: Alam, Khairul, et al.
Published: (2026)
MELT: Mining Effective Lightweight Transformations from Pull Requests
by: Ramos, Daniel, et al.
Published: (2023)
by: Ramos, Daniel, et al.
Published: (2023)
Spec-Driven Development:From Code to Contract in the Age of AI Coding Assistants
by: Piskala, Deepak Babu
Published: (2026)
by: Piskala, Deepak Babu
Published: (2026)
On Wasted Contributions: Understanding the Dynamics of Contributor-Abandoned Pull Requests
by: Khatoonabadi, SayedHassan, et al.
Published: (2021)
by: Khatoonabadi, SayedHassan, et al.
Published: (2021)
Predicting the First Response Latency of Maintainers and Contributors in Pull Requests
by: Khatoonabadi, SayedHassan, et al.
Published: (2023)
by: Khatoonabadi, SayedHassan, et al.
Published: (2023)
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
by: Lam, Man Ho, et al.
Published: (2026)
by: Lam, Man Ho, et al.
Published: (2026)
On The Impact of Merge Request Deviations on Code Review Practices
by: Kansab, Samah, et al.
Published: (2025)
by: Kansab, Samah, et al.
Published: (2025)
How AI Coding Agents Modify Code: A Large-Scale Study of GitHub Pull Requests
by: Ogenrwot, Daniel, et al.
Published: (2026)
by: Ogenrwot, Daniel, et al.
Published: (2026)
Quality Gatekeepers: Investigating the Effects ofCode Review Bots on Pull Request Activities
by: Wessel, Mairieli, et al.
Published: (2021)
by: Wessel, Mairieli, et al.
Published: (2021)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
by: Garg, Spandan, et al.
Published: (2025)
by: Garg, Spandan, et al.
Published: (2025)
SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
by: Xu, Jingxuan, et al.
Published: (2025)
by: Xu, Jingxuan, et al.
Published: (2025)
Code Review Agent Benchmark
by: Zhang, Yuntong, et al.
Published: (2026)
by: Zhang, Yuntong, et al.
Published: (2026)
AI-Assisted Code Review as a Scaffold for Code Quality and Self-Regulated Learning: An Experience Report
by: Oliveira, Eduardo, et al.
Published: (2026)
by: Oliveira, Eduardo, et al.
Published: (2026)
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
by: Guo, Lianghong, et al.
Published: (2025)
by: Guo, Lianghong, et al.
Published: (2025)
AutoFeedback: An LLM-based Framework for Efficient and Accurate API Request Generation
by: Liu, Huanxi, et al.
Published: (2024)
by: Liu, Huanxi, et al.
Published: (2024)
Ambiguity Resolution with Human Feedback for Code Writing Tasks
by: Nandan, Aditey, et al.
Published: (2025)
by: Nandan, Aditey, et al.
Published: (2025)
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
by: Cai, Songcheng, et al.
Published: (2026)
by: Cai, Songcheng, et al.
Published: (2026)
Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Unified Approach for Elevating Benchmark Quality
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
SWE-Bench-CL: Continual Learning for Coding Agents
by: Joshi, Thomas, et al.
Published: (2025)
by: Joshi, Thomas, et al.
Published: (2025)
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
by: Fan, Zhiyu, et al.
Published: (2025)
by: Fan, Zhiyu, et al.
Published: (2025)
Resolving Java Code Repository Issues with iSWE Agent
by: Ganhotra, Jatin, et al.
Published: (2026)
by: Ganhotra, Jatin, et al.
Published: (2026)
How Do Developers Use Code Suggestions in Pull Request Reviews?
by: Bouraffa, Abir, et al.
Published: (2025)
by: Bouraffa, Abir, et al.
Published: (2025)
Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
by: Zan, Daoguang, et al.
Published: (2025)
by: Zan, Daoguang, et al.
Published: (2025)
Rethinking Code Review in the Age of AI: A Vision for Agentic Code Review
by: Kamalı, Hüseyin Özgür, et al.
Published: (2026)
by: Kamalı, Hüseyin Özgür, et al.
Published: (2026)
APEX-SWE
by: Kottamasu, Abhi, et al.
Published: (2026)
by: Kottamasu, Abhi, et al.
Published: (2026)
AI-Assisted Assessment of Coding Practices in Modern Code Review
by: Vijayvergiya, Manushree, et al.
Published: (2024)
by: Vijayvergiya, Manushree, et al.
Published: (2024)
SWE-chat: Coding Agent Interactions From Real Users in the Wild
by: Baumann, Joachim, et al.
Published: (2026)
by: Baumann, Joachim, et al.
Published: (2026)
On the Footprints of Reviewer Bots Feedback on Agentic Pull Requests in OSS GitHub Repositories
by: Fatima, Syeda Kaneez, et al.
Published: (2026)
by: Fatima, Syeda Kaneez, et al.
Published: (2026)
Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review
by: Paul, Debalina Ghosh, et al.
Published: (2024)
by: Paul, Debalina Ghosh, et al.
Published: (2024)
Similar Items
-
Analyzing Message-Code Inconsistency in AI Coding Agent-Authored Pull Requests
by: Gong, Jingzhi, et al.
Published: (2026) -
Quality and Security Signals in AI-Generated Python Refactoring Pull Requests
by: Almukhtar, Mohamed, et al.
Published: (2026) -
More Code, Less Reuse: Investigating Code Quality and Reviewer Sentiment towards AI-generated Pull Requests
by: Huang, Haoming, et al.
Published: (2026) -
How AI Coding Agents Communicate: A Study of Pull Request Description Characteristics and Human Review Responses
by: Watanabe, Kan, et al.
Published: (2026) -
When AI Teammates Meet Code Review: Collaboration Signals Shaping the Integration of Agent-Authored Pull Requests
by: Nachuma, Costain, et al.
Published: (2026)