Characterizing the Failure Modes of LLMs in Resolving Real-World GitHub Issues
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Yanjie, Huang, Yian, Wang, Guancheng, Chen, Junjie, Liu, Hui, Briand, Lionel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
by: Jimenez, Carlos E., et al.
Published: (2023)
by: Jimenez, Carlos E., et al.
Published: (2023)
GitHub Proxy Server: A tool for supporting massive data collection on GitHub
by: Borges, Hudson Silva, et al.
Published: (2025)
by: Borges, Hudson Silva, et al.
Published: (2025)
SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
by: Zan, Daoguang, et al.
Published: (2024)
by: Zan, Daoguang, et al.
Published: (2024)
Classifying Issues in Open-source GitHub Repositories
by: Raaj, Amir Hossain, et al.
Published: (2025)
by: Raaj, Amir Hossain, et al.
Published: (2025)
Evaluating Software Development Agents: Patch Patterns, Code Quality, and Issue Complexity in Real-World GitHub Scenarios
by: Chen, Zhi, et al.
Published: (2024)
by: Chen, Zhi, et al.
Published: (2024)
OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution
by: Guo, Lianghong, et al.
Published: (2025)
by: Guo, Lianghong, et al.
Published: (2025)
Guidelines for Developing Bots for GitHub
by: Wessel, Mairieli, et al.
Published: (2022)
by: Wessel, Mairieli, et al.
Published: (2022)
Prioritising GitHub Priority Labels
by: Caddy, James, et al.
Published: (2024)
by: Caddy, James, et al.
Published: (2024)
"My GitHub Sponsors profile is live!" Investigating the Impact of Twitter/X Mentions on GitHub Sponsors
by: Fan, Youmei, et al.
Published: (2024)
by: Fan, Youmei, et al.
Published: (2024)
Beyond the YAML File: Understanding Real-World GitHub Actions Workflow Adoption
by: Khatami, Ali, et al.
Published: (2026)
by: Khatami, Ali, et al.
Published: (2026)
Immersion in the GitHub Universe: Scaling Coding Agents to Mastery
by: Zhao, Jiale, et al.
Published: (2026)
by: Zhao, Jiale, et al.
Published: (2026)
Hallucination to Consensus: Multi-Agent LLMs for End-to-End JUnit Test Generation
by: Xu, Qinghua, et al.
Published: (2025)
by: Xu, Qinghua, et al.
Published: (2025)
An Empirical Study of ChatGPT-Related Projects and Their Issues on GitHub
by: Lin, Zheng, et al.
Published: (2024)
by: Lin, Zheng, et al.
Published: (2024)
Can GitHub Issues Help in App Review Classifications?
by: Abedini, Yasaman, et al.
Published: (2023)
by: Abedini, Yasaman, et al.
Published: (2023)
What Makes a GitHub Issue Ready for Copilot?
by: Sayagh, Mohammed
Published: (2025)
by: Sayagh, Mohammed
Published: (2025)
The Impact of Sanctions on GitHub Developers and Activities
by: Fan, Youmei, et al.
Published: (2024)
by: Fan, Youmei, et al.
Published: (2024)
Fingerprinting AI Coding Agents on GitHub
by: Ghaleb, Taher A.
Published: (2026)
by: Ghaleb, Taher A.
Published: (2026)
Explaining GitHub Actions Failures with Large Language Models: Challenges, Insights, and Limitations
by: Valenzuela-Toledo, Pablo, et al.
Published: (2025)
by: Valenzuela-Toledo, Pablo, et al.
Published: (2025)
Characterizing and Modeling the GitHub Security Advisories Review Pipeline
by: Segal, Claudio, et al.
Published: (2026)
by: Segal, Claudio, et al.
Published: (2026)
GitHub Marketplace for Automation and Innovation in Software Production
by: Saroar, SK Golam, et al.
Published: (2024)
by: Saroar, SK Golam, et al.
Published: (2024)
Introducing Traceability in GitHub for Medical Software Development
by: Stirbu, Vlad, et al.
Published: (2021)
by: Stirbu, Vlad, et al.
Published: (2021)
The Landscape of Toxicity: An Empirical Investigation of Toxicity on GitHub
by: Sarker, Jaydeb, et al.
Published: (2025)
by: Sarker, Jaydeb, et al.
Published: (2025)
Chaos Engineering in the Wild: Findings from GitHub
by: Owotogbe, Joshua, et al.
Published: (2025)
by: Owotogbe, Joshua, et al.
Published: (2025)
An Empirical Study of the Evolution of GitHub Actions Workflows
by: Mazrae, Pooya Rostami, et al.
Published: (2026)
by: Mazrae, Pooya Rostami, et al.
Published: (2026)
GitHub Actions: The Impact on the Pull Request Process
by: Wessel, Mairieli, et al.
Published: (2022)
by: Wessel, Mairieli, et al.
Published: (2022)
Agentic Much? Adoption of Coding Agents on GitHub
by: Robbes, Romain, et al.
Published: (2026)
by: Robbes, Romain, et al.
Published: (2026)
Understanding and Predicting Derailment in Toxic Conversations on GitHub
by: Imran, Mia Mohammad, et al.
Published: (2025)
by: Imran, Mia Mohammad, et al.
Published: (2025)
Visual Analysis of GitHub Issues to Gain Insights
by: Proma, Rifat Ara, et al.
Published: (2024)
by: Proma, Rifat Ara, et al.
Published: (2024)
IssueGuard: Real-Time Secret Leak Prevention Tool for GitHub Issue Reports
by: Rahman, Md Nafiu, et al.
Published: (2026)
by: Rahman, Md Nafiu, et al.
Published: (2026)
Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub Scenarios
by: Chen, Zhi, et al.
Published: (2025)
by: Chen, Zhi, et al.
Published: (2025)
Test Code Review in the Era of GitHub Actions: A Replication Study
by: Sun, Hui, et al.
Published: (2026)
by: Sun, Hui, et al.
Published: (2026)
Testing in the Evolving World of DL Systems:Insights from Python GitHub Projects
by: Ali, Qurban, et al.
Published: (2024)
by: Ali, Qurban, et al.
Published: (2024)
Transforming Software Development: Evaluating the Efficiency and Challenges of GitHub Copilot in Real-World Projects
by: Pandey, Ruchika, et al.
Published: (2024)
by: Pandey, Ruchika, et al.
Published: (2024)
GitBug-Actions: Building Reproducible Bug-Fix Benchmarks with GitHub Actions
by: Saavedra, Nuno, et al.
Published: (2023)
by: Saavedra, Nuno, et al.
Published: (2023)
On the GitHub Actions Language: Usage, Evolution, and Workflow Reliability
by: Bardsiri, Aref Talebzadeh, et al.
Published: (2026)
by: Bardsiri, Aref Talebzadeh, et al.
Published: (2026)
Designing for Cognitive Diversity: Improving the GitHub Experience for Newcomers
by: Santos, Italo, et al.
Published: (2023)
by: Santos, Italo, et al.
Published: (2023)
Weaponizing the Commons: A Taxonomy and Detection Framework of Abuse on GitHub
by: Cheng, Yuli, et al.
Published: (2026)
by: Cheng, Yuli, et al.
Published: (2026)
Where Is Self-admitted Code Generated by Large Language Models on GitHub?
by: Yu, Xiao, et al.
Published: (2024)
by: Yu, Xiao, et al.
Published: (2024)
Mutation-Guided Unit Test Generation with a Large Language Model
by: Wang, Guancheng, et al.
Published: (2025)
by: Wang, Guancheng, et al.
Published: (2025)
Can LLMs Write CI? A Study on Automatic Generation of GitHub Actions Configurations
by: Ghaleb, Taher A., et al.
Published: (2025)
by: Ghaleb, Taher A., et al.
Published: (2025)
Similar Items
-
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
by: Jimenez, Carlos E., et al.
Published: (2023) -
GitHub Proxy Server: A tool for supporting massive data collection on GitHub
by: Borges, Hudson Silva, et al.
Published: (2025) -
SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
by: Zan, Daoguang, et al.
Published: (2024) -
Classifying Issues in Open-source GitHub Repositories
by: Raaj, Amir Hossain, et al.
Published: (2025) -
Evaluating Software Development Agents: Patch Patterns, Code Quality, and Issue Complexity in Real-World GitHub Scenarios
by: Chen, Zhi, et al.
Published: (2024)