How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Ningzhi, Chen, Chaoran, Xu, Gelei, Shi, Yiyu, Huang, Yu, McMillan, Collin, Dong, Tao, Li, Toby Jia-Jun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Programming by Chat: A Large-Scale Behavioral Analysis of 11,579 Real-World AI-Assisted IDE Sessions
by: Tang, Ningzhi, et al.
Published: (2026)
by: Tang, Ningzhi, et al.
Published: (2026)
NaturalEdit: Code Modification through Direct Interaction with Adaptive Natural Language Representation
by: Tang, Ningzhi, et al.
Published: (2025)
by: Tang, Ningzhi, et al.
Published: (2025)
Exploring Direct Instruction and Summary-Mediated Prompting in LLM-Assisted Code Modification
by: Tang, Ningzhi, et al.
Published: (2025)
by: Tang, Ningzhi, et al.
Published: (2025)
Programmer Visual Attention During Context-Aware Code Summarization
by: Wallace, Robert, et al.
Published: (2024)
by: Wallace, Robert, et al.
Published: (2024)
A Study on Developer Behaviors for Validating and Repairing LLM-Generated Code Using Eye Tracking and IDE Actions
by: Tang, Ningzhi, et al.
Published: (2024)
by: Tang, Ningzhi, et al.
Published: (2024)
APIDA-Chat: Structured Synthesis of API Search Dialogues to Bootstrap Conversational Agents
by: Eberhart, Zachary, et al.
Published: (2025)
by: Eberhart, Zachary, et al.
Published: (2025)
CMind: An AI Agent for Localizing C Memory Bugs
by: Su, Chia-Yi, et al.
Published: (2026)
by: Su, Chia-Yi, et al.
Published: (2026)
Do Code LLMs Do Static Analysis?
by: Su, Chia-Yi, et al.
Published: (2025)
by: Su, Chia-Yi, et al.
Published: (2025)
Context-aware Code Summary Generation
by: Su, Chia-Yi, et al.
Published: (2024)
by: Su, Chia-Yi, et al.
Published: (2024)
Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables
by: McMillan, Damon
Published: (2026)
by: McMillan, Damon
Published: (2026)
Distilled GPT for Source Code Summarization
by: Su, Chia-Yi, et al.
Published: (2023)
by: Su, Chia-Yi, et al.
Published: (2023)
Semantic Similarity Loss for Neural Source Code Summarization
by: Su, Chia-Yi, et al.
Published: (2023)
by: Su, Chia-Yi, et al.
Published: (2023)
Semantic similarity loss for neural source code summarization
by: Chia‐Yi Su, et al.
Published: (2024)
by: Chia‐Yi Su, et al.
Published: (2024)
CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges
by: Zhang, Kechi, et al.
Published: (2024)
by: Zhang, Kechi, et al.
Published: (2024)
Automated Generation of Accurate Privacy Captions From Android Source Code Using Large Language Models
by: Jain, Vijayanta, et al.
Published: (2026)
by: Jain, Vijayanta, et al.
Published: (2026)
AI-Mediated Code Comment Improvement
by: Dhakal, Maria, et al.
Published: (2025)
by: Dhakal, Maria, et al.
Published: (2025)
EyeTrans: Merging Human and Machine Attention for Neural Code Summarization
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Do Machines and Humans Focus on Similar Code? Exploring Explainability of Large Language Models in Code Summarization
by: Li, Jiliang, et al.
Published: (2024)
by: Li, Jiliang, et al.
Published: (2024)
Semantic Neighborhood Density and Eye Gaze Time in Human Programmer Attention
by: Wallace, Robert, et al.
Published: (2026)
by: Wallace, Robert, et al.
Published: (2026)
Which Code Statements Implement Privacy Behaviors in Android Applications?
by: Su, Chia-Yi, et al.
Published: (2025)
by: Su, Chia-Yi, et al.
Published: (2025)
An Empirical Study of Agent Skills for Healthcare: Practice, Gaps, and Governance
by: Xu, Gelei, et al.
Published: (2026)
by: Xu, Gelei, et al.
Published: (2026)
Human Attention During Localization of Memory Bugs in C Programs
by: Smith, Emory, et al.
Published: (2025)
by: Smith, Emory, et al.
Published: (2025)
Coherence Collapse: Diagnosing Why Code Agents Fail After Reaching the Right Code
by: Kim, Myeongsoo, et al.
Published: (2026)
by: Kim, Myeongsoo, et al.
Published: (2026)
An Empirical Study of Proactive Coding Assistants in Real-World Software Development
by: Li, Lehui, et al.
Published: (2026)
by: Li, Lehui, et al.
Published: (2026)
GraphCodeAgent: Dual Graph-Guided LLM Agent for Retrieval-Augmented Repo-Level Code Generation
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
by: Ehsani, Ramtin, et al.
Published: (2026)
by: Ehsani, Ramtin, et al.
Published: (2026)
Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases
by: Wong, Sherman, et al.
Published: (2025)
by: Wong, Sherman, et al.
Published: (2025)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
by: Zhao, Songwen, et al.
Published: (2025)
by: Zhao, Songwen, et al.
Published: (2025)
SWE-chat: Coding Agent Interactions From Real Users in the Wild
by: Baumann, Joachim, et al.
Published: (2026)
by: Baumann, Joachim, et al.
Published: (2026)
What Makes a Good LLM Agent for Real-world Penetration Testing?
by: Deng, Gelei, et al.
Published: (2026)
by: Deng, Gelei, et al.
Published: (2026)
Evaluating Software Development Agents: Patch Patterns, Code Quality, and Issue Complexity in Real-World GitHub Scenarios
by: Chen, Zhi, et al.
Published: (2024)
by: Chen, Zhi, et al.
Published: (2024)
Immersion in the GitHub Universe: Scaling Coding Agents to Mastery
by: Zhao, Jiale, et al.
Published: (2026)
by: Zhao, Jiale, et al.
Published: (2026)
When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling
by: Islam, Niful, et al.
Published: (2026)
by: Islam, Niful, et al.
Published: (2026)
AI IDEs or Autonomous Agents? Measuring the Impact of Coding Agents on Software Development
by: Agarwal, Shyam, et al.
Published: (2026)
by: Agarwal, Shyam, et al.
Published: (2026)
Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub Scenarios
by: Chen, Zhi, et al.
Published: (2025)
by: Chen, Zhi, et al.
Published: (2025)
CodePori: Large-Scale System for Autonomous Software Development Using Multi-Agent Technology
by: Rasheed, Zeeshan, et al.
Published: (2024)
by: Rasheed, Zeeshan, et al.
Published: (2024)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
by: Mündler, Niels, et al.
Published: (2024)
by: Mündler, Niels, et al.
Published: (2024)
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
by: Pereira, Kristen, et al.
Published: (2026)
by: Pereira, Kristen, et al.
Published: (2026)
Similar Items
-
Programming by Chat: A Large-Scale Behavioral Analysis of 11,579 Real-World AI-Assisted IDE Sessions
by: Tang, Ningzhi, et al.
Published: (2026) -
NaturalEdit: Code Modification through Direct Interaction with Adaptive Natural Language Representation
by: Tang, Ningzhi, et al.
Published: (2025) -
Exploring Direct Instruction and Summary-Mediated Prompting in LLM-Assisted Code Modification
by: Tang, Ningzhi, et al.
Published: (2025) -
Programmer Visual Attention During Context-Aware Code Summarization
by: Wallace, Robert, et al.
Published: (2024) -
A Study on Developer Behaviors for Validating and Repairing LLM-Generated Code Using Eye Tracking and IDE Actions
by: Tang, Ningzhi, et al.
Published: (2024)