Coherence Collapse: Diagnosing Why Code Agents Fail After Reaching the Right Code
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Myeongsoo, Wang, Dingmin, Cui, Siwei, Farmahinifarahani, Farima, Zhuo, Terry Yue, Garg, Shweta, Ray, Baishakhi, Mukherjee, Rajdeep, Kumar, Varun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CodeAssistBench (CAB): Dataset & Benchmarking for Multi-turn Chat-Based Code Assistance
by: Kim, Myeongsoo, et al.
Published: (2025)
by: Kim, Myeongsoo, et al.
Published: (2025)
CODESTRUCT: Code Agents over Structured Action Spaces
by: Kim, Myeongsoo, et al.
Published: (2026)
by: Kim, Myeongsoo, et al.
Published: (2026)
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
by: Haque, Mirazul, et al.
Published: (2025)
by: Haque, Mirazul, et al.
Published: (2025)
Cyber-Zero: Training Cybersecurity Agents without Runtime
by: Zhuo, Terry Yue, et al.
Published: (2025)
by: Zhuo, Terry Yue, et al.
Published: (2025)
Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
by: Zhuo, Terry Yue, et al.
Published: (2025)
by: Zhuo, Terry Yue, et al.
Published: (2025)
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
by: Roh, Jaechul, et al.
Published: (2025)
by: Roh, Jaechul, et al.
Published: (2025)
On Mitigating Code LLM Hallucinations with API Documentation
by: Jain, Nihal, et al.
Published: (2024)
by: Jain, Nihal, et al.
Published: (2024)
ICE-Score: Instructing Large Language Models to Evaluate Code
by: Zhuo, Terry Yue
Published: (2023)
by: Zhuo, Terry Yue
Published: (2023)
CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation
by: Peng, Jinjun, et al.
Published: (2025)
by: Peng, Jinjun, et al.
Published: (2025)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
SpecTra: Enhancing the Code Translation Ability of Language Models by Generating Multi-Modal Specifications
by: Nitin, Vikram, et al.
Published: (2024)
by: Nitin, Vikram, et al.
Published: (2024)
CYCLE: Learning to Self-Refine the Code Generation
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
EditLord: Learning Code Transformation Rules for Code Editing
by: Li, Weichen, et al.
Published: (2025)
by: Li, Weichen, et al.
Published: (2025)
Automated Code Editing with Search-Generate-Modify
by: Liu, Changshu, et al.
Published: (2023)
by: Liu, Changshu, et al.
Published: (2023)
CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning
by: Roy, Monoshi Kumar, et al.
Published: (2025)
by: Roy, Monoshi Kumar, et al.
Published: (2025)
LeDex: Training LLMs to Better Self-Debug and Explain Code
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
Substance Beats Style: Why Beginning Students Fail to Code with LLMs
by: Lucchetti, Francesca, et al.
Published: (2024)
by: Lucchetti, Francesca, et al.
Published: (2024)
CodeScout: Contextual Problem Statement Enhancement for Software Agents
by: Suri, Manan, et al.
Published: (2026)
by: Suri, Manan, et al.
Published: (2026)
Revealing degradation mechanisms in YSZ ceramics through machine learning-guided aging and multiscale characterization
by: Garg, Prachi, et al.
Published: (2025)
by: Garg, Prachi, et al.
Published: (2025)
Code Quality Analysis of Translations from C to Rust
by: Tadesse, Biruk, et al.
Published: (2026)
by: Tadesse, Biruk, et al.
Published: (2026)
Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action
by: Pujar, Saurabh, et al.
Published: (2025)
by: Pujar, Saurabh, et al.
Published: (2025)
PrivCode: When Code Generation Meets Differential Privacy
by: Liu, Zheng, et al.
Published: (2025)
by: Liu, Zheng, et al.
Published: (2025)
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
Why Is Exclusivity in Broadcasting Rights Prevalent and Why Does Simple Regulation Fail?
by: David Martimort, et al.
Published: (2025)
by: David Martimort, et al.
Published: (2025)
Why Prototypes Collapse: Diagnosing and Preventing Partial Collapse in Prototypical Self-Supervised Learning
by: Arteaga, Gabriel Y., et al.
Published: (2025)
by: Arteaga, Gabriel Y., et al.
Published: (2025)
Why Self-Awareness Fails: Coherence, Defense, and the Structure of Repair
by: Jovanovic, Vladisav
Published: (2026)
by: Jovanovic, Vladisav
Published: (2026)
Cracking the Code of Arctic Sea Ice: Why Models Fail to Predict Its Retreat?
by: Gou, Ruijian, et al.
Published: (2025)
by: Gou, Ruijian, et al.
Published: (2025)
The Observability Gap: Why Output-Level Human Feedback Fails for LLM Coding Agents
by: Wang, Yinghao, et al.
Published: (2026)
by: Wang, Yinghao, et al.
Published: (2026)
CodeSSM: Towards State Space Models for Code Understanding
by: Verma, Shweta, et al.
Published: (2025)
by: Verma, Shweta, et al.
Published: (2025)
CodeFort: Robust Training for Code Generation Models
by: Zhang, Yuhao, et al.
Published: (2024)
by: Zhang, Yuhao, et al.
Published: (2024)
Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalisation
by: Pacela, Vitória Barin, et al.
Published: (2026)
by: Pacela, Vitória Barin, et al.
Published: (2026)
Robustness, Security, Privacy, Explainability, Efficiency, and Usability of Large Language Models for Code
by: Yang, Zhou, et al.
Published: (2024)
by: Yang, Zhou, et al.
Published: (2024)
CodeArena: A Collective Evaluation Platform for LLM Code Generation
by: Du, Mingzhe, et al.
Published: (2025)
by: Du, Mingzhe, et al.
Published: (2025)
XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts
by: Ding, Yifeng, et al.
Published: (2024)
by: Ding, Yifeng, et al.
Published: (2024)
Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChain
by: Min, Marcus J., et al.
Published: (2023)
by: Min, Marcus J., et al.
Published: (2023)
Coherence manipulation in asymmetry and thermodynamics
by: Kondra, Tulja Varun, et al.
Published: (2023)
by: Kondra, Tulja Varun, et al.
Published: (2023)
Colour Codes Reach Surface Code Performance using Vibe Decoding
by: Koutsioumpas, Stergios, et al.
Published: (2025)
by: Koutsioumpas, Stergios, et al.
Published: (2025)
Machine learning–guided optimization of coercive field in Al 1− x Sc x N thin films for nonvolatile memory
by: Shaon Das, et al.
Published: (2024)
by: Shaon Das, et al.
Published: (2024)
Vulnerability Detection with Code Language Models: How Far Are We?
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
Chain-of-Thought in Neural Code Generation: From and For Lightweight Language Models
by: Yang, Guang, et al.
Published: (2023)
by: Yang, Guang, et al.
Published: (2023)
Similar Items
-
CodeAssistBench (CAB): Dataset & Benchmarking for Multi-turn Chat-Based Code Assistance
by: Kim, Myeongsoo, et al.
Published: (2025) -
CODESTRUCT: Code Agents over Structured Action Spaces
by: Kim, Myeongsoo, et al.
Published: (2026) -
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
by: Haque, Mirazul, et al.
Published: (2025) -
Cyber-Zero: Training Cybersecurity Agents without Runtime
by: Zhuo, Terry Yue, et al.
Published: (2025) -
Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
by: Zhuo, Terry Yue, et al.
Published: (2025)