Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency
Fuente:
arXiv
Saved in:
| Main Authors: | Ashrafi, Nazmus, Bouktif, Salah, Mediani, Mohammed |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval
by: Ashrafi, Nazmus
Published: (2026)
by: Ashrafi, Nazmus
Published: (2026)
Improving Code Reviewer Recommendation: Accuracy, Latency, Workload, and Bystanders
by: Rigby, Peter C., et al.
Published: (2023)
by: Rigby, Peter C., et al.
Published: (2023)
DebugTA: An LLM-Based Agent for Simplifying Debugging and Teaching in Programming Education
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
Multi-Agent Code-Orchestrated Generation for Reliable Infrastructure-as-Code
by: Khan, Rana Nameer Hussain, et al.
Published: (2025)
by: Khan, Rana Nameer Hussain, et al.
Published: (2025)
CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging
by: Islam, Md. Ashraful, et al.
Published: (2025)
by: Islam, Md. Ashraful, et al.
Published: (2025)
Structural Verification for Reliable EDA Code Generation without Tool-in-the-Loop Debugging
by: Jayasuriya, Dinithi, et al.
Published: (2026)
by: Jayasuriya, Dinithi, et al.
Published: (2026)
TraceCoder: A Trace-Driven Multi-Agent Framework for Automated Debugging of LLM-Generated Code
by: Huang, Jiangping, et al.
Published: (2026)
by: Huang, Jiangping, et al.
Published: (2026)
DebugRepair: Enhancing LLM-Based Automated Program Repair via Self-Directed Debugging
by: Wu, Linhao, et al.
Published: (2026)
by: Wu, Linhao, et al.
Published: (2026)
Debugging Performance Issues in WebAssembly Runtimes via Mutation-based Inference
by: Zeng, Ruiying, et al.
Published: (2026)
by: Zeng, Ruiying, et al.
Published: (2026)
Requirements Development and Formalization for Reliable Code Generation: A Multi-Agent Vision
by: Lu, Xu, et al.
Published: (2025)
by: Lu, Xu, et al.
Published: (2025)
Leveraging Print Debugging to Improve Code Generation in Large Language Models
by: Hu, Xueyu, et al.
Published: (2024)
by: Hu, Xueyu, et al.
Published: (2024)
Coding with Eyes: Visual Feedback Unlocks Reliable GUI Code Generating and Debugging
by: Liu, Zhilin, et al.
Published: (2026)
by: Liu, Zhilin, et al.
Published: (2026)
Debug2Fix: Can Interactive Debugging Help Coding Agents Fix More Bugs?
by: Garg, Spandan, et al.
Published: (2026)
by: Garg, Spandan, et al.
Published: (2026)
Balancing Latency and Accuracy of Code Completion via Local-Cloud Model Cascading
by: Lu, Hanzhen, et al.
Published: (2026)
by: Lu, Hanzhen, et al.
Published: (2026)
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
by: Ma, Ming, et al.
Published: (2025)
by: Ma, Ming, et al.
Published: (2025)
A Generic and Efficient Python Runtime Verification System and its Large-scale Evaluation
by: Shen, Zhuohang, et al.
Published: (2025)
by: Shen, Zhuohang, et al.
Published: (2025)
UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging
by: Lee, Cheryl, et al.
Published: (2024)
by: Lee, Cheryl, et al.
Published: (2024)
COAST: Enhancing the Code Debugging Ability of LLMs through Communicative Agent Based Data Synthesis
by: Yang, Weiqing, et al.
Published: (2024)
by: Yang, Weiqing, et al.
Published: (2024)
Is LLM-Generated Code More Maintainable \& Reliable than Human-Written Code?
by: Molison, Alfred Santa, et al.
Published: (2025)
by: Molison, Alfred Santa, et al.
Published: (2025)
Correctness isnt Efficiency: Runtime Memory Divergence in LLM-Generated Code
by: Rajput, Prateek, et al.
Published: (2026)
by: Rajput, Prateek, et al.
Published: (2026)
The Debugging Decay Index: Rethinking Debugging Strategies for Code LLMs
by: Adnan, Muntasir, et al.
Published: (2025)
by: Adnan, Muntasir, et al.
Published: (2025)
An Empirical Study of Interaction Smells in Multi-Turn Human-LLM Collaborative Code Generation
by: Zhang, Binquan, et al.
Published: (2026)
by: Zhang, Binquan, et al.
Published: (2026)
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
A Systematic Survey on Debugging Techniques for Machine Learning Systems
by: Nguyen, Thanh-Dat, et al.
Published: (2025)
by: Nguyen, Thanh-Dat, et al.
Published: (2025)
A Systematic Mapping Study on the Debugging of Autonomous Driving Systems
by: Shaw, Nathan, et al.
Published: (2026)
by: Shaw, Nathan, et al.
Published: (2026)
LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review
by: Rasheeda, Zeeshan, et al.
Published: (2026)
by: Rasheeda, Zeeshan, et al.
Published: (2026)
Revisit Self-Debugging with Self-Generated Tests for Code Generation
by: Chen, Xiancai, et al.
Published: (2025)
by: Chen, Xiancai, et al.
Published: (2025)
BugSpotter: Automated Generation of Code Debugging Exercises
by: Pădurean, Victor-Alexandru, et al.
Published: (2024)
by: Pădurean, Victor-Alexandru, et al.
Published: (2024)
CodeCoR: An LLM-Based Self-Reflective Multi-Agent Framework for Code Generation
by: Pan, Ruwei, et al.
Published: (2025)
by: Pan, Ruwei, et al.
Published: (2025)
RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models
by: Liu, Jingjing, et al.
Published: (2025)
by: Liu, Jingjing, et al.
Published: (2025)
MANTRA: Enhancing Automated Method-Level Refactoring with Contextual RAG and Multi-Agent LLM Collaboration
by: Xu, Yisen, et al.
Published: (2025)
by: Xu, Yisen, et al.
Published: (2025)
Timing Analysis Agent: Autonomous Multi-Corner Multi-Mode (MCMM) Timing Debugging with Timing Debug Relation Graph
by: Nainani, Jatin, et al.
Published: (2025)
by: Nainani, Jatin, et al.
Published: (2025)
Personalization of Code Readability Evaluation Based on LLM Using Collaborative Filtering
by: Hiraki, Buntaro, et al.
Published: (2024)
by: Hiraki, Buntaro, et al.
Published: (2024)
Debugging and Runtime Analysis of Neural Networks with VLMs (A Case Study)
by: Hu, Boyue Caroline, et al.
Published: (2025)
by: Hu, Boyue Caroline, et al.
Published: (2025)
Guided Debugging of Auto-Translated Code Using Differential Testing
by: Wu, Shengnan, et al.
Published: (2025)
by: Wu, Shengnan, et al.
Published: (2025)
AP2O-Coder: Adaptively Progressive Preference Optimization for Reducing Compilation and Runtime Errors in LLM-Generated Code
by: Zhang, Jianqing, et al.
Published: (2025)
by: Zhang, Jianqing, et al.
Published: (2025)
Large Language Model Guided Self-Debugging Code Generation
by: Adnan, Muntasir, et al.
Published: (2025)
by: Adnan, Muntasir, et al.
Published: (2025)
Agent That Debugs: Dynamic State-Guided Vulnerability Repair
by: Liu, Zhengyao, et al.
Published: (2025)
by: Liu, Zhengyao, et al.
Published: (2025)
Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis
by: Xiang, Jiahong, et al.
Published: (2026)
by: Xiang, Jiahong, et al.
Published: (2026)
CodeArena: A Collective Evaluation Platform for LLM Code Generation
by: Du, Mingzhe, et al.
Published: (2025)
by: Du, Mingzhe, et al.
Published: (2025)
Similar Items
-
How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval
by: Ashrafi, Nazmus
Published: (2026) -
Improving Code Reviewer Recommendation: Accuracy, Latency, Workload, and Bystanders
by: Rigby, Peter C., et al.
Published: (2023) -
DebugTA: An LLM-Based Agent for Simplifying Debugging and Teaching in Programming Education
by: Fu, Lingyue, et al.
Published: (2025) -
Multi-Agent Code-Orchestrated Generation for Reliable Infrastructure-as-Code
by: Khan, Rana Nameer Hussain, et al.
Published: (2025) -
CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging
by: Islam, Md. Ashraful, et al.
Published: (2025)