How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Arimbur, Johin Johny |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ReCode: Improving LLM-based Code Repair with Fine-Grained Retrieval-Augmented Generation
par: Zhao, Yicong, et autres
Publié: (2025)
par: Zhao, Yicong, et autres
Publié: (2025)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
par: Zhang, Wentao, et autres
Publié: (2026)
par: Zhang, Wentao, et autres
Publié: (2026)
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
par: Xu, Qiao, et autres
Publié: (2026)
par: Xu, Qiao, et autres
Publié: (2026)
NARRepair: Non-Autoregressive Code Generation Model for Automatic Program Repair
par: Yang, Zhenyu, et autres
Publié: (2024)
par: Yang, Zhenyu, et autres
Publié: (2024)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
par: Qiu, Ruizhong, et autres
Publié: (2024)
par: Qiu, Ruizhong, et autres
Publié: (2024)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
par: Orlanski, Gabriel, et autres
Publié: (2026)
par: Orlanski, Gabriel, et autres
Publié: (2026)
LLM-Powered Code Vulnerability Repair with Reinforcement Learning and Semantic Reward
par: Islam, Nafis Tanveer, et autres
Publié: (2024)
par: Islam, Nafis Tanveer, et autres
Publié: (2024)
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
par: Cui, Yi
Publié: (2025)
par: Cui, Yi
Publié: (2025)
MLDebugging: Towards Benchmarking Code Debugging Across Multi-Library Scenarios
par: Huang, Jinyang, et autres
Publié: (2025)
par: Huang, Jinyang, et autres
Publié: (2025)
Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
par: Singha, Ananya, et autres
Publié: (2025)
par: Singha, Ananya, et autres
Publié: (2025)
A Study on the Impact of Fault localization Granularity for Repository-Scale Code Repair Tasks
par: Townsend, Joseph, et autres
Publié: (2026)
par: Townsend, Joseph, et autres
Publié: (2026)
How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study
par: Velasco, Alejandro, et autres
Publié: (2024)
par: Velasco, Alejandro, et autres
Publié: (2024)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
par: Guo, Chengquan, et autres
Publié: (2024)
par: Guo, Chengquan, et autres
Publié: (2024)
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
par: Trinkenreich, Bianca, et autres
Publié: (2026)
par: Trinkenreich, Bianca, et autres
Publié: (2026)
RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation
par: Liang, Linxi, et autres
Publié: (2025)
par: Liang, Linxi, et autres
Publié: (2025)
Benchmarking Large Language Models for ABAP Code Generation: An Empirical Study on Iterative Improvement by Compiler Feedback
par: Wallraven, Stephan, et autres
Publié: (2026)
par: Wallraven, Stephan, et autres
Publié: (2026)
LLM Test Generation via Iterative Hybrid Program Analysis
par: Gu, Sijia, et autres
Publié: (2025)
par: Gu, Sijia, et autres
Publié: (2025)
From Defects to Demands: A Unified, Iterative, and Heuristically Guided LLM-Based Framework for Automated Software Repair and Requirement Realization
par: Alex, et autres
Publié: (2024)
par: Alex, et autres
Publié: (2024)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
par: Pan, Zhiyuan, et autres
Publié: (2025)
par: Pan, Zhiyuan, et autres
Publié: (2025)
Revisit Self-Debugging with Self-Generated Tests for Code Generation
par: Chen, Xiancai, et autres
Publié: (2025)
par: Chen, Xiancai, et autres
Publié: (2025)
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
par: Bouzenia, Islem, et autres
Publié: (2024)
par: Bouzenia, Islem, et autres
Publié: (2024)
Large Language Model Guided Self-Debugging Code Generation
par: Adnan, Muntasir, et autres
Publié: (2025)
par: Adnan, Muntasir, et autres
Publié: (2025)
Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair
par: de-Fitero-Dominguez, David, et autres
Publié: (2025)
par: de-Fitero-Dominguez, David, et autres
Publié: (2025)
RepoRepair: Leveraging Code Documentation for Repository-Level Automated Program Repair
par: Pan, Zhongqiang, et autres
Publié: (2026)
par: Pan, Zhongqiang, et autres
Publié: (2026)
Automated Repair of AI Code with Large Language Models and Formal Verification
par: Charalambous, Yiannis, et autres
Publié: (2024)
par: Charalambous, Yiannis, et autres
Publié: (2024)
Investigating The Smells of LLM Generated Code
par: Paul, Debalina Ghosh, et autres
Publié: (2025)
par: Paul, Debalina Ghosh, et autres
Publié: (2025)
Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
par: Zhang, Zhe, et autres
Publié: (2025)
par: Zhang, Zhe, et autres
Publié: (2025)
LLM Code Customization with Visual Results: A Benchmark on TikZ
par: Reux, Charly, et autres
Publié: (2025)
par: Reux, Charly, et autres
Publié: (2025)
SimdBench: Benchmarking Large Language Models for SIMD-Intrinsic Code Generation
par: He, Yibo, et autres
Publié: (2025)
par: He, Yibo, et autres
Publié: (2025)
Insights from Benchmarking Frontier Language Models on Web App Code Generation
par: Cui, Yi
Publié: (2024)
par: Cui, Yi
Publié: (2024)
ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization
par: Huang, Yixu, et autres
Publié: (2026)
par: Huang, Yixu, et autres
Publié: (2026)
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
par: Diehl, Patrick, et autres
Publié: (2025)
par: Diehl, Patrick, et autres
Publié: (2025)
INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair
par: Wang, Hanbin, et autres
Publié: (2023)
par: Wang, Hanbin, et autres
Publié: (2023)
Code Copycat Conundrum: Demystifying Repetition in LLM-based Code Generation
par: Liu, Mingwei, et autres
Publié: (2025)
par: Liu, Mingwei, et autres
Publié: (2025)
Benchmarking Correctness and Security in Multi-Turn Code Generation
par: Rawal, Ruchit, et autres
Publié: (2025)
par: Rawal, Ruchit, et autres
Publié: (2025)
Lyra: A Benchmark for Turducken-Style Code Generation
par: Liang, Qingyuan, et autres
Publié: (2021)
par: Liang, Qingyuan, et autres
Publié: (2021)
Automated Benchmark Generation for Repository-Level Coding Tasks
par: Vergopoulos, Konstantinos, et autres
Publié: (2025)
par: Vergopoulos, Konstantinos, et autres
Publié: (2025)
Is Self-Repair a Silver Bullet for Code Generation?
par: Olausson, Theo X., et autres
Publié: (2023)
par: Olausson, Theo X., et autres
Publié: (2023)
Uncertainty Quantification for LLM-based Code Generation
par: Xu, Senrong, et autres
Publié: (2026)
par: Xu, Senrong, et autres
Publié: (2026)
CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models
par: Padwal, Vedant
Publié: (2026)
par: Padwal, Vedant
Publié: (2026)
Documents similaires
-
ReCode: Improving LLM-based Code Repair with Fine-Grained Retrieval-Augmented Generation
par: Zhao, Yicong, et autres
Publié: (2025) -
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
par: Zhang, Wentao, et autres
Publié: (2026) -
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
par: Xu, Qiao, et autres
Publié: (2026) -
NARRepair: Non-Autoregressive Code Generation Model for Automatic Program Repair
par: Yang, Zhenyu, et autres
Publié: (2024) -
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
par: Qiu, Ruizhong, et autres
Publié: (2024)