How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks
Fuente:
arXiv
Guardado en:
| Autor principal: | Arimbur, Johin Johny |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ReCode: Improving LLM-based Code Repair with Fine-Grained Retrieval-Augmented Generation
por: Zhao, Yicong, et al.
Publicado: (2025)
por: Zhao, Yicong, et al.
Publicado: (2025)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
por: Zhang, Wentao, et al.
Publicado: (2026)
por: Zhang, Wentao, et al.
Publicado: (2026)
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
por: Xu, Qiao, et al.
Publicado: (2026)
por: Xu, Qiao, et al.
Publicado: (2026)
NARRepair: Non-Autoregressive Code Generation Model for Automatic Program Repair
por: Yang, Zhenyu, et al.
Publicado: (2024)
por: Yang, Zhenyu, et al.
Publicado: (2024)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
por: Qiu, Ruizhong, et al.
Publicado: (2024)
por: Qiu, Ruizhong, et al.
Publicado: (2024)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
por: Orlanski, Gabriel, et al.
Publicado: (2026)
por: Orlanski, Gabriel, et al.
Publicado: (2026)
LLM-Powered Code Vulnerability Repair with Reinforcement Learning and Semantic Reward
por: Islam, Nafis Tanveer, et al.
Publicado: (2024)
por: Islam, Nafis Tanveer, et al.
Publicado: (2024)
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
por: Cui, Yi
Publicado: (2025)
por: Cui, Yi
Publicado: (2025)
MLDebugging: Towards Benchmarking Code Debugging Across Multi-Library Scenarios
por: Huang, Jinyang, et al.
Publicado: (2025)
por: Huang, Jinyang, et al.
Publicado: (2025)
Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
por: Singha, Ananya, et al.
Publicado: (2025)
por: Singha, Ananya, et al.
Publicado: (2025)
A Study on the Impact of Fault localization Granularity for Repository-Scale Code Repair Tasks
por: Townsend, Joseph, et al.
Publicado: (2026)
por: Townsend, Joseph, et al.
Publicado: (2026)
How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study
por: Velasco, Alejandro, et al.
Publicado: (2024)
por: Velasco, Alejandro, et al.
Publicado: (2024)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
por: Guo, Chengquan, et al.
Publicado: (2024)
por: Guo, Chengquan, et al.
Publicado: (2024)
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
por: Trinkenreich, Bianca, et al.
Publicado: (2026)
por: Trinkenreich, Bianca, et al.
Publicado: (2026)
RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation
por: Liang, Linxi, et al.
Publicado: (2025)
por: Liang, Linxi, et al.
Publicado: (2025)
Benchmarking Large Language Models for ABAP Code Generation: An Empirical Study on Iterative Improvement by Compiler Feedback
por: Wallraven, Stephan, et al.
Publicado: (2026)
por: Wallraven, Stephan, et al.
Publicado: (2026)
LLM Test Generation via Iterative Hybrid Program Analysis
por: Gu, Sijia, et al.
Publicado: (2025)
por: Gu, Sijia, et al.
Publicado: (2025)
From Defects to Demands: A Unified, Iterative, and Heuristically Guided LLM-Based Framework for Automated Software Repair and Requirement Realization
por: Alex, et al.
Publicado: (2024)
por: Alex, et al.
Publicado: (2024)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
por: Pan, Zhiyuan, et al.
Publicado: (2025)
por: Pan, Zhiyuan, et al.
Publicado: (2025)
Revisit Self-Debugging with Self-Generated Tests for Code Generation
por: Chen, Xiancai, et al.
Publicado: (2025)
por: Chen, Xiancai, et al.
Publicado: (2025)
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
por: Bouzenia, Islem, et al.
Publicado: (2024)
por: Bouzenia, Islem, et al.
Publicado: (2024)
Large Language Model Guided Self-Debugging Code Generation
por: Adnan, Muntasir, et al.
Publicado: (2025)
por: Adnan, Muntasir, et al.
Publicado: (2025)
Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair
por: de-Fitero-Dominguez, David, et al.
Publicado: (2025)
por: de-Fitero-Dominguez, David, et al.
Publicado: (2025)
RepoRepair: Leveraging Code Documentation for Repository-Level Automated Program Repair
por: Pan, Zhongqiang, et al.
Publicado: (2026)
por: Pan, Zhongqiang, et al.
Publicado: (2026)
Automated Repair of AI Code with Large Language Models and Formal Verification
por: Charalambous, Yiannis, et al.
Publicado: (2024)
por: Charalambous, Yiannis, et al.
Publicado: (2024)
Investigating The Smells of LLM Generated Code
por: Paul, Debalina Ghosh, et al.
Publicado: (2025)
por: Paul, Debalina Ghosh, et al.
Publicado: (2025)
Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
por: Zhang, Zhe, et al.
Publicado: (2025)
por: Zhang, Zhe, et al.
Publicado: (2025)
LLM Code Customization with Visual Results: A Benchmark on TikZ
por: Reux, Charly, et al.
Publicado: (2025)
por: Reux, Charly, et al.
Publicado: (2025)
SimdBench: Benchmarking Large Language Models for SIMD-Intrinsic Code Generation
por: He, Yibo, et al.
Publicado: (2025)
por: He, Yibo, et al.
Publicado: (2025)
Insights from Benchmarking Frontier Language Models on Web App Code Generation
por: Cui, Yi
Publicado: (2024)
por: Cui, Yi
Publicado: (2024)
ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization
por: Huang, Yixu, et al.
Publicado: (2026)
por: Huang, Yixu, et al.
Publicado: (2026)
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
por: Diehl, Patrick, et al.
Publicado: (2025)
por: Diehl, Patrick, et al.
Publicado: (2025)
INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair
por: Wang, Hanbin, et al.
Publicado: (2023)
por: Wang, Hanbin, et al.
Publicado: (2023)
Code Copycat Conundrum: Demystifying Repetition in LLM-based Code Generation
por: Liu, Mingwei, et al.
Publicado: (2025)
por: Liu, Mingwei, et al.
Publicado: (2025)
Benchmarking Correctness and Security in Multi-Turn Code Generation
por: Rawal, Ruchit, et al.
Publicado: (2025)
por: Rawal, Ruchit, et al.
Publicado: (2025)
Lyra: A Benchmark for Turducken-Style Code Generation
por: Liang, Qingyuan, et al.
Publicado: (2021)
por: Liang, Qingyuan, et al.
Publicado: (2021)
Automated Benchmark Generation for Repository-Level Coding Tasks
por: Vergopoulos, Konstantinos, et al.
Publicado: (2025)
por: Vergopoulos, Konstantinos, et al.
Publicado: (2025)
Is Self-Repair a Silver Bullet for Code Generation?
por: Olausson, Theo X., et al.
Publicado: (2023)
por: Olausson, Theo X., et al.
Publicado: (2023)
Uncertainty Quantification for LLM-based Code Generation
por: Xu, Senrong, et al.
Publicado: (2026)
por: Xu, Senrong, et al.
Publicado: (2026)
CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models
por: Padwal, Vedant
Publicado: (2026)
por: Padwal, Vedant
Publicado: (2026)
Ejemplares similares
-
ReCode: Improving LLM-based Code Repair with Fine-Grained Retrieval-Augmented Generation
por: Zhao, Yicong, et al.
Publicado: (2025) -
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
por: Zhang, Wentao, et al.
Publicado: (2026) -
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
por: Xu, Qiao, et al.
Publicado: (2026) -
NARRepair: Non-Autoregressive Code Generation Model for Automatic Program Repair
por: Yang, Zhenyu, et al.
Publicado: (2024) -
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
por: Qiu, Ruizhong, et al.
Publicado: (2024)