From Empirical Evaluation to Context-Aware Enhancement: Repairing Regression Errors with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ho, Anh, Le-Cong, Thanh, Le, Bach, Rizkallah, Christine |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
by: Le-Cong, Thanh, et al.
Published: (2024)
by: Le-Cong, Thanh, et al.
Published: (2024)
Memory-Efficient Large Language Models for Program Repair with Semantic-Guided Patch Generation
by: Le-Cong, Thanh, et al.
Published: (2024)
by: Le-Cong, Thanh, et al.
Published: (2024)
Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
by: Le-Cong, Thanh, et al.
Published: (2025)
by: Le-Cong, Thanh, et al.
Published: (2025)
Perish or Flourish? A Holistic Evaluation of Large Language Models for Code Generation in Functional Programming
by: Lang, Nguyet-Anh H., et al.
Published: (2026)
by: Lang, Nguyet-Anh H., et al.
Published: (2026)
Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
by: Singha, Ananya, et al.
Published: (2025)
by: Singha, Ananya, et al.
Published: (2025)
Comparison of Static Application Security Testing Tools and Large Language Models for Repo-level Vulnerability Detection
by: Zhou, Xin, et al.
Published: (2024)
by: Zhou, Xin, et al.
Published: (2024)
Holistic Evaluation of State-of-the-Art LLMs for Code Generation
by: Zhang, Le, et al.
Published: (2025)
by: Zhang, Le, et al.
Published: (2025)
Evaluating the Generalizability of LLMs in Automated Program Repair
by: Li, Fengjie, et al.
Published: (2025)
by: Li, Fengjie, et al.
Published: (2025)
Do LLMs Consider Security? An Empirical Study on Responses to Programming Questions
by: Sajadi, Amirali, et al.
Published: (2025)
by: Sajadi, Amirali, et al.
Published: (2025)
An Empirical Study on Self-correcting Large Language Models for Data Science Code Generation
by: Quoc, Thai Tang, et al.
Published: (2024)
by: Quoc, Thai Tang, et al.
Published: (2024)
Using LLMs in Software Requirements Specifications: An Empirical Evaluation
by: Krishna, Madhava, et al.
Published: (2024)
by: Krishna, Madhava, et al.
Published: (2024)
How Do Semantically Equivalent Code Transformations Impact Membership Inference on LLMs for Code?
by: Yang, Hua, et al.
Published: (2025)
by: Yang, Hua, et al.
Published: (2025)
Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair
by: de-Fitero-Dominguez, David, et al.
Published: (2025)
by: de-Fitero-Dominguez, David, et al.
Published: (2025)
Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
by: Wang, Shufan, et al.
Published: (2025)
by: Wang, Shufan, et al.
Published: (2025)
Multi-Task Program Error Repair and Explanatory Diagnosis
by: Xu, Zhenyu, et al.
Published: (2024)
by: Xu, Zhenyu, et al.
Published: (2024)
An Empirical Study of the Imbalance Issue in Software Vulnerability Detection
by: Guo, Yuejun, et al.
Published: (2026)
by: Guo, Yuejun, et al.
Published: (2026)
RisConFix: LLM-based Automated Repair of Risk-Prone Drone Configurations
by: Han, Liping, et al.
Published: (2025)
by: Han, Liping, et al.
Published: (2025)
An Experience Report on Regression-Free Repair of Deep Neural Network Model
by: Nakagawa, Takao, et al.
Published: (2025)
by: Nakagawa, Takao, et al.
Published: (2025)
Understanding Automated Program Repair Agents Through the Lens of Traceability: An Empirical Study
by: Ceka, Ira, et al.
Published: (2025)
by: Ceka, Ira, et al.
Published: (2025)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Improving MPI Error Detection and Repair with Large Language Models and Bug References
by: Piersall, Scott, et al.
Published: (2026)
by: Piersall, Scott, et al.
Published: (2026)
On the Role of Fault Localization Context for LLM-Based Program Repair
by: Sepidband, Melika, et al.
Published: (2026)
by: Sepidband, Melika, et al.
Published: (2026)
MergeRepair: An Exploratory Study on Merging Task-Specific Adapters in Code LLMs for Automated Program Repair
by: Dehghan, Meghdad, et al.
Published: (2024)
by: Dehghan, Meghdad, et al.
Published: (2024)
IRepair: An Intent-Aware Approach to Repair Data-Driven Errors in Large Language Models
by: Imtiaz, Sayem Mohammad, et al.
Published: (2025)
by: Imtiaz, Sayem Mohammad, et al.
Published: (2025)
On the Impacts of Contexts on Repository-Level Code Generation
by: Hai, Nam Le, et al.
Published: (2024)
by: Hai, Nam Le, et al.
Published: (2024)
HAFixAgent: History-Aware Program Repair Agent
by: Shi, Yu, et al.
Published: (2025)
by: Shi, Yu, et al.
Published: (2025)
ComBench: A Repo-level Real-world Benchmark for Compilation Error Repair
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles
by: Vinh, Nguyen Phu, et al.
Published: (2025)
by: Vinh, Nguyen Phu, et al.
Published: (2025)
An Empirical Study on Capability of Large Language Models in Understanding Code Semantics
by: Nguyen, Thu-Trang, et al.
Published: (2024)
by: Nguyen, Thu-Trang, et al.
Published: (2024)
Evaluating Agent-based Program Repair at Google
by: Rondon, Pat, et al.
Published: (2025)
by: Rondon, Pat, et al.
Published: (2025)
Coding in a Bubble? Evaluating LLMs in Resolving Context Adaptation Bugs During Code Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026)
by: Zhang, Tanghaoran, et al.
Published: (2026)
Identifying Helpful Context for LLM-based Vulnerability Repair: A Preliminary Study
by: Antal, Gábor, et al.
Published: (2025)
by: Antal, Gábor, et al.
Published: (2025)
ReCatcher: Towards LLMs Regression Testing for Code Generation
by: Abbassi, Altaf Allah, et al.
Published: (2025)
by: Abbassi, Altaf Allah, et al.
Published: (2025)
Boosting Source Code Learning with Text-Oriented Data Augmentation: An Empirical Study
by: Dong, Zeming, et al.
Published: (2023)
by: Dong, Zeming, et al.
Published: (2023)
FailureMem: A Failure-Aware Multimodal Framework for Autonomous Software Repair
by: Ma, Ruize, et al.
Published: (2026)
by: Ma, Ruize, et al.
Published: (2026)
When Fine-Tuning LLMs Meets Data Privacy: An Empirical Study of Federated Learning in LLM-Based Program Repair
by: Luo, Wenqiang, et al.
Published: (2024)
by: Luo, Wenqiang, et al.
Published: (2024)
A Study on the Impact of Fault localization Granularity for Repository-Scale Code Repair Tasks
by: Townsend, Joseph, et al.
Published: (2026)
by: Townsend, Joseph, et al.
Published: (2026)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
From Benchmark Data To Applicable Program Repair: An Experience Report
by: Chandramohan, Mahinthan, et al.
Published: (2025)
by: Chandramohan, Mahinthan, et al.
Published: (2025)
An Empirically-grounded tool for Automatic Prompt Linting and Repair: A Case Study on Bias, Vulnerability, and Optimization in Developer Prompts
by: Rzig, Dhia Elhaq, et al.
Published: (2025)
by: Rzig, Dhia Elhaq, et al.
Published: (2025)
Similar Items
-
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
by: Le-Cong, Thanh, et al.
Published: (2024) -
Memory-Efficient Large Language Models for Program Repair with Semantic-Guided Patch Generation
by: Le-Cong, Thanh, et al.
Published: (2024) -
Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
by: Le-Cong, Thanh, et al.
Published: (2025) -
Perish or Flourish? A Holistic Evaluation of Large Language Models for Code Generation in Functional Programming
by: Lang, Nguyet-Anh H., et al.
Published: (2026) -
Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
by: Singha, Ananya, et al.
Published: (2025)