ReCatcher: Towards LLMs Regression Testing for Code Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Abbassi, Altaf Allah, Da Silva, Leuson, Nikanjam, Amin, Khomh, Foutse |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Taxonomy of Inefficiencies in LLM-Generated Python Code
by: Abbassi, Altaf Allah, et al.
Published: (2025)
by: Abbassi, Altaf Allah, et al.
Published: (2025)
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
by: Majdinasab, Vahid, et al.
Published: (2025)
by: Majdinasab, Vahid, et al.
Published: (2025)
LLMs and Stack Overflow Discussions: Reliability, Impact, and Challenges
by: Da Silva, Leuson, et al.
Published: (2024)
by: Da Silva, Leuson, et al.
Published: (2024)
DeepCodeProbe: Towards Understanding What Models Trained on Code Learn
by: Majdinasab, Vahid, et al.
Published: (2024)
by: Majdinasab, Vahid, et al.
Published: (2024)
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models
by: Tambon, Florian, et al.
Published: (2024)
by: Tambon, Florian, et al.
Published: (2024)
Continuously Learning Bug Locations
by: Mindom, Paulina Stevia Nouwou, et al.
Published: (2024)
by: Mindom, Paulina Stevia Nouwou, et al.
Published: (2024)
Bugs in Large Language Models Generated Code: An Empirical Study
by: Tambon, Florian, et al.
Published: (2024)
by: Tambon, Florian, et al.
Published: (2024)
Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code
by: Majdinasab, Vahid, et al.
Published: (2024)
by: Majdinasab, Vahid, et al.
Published: (2024)
Refactoring with LLMs: Bridging Human Expertise and Machine Understanding
by: Piao, Yonnel Chen Kuang, et al.
Published: (2025)
by: Piao, Yonnel Chen Kuang, et al.
Published: (2025)
Trimming the Risk: Towards Reliable Continuous Training for Deep Learning Inspection Systems
by: Abbassi, Altaf Allah, et al.
Published: (2024)
by: Abbassi, Altaf Allah, et al.
Published: (2024)
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
by: Ngassom, Sylvain Kouemo, et al.
Published: (2024)
by: Ngassom, Sylvain Kouemo, et al.
Published: (2024)
Toward Debugging Deep Reinforcement Learning Programs with RLExplorer
by: Bouchoucha, Rached, et al.
Published: (2024)
by: Bouchoucha, Rached, et al.
Published: (2024)
An Empirical Study of Policy-as-Code Adoption in Open-Source Software Projects
by: Foalem, Patrick Loic, et al.
Published: (2026)
by: Foalem, Patrick Loic, et al.
Published: (2026)
Exploring Security Practices in Infrastructure as Code: An Empirical Study
by: Verdet, Alexandre, et al.
Published: (2023)
by: Verdet, Alexandre, et al.
Published: (2023)
A Survey of Bugs in AI-Generated Code
by: Gao, Ruofan, et al.
Published: (2025)
by: Gao, Ruofan, et al.
Published: (2025)
Performance Smells in ML and Non-ML Python Projects: A Comparative Study
by: Belias, François, et al.
Published: (2025)
by: Belias, François, et al.
Published: (2025)
Mitigating False Positives in Static Memory Safety Analysis of Rust Programs via Reinforcement Learning
by: P, Akilesh, et al.
Published: (2026)
by: P, Akilesh, et al.
Published: (2026)
PathOCL: Path-Based Prompt Augmentation for OCL Generation with GPT-4
by: Abukhalaf, Seif, et al.
Published: (2024)
by: Abukhalaf, Seif, et al.
Published: (2024)
Fault Localization in Deep Learning-based Software: A System-level Approach
by: Morovati, Mohammad Mehdi, et al.
Published: (2024)
by: Morovati, Mohammad Mehdi, et al.
Published: (2024)
Empirical Characterization of Logging Smells in Machine Learning Code
by: Foalem, Patrick Loic, et al.
Published: (2026)
by: Foalem, Patrick Loic, et al.
Published: (2026)
Empirical Characterization of Logging Smells in Machine Learning Code
by: Foalem, Patrick Loic, et al.
Published: (2026)
by: Foalem, Patrick Loic, et al.
Published: (2026)
RefAgent: A Multi-agent LLM-based Framework for Automatic Software Refactoring
by: Oueslati, Khouloud, et al.
Published: (2025)
by: Oueslati, Khouloud, et al.
Published: (2025)
An Efficient Model Maintenance Approach for MLOps
by: Majidi, Forough, et al.
Published: (2024)
by: Majidi, Forough, et al.
Published: (2024)
Machine Learning Robustness: A Primer
by: Braiek, Houssem Ben, et al.
Published: (2024)
by: Braiek, Houssem Ben, et al.
Published: (2024)
SDLog: A Deep Learning Framework for Detecting Sensitive Information in Software Logs
by: Aghili, Roozbeh, et al.
Published: (2025)
by: Aghili, Roozbeh, et al.
Published: (2025)
Logging Requirement for Continuous Auditing of Responsible Machine Learning-based Applications
by: Foalem, Patrick Loic, et al.
Published: (2025)
by: Foalem, Patrick Loic, et al.
Published: (2025)
Common Challenges of Deep Reinforcement Learning Applications Development: An Empirical Study
by: Morovati, Mohammad Mehdi, et al.
Published: (2023)
by: Morovati, Mohammad Mehdi, et al.
Published: (2023)
What Information Contributes to Log-based Anomaly Detection? Insights from a Configurable Transformer-Based Approach
by: Wu, Xingfang, et al.
Published: (2024)
by: Wu, Xingfang, et al.
Published: (2024)
Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
by: Shah, Mehil B, et al.
Published: (2025)
by: Shah, Mehil B, et al.
Published: (2025)
Adversarial Attack Classification and Robustness Testing for Large Language Models for Code
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Impact of LLM-based Review Comment Generation in Practice: A Mixed Open-/Closed-source User Study
by: Olewicki, Doriane, et al.
Published: (2024)
by: Olewicki, Doriane, et al.
Published: (2024)
BloomAPR: A Bloom's Taxonomy-based Framework for Assessing the Capabilities of LLM-Powered APR Solutions
by: Ma, Yinghang, et al.
Published: (2025)
by: Ma, Yinghang, et al.
Published: (2025)
Mock Deep Testing: Toward Separate Development of Data and Models for Deep Learning
by: Manke, Ruchira, et al.
Published: (2025)
by: Manke, Ruchira, et al.
Published: (2025)
Harnessing the Power of LLMs: Automating Unit Test Generation for High-Performance Computing
by: Karanjai, Rabimba, et al.
Published: (2024)
by: Karanjai, Rabimba, et al.
Published: (2024)
Structural Anchors and Reasoning Fragility:Understanding CoT Robustness in LLM4Code
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Helping LLMs Improve Code Generation Using Feedback from Testing and Static Analysis
by: Dolcetti, Greta, et al.
Published: (2024)
by: Dolcetti, Greta, et al.
Published: (2024)
Quality Issues in Machine Learning Software Systems
by: Côté, Pierre-Olivier, et al.
Published: (2023)
by: Côté, Pierre-Olivier, et al.
Published: (2023)
Klear-CodeTest: Scalable Test Case Generation for Code Reinforcement Learning
by: Fu, Jia, et al.
Published: (2025)
by: Fu, Jia, et al.
Published: (2025)
Can LLM Generate Regression Tests for Software Commits?
by: Liu, Jing, et al.
Published: (2025)
by: Liu, Jing, et al.
Published: (2025)
CoRe: Benchmarking LLMs Code Reasoning Capabilities through Static Analysis Tasks
by: Xie, Danning, et al.
Published: (2025)
by: Xie, Danning, et al.
Published: (2025)
Similar Items
-
A Taxonomy of Inefficiencies in LLM-Generated Python Code
by: Abbassi, Altaf Allah, et al.
Published: (2025) -
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
by: Majdinasab, Vahid, et al.
Published: (2025) -
LLMs and Stack Overflow Discussions: Reliability, Impact, and Challenges
by: Da Silva, Leuson, et al.
Published: (2024) -
DeepCodeProbe: Towards Understanding What Models Trained on Code Learn
by: Majdinasab, Vahid, et al.
Published: (2024) -
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models
by: Tambon, Florian, et al.
Published: (2024)