Design choices made by LLM-based test generators prevent them from finding bugs
Fuente:
arXiv
Saved in:
| Main Authors: | Mathews, Noble Saji, Nagappan, Meiyappan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Test-Driven Development for Code Generation
by: Mathews, Noble Saji, et al.
Published: (2024)
by: Mathews, Noble Saji, et al.
Published: (2024)
Is Your Automated Software Engineer Trustworthy?
by: Mathews, Noble Saji, et al.
Published: (2025)
by: Mathews, Noble Saji, et al.
Published: (2025)
AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing Tests
by: Khatib, Lara, et al.
Published: (2025)
by: Khatib, Lara, et al.
Published: (2025)
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
by: Prathifkumar, Thanosan, et al.
Published: (2025)
by: Prathifkumar, Thanosan, et al.
Published: (2025)
LLbezpeky: Leveraging Large Language Models for Vulnerability Detection
by: Mathews, Noble Saji, et al.
Published: (2024)
by: Mathews, Noble Saji, et al.
Published: (2024)
FuzzSlice: Pruning False Positives in Static Analysis Warnings Through Function-Level Fuzzing
by: Murali, Aniruddhan, et al.
Published: (2024)
by: Murali, Aniruddhan, et al.
Published: (2024)
RLocator: Reinforcement Learning for Bug Localization
by: Chakraborty, Partha, et al.
Published: (2023)
by: Chakraborty, Partha, et al.
Published: (2023)
Aligning Programming Language and Natural Language: Exploring Design Choices in Multi-Modal Transformer-Based Embedding for Bug Localization
by: Chakraborty, Partha, et al.
Published: (2024)
by: Chakraborty, Partha, et al.
Published: (2024)
Understanding the Human-LLM Dynamic: A Literature Survey of LLM Use in Programming Tasks
by: Etsenake, Deborah, et al.
Published: (2024)
by: Etsenake, Deborah, et al.
Published: (2024)
Whodunit: Classifying Code as Human Authored or GPT-4 Generated -- A case study on CodeChef problems
by: Idialu, Oseremen Joy, et al.
Published: (2024)
by: Idialu, Oseremen Joy, et al.
Published: (2024)
Is GitHub's Copilot as Bad as Humans at Introducing Vulnerabilities in Code?
by: Asare, Owura, et al.
Published: (2022)
by: Asare, Owura, et al.
Published: (2022)
A User-centered Security Evaluation of Copilot
by: Asare, Owura, et al.
Published: (2023)
by: Asare, Owura, et al.
Published: (2023)
Examining LLMs Ability to Summarize Code Through Mutation-Analysis
by: Khatib, Lara, et al.
Published: (2026)
by: Khatib, Lara, et al.
Published: (2026)
GenAI-based test case generation and execution in SDV platform
by: Zyberaj, Denesa, et al.
Published: (2025)
by: Zyberaj, Denesa, et al.
Published: (2025)
Do AI models help produce verified bug fixes?
by: Huang, Li, et al.
Published: (2025)
by: Huang, Li, et al.
Published: (2025)
Measuring the Runtime Performance of C++ Code Written by Humans using GitHub Copilot
by: Erhabor, Daniel, et al.
Published: (2023)
by: Erhabor, Daniel, et al.
Published: (2023)
BLAZE: Cross-Language and Cross-Project Bug Localization via Dynamic Chunking and Hard Example Learning
by: Chakraborty, Partha, et al.
Published: (2024)
by: Chakraborty, Partha, et al.
Published: (2024)
Revisiting the Performance of Deep Learning-Based Vulnerability Detection on Realistic Datasets
by: Chakraborty, Partha, et al.
Published: (2024)
by: Chakraborty, Partha, et al.
Published: (2024)
Empirical evaluation of LLMs in predicting fixes of Configuration bugs in Smart Home System
by: Monisha, Sheikh Moonwara Anjum, et al.
Published: (2025)
by: Monisha, Sheikh Moonwara Anjum, et al.
Published: (2025)
Canonical Intermediate Representation for LLM-based optimization problem formulation and code generation
by: Lyu, Zhongyuan, et al.
Published: (2026)
by: Lyu, Zhongyuan, et al.
Published: (2026)
Towards LLM-generated explanations for Component-based Knowledge Graph Question Answering Systems
by: Schiese, Dennis, et al.
Published: (2025)
by: Schiese, Dennis, et al.
Published: (2025)
An empirical study of LoRA-based fine-tuning of large language models for automated test case generation
by: Moradi, Milad, et al.
Published: (2026)
by: Moradi, Milad, et al.
Published: (2026)
Designing LLM-based Multi-Agent Systems for Software Engineering Tasks: Quality Attributes, Design Patterns and Rationale
by: Cai, Yangxiao, et al.
Published: (2025)
by: Cai, Yangxiao, et al.
Published: (2025)
CMSA algorithm for solving the prioritized pairwise test data generation problem in software product lines
by: Ferrer, Javier, et al.
Published: (2024)
by: Ferrer, Javier, et al.
Published: (2024)
Assessing LLM code generation quality through path planning tasks
by: Chen, Wanyi, et al.
Published: (2025)
by: Chen, Wanyi, et al.
Published: (2025)
Quality Assurance of LLM-generated Code: Addressing Non-Functional Quality Characteristics
by: Sun, Xin, et al.
Published: (2025)
by: Sun, Xin, et al.
Published: (2025)
AI-assisted Code Authoring at Scale: Fine-tuning, deploying, and mixed methods evaluation
by: Murali, Vijayaraghavan, et al.
Published: (2023)
by: Murali, Vijayaraghavan, et al.
Published: (2023)
Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution
by: Yu, Kai, et al.
Published: (2026)
by: Yu, Kai, et al.
Published: (2026)
Adopting RAG for LLM-Aided Future Vehicle Design
by: Zolfaghari, Vahid, et al.
Published: (2024)
by: Zolfaghari, Vahid, et al.
Published: (2024)
Designing Adaptive Digital Nudging Systems with LLM-Driven Reasoning
by: Santilli, Tiziano, et al.
Published: (2026)
by: Santilli, Tiziano, et al.
Published: (2026)
LLM-Empowered Functional Safety and Security by Design in Automotive Systems
by: Petrovic, Nenad, et al.
Published: (2026)
by: Petrovic, Nenad, et al.
Published: (2026)
DeepSample: DNN sampling-based testing for operational accuracy assessment
by: Guerriero, Antonio, et al.
Published: (2024)
by: Guerriero, Antonio, et al.
Published: (2024)
Evaluating the effectiveness of LLM-based interoperability
by: Falcão, Rodrigo, et al.
Published: (2025)
by: Falcão, Rodrigo, et al.
Published: (2025)
LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile Apps
by: Zhao, Shanhui, et al.
Published: (2025)
by: Zhao, Shanhui, et al.
Published: (2025)
Uncertainty Quantification for LLM-based Code Generation
by: Xu, Senrong, et al.
Published: (2026)
by: Xu, Senrong, et al.
Published: (2026)
LLM-based Iterative Approach to Metamodeling in Automotive
by: Petrovic, Nenad, et al.
Published: (2025)
by: Petrovic, Nenad, et al.
Published: (2025)
NeSy is alive and well: A LLM-driven symbolic approach for better code comment data generation and classification
by: Akl, Hanna Abi
Published: (2024)
by: Akl, Hanna Abi
Published: (2024)
The Tyranny of Possibilities in the Design of Task-Oriented LLM Systems: A Scoping Survey
by: Dhamani, Dhruv, et al.
Published: (2023)
by: Dhamani, Dhruv, et al.
Published: (2023)
Bias Testing and Mitigation in LLM-based Code Generation
by: Huang, Dong, et al.
Published: (2023)
by: Huang, Dong, et al.
Published: (2023)
Evolving Excellence: Automated Optimization of LLM-based Agents
by: Brookes, Paul, et al.
Published: (2025)
by: Brookes, Paul, et al.
Published: (2025)
Similar Items
-
Test-Driven Development for Code Generation
by: Mathews, Noble Saji, et al.
Published: (2024) -
Is Your Automated Software Engineer Trustworthy?
by: Mathews, Noble Saji, et al.
Published: (2025) -
AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing Tests
by: Khatib, Lara, et al.
Published: (2025) -
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
by: Prathifkumar, Thanosan, et al.
Published: (2025) -
LLbezpeky: Leveraging Large Language Models for Vulnerability Detection
by: Mathews, Noble Saji, et al.
Published: (2024)