AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing Tests
Fuente:
arXiv
Saved in:
| Main Authors: | Khatib, Lara, Mathews, Noble Saji, Nagappan, Meiyappan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Test-Driven Development for Code Generation
by: Mathews, Noble Saji, et al.
Published: (2024)
by: Mathews, Noble Saji, et al.
Published: (2024)
Design choices made by LLM-based test generators prevent them from finding bugs
by: Mathews, Noble Saji, et al.
Published: (2024)
by: Mathews, Noble Saji, et al.
Published: (2024)
Is Your Automated Software Engineer Trustworthy?
by: Mathews, Noble Saji, et al.
Published: (2025)
by: Mathews, Noble Saji, et al.
Published: (2025)
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
by: Prathifkumar, Thanosan, et al.
Published: (2025)
by: Prathifkumar, Thanosan, et al.
Published: (2025)
FuzzSlice: Pruning False Positives in Static Analysis Warnings Through Function-Level Fuzzing
by: Murali, Aniruddhan, et al.
Published: (2024)
by: Murali, Aniruddhan, et al.
Published: (2024)
Examining LLMs Ability to Summarize Code Through Mutation-Analysis
by: Khatib, Lara, et al.
Published: (2026)
by: Khatib, Lara, et al.
Published: (2026)
LLbezpeky: Leveraging Large Language Models for Vulnerability Detection
by: Mathews, Noble Saji, et al.
Published: (2024)
by: Mathews, Noble Saji, et al.
Published: (2024)
RLocator: Reinforcement Learning for Bug Localization
by: Chakraborty, Partha, et al.
Published: (2023)
by: Chakraborty, Partha, et al.
Published: (2023)
BLAZE: Cross-Language and Cross-Project Bug Localization via Dynamic Chunking and Hard Example Learning
by: Chakraborty, Partha, et al.
Published: (2024)
by: Chakraborty, Partha, et al.
Published: (2024)
Understanding the Human-LLM Dynamic: A Literature Survey of LLM Use in Programming Tasks
by: Etsenake, Deborah, et al.
Published: (2024)
by: Etsenake, Deborah, et al.
Published: (2024)
Whodunit: Classifying Code as Human Authored or GPT-4 Generated -- A case study on CodeChef problems
by: Idialu, Oseremen Joy, et al.
Published: (2024)
by: Idialu, Oseremen Joy, et al.
Published: (2024)
Aligning Programming Language and Natural Language: Exploring Design Choices in Multi-Modal Transformer-Based Embedding for Bug Localization
by: Chakraborty, Partha, et al.
Published: (2024)
by: Chakraborty, Partha, et al.
Published: (2024)
Is GitHub's Copilot as Bad as Humans at Introducing Vulnerabilities in Code?
by: Asare, Owura, et al.
Published: (2022)
by: Asare, Owura, et al.
Published: (2022)
A User-centered Security Evaluation of Copilot
by: Asare, Owura, et al.
Published: (2023)
by: Asare, Owura, et al.
Published: (2023)
Measuring the Runtime Performance of C++ Code Written by Humans using GitHub Copilot
by: Erhabor, Daniel, et al.
Published: (2023)
by: Erhabor, Daniel, et al.
Published: (2023)
Understanding Bug-Reproducing Tests: A First Empirical Study
by: Hora, Andre, et al.
Published: (2026)
by: Hora, Andre, et al.
Published: (2026)
SAGA: Summarization-Guided Assert Statement Generation
by: Zhang, Yuwei, et al.
Published: (2023)
by: Zhang, Yuwei, et al.
Published: (2023)
GitBug-Java: A Reproducible Benchmark of Recent Java Bugs
by: Silva, André, et al.
Published: (2024)
by: Silva, André, et al.
Published: (2024)
SysPro: Reproducing System-level Concurrency Bugs from Bug Reports
by: Zaman, Tarannum Shaila, et al.
Published: (2026)
by: Zaman, Tarannum Shaila, et al.
Published: (2026)
TreeMind: Automatically Reproducing Android Bug Reports via LLM-empowered Monte Carlo Tree Search
by: Chen, Zhengyu, et al.
Published: (2025)
by: Chen, Zhengyu, et al.
Published: (2025)
Testing Refactoring Engine via Historical Bug Report driven LLM
by: Wang, Haibo, et al.
Published: (2025)
by: Wang, Haibo, et al.
Published: (2025)
GitBug-Actions: Building Reproducible Bug-Fix Benchmarks with GitHub Actions
by: Saavedra, Nuno, et al.
Published: (2023)
by: Saavedra, Nuno, et al.
Published: (2023)
CodeSAM: Source Code Representation Learning by Infusing Self-Attention with Multi-Code-View Graphs
by: Mathai, Alex, et al.
Published: (2024)
by: Mathai, Alex, et al.
Published: (2024)
Uncovering Business Logic Bugs via Semantics-Driven Unit Test Generation
by: Yang, Chen, et al.
Published: (2026)
by: Yang, Chen, et al.
Published: (2026)
AssertCoder: LLM-Based Assertion Generation via Multimodal Specification Extraction
by: Tian, Enyuan, et al.
Published: (2025)
by: Tian, Enyuan, et al.
Published: (2025)
Checker Bug Detection and Repair in Deep Learning Libraries
by: Harzevili, Nima Shiri, et al.
Published: (2024)
by: Harzevili, Nima Shiri, et al.
Published: (2024)
LLM-Powered Test Case Generation for Detecting Bugs in Plausible Programs
by: Liu, Kaibo, et al.
Published: (2024)
by: Liu, Kaibo, et al.
Published: (2024)
LLM-Powered Silent Bug Fuzzing in Deep Learning Libraries via Versatile and Controlled Bug Transfer
by: Zhang, Kunpeng, et al.
Published: (2026)
by: Zhang, Kunpeng, et al.
Published: (2026)
Chat-like Asserts Prediction with the Support of Large Language Model
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
BugForge: Constructing and Utilizing DBMS Bug Repository to Enhance DBMS Testing
by: Li, Dawei, et al.
Published: (2026)
by: Li, Dawei, et al.
Published: (2026)
AddressWatcher: Sanitizer-Based Localization of Memory Leak Fixes
by: Murali, Aniruddhan, et al.
Published: (2024)
by: Murali, Aniruddhan, et al.
Published: (2024)
Finding XPath Bugs in XML Document Processors via Differential Testing
by: Li, Shuxin, et al.
Published: (2024)
by: Li, Shuxin, et al.
Published: (2024)
Enriching Automatic Test Case Generation by Extracting Relevant Test Inputs from Bug Reports
by: Ouédraogo, Wendkûuni C., et al.
Published: (2023)
by: Ouédraogo, Wendkûuni C., et al.
Published: (2023)
AutoAssert 1: A LoRA Fine-Tuned LLM Model for Efficient Automated Assertion Generation
by: Zhong, Yi, et al.
Published: (2025)
by: Zhong, Yi, et al.
Published: (2025)
Issue2Test: Generating Reproducing Test Cases from Issue Reports
by: Nashid, Noor, et al.
Published: (2025)
by: Nashid, Noor, et al.
Published: (2025)
Towards Enhancing the Reproducibility of Deep Learning Bugs: An Empirical Study
by: Shah, Mehil B., et al.
Published: (2024)
by: Shah, Mehil B., et al.
Published: (2024)
How is Testing Related to Single Statement Bugs?
by: Rahman, Habibur, et al.
Published: (2024)
by: Rahman, Habibur, et al.
Published: (2024)
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
by: Liu, Steven, et al.
Published: (2026)
by: Liu, Steven, et al.
Published: (2026)
BugsInPy: A Database of Existing Bugs in Python Programs to Enable Controlled Testing and Debugging Studies
by: Widyasari, Ratnadira, et al.
Published: (2024)
by: Widyasari, Ratnadira, et al.
Published: (2024)
Towards Understanding Barriers and Mitigation Strategies of Software Engineers with Non-traditional Educational and Occupational Backgrounds
by: Barnes, Tavian, et al.
Published: (2022)
by: Barnes, Tavian, et al.
Published: (2022)
Similar Items
-
Test-Driven Development for Code Generation
by: Mathews, Noble Saji, et al.
Published: (2024) -
Design choices made by LLM-based test generators prevent them from finding bugs
by: Mathews, Noble Saji, et al.
Published: (2024) -
Is Your Automated Software Engineer Trustworthy?
by: Mathews, Noble Saji, et al.
Published: (2025) -
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
by: Prathifkumar, Thanosan, et al.
Published: (2025) -
FuzzSlice: Pruning False Positives in Static Analysis Warnings Through Function-Level Fuzzing
by: Murali, Aniruddhan, et al.
Published: (2024)