AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing Tests

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Khatib, Lara, Mathews, Noble Saji, Nagappan, Meiyappan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915714779054080
author Khatib, Lara
Mathews, Noble Saji
Nagappan, Meiyappan
author_facet Khatib, Lara
Mathews, Noble Saji
Nagappan, Meiyappan
contents Bug reproduction is critical in the software debugging and repair process, yet the majority of bugs in open-source and industrial settings lack executable tests to reproduce them at the time they are reported, making diagnosis and resolution more difficult and time-consuming. To address this challenge, we introduce AssertFlip, a novel technique for automatically generating Bug Reproducible Tests (BRTs) using large language models (LLMs). Unlike existing methods that attempt direct generation of failing tests, AssertFlip first generates passing tests on the buggy behaviour and then inverts these tests to fail when the bug is present. We hypothesize that LLMs are better at writing passing tests than ones that crash or fail on purpose. Our results show that AssertFlip outperforms all known techniques in the leaderboard of SWT-Bench, a benchmark curated for BRTs. Specifically, AssertFlip achieves a fail-to-pass success rate of 43.6% on the SWT-Bench-Verified subset.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17542
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing Tests
Khatib, Lara
Mathews, Noble Saji
Nagappan, Meiyappan
Software Engineering
Bug reproduction is critical in the software debugging and repair process, yet the majority of bugs in open-source and industrial settings lack executable tests to reproduce them at the time they are reported, making diagnosis and resolution more difficult and time-consuming. To address this challenge, we introduce AssertFlip, a novel technique for automatically generating Bug Reproducible Tests (BRTs) using large language models (LLMs). Unlike existing methods that attempt direct generation of failing tests, AssertFlip first generates passing tests on the buggy behaviour and then inverts these tests to fail when the bug is present. We hypothesize that LLMs are better at writing passing tests than ones that crash or fail on purpose. Our results show that AssertFlip outperforms all known techniques in the leaderboard of SWT-Bench, a benchmark curated for BRTs. Specifically, AssertFlip achieves a fail-to-pass success rate of 43.6% on the SWT-Bench-Verified subset.
title AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing Tests
topic Software Engineering
url https://arxiv.org/abs/2507.17542