Are Benchmark Tests Strong Enough? Mutation-Guided Diagnosis and Augmentation of Regression Suites
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Chenglin, Xu, Yisen, Wang, Zehao, Tan, Shin Hwei, Tse-Hsun, Chen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Historical Patches to Repair Plans: Outcome-Conditioned Reasoning for Repository-Level Program Repair
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
von: Xu, Yisen, et al.
Veröffentlicht: (2026)
von: Xu, Yisen, et al.
Veröffentlicht: (2026)
Automated Harmfulness Testing for Code Large Language Models
von: Tan, Honghao, et al.
Veröffentlicht: (2025)
von: Tan, Honghao, et al.
Veröffentlicht: (2025)
Guiding ChatGPT to Fix Web UI Tests via Explanation-Consistency Checking
von: Xu, Zhuolin, et al.
Veröffentlicht: (2023)
von: Xu, Zhuolin, et al.
Veröffentlicht: (2023)
Testing Refactoring Engine via Historical Bug Report driven LLM
von: Wang, Haibo, et al.
Veröffentlicht: (2025)
von: Wang, Haibo, et al.
Veröffentlicht: (2025)
What Makes Code Generation Ethically Sourced?
von: Xu, Zhuolin, et al.
Veröffentlicht: (2025)
von: Xu, Zhuolin, et al.
Veröffentlicht: (2025)
On Rank Aggregating Test Prioritizations
von: Mondal, Shouvick, et al.
Veröffentlicht: (2024)
von: Mondal, Shouvick, et al.
Veröffentlicht: (2024)
MANTRA: Enhancing Automated Method-Level Refactoring with Contextual RAG and Multi-Agent LLM Collaboration
von: Xu, Yisen, et al.
Veröffentlicht: (2025)
von: Xu, Yisen, et al.
Veröffentlicht: (2025)
Ethics Testing: Proactive Identification of Generative AI System Harms
von: Tan, Shin Hwei, et al.
Veröffentlicht: (2026)
von: Tan, Shin Hwei, et al.
Veröffentlicht: (2026)
LLM-Guided Issue Generation from Uncovered Code Segments
von: Pressato, Diany, et al.
Veröffentlicht: (2026)
von: Pressato, Diany, et al.
Veröffentlicht: (2026)
Identifying Performance-Sensitive Configurations in Software Systems through Code Analysis with LLM Agents
von: Wang, Zehao, et al.
Veröffentlicht: (2024)
von: Wang, Zehao, et al.
Veröffentlicht: (2024)
CODEPROMPTZIP: Code-specific Prompt Compression for Retrieval-Augmented Generation in Coding Tasks with LMs
von: He, Pengfei, et al.
Veröffentlicht: (2025)
von: He, Pengfei, et al.
Veröffentlicht: (2025)
Structured Safety Auditing for Balancing Code Correctness and Content Safety in LLM-Generated Code
von: Tan, Honghao, et al.
Veröffentlicht: (2026)
von: Tan, Honghao, et al.
Veröffentlicht: (2026)
Towards Structured, State-Aware, and Execution-Grounded Reasoning for Software Engineering Agents
von: Tse-Hsun, et al.
Veröffentlicht: (2026)
von: Tse-Hsun, et al.
Veröffentlicht: (2026)
Efficient Incremental Code Coverage Analysis for Regression Test Suites
von: Wang, Jiale Amber, et al.
Veröffentlicht: (2024)
von: Wang, Jiale Amber, et al.
Veröffentlicht: (2024)
Empowering AIOps: Leveraging Large Language Models for IT Operations Management
von: Vitui, Arthur, et al.
Veröffentlicht: (2025)
von: Vitui, Arthur, et al.
Veröffentlicht: (2025)
Investigating Code Reuse in Software Redesign: A Case Study
von: Zhang, Xiaowen, et al.
Veröffentlicht: (2026)
von: Zhang, Xiaowen, et al.
Veröffentlicht: (2026)
Dissecting Bug Triggers and Failure Modes in Modern Agentic Frameworks: An Empirical Study
von: Zhang, Xiaowen, et al.
Veröffentlicht: (2026)
von: Zhang, Xiaowen, et al.
Veröffentlicht: (2026)
A Survey of Code Review Benchmarks and Evaluation Practices in Pre-LLM and LLM Era
von: Khan, Taufiqul Islam, et al.
Veröffentlicht: (2026)
von: Khan, Taufiqul Islam, et al.
Veröffentlicht: (2026)
CI-Repair-Bench: A Repository-Aware Benchmark for Automated Patch Validation via CI Workflows
von: Muna, Rabeya Khatun, et al.
Veröffentlicht: (2026)
von: Muna, Rabeya Khatun, et al.
Veröffentlicht: (2026)
Screencast-Based Analysis of User-Perceived GUI Responsiveness
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
SLICET5: Static Program Slicing using Language Models with Copy Mechanism and Constrained Decoding
von: He, Pengfei, et al.
Veröffentlicht: (2025)
von: He, Pengfei, et al.
Veröffentlicht: (2025)
An Empirical Study of Refactoring Engine Bugs
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
Regression Test Suite for Payment Switch using jPOS
von: Sardesai, Atharv, et al.
Veröffentlicht: (2022)
von: Sardesai, Atharv, et al.
Veröffentlicht: (2022)
Studying the Impact of Early Test Termination Due to Assertion Failure on Code Coverage and Spectrum-based Fault Localization
von: Uddin, Md. Ashraf, et al.
Veröffentlicht: (2025)
von: Uddin, Md. Ashraf, et al.
Veröffentlicht: (2025)
MobileUPReg: Identifying User-Perceived Performance Regressions in Mobile OS Versions
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
Studying and Benchmarking Large Language Models For Log Level Suggestion
von: Heng, Yi Wen, et al.
Veröffentlicht: (2024)
von: Heng, Yi Wen, et al.
Veröffentlicht: (2024)
PopSweeper: Automatically Detecting and Resolving App-Blocking Pop-Ups to Assist Automated Mobile GUI Testing
von: Guo, Linqiang, et al.
Veröffentlicht: (2024)
von: Guo, Linqiang, et al.
Veröffentlicht: (2024)
Understanding and Detecting Annotation-Induced Faults of Static Analyzers
von: Zhang, Huaien, et al.
Veröffentlicht: (2024)
von: Zhang, Huaien, et al.
Veröffentlicht: (2024)
Mutation-Guided Unit Test Generation with a Large Language Model
von: Wang, Guancheng, et al.
Veröffentlicht: (2025)
von: Wang, Guancheng, et al.
Veröffentlicht: (2025)
Moving beyond Deletions: Program Simplification via Diverse Program Transformations
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
SBEST: Spectrum-Based Fault Localization Without Fault-Triggering Tests
von: Rafi, Md Nakhla, et al.
Veröffentlicht: (2024)
von: Rafi, Md Nakhla, et al.
Veröffentlicht: (2024)
Evaluating the Effectiveness and Efficiency of Demonstration Retrievers in RAG for Coding Tasks
von: He, Pengfei, et al.
Veröffentlicht: (2024)
von: He, Pengfei, et al.
Veröffentlicht: (2024)
MIST-RL: Mutation-based Incremental Suite Testing via Reinforcement Learning
von: Zhu, Sicheng, et al.
Veröffentlicht: (2026)
von: Zhu, Sicheng, et al.
Veröffentlicht: (2026)
Back to the Future! Studying Data Cleanness in Defects4J and its Impact on Fault Localization
von: Rafi, Md Nakhla, et al.
Veröffentlicht: (2023)
von: Rafi, Md Nakhla, et al.
Veröffentlicht: (2023)
An Empirical Study on the Characteristics of Database Access Bugs in Java Applications
von: Liu, Wei, et al.
Veröffentlicht: (2024)
von: Liu, Wei, et al.
Veröffentlicht: (2024)
An Empirical Study of False Negatives and Positives of Static Code Analyzers From the Perspective of Historical Issues
von: Cui, Han, et al.
Veröffentlicht: (2024)
von: Cui, Han, et al.
Veröffentlicht: (2024)
A Comprehensive Framework for Evaluating API-oriented Code Generation in Large Language Models
von: Wu, Yixi, et al.
Veröffentlicht: (2024)
von: Wu, Yixi, et al.
Veröffentlicht: (2024)
Towards a Fault-Injection Benchmarking Suite
von: Wang, Tianhao, et al.
Veröffentlicht: (2024)
von: Wang, Tianhao, et al.
Veröffentlicht: (2024)
Evaluating Software Process Models for Multi-Agent Class-Level Code Generation
von: Shafin, Wasique Islam, et al.
Veröffentlicht: (2025)
von: Shafin, Wasique Islam, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
From Historical Patches to Repair Plans: Outcome-Conditioned Reasoning for Repository-Level Program Repair
von: Li, Chenglin, et al.
Veröffentlicht: (2026) -
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
von: Xu, Yisen, et al.
Veröffentlicht: (2026) -
Automated Harmfulness Testing for Code Large Language Models
von: Tan, Honghao, et al.
Veröffentlicht: (2025) -
Guiding ChatGPT to Fix Web UI Tests via Explanation-Consistency Checking
von: Xu, Zhuolin, et al.
Veröffentlicht: (2023) -
Testing Refactoring Engine via Historical Bug Report driven LLM
von: Wang, Haibo, et al.
Veröffentlicht: (2025)