Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Meiziniu, Li, Dongze, Liu, Jianmeng, Cao, Jialun, Tian, Yongqiang, Cheung, Shing-Chi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
por: Li, Meiziniu, et al.
Publicado: (2022)
por: Li, Meiziniu, et al.
Publicado: (2022)
Demystifying the Silence of Correctness Bugs in PyTorch Compiler
por: Li, Meiziniu, et al.
Publicado: (2026)
por: Li, Meiziniu, et al.
Publicado: (2026)
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
por: Yu, Boxi, et al.
Publicado: (2026)
por: Yu, Boxi, et al.
Publicado: (2026)
SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
por: Ahmed, Syed Yusuf, et al.
Publicado: (2026)
por: Ahmed, Syed Yusuf, et al.
Publicado: (2026)
Automated Generation of Issue-Reproducing Tests by Combining LLMs and Search-Based Testing
por: Kitsios, Konstantinos, et al.
Publicado: (2025)
por: Kitsios, Konstantinos, et al.
Publicado: (2025)
From Untestable to Testable: Metamorphic Testing in the Age of LLMs
por: Terragni, Valerio
Publicado: (2026)
por: Terragni, Valerio
Publicado: (2026)
E-Test: E'er-Improving Test Suites
por: Qiu, Ketai, et al.
Publicado: (2025)
por: Qiu, Ketai, et al.
Publicado: (2025)
Towards Explainable Test Case Prioritisation with Learning-to-Rank Models
por: Ramírez, Aurora, et al.
Publicado: (2024)
por: Ramírez, Aurora, et al.
Publicado: (2024)
Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models
por: More, Riddhi, et al.
Publicado: (2025)
por: More, Riddhi, et al.
Publicado: (2025)
An Analysis of LLM Fine-Tuning and Few-Shot Learning for Flaky Test Detection and Classification
por: More, Riddhi, et al.
Publicado: (2025)
por: More, Riddhi, et al.
Publicado: (2025)
Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate
por: Wu, Jiaqing, et al.
Publicado: (2026)
por: Wu, Jiaqing, et al.
Publicado: (2026)
Review Beats Planning: Dual-Model Interaction Patterns for Code Synthesis
por: Miller, Jan
Publicado: (2026)
por: Miller, Jan
Publicado: (2026)
On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning Implementations
por: Hundal, Rajdeep Singh, et al.
Publicado: (2025)
por: Hundal, Rajdeep Singh, et al.
Publicado: (2025)
MeDeT: Medical Device Digital Twins Creation with Few-shot Meta-learning
por: Sartaj, Hassan, et al.
Publicado: (2024)
por: Sartaj, Hassan, et al.
Publicado: (2024)
PALM: Path-aware LLM-based Test Generation with Comprehension
por: Wu, Yaoxuan, et al.
Publicado: (2025)
por: Wu, Yaoxuan, et al.
Publicado: (2025)
A Comprehensive Study on Large Language Models for Mutation Testing
por: Wang, Bo, et al.
Publicado: (2024)
por: Wang, Bo, et al.
Publicado: (2024)
An LSTM-based Test Selection Method for Self-Driving Cars
por: Güllü, Ali, et al.
Publicado: (2025)
por: Güllü, Ali, et al.
Publicado: (2025)
Multi-modal Traffic Scenario Generation for Autonomous Driving System Testing
por: Tu, Zhi, et al.
Publicado: (2025)
por: Tu, Zhi, et al.
Publicado: (2025)
A Catalog of Transformations to Remove Smells From Natural Language Tests
por: Aranda, Manoel, et al.
Publicado: (2024)
por: Aranda, Manoel, et al.
Publicado: (2024)
CASCADE: Detecting Inconsistencies between Code and Documentation with Automatic Test Generation
por: Kiecker, Tobias, et al.
Publicado: (2026)
por: Kiecker, Tobias, et al.
Publicado: (2026)
Automated Test Generation from Program Documentation Encoded in Code Comments
por: Denaro, Giovanni, et al.
Publicado: (2025)
por: Denaro, Giovanni, et al.
Publicado: (2025)
AgentAssay: Token-Efficient Regression Testing for Non-Deterministic AI Agent Workflows
por: Bhardwaj, Varun Pratap
Publicado: (2026)
por: Bhardwaj, Varun Pratap
Publicado: (2026)
A Systematic Approach for Assessing Large Language Models' Test Case Generation Capability
por: Chang, Hung-Fu, et al.
Publicado: (2025)
por: Chang, Hung-Fu, et al.
Publicado: (2025)
How Do Developers Structure Unit Test Cases? An Empirical Study from the "AAA" Perspective
por: Wei, Chenhao, et al.
Publicado: (2024)
por: Wei, Chenhao, et al.
Publicado: (2024)
Understanding and Detecting Flaky Builds in GitHub Actions
por: Ge, Wenhao, et al.
Publicado: (2026)
por: Ge, Wenhao, et al.
Publicado: (2026)
Tests4Py: A Benchmark for System Testing
por: Smytzek, Marius, et al.
Publicado: (2023)
por: Smytzek, Marius, et al.
Publicado: (2023)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
por: Liu, Shunyu, et al.
Publicado: (2025)
por: Liu, Shunyu, et al.
Publicado: (2025)
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
por: Ahmed, Sheikh Nazib, et al.
Publicado: (2026)
por: Ahmed, Sheikh Nazib, et al.
Publicado: (2026)
Addressing Data Leakage in HumanEval Using Combinatorial Test Design
por: Bradbury, Jeremy S., et al.
Publicado: (2024)
por: Bradbury, Jeremy S., et al.
Publicado: (2024)
VLM-Fuzz: Vision Language Model Assisted Recursive Depth-first Search Exploration for Effective UI Testing of Android Apps
por: Demissie, Biniam Fisseha, et al.
Publicado: (2025)
por: Demissie, Biniam Fisseha, et al.
Publicado: (2025)
Benchmarking Generative AI Models for Deep Learning Test Input Generation
por: Maryam, et al.
Publicado: (2024)
por: Maryam, et al.
Publicado: (2024)
Which Combination of Test Metrics Can Predict Success of a Software Project? A Case Study in a Year-Long Project Course
por: Filipovic, Marina, et al.
Publicado: (2024)
por: Filipovic, Marina, et al.
Publicado: (2024)
LLM Based Input Space Partitioning Testing for Library APIs
por: Li, Jiageng, et al.
Publicado: (2024)
por: Li, Jiageng, et al.
Publicado: (2024)
LSPFuzz: Hunting Bugs in Language Servers
por: Zhu, Hengcheng, et al.
Publicado: (2025)
por: Zhu, Hengcheng, et al.
Publicado: (2025)
Toward Efficient Testing of Graph Neural Networks via Test Input Prioritization
por: Yang, Lichen, et al.
Publicado: (2025)
por: Yang, Lichen, et al.
Publicado: (2025)
VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs
por: Daneshvar, Seyed Shayan, et al.
Publicado: (2024)
por: Daneshvar, Seyed Shayan, et al.
Publicado: (2024)
Flow-of-Action: SOP Enhanced LLM-Based Multi-Agent System for Root Cause Analysis
por: Pei, Changhua, et al.
Publicado: (2025)
por: Pei, Changhua, et al.
Publicado: (2025)
From Feedback to Failure: Automated Android Performance Issue Reproduction
por: Li, Zhengquan, et al.
Publicado: (2025)
por: Li, Zhengquan, et al.
Publicado: (2025)
Synthesizing Performance Constraints for Evaluating and Improving Code Efficiency
por: Yang, Jun, et al.
Publicado: (2025)
por: Yang, Jun, et al.
Publicado: (2025)
Auto-repair without test cases: How LLMs fix compilation errors in large industrial embedded code
por: Fu, Han, et al.
Publicado: (2025)
por: Fu, Han, et al.
Publicado: (2025)
Ejemplares similares
-
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
por: Li, Meiziniu, et al.
Publicado: (2022) -
Demystifying the Silence of Correctness Bugs in PyTorch Compiler
por: Li, Meiziniu, et al.
Publicado: (2026) -
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
por: Yu, Boxi, et al.
Publicado: (2026) -
SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
por: Ahmed, Syed Yusuf, et al.
Publicado: (2026) -
Automated Generation of Issue-Reproducing Tests by Combining LLMs and Search-Based Testing
por: Kitsios, Konstantinos, et al.
Publicado: (2025)