Beyond Accuracy: An Empirical Study on Unit Testing in Open-source Deep Learning Projects
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Han, Yu, Sijia, Chen, Chunyang, Turhan, Burak, Zhu, Xiaodong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chat-like Asserts Prediction with the Support of Large Language Model
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Empowering AI to Generate Better AI Code: Guided Generation of Deep Learning Projects with LLMs
by: Xie, Chen, et al.
Published: (2025)
by: Xie, Chen, et al.
Published: (2025)
An Empirical Study of OpenAI API Discussions on Stack Overflow
by: Chen, Xiang, et al.
Published: (2025)
by: Chen, Xiang, et al.
Published: (2025)
Navigating Fairness: Practitioners' Understanding, Challenges, and Strategies in AI/ML Development
by: Pant, Aastha, et al.
Published: (2024)
by: Pant, Aastha, et al.
Published: (2024)
Software Dependencies 2.0: An Empirical Study of Reuse and Integration of Pre-Trained Models in Open-Source Projects
by: Yasmin, Jerin, et al.
Published: (2025)
by: Yasmin, Jerin, et al.
Published: (2025)
Clarifying Semantics of In-Context Examples for Unit Test Generation
by: Yang, Chen, et al.
Published: (2025)
by: Yang, Chen, et al.
Published: (2025)
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
Leveraging Large Language Models for Command Injection Vulnerability Analysis in Python: An Empirical Study on Popular Open-Source Projects
by: Wang, Yuxuan, et al.
Published: (2025)
by: Wang, Yuxuan, et al.
Published: (2025)
Learning to Generate Unit Test via Adversarial Reinforcement Learning
by: Lee, Dongjun, et al.
Published: (2025)
by: Lee, Dongjun, et al.
Published: (2025)
Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study
by: Storhaug, André, et al.
Published: (2024)
by: Storhaug, André, et al.
Published: (2024)
AI-Assisted Unit Test Writing and Test-Driven Code Refactoring: A Case Study
by: Smolic, Ema, et al.
Published: (2026)
by: Smolic, Ema, et al.
Published: (2026)
LLM Test Generation via Iterative Hybrid Program Analysis
by: Gu, Sijia, et al.
Published: (2025)
by: Gu, Sijia, et al.
Published: (2025)
LLM-Based Robustness Testing of Microservice Applications: An Empirical Study
by: Tigulla, Hrushitha Goud, et al.
Published: (2026)
by: Tigulla, Hrushitha Goud, et al.
Published: (2026)
Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
by: Wang, Shufan, et al.
Published: (2025)
by: Wang, Shufan, et al.
Published: (2025)
An Empirical Study of Challenges in Machine Learning Asset Management
by: Zhao, Zhimin, et al.
Published: (2024)
by: Zhao, Zhimin, et al.
Published: (2024)
Beyond Autoregression: An Empirical Study of Diffusion Large Language Models for Code Generation
by: Li, Chengze, et al.
Published: (2025)
by: Li, Chengze, et al.
Published: (2025)
Unit Testing in ASP Revisited: Language and Test-Driven Development Environment
by: Amendola, Giovanni, et al.
Published: (2024)
by: Amendola, Giovanni, et al.
Published: (2024)
EvoC2Rust: A Skeleton-guided Framework for Project-Level C-to-Rust Translation
by: Wang, Chaofan, et al.
Published: (2025)
by: Wang, Chaofan, et al.
Published: (2025)
An Empirical Study of Fault Localisation Techniques for Deep Learning
by: Humbatova, Nargiz, et al.
Published: (2024)
by: Humbatova, Nargiz, et al.
Published: (2024)
Automated Unit Test Case Generation: A Systematic Literature Review
by: Wang, Jason, et al.
Published: (2025)
by: Wang, Jason, et al.
Published: (2025)
Domain Adaptation for Code Model-based Unit Test Case Generation
by: Shin, Jiho, et al.
Published: (2023)
by: Shin, Jiho, et al.
Published: (2023)
Deep Learning Library Testing: Definition, Methods and Challenges
by: Zhang, Xiaoyu, et al.
Published: (2024)
by: Zhang, Xiaoyu, et al.
Published: (2024)
LLMs for Automated Unit Test Generation and Assessment in Java: The AgoneTest Framework
by: Lops, Andrea, et al.
Published: (2025)
by: Lops, Andrea, et al.
Published: (2025)
Understanding the Helpfulness of Stale Bot for Pull-based Development: An Empirical Study of 20 Large Open-Source Projects
by: Khatoonabadi, SayedHassan, et al.
Published: (2023)
by: Khatoonabadi, SayedHassan, et al.
Published: (2023)
ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization
by: Huang, Yixu, et al.
Published: (2026)
by: Huang, Yixu, et al.
Published: (2026)
An Empirical Investigation of Pre-Trained Deep Learning Model Reuse in the Scientific Process
by: Synovic, Nicholas M., et al.
Published: (2026)
by: Synovic, Nicholas M., et al.
Published: (2026)
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
by: Huang, Donghao, et al.
Published: (2025)
by: Huang, Donghao, et al.
Published: (2025)
Beyond Accuracy: LLM Variability in Evidence Screening for Software Engineering SLRs
by: Hida, Gilberto Sussumu, et al.
Published: (2026)
by: Hida, Gilberto Sussumu, et al.
Published: (2026)
Beyond Accuracy: Characterizing Code Comprehension Capabilities in (Large) Language Models
by: Mächtle, Felix, et al.
Published: (2026)
by: Mächtle, Felix, et al.
Published: (2026)
TAM-Eval: Evaluating LLMs for Automated Unit Test Maintenance
by: Bruches, Elena, et al.
Published: (2026)
by: Bruches, Elena, et al.
Published: (2026)
Leveraging GPT-4 for Vulnerability-Witnessing Unit Test Generation
by: Antal, Gábor, et al.
Published: (2025)
by: Antal, Gábor, et al.
Published: (2025)
An Empirical Study of Agent Developer Practices in AI Agent Frameworks
by: Wang, Yanlin, et al.
Published: (2025)
by: Wang, Yanlin, et al.
Published: (2025)
Demystifying Issues, Causes and Solutions in LLM Open-Source Projects
by: Cai, Yangxiao, et al.
Published: (2024)
by: Cai, Yangxiao, et al.
Published: (2024)
Static Program Analysis Guided LLM Based Unit Test Generation
by: Roychowdhury, Sujoy, et al.
Published: (2025)
by: Roychowdhury, Sujoy, et al.
Published: (2025)
SPARC: Scenario Planning and Reasoning for Automated C Unit Test Generation
by: Chowdhury, Jaid Monwar, et al.
Published: (2026)
by: Chowdhury, Jaid Monwar, et al.
Published: (2026)
Leveraging Large Language Models for Enhancing the Understandability of Generated Unit Tests
by: Deljouyi, Amirhossein, et al.
Published: (2024)
by: Deljouyi, Amirhossein, et al.
Published: (2024)
LAUDE: LLM-Assisted Unit Test Generation and Debugging of Hardware DEsigns
by: Nandal, Deeksha, et al.
Published: (2026)
by: Nandal, Deeksha, et al.
Published: (2026)
Rethinking Diversity in Deep Neural Network Testing
by: Wang, Zi, et al.
Published: (2023)
by: Wang, Zi, et al.
Published: (2023)
A System for Automated Unit Test Generation Using Large Language Models and Assessment of Generated Test Suites
by: Lops, Andrea, et al.
Published: (2024)
by: Lops, Andrea, et al.
Published: (2024)
An Empirical Study of AI Techniques in Mobile Applications
by: Li, Yinghua, et al.
Published: (2022)
by: Li, Yinghua, et al.
Published: (2022)
Similar Items
-
Chat-like Asserts Prediction with the Support of Large Language Model
by: Wang, Han, et al.
Published: (2024) -
Empowering AI to Generate Better AI Code: Guided Generation of Deep Learning Projects with LLMs
by: Xie, Chen, et al.
Published: (2025) -
An Empirical Study of OpenAI API Discussions on Stack Overflow
by: Chen, Xiang, et al.
Published: (2025) -
Navigating Fairness: Practitioners' Understanding, Challenges, and Strategies in AI/ML Development
by: Pant, Aastha, et al.
Published: (2024) -
Software Dependencies 2.0: An Empirical Study of Reuse and Integration of Pre-Trained Models in Open-Source Projects
by: Yasmin, Jerin, et al.
Published: (2025)