Deep Learning Framework Testing via Model Mutation: How Far Are We?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mu, Yanzhou, Wang, Rong, Zhai, Juan, Fang, Chunrong, Chen, Xiang, Peng, Zhiyuan, Yang, Peiran, Qian, Ruixiang, Yang, Shaoyu, Chen, Zhenyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912467934773248
author Mu, Yanzhou
Wang, Rong
Zhai, Juan
Fang, Chunrong
Chen, Xiang
Peng, Zhiyuan
Yang, Peiran
Qian, Ruixiang
Yang, Shaoyu
Chen, Zhenyu
author_facet Mu, Yanzhou
Wang, Rong
Zhai, Juan
Fang, Chunrong
Chen, Xiang
Peng, Zhiyuan
Yang, Peiran
Qian, Ruixiang
Yang, Shaoyu
Chen, Zhenyu
contents Deep Learning (DL) frameworks are a fundamental component of DL development. Therefore, the detection of DL framework defects is important and challenging. As one of the most widely adopted DL testing techniques, model mutation has recently gained significant attention. In this study, we revisit the defect detection ability of existing mutation-based testing methods and investigate the factors that influence their effectiveness. To begin with, we reviewed existing methods and observed that many of them mutate DL models (e.g., changing their parameters) without any customization, ignoring the unique challenges in framework testing. Another issue with these methods is their limited effectiveness, characterized by a high rate of false positives caused by illegal mutations arising from the use of generic, non-customized mutation operators. Moreover, we tracked the defects identified by these methods and discovered that most of them were ignored by developers. Motivated by these observations, we investigate the effectiveness of existing mutation-based testing methods in detecting important defects that have been authenticated by framework developers. We begin by collecting defect reports from three popular frameworks and classifying them based on framework developers' ratings to build a comprehensive dataset. We then perform an in-depth analysis to uncover valuable insights. Based on our findings, we propose optimization strategies to address the shortcomings of existing approaches. Following these optimizations, we identified seven new defects, four of which were confirmed by developers as high-priority issues, with three resolved. In summary, we identified 39 unique defects across just 23 models, of which 31 were confirmed by developers, and eight have been fixed.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17638
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deep Learning Framework Testing via Model Mutation: How Far Are We?
Mu, Yanzhou
Wang, Rong
Zhai, Juan
Fang, Chunrong
Chen, Xiang
Peng, Zhiyuan
Yang, Peiran
Qian, Ruixiang
Yang, Shaoyu
Chen, Zhenyu
Software Engineering
Deep Learning (DL) frameworks are a fundamental component of DL development. Therefore, the detection of DL framework defects is important and challenging. As one of the most widely adopted DL testing techniques, model mutation has recently gained significant attention. In this study, we revisit the defect detection ability of existing mutation-based testing methods and investigate the factors that influence their effectiveness. To begin with, we reviewed existing methods and observed that many of them mutate DL models (e.g., changing their parameters) without any customization, ignoring the unique challenges in framework testing. Another issue with these methods is their limited effectiveness, characterized by a high rate of false positives caused by illegal mutations arising from the use of generic, non-customized mutation operators. Moreover, we tracked the defects identified by these methods and discovered that most of them were ignored by developers. Motivated by these observations, we investigate the effectiveness of existing mutation-based testing methods in detecting important defects that have been authenticated by framework developers. We begin by collecting defect reports from three popular frameworks and classifying them based on framework developers' ratings to build a comprehensive dataset. We then perform an in-depth analysis to uncover valuable insights. Based on our findings, we propose optimization strategies to address the shortcomings of existing approaches. Following these optimizations, we identified seven new defects, four of which were confirmed by developers as high-priority issues, with three resolved. In summary, we identified 39 unique defects across just 23 models, of which 31 were confirmed by developers, and eight have been fixed.
title Deep Learning Framework Testing via Model Mutation: How Far Are We?
topic Software Engineering
url https://arxiv.org/abs/2506.17638