Deep Learning Framework Testing via Heuristic Guidance Based on Multiple Model Measurements

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zou, Yinglong, Zhai, Juan, Fang, Chunrong, Mu, Yanzhou, Liu, Jiawei, Chen, Zhenyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912661923430400
author Zou, Yinglong
Zhai, Juan
Fang, Chunrong
Mu, Yanzhou
Liu, Jiawei
Chen, Zhenyu
author_facet Zou, Yinglong
Zhai, Juan
Fang, Chunrong
Mu, Yanzhou
Liu, Jiawei
Chen, Zhenyu
contents Deep learning frameworks serve as the foundation for developing and deploying deep learning applications. To enhance the quality of deep learning frameworks, researchers have proposed numerous testing methods using deep learning models as test inputs. However, existing methods predominantly measure model bug detection effectiveness as heuristic indicators, presenting three critical limitations. Firstly, existing methods fail to quantitatively measure model's operator combination variety, potentially missing critical operator combinations that could trigger framework bugs. Secondly, existing methods neglect measuring and heuristically guiding the model execution time, resulting in the omission of numerous models potential for detecting more framework bugs within limited testing time. Thirdly, existing methods overlook correlation between different model measurements, relying simply on single-indicator heuristic guidance without considering their trade-offs. To overcome these limitations, we propose DLMMM, the first deep learning framework testing method to include multiple model measurements into heuristic guidance and fuse these measurements to achieve their trade-offs. DLMMM firstly quantitatively measures model's bug detection performance, operator combination variety, and model execution time. After that, DLMMM fuses these measurements based on their correlation to achieve their trade-offs. To further enhance testing effectiveness, DLMMM designs multi-level heuristic guidance for test input model generation. We apply DLMMM to test three widely used deep learning frameworks (including TensorFlow, PyTorch, and MindSpore). The experimental results show that DLMMM outperforms state-of-the-art methods in effectiveness and efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15181
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deep Learning Framework Testing via Heuristic Guidance Based on Multiple Model Measurements
Zou, Yinglong
Zhai, Juan
Fang, Chunrong
Mu, Yanzhou
Liu, Jiawei
Chen, Zhenyu
Software Engineering
Deep learning frameworks serve as the foundation for developing and deploying deep learning applications. To enhance the quality of deep learning frameworks, researchers have proposed numerous testing methods using deep learning models as test inputs. However, existing methods predominantly measure model bug detection effectiveness as heuristic indicators, presenting three critical limitations. Firstly, existing methods fail to quantitatively measure model's operator combination variety, potentially missing critical operator combinations that could trigger framework bugs. Secondly, existing methods neglect measuring and heuristically guiding the model execution time, resulting in the omission of numerous models potential for detecting more framework bugs within limited testing time. Thirdly, existing methods overlook correlation between different model measurements, relying simply on single-indicator heuristic guidance without considering their trade-offs. To overcome these limitations, we propose DLMMM, the first deep learning framework testing method to include multiple model measurements into heuristic guidance and fuse these measurements to achieve their trade-offs. DLMMM firstly quantitatively measures model's bug detection performance, operator combination variety, and model execution time. After that, DLMMM fuses these measurements based on their correlation to achieve their trade-offs. To further enhance testing effectiveness, DLMMM designs multi-level heuristic guidance for test input model generation. We apply DLMMM to test three widely used deep learning frameworks (including TensorFlow, PyTorch, and MindSpore). The experimental results show that DLMMM outperforms state-of-the-art methods in effectiveness and efficiency.
title Deep Learning Framework Testing via Heuristic Guidance Based on Multiple Model Measurements
topic Software Engineering
url https://arxiv.org/abs/2507.15181