AutoTestForge: A Multidimensional Automated Testing Framework for Natural Language Processing Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xing, Hengrui, Tian, Cong, Zhao, Liang, Ma, Zhi, Wang, WenSheng, Zhang, Nan, Huang, Chao, Duan, Zhenhua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916645807587328
author Xing, Hengrui
Tian, Cong
Zhao, Liang
Ma, Zhi
Wang, WenSheng
Zhang, Nan
Huang, Chao
Duan, Zhenhua
author_facet Xing, Hengrui
Tian, Cong
Zhao, Liang
Ma, Zhi
Wang, WenSheng
Zhang, Nan
Huang, Chao
Duan, Zhenhua
contents In recent years, the application of behavioral testing in Natural Language Processing (NLP) model evaluation has experienced a remarkable and substantial growth. However, the existing methods continue to be restricted by the requirements for manual labor and the limited scope of capability assessment. To address these limitations, we introduce AutoTestForge, an automated and multidimensional testing framework for NLP models in this paper. Within AutoTestForge, through the utilization of Large Language Models (LLMs) to automatically generate test templates and instantiate them, manual involvement is significantly reduced. Additionally, a mechanism for the validation of test case labels based on differential testing is implemented which makes use of a multi-model voting system to guarantee the quality of test cases. The framework also extends the test suite across three dimensions, taxonomy, fairness, and robustness, offering a comprehensive evaluation of the capabilities of NLP models. This expansion enables a more in-depth and thorough assessment of the models, providing valuable insights into their strengths and weaknesses. A comprehensive evaluation across sentiment analysis (SA) and semantic textual similarity (STS) tasks demonstrates that AutoTestForge consistently outperforms existing datasets and testing tools, achieving higher error detection rates (an average of $30.89\%$ for SA and $34.58\%$ for STS). Moreover, different generation strategies exhibit stable effectiveness, with error detection rates ranging from $29.03\% - 36.82\%$.
format Preprint
id arxiv_https___arxiv_org_abs_2503_05102
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AutoTestForge: A Multidimensional Automated Testing Framework for Natural Language Processing Models
Xing, Hengrui
Tian, Cong
Zhao, Liang
Ma, Zhi
Wang, WenSheng
Zhang, Nan
Huang, Chao
Duan, Zhenhua
Software Engineering
Computation and Language
Cryptography and Security
In recent years, the application of behavioral testing in Natural Language Processing (NLP) model evaluation has experienced a remarkable and substantial growth. However, the existing methods continue to be restricted by the requirements for manual labor and the limited scope of capability assessment. To address these limitations, we introduce AutoTestForge, an automated and multidimensional testing framework for NLP models in this paper. Within AutoTestForge, through the utilization of Large Language Models (LLMs) to automatically generate test templates and instantiate them, manual involvement is significantly reduced. Additionally, a mechanism for the validation of test case labels based on differential testing is implemented which makes use of a multi-model voting system to guarantee the quality of test cases. The framework also extends the test suite across three dimensions, taxonomy, fairness, and robustness, offering a comprehensive evaluation of the capabilities of NLP models. This expansion enables a more in-depth and thorough assessment of the models, providing valuable insights into their strengths and weaknesses. A comprehensive evaluation across sentiment analysis (SA) and semantic textual similarity (STS) tasks demonstrates that AutoTestForge consistently outperforms existing datasets and testing tools, achieving higher error detection rates (an average of $30.89\%$ for SA and $34.58\%$ for STS). Moreover, different generation strategies exhibit stable effectiveness, with error detection rates ranging from $29.03\% - 36.82\%$.
title AutoTestForge: A Multidimensional Automated Testing Framework for Natural Language Processing Models
topic Software Engineering
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2503.05102