RDI: An adversarial robustness evaluation metric for deep neural networks based on model statistical features

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Jialei, Zuo, Xingquan, Wang, Feiyang, Huang, Hai, Zhang, Tianle
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915304774303744
author Song, Jialei
Zuo, Xingquan
Wang, Feiyang
Huang, Hai
Zhang, Tianle
author_facet Song, Jialei
Zuo, Xingquan
Wang, Feiyang
Huang, Hai
Zhang, Tianle
contents Deep neural networks (DNNs) are highly susceptible to adversarial samples, raising concerns about their reliability in safety-critical tasks. Currently, methods of evaluating adversarial robustness are primarily categorized into attack-based and certified robustness evaluation approaches. The former not only relies on specific attack algorithms but also is highly time-consuming, while the latter due to its analytical nature, is typically difficult to implement for large and complex models. A few studies evaluate model robustness based on the model's decision boundary, but they suffer from low evaluation accuracy. To address the aforementioned issues, we propose a novel adversarial robustness evaluation metric, Robustness Difference Index (RDI), which is based on model statistical features. RDI draws inspiration from clustering evaluation by analyzing the intra-class and inter-class distances of feature vectors separated by the decision boundary to quantify model robustness. It is attack-independent and has high computational efficiency. Experiments show that, RDI demonstrates a stronger correlation with the gold-standard adversarial robustness metric of attack success rate (ASR). The average computation time of RDI is only 1/30 of the evaluation method based on the PGD attack. Our open-source code is available at: https://github.com/BUPTAIOC/RDI.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18556
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RDI: An adversarial robustness evaluation metric for deep neural networks based on model statistical features
Song, Jialei
Zuo, Xingquan
Wang, Feiyang
Huang, Hai
Zhang, Tianle
Machine Learning
Artificial Intelligence
Deep neural networks (DNNs) are highly susceptible to adversarial samples, raising concerns about their reliability in safety-critical tasks. Currently, methods of evaluating adversarial robustness are primarily categorized into attack-based and certified robustness evaluation approaches. The former not only relies on specific attack algorithms but also is highly time-consuming, while the latter due to its analytical nature, is typically difficult to implement for large and complex models. A few studies evaluate model robustness based on the model's decision boundary, but they suffer from low evaluation accuracy. To address the aforementioned issues, we propose a novel adversarial robustness evaluation metric, Robustness Difference Index (RDI), which is based on model statistical features. RDI draws inspiration from clustering evaluation by analyzing the intra-class and inter-class distances of feature vectors separated by the decision boundary to quantify model robustness. It is attack-independent and has high computational efficiency. Experiments show that, RDI demonstrates a stronger correlation with the gold-standard adversarial robustness metric of attack success rate (ASR). The average computation time of RDI is only 1/30 of the evaluation method based on the PGD attack. Our open-source code is available at: https://github.com/BUPTAIOC/RDI.
title RDI: An adversarial robustness evaluation metric for deep neural networks based on model statistical features
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.18556