Probing Network Decisions: Capturing Uncertainties and Unveiling Vulnerabilities Without Label Information

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Joung, Youngju, Lee, Sehyun, Choi, Jaesik
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917953399685120
author Joung, Youngju
Lee, Sehyun
Choi, Jaesik
author_facet Joung, Youngju
Lee, Sehyun
Choi, Jaesik
contents To improve trust and transparency, it is crucial to be able to interpret the decisions of Deep Neural classifiers (DNNs). Instance-level examinations, such as attribution techniques, are commonly employed to interpret the model decisions. However, when interpreting misclassified decisions, human intervention may be required. Analyzing the attribu tions across each class within one instance can be particularly labor intensive and influenced by the bias of the human interpreter. In this paper, we present a novel framework to uncover the weakness of the classifier via counterfactual examples. A prober is introduced to learn the correctness of the classifier's decision in terms of binary code-hit or miss. It enables the creation of the counterfactual example concerning the prober's decision. We test the performance of our prober's misclassification detection and verify its effectiveness on the image classification benchmark datasets. Furthermore, by generating counterfactuals that penetrate the prober, we demonstrate that our framework effectively identifies vulnerabilities in the target classifier without relying on label information on the MNIST dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2503_09068
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Probing Network Decisions: Capturing Uncertainties and Unveiling Vulnerabilities Without Label Information
Joung, Youngju
Lee, Sehyun
Choi, Jaesik
Machine Learning
Artificial Intelligence
Cryptography and Security
To improve trust and transparency, it is crucial to be able to interpret the decisions of Deep Neural classifiers (DNNs). Instance-level examinations, such as attribution techniques, are commonly employed to interpret the model decisions. However, when interpreting misclassified decisions, human intervention may be required. Analyzing the attribu tions across each class within one instance can be particularly labor intensive and influenced by the bias of the human interpreter. In this paper, we present a novel framework to uncover the weakness of the classifier via counterfactual examples. A prober is introduced to learn the correctness of the classifier's decision in terms of binary code-hit or miss. It enables the creation of the counterfactual example concerning the prober's decision. We test the performance of our prober's misclassification detection and verify its effectiveness on the image classification benchmark datasets. Furthermore, by generating counterfactuals that penetrate the prober, we demonstrate that our framework effectively identifies vulnerabilities in the target classifier without relying on label information on the MNIST dataset.
title Probing Network Decisions: Capturing Uncertainties and Unveiling Vulnerabilities Without Label Information
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2503.09068