Saved in:
Bibliographic Details
Main Authors: Su, Qin, Luo, Tie
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.01605
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918367819988992
author Su, Qin
Luo, Tie
author_facet Su, Qin
Luo, Tie
contents Vision Transformers (ViTs) achieve strong performance in visual recognition, yet their decision-making remains difficult to interpret. We propose BiCAM, a bidirectional class activation mapping method that captures both supportive (positive) and suppressive (negative) contributions to model predictions. Unlike prior CAM-based approaches that discard negative signals, BiCAM preserves signed attributions to produce more complete and contrastive explanations. BiCAM further introduces a Positive-to-Negative Ratio (PNR) that summarizes attribution balance and enables lightweight detection of adversarial examples without retraining. Across ImageNet, VOC, and COCO, BiCAM improves localization and faithfulness while remaining computationally efficient. It generalizes to multiple ViT variants, including DeiT and Swin. These results suggest the importance of modeling both supportive and suppressive evidence for interpreting transformer-based vision models.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01605
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle What Helps---and What Hurts: Bidirectional Explanations for Vision Transformers
Su, Qin
Luo, Tie
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Vision Transformers (ViTs) achieve strong performance in visual recognition, yet their decision-making remains difficult to interpret. We propose BiCAM, a bidirectional class activation mapping method that captures both supportive (positive) and suppressive (negative) contributions to model predictions. Unlike prior CAM-based approaches that discard negative signals, BiCAM preserves signed attributions to produce more complete and contrastive explanations. BiCAM further introduces a Positive-to-Negative Ratio (PNR) that summarizes attribution balance and enables lightweight detection of adversarial examples without retraining. Across ImageNet, VOC, and COCO, BiCAM improves localization and faithfulness while remaining computationally efficient. It generalizes to multiple ViT variants, including DeiT and Swin. These results suggest the importance of modeling both supportive and suppressive evidence for interpreting transformer-based vision models.
title What Helps---and What Hurts: Bidirectional Explanations for Vision Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.01605