Transformer Interpretability from Perspective of Attention and Gradient

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cui, Yongjin, Fan, Xiaohui, Chen, Huajun
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913115591933952
author Cui, Yongjin
Fan, Xiaohui
Chen, Huajun
author_facet Cui, Yongjin
Fan, Xiaohui
Chen, Huajun
contents Although researchers' attention is more focused on the performance of Transformer models, the interpretation of Transformer can never be ignored. Gradient is widely utilized in Transformer interpretation. From the perspective of attention and gradient, we conduct an in-depth study of Transformer interpretation and propose a method to achieve it by guiding the gradient direction, or more precisely, the attention direction. The method enables more comprehensive interpretation of feature regions, offers detail interpretation, and helps to better understand Transformer mechanism. Leveraging the difference in how Vision Transformer (ViT) and humans perceive images, we alter the class of an image in a way that is almost imperceptible to the human eye. This class rewriting phenomenon may potentially pose security risks in certain scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11392
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Transformer Interpretability from Perspective of Attention and Gradient
Cui, Yongjin
Fan, Xiaohui
Chen, Huajun
Artificial Intelligence
Although researchers' attention is more focused on the performance of Transformer models, the interpretation of Transformer can never be ignored. Gradient is widely utilized in Transformer interpretation. From the perspective of attention and gradient, we conduct an in-depth study of Transformer interpretation and propose a method to achieve it by guiding the gradient direction, or more precisely, the attention direction. The method enables more comprehensive interpretation of feature regions, offers detail interpretation, and helps to better understand Transformer mechanism. Leveraging the difference in how Vision Transformer (ViT) and humans perceive images, we alter the class of an image in a way that is almost imperceptible to the human eye. This class rewriting phenomenon may potentially pose security risks in certain scenarios.
title Transformer Interpretability from Perspective of Attention and Gradient
topic Artificial Intelligence
url https://arxiv.org/abs/2605.11392