A Close Look at Decomposition-based XAI-Methods for Transformer Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arras, Leila, Puri, Bruno, Kahardipraja, Patrick, Lapuschkin, Sebastian, Samek, Wojciech
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917933912948736
author Arras, Leila
Puri, Bruno
Kahardipraja, Patrick
Lapuschkin, Sebastian
Samek, Wojciech
author_facet Arras, Leila
Puri, Bruno
Kahardipraja, Patrick
Lapuschkin, Sebastian
Samek, Wojciech
contents Various XAI attribution methods have been recently proposed for the transformer architecture, allowing for insights into the decision-making process of large language models by assigning importance scores to input tokens and intermediate representations. One class of methods that seems very promising in this direction includes decomposition-based approaches, i.e., XAI-methods that redistribute the model's prediction logit through the network, as this value is directly related to the prediction. In the previous literature we note though that two prominent methods of this category, namely ALTI-Logit and LRP, have not yet been analyzed in juxtaposition and hence we propose to close this gap by conducting a careful quantitative evaluation w.r.t. ground truth annotations on a subject-verb agreement task, as well as various qualitative inspections, using BERT, GPT-2 and LLaMA-3 as a testbed. Along the way we compare and extend the ALTI-Logit and LRP methods, including the recently proposed AttnLRP variant, from an algorithmic and implementation perspective. We further incorporate in our benchmark two widely-used gradient-based attribution techniques. Finally, we make our carefullly constructed benchmark dataset for evaluating attributions on language models, as well as our code, publicly available in order to foster evaluation of XAI-methods on a well-defined common ground.
format Preprint
id arxiv_https___arxiv_org_abs_2502_15886
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Close Look at Decomposition-based XAI-Methods for Transformer Language Models
Arras, Leila
Puri, Bruno
Kahardipraja, Patrick
Lapuschkin, Sebastian
Samek, Wojciech
Computation and Language
Various XAI attribution methods have been recently proposed for the transformer architecture, allowing for insights into the decision-making process of large language models by assigning importance scores to input tokens and intermediate representations. One class of methods that seems very promising in this direction includes decomposition-based approaches, i.e., XAI-methods that redistribute the model's prediction logit through the network, as this value is directly related to the prediction. In the previous literature we note though that two prominent methods of this category, namely ALTI-Logit and LRP, have not yet been analyzed in juxtaposition and hence we propose to close this gap by conducting a careful quantitative evaluation w.r.t. ground truth annotations on a subject-verb agreement task, as well as various qualitative inspections, using BERT, GPT-2 and LLaMA-3 as a testbed. Along the way we compare and extend the ALTI-Logit and LRP methods, including the recently proposed AttnLRP variant, from an algorithmic and implementation perspective. We further incorporate in our benchmark two widely-used gradient-based attribution techniques. Finally, we make our carefullly constructed benchmark dataset for evaluating attributions on language models, as well as our code, publicly available in order to foster evaluation of XAI-methods on a well-defined common ground.
title A Close Look at Decomposition-based XAI-Methods for Transformer Language Models
topic Computation and Language
url https://arxiv.org/abs/2502.15886