Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Rongyuan, Zhang, Jue, Li, Zhuozhao, Lin, Qingwei, Rajmohan, Saravan, Zhang, Dongmei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915944874377216
author Tan, Rongyuan
Zhang, Jue
Li, Zhuozhao
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
author_facet Tan, Rongyuan
Zhang, Jue
Li, Zhuozhao
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
contents Interpretability tools are increasingly used to analyze failures of Large Language Models (LLMs), yet prior work largely focuses on short prompts or toy settings, leaving their behavior on commonly used benchmarks underexplored. To address this gap, we study contrastive, LRP-based attribution as a practical tool for analyzing LLM failures in realistic settings. We formulate failure analysis as \textit{contrastive attribution}, attributing the logit difference between an incorrect output token and a correct alternative to input tokens and internal model states, and introduce an efficient extension that enables construction of cross-layer attribution graphs for long-context inputs. Using this framework, we conduct a systematic empirical study across benchmarks, comparing attribution patterns across datasets, model sizes, and training checkpoints. Our results show that this token-level contrastive attribution can yield informative signals in some failure cases, but is not universally applicable, highlighting both its utility and its limitations for realistic LLM failure analysis. Our code is available at: https://aka.ms/Debug-XAI.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17761
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
Tan, Rongyuan
Zhang, Jue
Li, Zhuozhao
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
Artificial Intelligence
Computation and Language
Interpretability tools are increasingly used to analyze failures of Large Language Models (LLMs), yet prior work largely focuses on short prompts or toy settings, leaving their behavior on commonly used benchmarks underexplored. To address this gap, we study contrastive, LRP-based attribution as a practical tool for analyzing LLM failures in realistic settings. We formulate failure analysis as \textit{contrastive attribution}, attributing the logit difference between an incorrect output token and a correct alternative to input tokens and internal model states, and introduce an efficient extension that enables construction of cross-layer attribution graphs for long-context inputs. Using this framework, we conduct a systematic empirical study across benchmarks, comparing attribution patterns across datasets, model sizes, and training checkpoints. Our results show that this token-level contrastive attribution can yield informative signals in some failure cases, but is not universally applicable, highlighting both its utility and its limitations for realistic LLM failure analysis. Our code is available at: https://aka.ms/Debug-XAI.
title Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.17761