SPEX: Scaling Feature Interaction Explanations for LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kang, Justin Singh, Butler, Landon, Agarwal, Abhineet, Erginbas, Yigit Efe, Pedarsani, Ramtin, Ramchandran, Kannan, Yu, Bin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915161319669760
author Kang, Justin Singh
Butler, Landon
Agarwal, Abhineet
Erginbas, Yigit Efe
Pedarsani, Ramtin
Ramchandran, Kannan
Yu, Bin
author_facet Kang, Justin Singh
Butler, Landon
Agarwal, Abhineet
Erginbas, Yigit Efe
Pedarsani, Ramtin
Ramchandran, Kannan
Yu, Bin
contents Large language models (LLMs) have revolutionized machine learning due to their ability to capture complex interactions between input features. Popular post-hoc explanation methods like SHAP provide marginal feature attributions, while their extensions to interaction importances only scale to small input lengths ($\approx 20$). We propose Spectral Explainer (SPEX), a model-agnostic interaction attribution algorithm that efficiently scales to large input lengths ($\approx 1000)$. SPEX exploits underlying natural sparsity among interactions -- common in real-world data -- and applies a sparse Fourier transform using a channel decoding algorithm to efficiently identify important interactions. We perform experiments across three difficult long-context datasets that require LLMs to utilize interactions between inputs to complete the task. For large inputs, SPEX outperforms marginal attribution methods by up to 20% in terms of faithfully reconstructing LLM outputs. Further, SPEX successfully identifies key features and interactions that strongly influence model output. For one of our datasets, HotpotQA, SPEX provides interactions that align with human annotations. Finally, we use our model-agnostic approach to generate explanations to demonstrate abstract reasoning in closed-source LLMs (GPT-4o mini) and compositional reasoning in vision-language models.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13870
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SPEX: Scaling Feature Interaction Explanations for LLMs
Kang, Justin Singh
Butler, Landon
Agarwal, Abhineet
Erginbas, Yigit Efe
Pedarsani, Ramtin
Ramchandran, Kannan
Yu, Bin
Machine Learning
Artificial Intelligence
Computation and Language
Information Theory
Large language models (LLMs) have revolutionized machine learning due to their ability to capture complex interactions between input features. Popular post-hoc explanation methods like SHAP provide marginal feature attributions, while their extensions to interaction importances only scale to small input lengths ($\approx 20$). We propose Spectral Explainer (SPEX), a model-agnostic interaction attribution algorithm that efficiently scales to large input lengths ($\approx 1000)$. SPEX exploits underlying natural sparsity among interactions -- common in real-world data -- and applies a sparse Fourier transform using a channel decoding algorithm to efficiently identify important interactions. We perform experiments across three difficult long-context datasets that require LLMs to utilize interactions between inputs to complete the task. For large inputs, SPEX outperforms marginal attribution methods by up to 20% in terms of faithfully reconstructing LLM outputs. Further, SPEX successfully identifies key features and interactions that strongly influence model output. For one of our datasets, HotpotQA, SPEX provides interactions that align with human annotations. Finally, we use our model-agnostic approach to generate explanations to demonstrate abstract reasoning in closed-source LLMs (GPT-4o mini) and compositional reasoning in vision-language models.
title SPEX: Scaling Feature Interaction Explanations for LLMs
topic Machine Learning
Artificial Intelligence
Computation and Language
Information Theory
url https://arxiv.org/abs/2502.13870