Saved in:
Bibliographic Details
Main Authors: Naudot, Filip, Sundqvist, Tobias, Kampik, Timotheus
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.01311
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918182510395392
author Naudot, Filip
Sundqvist, Tobias
Kampik, Timotheus
author_facet Naudot, Filip
Sundqvist, Tobias
Kampik, Timotheus
contents Feature attribution methods help make machine learning-based inference explainable by determining how much one or several features have contributed to a model's output. A particularly popular attribution method is based on the Shapley value from cooperative game theory, a measure that guarantees the satisfaction of several desirable principles, assuming deterministic inference. We apply the Shapley value to feature attribution in large language model (LLM)-based decision support systems, where inference is, by design, stochastic (non-deterministic). We then demonstrate when we can and cannot guarantee Shapley value principle satisfaction across different implementation variants applied to LLM-based decision support, and analyze how the stochastic nature of LLMs affects these guarantees. We also highlight trade-offs between explainable inference speed, agreement with exact Shapley value attributions, and principle attainment.
format Preprint
id arxiv_https___arxiv_org_abs_2511_01311
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle llmSHAP: A Principled Approach to LLM Explainability
Naudot, Filip
Sundqvist, Tobias
Kampik, Timotheus
Artificial Intelligence
Feature attribution methods help make machine learning-based inference explainable by determining how much one or several features have contributed to a model's output. A particularly popular attribution method is based on the Shapley value from cooperative game theory, a measure that guarantees the satisfaction of several desirable principles, assuming deterministic inference. We apply the Shapley value to feature attribution in large language model (LLM)-based decision support systems, where inference is, by design, stochastic (non-deterministic). We then demonstrate when we can and cannot guarantee Shapley value principle satisfaction across different implementation variants applied to LLM-based decision support, and analyze how the stochastic nature of LLMs affects these guarantees. We also highlight trade-offs between explainable inference speed, agreement with exact Shapley value attributions, and principle attainment.
title llmSHAP: A Principled Approach to LLM Explainability
topic Artificial Intelligence
url https://arxiv.org/abs/2511.01311