Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.01311 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918182510395392 |
|---|---|
| author | Naudot, Filip Sundqvist, Tobias Kampik, Timotheus |
| author_facet | Naudot, Filip Sundqvist, Tobias Kampik, Timotheus |
| contents | Feature attribution methods help make machine learning-based inference explainable by determining how much one or several features have contributed to a model's output. A particularly popular attribution method is based on the Shapley value from cooperative game theory, a measure that guarantees the satisfaction of several desirable principles, assuming deterministic inference. We apply the Shapley value to feature attribution in large language model (LLM)-based decision support systems, where inference is, by design, stochastic (non-deterministic). We then demonstrate when we can and cannot guarantee Shapley value principle satisfaction across different implementation variants applied to LLM-based decision support, and analyze how the stochastic nature of LLMs affects these guarantees. We also highlight trade-offs between explainable inference speed, agreement with exact Shapley value attributions, and principle attainment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_01311 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | llmSHAP: A Principled Approach to LLM Explainability Naudot, Filip Sundqvist, Tobias Kampik, Timotheus Artificial Intelligence Feature attribution methods help make machine learning-based inference explainable by determining how much one or several features have contributed to a model's output. A particularly popular attribution method is based on the Shapley value from cooperative game theory, a measure that guarantees the satisfaction of several desirable principles, assuming deterministic inference. We apply the Shapley value to feature attribution in large language model (LLM)-based decision support systems, where inference is, by design, stochastic (non-deterministic). We then demonstrate when we can and cannot guarantee Shapley value principle satisfaction across different implementation variants applied to LLM-based decision support, and analyze how the stochastic nature of LLMs affects these guarantees. We also highlight trade-offs between explainable inference speed, agreement with exact Shapley value attributions, and principle attainment. |
| title | llmSHAP: A Principled Approach to LLM Explainability |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2511.01311 |