Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Silva, Inês Oliveira e, Jesus, Sérgio, Perez, Iker, Ribeiro, Rita P., Soares, Carlos, Ferreira, Hugo, Bizarro, Pedro
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917433643630592
author Silva, Inês Oliveira e
Jesus, Sérgio
Perez, Iker
Ribeiro, Rita P.
Soares, Carlos
Ferreira, Hugo
Bizarro, Pedro
author_facet Silva, Inês Oliveira e
Jesus, Sérgio
Perez, Iker
Ribeiro, Rita P.
Soares, Carlos
Ferreira, Hugo
Bizarro, Pedro
contents Shapley values are a cornerstone of explainable AI, yet their proliferation into competing formulations has created a fragmented landscape with little consensus on practical deployment. While theoretical differences are well-documented, evaluation remains reliant on quantitative proxies whose alignment with human utility is unverified. In this work, we use a unified amortized framework to isolate semantic differences between eight Shapley variants under the low-latency constraints of operational risk workflows. We conduct a large-scale empirical evaluation across four risk datasets and a realistic fraud-detection environment involving professional analysts and 3,735 case reviews. Our results reveal a fundamental misalignment: standard quantitative metrics, such as sparsity and faithfulness, are decoupled from human-perceived clarity and decision utility. Furthermore, while no formulation improved objective analyst performance, explanations consistently increased decision confidence, signaling a critical risk of automation bias in high-stakes settings. These findings suggest that current evaluation proxies are insufficient for predicting downstream human impact, and we provide evidence-based guidance for selecting formulations and metrics in operational decision systems.
format Preprint
id arxiv_https___arxiv_org_abs_2604_22662
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings
Silva, Inês Oliveira e
Jesus, Sérgio
Perez, Iker
Ribeiro, Rita P.
Soares, Carlos
Ferreira, Hugo
Bizarro, Pedro
Machine Learning
Artificial Intelligence
Human-Computer Interaction
Shapley values are a cornerstone of explainable AI, yet their proliferation into competing formulations has created a fragmented landscape with little consensus on practical deployment. While theoretical differences are well-documented, evaluation remains reliant on quantitative proxies whose alignment with human utility is unverified. In this work, we use a unified amortized framework to isolate semantic differences between eight Shapley variants under the low-latency constraints of operational risk workflows. We conduct a large-scale empirical evaluation across four risk datasets and a realistic fraud-detection environment involving professional analysts and 3,735 case reviews. Our results reveal a fundamental misalignment: standard quantitative metrics, such as sparsity and faithfulness, are decoupled from human-perceived clarity and decision utility. Furthermore, while no formulation improved objective analyst performance, explanations consistently increased decision confidence, signaling a critical risk of automation bias in high-stakes settings. These findings suggest that current evaluation proxies are insufficient for predicting downstream human impact, and we provide evidence-based guidance for selecting formulations and metrics in operational decision systems.
title Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings
topic Machine Learning
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2604.22662