Saved in:
Bibliographic Details
Main Authors: Stradiotti, Luca, Pesenti, Dario, Teso, Stefano, Davis, Jesse
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.12900
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910056744747008
author Stradiotti, Luca
Pesenti, Dario
Teso, Stefano
Davis, Jesse
author_facet Stradiotti, Luca
Pesenti, Dario
Teso, Stefano
Davis, Jesse
contents Learning to Reject (LtR) frameworks allow ML models to abstain from uncertain predictions and promote user trust. However, since current LtR strategies focus solely on predictive performance, they completely neglect explanation quality. Low-quality explanations -- whether they inaccurately reflect the model's reasoning or fail to satisfy users -- can severely compromise trust assessments and induce over-reliance on incorrect predictions. We argue that models should abstain from making a prediction when they cannot offer a satisfactory explanation for it and introduce a framework for learning to reject low-quality explanations (LtX) in which predictors are equipped with a rejector that evaluates the explanation quality. Focusing on popular attribution techniques, we propose REX (REjector of low-quality eXplanations), which learns a rejector from explanation quality labels combining machine-side judgments with explicit human annotations to assess explanation quality. Our empirical evaluation demonstrates that \method outperforms popular LtR strategies and baselines relying on isolated explanation metrics. Finally, to support future research, we publicly release a novel, larger-scale dataset of 1050 human-annotated machine explanations.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12900
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Knowing What You Cannot Explain: Learning to Reject Low-Quality Explanations
Stradiotti, Luca
Pesenti, Dario
Teso, Stefano
Davis, Jesse
Machine Learning
Learning to Reject (LtR) frameworks allow ML models to abstain from uncertain predictions and promote user trust. However, since current LtR strategies focus solely on predictive performance, they completely neglect explanation quality. Low-quality explanations -- whether they inaccurately reflect the model's reasoning or fail to satisfy users -- can severely compromise trust assessments and induce over-reliance on incorrect predictions. We argue that models should abstain from making a prediction when they cannot offer a satisfactory explanation for it and introduce a framework for learning to reject low-quality explanations (LtX) in which predictors are equipped with a rejector that evaluates the explanation quality. Focusing on popular attribution techniques, we propose REX (REjector of low-quality eXplanations), which learns a rejector from explanation quality labels combining machine-side judgments with explicit human annotations to assess explanation quality. Our empirical evaluation demonstrates that \method outperforms popular LtR strategies and baselines relying on isolated explanation metrics. Finally, to support future research, we publicly release a novel, larger-scale dataset of 1050 human-annotated machine explanations.
title Knowing What You Cannot Explain: Learning to Reject Low-Quality Explanations
topic Machine Learning
url https://arxiv.org/abs/2507.12900