Deferring Concept Bottleneck Models: Learning to Defer Interventions to Inaccurate Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pugnana, Andrea, Massidda, Riccardo, Giannini, Francesco, Barbiero, Pietro, Zarlenga, Mateo Espinosa, Pellungrini, Roberto, Dominici, Gabriele, Giannotti, Fosca, Bacciu, Davide
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908276105412608
author Pugnana, Andrea
Massidda, Riccardo
Giannini, Francesco
Barbiero, Pietro
Zarlenga, Mateo Espinosa
Pellungrini, Roberto
Dominici, Gabriele
Giannotti, Fosca
Bacciu, Davide
author_facet Pugnana, Andrea
Massidda, Riccardo
Giannini, Francesco
Barbiero, Pietro
Zarlenga, Mateo Espinosa
Pellungrini, Roberto
Dominici, Gabriele
Giannotti, Fosca
Bacciu, Davide
contents Concept Bottleneck Models (CBMs) are machine learning models that improve interpretability by grounding their predictions on human-understandable concepts, allowing for targeted interventions in their decision-making process. However, when intervened on, CBMs assume the availability of humans that can identify the need to intervene and always provide correct interventions. Both assumptions are unrealistic and impractical, considering labor costs and human error-proneness. In contrast, Learning to Defer (L2D) extends supervised learning by allowing machine learning models to identify cases where a human is more likely to be correct than the model, thus leading to deferring systems with improved performance. In this work, we gain inspiration from L2D and propose Deferring CBMs (DCBMs), a novel framework that allows CBMs to learn when an intervention is needed. To this end, we model DCBMs as a composition of deferring systems and derive a consistent L2D loss to train them. Moreover, by relying on a CBM architecture, DCBMs can explain why defer occurs on the final task. Our results show that DCBMs achieve high predictive performance and interpretability at the cost of deferring more to humans.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16199
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deferring Concept Bottleneck Models: Learning to Defer Interventions to Inaccurate Experts
Pugnana, Andrea
Massidda, Riccardo
Giannini, Francesco
Barbiero, Pietro
Zarlenga, Mateo Espinosa
Pellungrini, Roberto
Dominici, Gabriele
Giannotti, Fosca
Bacciu, Davide
Machine Learning
Concept Bottleneck Models (CBMs) are machine learning models that improve interpretability by grounding their predictions on human-understandable concepts, allowing for targeted interventions in their decision-making process. However, when intervened on, CBMs assume the availability of humans that can identify the need to intervene and always provide correct interventions. Both assumptions are unrealistic and impractical, considering labor costs and human error-proneness. In contrast, Learning to Defer (L2D) extends supervised learning by allowing machine learning models to identify cases where a human is more likely to be correct than the model, thus leading to deferring systems with improved performance. In this work, we gain inspiration from L2D and propose Deferring CBMs (DCBMs), a novel framework that allows CBMs to learn when an intervention is needed. To this end, we model DCBMs as a composition of deferring systems and derive a consistent L2D loss to train them. Moreover, by relying on a CBM architecture, DCBMs can explain why defer occurs on the final task. Our results show that DCBMs achieve high predictive performance and interpretability at the cost of deferring more to humans.
title Deferring Concept Bottleneck Models: Learning to Defer Interventions to Inaccurate Experts
topic Machine Learning
url https://arxiv.org/abs/2503.16199