Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909720130879488 |
|---|---|
| author | Zarlenga, Mateo Espinosa Dominici, Gabriele Barbiero, Pietro Shams, Zohreh Jamnik, Mateja |
| author_facet | Zarlenga, Mateo Espinosa Dominici, Gabriele Barbiero, Pietro Shams, Zohreh Jamnik, Mateja |
| contents | In this paper, we investigate how concept-based models (CMs) respond to out-of-distribution (OOD) inputs. CMs are interpretable neural architectures that first predict a set of high-level concepts (e.g., stripes, black) and then predict a task label from those concepts. In particular, we study the impact of concept interventions (i.e., operations where a human expert corrects a CM's mispredicted concepts at test time) on CMs' task predictions when inputs are OOD. Our analysis reveals a weakness in current state-of-the-art CMs, which we term leakage poisoning, that prevents them from properly improving their accuracy when intervened on for OOD inputs. To address this, we introduce MixCEM, a new CM that learns to dynamically exploit leaked information missing from its concepts only when this information is in-distribution. Our results across tasks with and without complete sets of concept annotations demonstrate that MixCEMs outperform strong baselines by significantly improving their accuracy for both in-distribution and OOD samples in the presence and absence of concept interventions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_17921 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts Zarlenga, Mateo Espinosa Dominici, Gabriele Barbiero, Pietro Shams, Zohreh Jamnik, Mateja Machine Learning Artificial Intelligence Cryptography and Security Human-Computer Interaction In this paper, we investigate how concept-based models (CMs) respond to out-of-distribution (OOD) inputs. CMs are interpretable neural architectures that first predict a set of high-level concepts (e.g., stripes, black) and then predict a task label from those concepts. In particular, we study the impact of concept interventions (i.e., operations where a human expert corrects a CM's mispredicted concepts at test time) on CMs' task predictions when inputs are OOD. Our analysis reveals a weakness in current state-of-the-art CMs, which we term leakage poisoning, that prevents them from properly improving their accuracy when intervened on for OOD inputs. To address this, we introduce MixCEM, a new CM that learns to dynamically exploit leaked information missing from its concepts only when this information is in-distribution. Our results across tasks with and without complete sets of concept annotations demonstrate that MixCEMs outperform strong baselines by significantly improving their accuracy for both in-distribution and OOD samples in the presence and absence of concept interventions. |
| title | Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts |
| topic | Machine Learning Artificial Intelligence Cryptography and Security Human-Computer Interaction |
| url | https://arxiv.org/abs/2504.17921 |