Shortcut Mitigation via Spurious-Positive Samples

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Le, Phuong Quynh, Schlötterer, Jörg, Sadiya, Sari, Roig, Gemma, Seifert, Christin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910216011907072
author Le, Phuong Quynh
Schlötterer, Jörg
Sadiya, Sari
Roig, Gemma
Seifert, Christin
author_facet Le, Phuong Quynh
Schlötterer, Jörg
Sadiya, Sari
Roig, Gemma
Seifert, Christin
contents Shortcut mitigation strategies commonly rely on training data annotations, group-balanced held-out data or the presence of all groups, i.e., all combinations of (spurious) attributes and classes, in the training data. However, these requirements are rarely met in practice. We instead propose a method for targeted model analysis to identify a small set of instances in which the model relies on spurious attributes. Using that set and following ``this feature should not be used for prediction'' reasoning, we identify highly relevant neurons in an intermediate layer and regularize their impact. This ensures that models learn to depend on informative features rather than being right for the wrong reasons, thereby improving robustness without requiring additional balanced held-out data or annotations.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13340
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Shortcut Mitigation via Spurious-Positive Samples
Le, Phuong Quynh
Schlötterer, Jörg
Sadiya, Sari
Roig, Gemma
Seifert, Christin
Machine Learning
Shortcut mitigation strategies commonly rely on training data annotations, group-balanced held-out data or the presence of all groups, i.e., all combinations of (spurious) attributes and classes, in the training data. However, these requirements are rarely met in practice. We instead propose a method for targeted model analysis to identify a small set of instances in which the model relies on spurious attributes. Using that set and following ``this feature should not be used for prediction'' reasoning, we identify highly relevant neurons in an intermediate layer and regularize their impact. This ensures that models learn to depend on informative features rather than being right for the wrong reasons, thereby improving robustness without requiring additional balanced held-out data or annotations.
title Shortcut Mitigation via Spurious-Positive Samples
topic Machine Learning
url https://arxiv.org/abs/2605.13340