ShortcutProbe: Probing Prediction Shortcuts for Learning Robust Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Guangtao, Ye, Wenqian, Zhang, Aidong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908410976403456
author Zheng, Guangtao
Ye, Wenqian
Zhang, Aidong
author_facet Zheng, Guangtao
Ye, Wenqian
Zhang, Aidong
contents Deep learning models often achieve high performance by inadvertently learning spurious correlations between targets and non-essential features. For example, an image classifier may identify an object via its background that spuriously correlates with it. This prediction behavior, known as spurious bias, severely degrades model performance on data that lacks the learned spurious correlations. Existing methods on spurious bias mitigation typically require a variety of data groups with spurious correlation annotations called group labels. However, group labels require costly human annotations and often fail to capture subtle spurious biases such as relying on specific pixels for predictions. In this paper, we propose a novel post hoc spurious bias mitigation framework without requiring group labels. Our framework, termed ShortcutProbe, identifies prediction shortcuts that reflect potential non-robustness in predictions in a given model's latent space. The model is then retrained to be invariant to the identified prediction shortcuts for improved robustness. We theoretically analyze the effectiveness of the framework and empirically demonstrate that it is an efficient and practical tool for improving a model's robustness to spurious bias on diverse datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13910
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ShortcutProbe: Probing Prediction Shortcuts for Learning Robust Models
Zheng, Guangtao
Ye, Wenqian
Zhang, Aidong
Machine Learning
Deep learning models often achieve high performance by inadvertently learning spurious correlations between targets and non-essential features. For example, an image classifier may identify an object via its background that spuriously correlates with it. This prediction behavior, known as spurious bias, severely degrades model performance on data that lacks the learned spurious correlations. Existing methods on spurious bias mitigation typically require a variety of data groups with spurious correlation annotations called group labels. However, group labels require costly human annotations and often fail to capture subtle spurious biases such as relying on specific pixels for predictions. In this paper, we propose a novel post hoc spurious bias mitigation framework without requiring group labels. Our framework, termed ShortcutProbe, identifies prediction shortcuts that reflect potential non-robustness in predictions in a given model's latent space. The model is then retrained to be invariant to the identified prediction shortcuts for improved robustness. We theoretically analyze the effectiveness of the framework and empirically demonstrate that it is an efficient and practical tool for improving a model's robustness to spurious bias on diverse datasets.
title ShortcutProbe: Probing Prediction Shortcuts for Learning Robust Models
topic Machine Learning
url https://arxiv.org/abs/2505.13910