Did Models Sufficient Learn? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Yannan, Chen, Ruoyu, Zeng, Bin, Wang, Wei, Liu, Shiming, Zhang, Qunli, Hu, Zheng, Wang, Laiyuan, Wang, Yaowei, Cao, Xiaochun
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915620170235904
author Chen, Yannan
Chen, Ruoyu
Zeng, Bin
Wang, Wei
Liu, Shiming
Zhang, Qunli
Hu, Zheng
Wang, Laiyuan
Wang, Yaowei
Cao, Xiaochun
author_facet Chen, Yannan
Chen, Ruoyu
Zeng, Bin
Wang, Wei
Liu, Shiming
Zhang, Qunli
Hu, Zheng
Wang, Laiyuan
Wang, Yaowei
Cao, Xiaochun
contents In current visual model training, models often rely on only limited sufficient causes for their predictions, which makes them sensitive to distribution shifts or the absence of key features. Attribution methods can accurately identify a model's critical regions. However, masking these areas to create counterfactuals often causes the model to misclassify the target, while humans can still easily recognize it. This divergence highlights that the model's learned dependencies may not be sufficiently causal. To address this issue, we propose Subset-Selected Counterfactual Augmentation (SS-CA), which integrates counterfactual explanations directly into the training process for targeted intervention. Building on the subset-selection-based LIMA attribution method, we develop Counterfactual LIMA to identify minimal spatial region sets whose removal can selectively alter model predictions. Leveraging these attributions, we introduce a data augmentation strategy that replaces the identified regions with natural background, and we train the model jointly on both augmented and original samples to mitigate incomplete causal learning. Extensive experiments across multiple ImageNet variants show that SS-CA improves generalization on in-distribution (ID) test data and achieves superior performance on out-of-distribution (OOD) benchmarks such as ImageNet-R and ImageNet-S. Under perturbations including noise, models trained with SS-CA also exhibit enhanced generalization, demonstrating that our approach effectively uses interpretability insights to correct model deficiencies and improve both performance and robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12100
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Did Models Sufficient Learn? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation
Chen, Yannan
Chen, Ruoyu
Zeng, Bin
Wang, Wei
Liu, Shiming
Zhang, Qunli
Hu, Zheng
Wang, Laiyuan
Wang, Yaowei
Cao, Xiaochun
Computer Vision and Pattern Recognition
In current visual model training, models often rely on only limited sufficient causes for their predictions, which makes them sensitive to distribution shifts or the absence of key features. Attribution methods can accurately identify a model's critical regions. However, masking these areas to create counterfactuals often causes the model to misclassify the target, while humans can still easily recognize it. This divergence highlights that the model's learned dependencies may not be sufficiently causal. To address this issue, we propose Subset-Selected Counterfactual Augmentation (SS-CA), which integrates counterfactual explanations directly into the training process for targeted intervention. Building on the subset-selection-based LIMA attribution method, we develop Counterfactual LIMA to identify minimal spatial region sets whose removal can selectively alter model predictions. Leveraging these attributions, we introduce a data augmentation strategy that replaces the identified regions with natural background, and we train the model jointly on both augmented and original samples to mitigate incomplete causal learning. Extensive experiments across multiple ImageNet variants show that SS-CA improves generalization on in-distribution (ID) test data and achieves superior performance on out-of-distribution (OOD) benchmarks such as ImageNet-R and ImageNet-S. Under perturbations including noise, models trained with SS-CA also exhibit enhanced generalization, demonstrating that our approach effectively uses interpretability insights to correct model deficiencies and improve both performance and robustness.
title Did Models Sufficient Learn? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.12100