Explaining Domain Shifts in Language: Concept erasing for Interpretable Image Classification

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zeng, Zequn, Su, Yudi, Sun, Jianqiao, Wen, Tiansheng, Zhang, Hao, Wang, Zhengjue, Chen, Bo, Liu, Hongwei, Ma, Jiawei
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916660694220800
author Zeng, Zequn
Su, Yudi
Sun, Jianqiao
Wen, Tiansheng
Zhang, Hao
Wang, Zhengjue
Chen, Bo
Liu, Hongwei
Ma, Jiawei
author_facet Zeng, Zequn
Su, Yudi
Sun, Jianqiao
Wen, Tiansheng
Zhang, Hao
Wang, Zhengjue
Chen, Bo
Liu, Hongwei
Ma, Jiawei
contents Concept-based models can map black-box representations to human-understandable concepts, which makes the decision-making process more transparent and then allows users to understand the reason behind predictions. However, domain-specific concepts often impact the final predictions, which subsequently undermine the model generalization capabilities, and prevent the model from being used in high-stake applications. In this paper, we propose a novel Language-guided Concept-Erasing (LanCE) framework. In particular, we empirically demonstrate that pre-trained vision-language models (VLMs) can approximate distinct visual domain shifts via domain descriptors while prompting large Language Models (LLMs) can easily simulate a wide range of descriptors of unseen visual domains. Then, we introduce a novel plug-in domain descriptor orthogonality (DDO) regularizer to mitigate the impact of these domain-specific concepts on the final predictions. Notably, the DDO regularizer is agnostic to the design of concept-based models and we integrate it into several prevailing models. Through evaluation of domain generalization on four standard benchmarks and three newly introduced benchmarks, we demonstrate that DDO can significantly improve the out-of-distribution (OOD) generalization over the previous state-of-the-art concept-based models.Our code is available at https://github.com/joeyz0z/LanCE.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18483
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Explaining Domain Shifts in Language: Concept erasing for Interpretable Image Classification
Zeng, Zequn
Su, Yudi
Sun, Jianqiao
Wen, Tiansheng
Zhang, Hao
Wang, Zhengjue
Chen, Bo
Liu, Hongwei
Ma, Jiawei
Computer Vision and Pattern Recognition
Concept-based models can map black-box representations to human-understandable concepts, which makes the decision-making process more transparent and then allows users to understand the reason behind predictions. However, domain-specific concepts often impact the final predictions, which subsequently undermine the model generalization capabilities, and prevent the model from being used in high-stake applications. In this paper, we propose a novel Language-guided Concept-Erasing (LanCE) framework. In particular, we empirically demonstrate that pre-trained vision-language models (VLMs) can approximate distinct visual domain shifts via domain descriptors while prompting large Language Models (LLMs) can easily simulate a wide range of descriptors of unseen visual domains. Then, we introduce a novel plug-in domain descriptor orthogonality (DDO) regularizer to mitigate the impact of these domain-specific concepts on the final predictions. Notably, the DDO regularizer is agnostic to the design of concept-based models and we integrate it into several prevailing models. Through evaluation of domain generalization on four standard benchmarks and three newly introduced benchmarks, we demonstrate that DDO can significantly improve the out-of-distribution (OOD) generalization over the previous state-of-the-art concept-based models.Our code is available at https://github.com/joeyz0z/LanCE.
title Explaining Domain Shifts in Language: Concept erasing for Interpretable Image Classification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.18483