LG-CAV: Train Any Concept Activation Vector with Language Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Qihan, Song, Jie, Xue, Mengqi, Zhang, Haofei, Hu, Bingde, Wang, Huiqiong, Jiang, Hao, Wang, Xingen, Song, Mingli
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912072317534208
author Huang, Qihan
Song, Jie
Xue, Mengqi
Zhang, Haofei
Hu, Bingde
Wang, Huiqiong
Jiang, Hao
Wang, Xingen
Song, Mingli
author_facet Huang, Qihan
Song, Jie
Xue, Mengqi
Zhang, Haofei
Hu, Bingde
Wang, Huiqiong
Jiang, Hao
Wang, Xingen
Song, Mingli
contents Concept activation vector (CAV) has attracted broad research interest in explainable AI, by elegantly attributing model predictions to specific concepts. However, the training of CAV often necessitates a large number of high-quality images, which are expensive to curate and thus limited to a predefined set of concepts. To address this issue, we propose Language-Guided CAV (LG-CAV) to harness the abundant concept knowledge within the certain pre-trained vision-language models (e.g., CLIP). This method allows training any CAV without labeled data, by utilizing the corresponding concept descriptions as guidance. To bridge the gap between vision-language model and the target model, we calculate the activation values of concept descriptions on a common pool of images (probe images) with vision-language model and utilize them as language guidance to train the LG-CAV. Furthermore, after training high-quality LG-CAVs related to all the predicted classes in the target model, we propose the activation sample reweighting (ASR), serving as a model correction technique, to improve the performance of the target model in return. Experiments on four datasets across nine architectures demonstrate that LG-CAV achieves significantly superior quality to previous CAV methods given any concept, and our model correction method achieves state-of-the-art performance compared to existing concept-based methods. Our code is available at https://github.com/hqhQAQ/LG-CAV.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10308
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LG-CAV: Train Any Concept Activation Vector with Language Guidance
Huang, Qihan
Song, Jie
Xue, Mengqi
Zhang, Haofei
Hu, Bingde
Wang, Huiqiong
Jiang, Hao
Wang, Xingen
Song, Mingli
Computer Vision and Pattern Recognition
Concept activation vector (CAV) has attracted broad research interest in explainable AI, by elegantly attributing model predictions to specific concepts. However, the training of CAV often necessitates a large number of high-quality images, which are expensive to curate and thus limited to a predefined set of concepts. To address this issue, we propose Language-Guided CAV (LG-CAV) to harness the abundant concept knowledge within the certain pre-trained vision-language models (e.g., CLIP). This method allows training any CAV without labeled data, by utilizing the corresponding concept descriptions as guidance. To bridge the gap between vision-language model and the target model, we calculate the activation values of concept descriptions on a common pool of images (probe images) with vision-language model and utilize them as language guidance to train the LG-CAV. Furthermore, after training high-quality LG-CAVs related to all the predicted classes in the target model, we propose the activation sample reweighting (ASR), serving as a model correction technique, to improve the performance of the target model in return. Experiments on four datasets across nine architectures demonstrate that LG-CAV achieves significantly superior quality to previous CAV methods given any concept, and our model correction method achieves state-of-the-art performance compared to existing concept-based methods. Our code is available at https://github.com/hqhQAQ/LG-CAV.
title LG-CAV: Train Any Concept Activation Vector with Language Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.10308