V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept Tokenizer

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: He, Hangzhou, Zhu, Lei, Zhang, Xinliang, Zeng, Shuang, Chen, Qian, Lu, Yanye
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918067142918144
author He, Hangzhou
Zhu, Lei
Zhang, Xinliang
Zeng, Shuang
Chen, Qian
Lu, Yanye
author_facet He, Hangzhou
Zhu, Lei
Zhang, Xinliang
Zeng, Shuang
Chen, Qian
Lu, Yanye
contents Concept Bottleneck Models (CBMs) offer inherent interpretability by initially translating images into human-comprehensible concepts, followed by a linear combination of these concepts for classification. However, the annotation of concepts for visual recognition tasks requires extensive expert knowledge and labor, constraining the broad adoption of CBMs. Recent approaches have leveraged the knowledge of large language models to construct concept bottlenecks, with multimodal models like CLIP subsequently mapping image features into the concept feature space for classification. Despite this, the concepts produced by language models can be verbose and may introduce non-visual attributes, which hurts accuracy and interpretability. In this study, we investigate to avoid these issues by constructing CBMs directly from multimodal models. To this end, we adopt common words as base concept vocabulary and leverage auxiliary unlabeled images to construct a Vision-to-Concept (V2C) tokenizer that can explicitly quantize images into their most relevant visual concepts, thus creating a vision-oriented concept bottleneck tightly coupled with the multimodal model. This leads to our V2C-CBM which is training efficient and interpretable with high accuracy. Our V2C-CBM has matched or outperformed LLM-supervised CBMs on various visual classification benchmarks, validating the efficacy of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04975
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept Tokenizer
He, Hangzhou
Zhu, Lei
Zhang, Xinliang
Zeng, Shuang
Chen, Qian
Lu, Yanye
Computer Vision and Pattern Recognition
Concept Bottleneck Models (CBMs) offer inherent interpretability by initially translating images into human-comprehensible concepts, followed by a linear combination of these concepts for classification. However, the annotation of concepts for visual recognition tasks requires extensive expert knowledge and labor, constraining the broad adoption of CBMs. Recent approaches have leveraged the knowledge of large language models to construct concept bottlenecks, with multimodal models like CLIP subsequently mapping image features into the concept feature space for classification. Despite this, the concepts produced by language models can be verbose and may introduce non-visual attributes, which hurts accuracy and interpretability. In this study, we investigate to avoid these issues by constructing CBMs directly from multimodal models. To this end, we adopt common words as base concept vocabulary and leverage auxiliary unlabeled images to construct a Vision-to-Concept (V2C) tokenizer that can explicitly quantize images into their most relevant visual concepts, thus creating a vision-oriented concept bottleneck tightly coupled with the multimodal model. This leads to our V2C-CBM which is training efficient and interpretable with high accuracy. Our V2C-CBM has matched or outperformed LLM-supervised CBMs on various visual classification benchmarks, validating the efficacy of our approach.
title V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept Tokenizer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.04975