WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jiang, Yiwen, Mehta, Deval, Yan, Siyuan, Shen, Yaling, Wang, Zimu, Ge, Zongyuan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914050741370880
author Jiang, Yiwen
Mehta, Deval
Yan, Siyuan
Shen, Yaling
Wang, Zimu
Ge, Zongyuan
author_facet Jiang, Yiwen
Mehta, Deval
Yan, Siyuan
Shen, Yaling
Wang, Zimu
Ge, Zongyuan
contents Multimodal Large Language Models (MLLMs) have shown promise in visual-textual reasoning, with Multimodal Chain-of-Thought (MCoT) prompting significantly enhancing interpretability. However, existing MCoT methods rely on rationale-rich datasets and largely focus on inter-object reasoning, overlooking the intra-object understanding crucial for image classification. To address this gap, we propose WISE, a Weak-supervision-guided Step-by-step Explanation method that augments any image classification dataset with MCoTs by reformulating the concept-based representations from Concept Bottleneck Models (CBMs) into concise, interpretable reasoning chains under weak supervision. Experiments across ten datasets show that our generated MCoTs not only improve interpretability by 37% but also lead to gains in classification accuracy when used to fine-tune MLLMs. Our work bridges concept-based interpretability and generative MCoT reasoning, providing a generalizable framework for enhancing MLLMs in fine-grained visual understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17740
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification
Jiang, Yiwen
Mehta, Deval
Yan, Siyuan
Shen, Yaling
Wang, Zimu
Ge, Zongyuan
Computer Vision and Pattern Recognition
Computation and Language
Multimodal Large Language Models (MLLMs) have shown promise in visual-textual reasoning, with Multimodal Chain-of-Thought (MCoT) prompting significantly enhancing interpretability. However, existing MCoT methods rely on rationale-rich datasets and largely focus on inter-object reasoning, overlooking the intra-object understanding crucial for image classification. To address this gap, we propose WISE, a Weak-supervision-guided Step-by-step Explanation method that augments any image classification dataset with MCoTs by reformulating the concept-based representations from Concept Bottleneck Models (CBMs) into concise, interpretable reasoning chains under weak supervision. Experiments across ten datasets show that our generated MCoTs not only improve interpretability by 37% but also lead to gains in classification accuracy when used to fine-tune MLLMs. Our work bridges concept-based interpretability and generative MCoT reasoning, providing a generalizable framework for enhancing MLLMs in fine-grained visual understanding.
title WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2509.17740