Grounded Chain-of-Thought for Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Qiong, Yang, Xiangcong, Zhou, Yiyi, Fang, Chenxin, Song, Baiyang, Sun, Xiaoshuai, Ji, Rongrong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908279529013248
author Wu, Qiong
Yang, Xiangcong
Zhou, Yiyi
Fang, Chenxin
Song, Baiyang
Sun, Xiaoshuai
Ji, Rongrong
author_facet Wu, Qiong
Yang, Xiangcong
Zhou, Yiyi
Fang, Chenxin
Song, Baiyang
Sun, Xiaoshuai
Ji, Rongrong
contents Despite great progress, existing multimodal large language models (MLLMs) are prone to visual hallucination, greatly impeding their trustworthy applications. In this paper, we study this problem from the perspective of visual-spatial reasoning, and propose a new learning task for MLLMs, termed Grounded Chain-of-Thought (GCoT). Different from recent visual CoT studies, which focus more on visual knowledge reasoning, GCoT is keen to helping MLLMs to recognize and ground the relevant visual cues step by step, thereby predicting the correct answer with grounding coordinates as the intuitive basis. To facilitate this task, we also carefully design and construct a dataset called multimodal grounded chain-of-thought (MM-GCoT) consisting of 24,022 GCoT examples for 5,033 images. Besides, a comprehensive consistency evaluation system is also introduced, including the metrics of answer accuracy, grounding accuracy and answer-grounding consistency. We further design and conduct a bunch of experiments on 12 advanced MLLMs, and reveal some notable findings: i. most MLLMs performs poorly on the consistency evaluation, indicating obvious visual hallucination; ii. visual hallucination is not directly related to the parameter size and general multimodal performance, i.e., a larger and stronger MLLM is not less affected by this issue. Lastly, we also demonstrate that the proposed dataset can help existing MLLMs to well cultivate their GCoT capability and reduce the inconsistent answering significantly. Moreover, their GCoT can be also generalized to exiting multimodal tasks, such as open-world QA and REC.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12799
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Grounded Chain-of-Thought for Multimodal Large Language Models
Wu, Qiong
Yang, Xiangcong
Zhou, Yiyi
Fang, Chenxin
Song, Baiyang
Sun, Xiaoshuai
Ji, Rongrong
Computer Vision and Pattern Recognition
Multimedia
Despite great progress, existing multimodal large language models (MLLMs) are prone to visual hallucination, greatly impeding their trustworthy applications. In this paper, we study this problem from the perspective of visual-spatial reasoning, and propose a new learning task for MLLMs, termed Grounded Chain-of-Thought (GCoT). Different from recent visual CoT studies, which focus more on visual knowledge reasoning, GCoT is keen to helping MLLMs to recognize and ground the relevant visual cues step by step, thereby predicting the correct answer with grounding coordinates as the intuitive basis. To facilitate this task, we also carefully design and construct a dataset called multimodal grounded chain-of-thought (MM-GCoT) consisting of 24,022 GCoT examples for 5,033 images. Besides, a comprehensive consistency evaluation system is also introduced, including the metrics of answer accuracy, grounding accuracy and answer-grounding consistency. We further design and conduct a bunch of experiments on 12 advanced MLLMs, and reveal some notable findings: i. most MLLMs performs poorly on the consistency evaluation, indicating obvious visual hallucination; ii. visual hallucination is not directly related to the parameter size and general multimodal performance, i.e., a larger and stronger MLLM is not less affected by this issue. Lastly, we also demonstrate that the proposed dataset can help existing MLLMs to well cultivate their GCoT capability and reduce the inconsistent answering significantly. Moreover, their GCoT can be also generalized to exiting multimodal tasks, such as open-world QA and REC.
title Grounded Chain-of-Thought for Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2503.12799