GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhan, Yufei, Wu, Ziheng, Zhu, Yousong, Xue, Rongkun, Luo, Ruipu, Chen, Zhenghao, Zhang, Can, Li, Yifan, He, Zhentao, Yang, Zheming, Tang, Ming, Qiu, Minghui, Wang, Jinqiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918042789740544
author Zhan, Yufei
Wu, Ziheng
Zhu, Yousong
Xue, Rongkun
Luo, Ruipu
Chen, Zhenghao
Zhang, Can
Li, Yifan
He, Zhentao
Yang, Zheming
Tang, Ming
Qiu, Minghui
Wang, Jinqiao
author_facet Zhan, Yufei
Wu, Ziheng
Zhu, Yousong
Xue, Rongkun
Luo, Ruipu
Chen, Zhenghao
Zhang, Can
Li, Yifan
He, Zhentao
Yang, Zheming
Tang, Ming
Qiu, Minghui
Wang, Jinqiao
contents Despite notable advancements in multimodal reasoning, leading Multimodal Large Language Models (MLLMs) still underperform on vision-centric multimodal reasoning tasks in general scenarios. This shortfall stems from their predominant reliance on logic- and knowledge-based slow thinking strategies, while effective for domains like math and science, fail to integrate visual information effectively during reasoning. Consequently, these models often fail to adequately ground visual cues, resulting in suboptimal performance in tasks that require multiple plausible visual interpretations and inferences. To address this, we present GThinker (General Thinker), a novel reasoning MLLM excelling in multimodal reasoning across general scenarios, mathematics, and science. GThinker introduces Cue-Rethinking, a flexible reasoning pattern that grounds inferences in visual cues and iteratively reinterprets these cues to resolve inconsistencies. Building on this pattern, we further propose a two-stage training pipeline, including pattern-guided cold start and incentive reinforcement learning, designed to enable multimodal reasoning capabilities across domains. Furthermore, to support the training, we construct GThinker-11K, comprising 7K high-quality, iteratively-annotated reasoning paths and 4K curated reinforcement learning samples, filling the data gap toward general multimodal reasoning. Extensive experiments demonstrate that GThinker achieves 81.5% on the challenging comprehensive multimodal reasoning benchmark M$^3$CoT, surpassing the latest O4-mini model. It also shows an average improvement of 2.1% on general scenario multimodal reasoning benchmarks, while maintaining on-par performance in mathematical reasoning compared to counterpart advanced reasoning models. The code, model, and data will be released soon at https://github.com/jefferyZhan/GThinker.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01078
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
Zhan, Yufei
Wu, Ziheng
Zhu, Yousong
Xue, Rongkun
Luo, Ruipu
Chen, Zhenghao
Zhang, Can
Li, Yifan
He, Zhentao
Yang, Zheming
Tang, Ming
Qiu, Minghui
Wang, Jinqiao
Computer Vision and Pattern Recognition
Artificial Intelligence
Despite notable advancements in multimodal reasoning, leading Multimodal Large Language Models (MLLMs) still underperform on vision-centric multimodal reasoning tasks in general scenarios. This shortfall stems from their predominant reliance on logic- and knowledge-based slow thinking strategies, while effective for domains like math and science, fail to integrate visual information effectively during reasoning. Consequently, these models often fail to adequately ground visual cues, resulting in suboptimal performance in tasks that require multiple plausible visual interpretations and inferences. To address this, we present GThinker (General Thinker), a novel reasoning MLLM excelling in multimodal reasoning across general scenarios, mathematics, and science. GThinker introduces Cue-Rethinking, a flexible reasoning pattern that grounds inferences in visual cues and iteratively reinterprets these cues to resolve inconsistencies. Building on this pattern, we further propose a two-stage training pipeline, including pattern-guided cold start and incentive reinforcement learning, designed to enable multimodal reasoning capabilities across domains. Furthermore, to support the training, we construct GThinker-11K, comprising 7K high-quality, iteratively-annotated reasoning paths and 4K curated reinforcement learning samples, filling the data gap toward general multimodal reasoning. Extensive experiments demonstrate that GThinker achieves 81.5% on the challenging comprehensive multimodal reasoning benchmark M$^3$CoT, surpassing the latest O4-mini model. It also shows an average improvement of 2.1% on general scenario multimodal reasoning benchmarks, while maintaining on-par performance in mathematical reasoning compared to counterpart advanced reasoning models. The code, model, and data will be released soon at https://github.com/jefferyZhan/GThinker.
title GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.01078