MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Shuhang, Zhang, Zhenrong, Hu, Pengfei, Ma, Jiefeng, Du, Jun, Wang, Qing, Zhang, Jianshu, Liu, Quan, Gao, Jianqing, Ma, Feng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908320448643072
author Liu, Shuhang
Zhang, Zhenrong
Hu, Pengfei
Ma, Jiefeng
Du, Jun
Wang, Qing
Zhang, Jianshu
Liu, Quan
Gao, Jianqing
Ma, Feng
author_facet Liu, Shuhang
Zhang, Zhenrong
Hu, Pengfei
Ma, Jiefeng
Du, Jun
Wang, Qing
Zhang, Jianshu
Liu, Quan
Gao, Jianqing
Ma, Feng
contents Visual language models (VLMs) have demonstrated strong performance across diverse multimodal reasoning tasks but still face challenges such as hallucinations, resulting in incorrect reasoning outcomes. Inspired by recent research on external feedback mechanisms in large language models (LLMs), we propose a multimodal actor-critic framework to enhance VLM reasoning capabilities. Specifically, the actor model generates step-by-step reasoning paths based on image and text inputs, while the critic model evaluates these reasoning paths and provides corrective feedback. The actor model iteratively refines its reasoning based on the feedback until the reasoning outcome is deemed satisfactory by the critic model. To reduce reliance on costly manual annotations, we introduce an automated method for constructing multimodal critique datasets. By leveraging Monte Carlo Tree Search (MCTS), we systematically guide the actor model to explore diverse reasoning paths. To obtain critique data for correcting erroneous reasoning steps, we prompt an annotator model to compare pairs of reasoning paths diverging from a shared ancestor node - one leading to a correct conclusion and the other to an incorrect one. This approach enables us to construct the MMC (MCTS-based Multimodal Critique) dataset, upon which we further develop a comprehensive training and inference pipeline. Extensive experiments conducted on several public benchmark datasets and mainstream VLMs demonstrate that our approach significantly improves the performance of VLM on complex multimodal reasoning tasks, underscoring its effectiveness and wide applicability.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11009
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
Liu, Shuhang
Zhang, Zhenrong
Hu, Pengfei
Ma, Jiefeng
Du, Jun
Wang, Qing
Zhang, Jianshu
Liu, Quan
Gao, Jianqing
Ma, Feng
Multimedia
Visual language models (VLMs) have demonstrated strong performance across diverse multimodal reasoning tasks but still face challenges such as hallucinations, resulting in incorrect reasoning outcomes. Inspired by recent research on external feedback mechanisms in large language models (LLMs), we propose a multimodal actor-critic framework to enhance VLM reasoning capabilities. Specifically, the actor model generates step-by-step reasoning paths based on image and text inputs, while the critic model evaluates these reasoning paths and provides corrective feedback. The actor model iteratively refines its reasoning based on the feedback until the reasoning outcome is deemed satisfactory by the critic model. To reduce reliance on costly manual annotations, we introduce an automated method for constructing multimodal critique datasets. By leveraging Monte Carlo Tree Search (MCTS), we systematically guide the actor model to explore diverse reasoning paths. To obtain critique data for correcting erroneous reasoning steps, we prompt an annotator model to compare pairs of reasoning paths diverging from a shared ancestor node - one leading to a correct conclusion and the other to an incorrect one. This approach enables us to construct the MMC (MCTS-based Multimodal Critique) dataset, upon which we further develop a comprehensive training and inference pipeline. Extensive experiments conducted on several public benchmark datasets and mainstream VLMs demonstrate that our approach significantly improves the performance of VLM on complex multimodal reasoning tasks, underscoring its effectiveness and wide applicability.
title MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
topic Multimedia
url https://arxiv.org/abs/2504.11009