Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Linhao, Ji, Xinguang, Liu, Yahui, Kong, Fanheng, Sun, Chenxi, Zhang, Jingyuan, Zhang, Hongzhi, W., V., Zhang, Fuzheng, Xiong, Deyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909647350267904
author Yu, Linhao
Ji, Xinguang
Liu, Yahui
Kong, Fanheng
Sun, Chenxi
Zhang, Jingyuan
Zhang, Hongzhi
W., V.
Zhang, Fuzheng
Xiong, Deyi
author_facet Yu, Linhao
Ji, Xinguang
Liu, Yahui
Kong, Fanheng
Sun, Chenxi
Zhang, Jingyuan
Zhang, Hongzhi
W., V.
Zhang, Fuzheng
Xiong, Deyi
contents Video captioning can be used to assess the video understanding capabilities of Multimodal Large Language Models (MLLMs). However, existing benchmarks and evaluation protocols suffer from crucial issues, such as inadequate or homogeneous creation of key points, exorbitant cost of data creation, and limited evaluation scopes. To address these issues, we propose an automatic framework, named AutoCaption, which leverages Monte Carlo Tree Search (MCTS) to construct numerous and diverse descriptive sentences (\textit{i.e.}, key points) that thoroughly represent video content in an iterative way. This iterative captioning strategy enables the continuous enhancement of video details such as actions, objects' attributes, environment details, etc. We apply AutoCaption to curate MCTS-VCB, a fine-grained video caption benchmark covering video details, thereby enabling a comprehensive evaluation of MLLMs on the video captioning task. We evaluate more than 20 open- and closed-source MLLMs of varying sizes on MCTS-VCB. Results show that MCTS-VCB can effectively and comprehensively evaluate the video captioning capability, with Gemini-1.5-Pro achieving the highest F1 score of 71.2. Interestingly, we fine-tune InternVL2.5-8B with the AutoCaption-generated data, which helps the model achieve an overall improvement of 25.0% on MCTS-VCB and 16.3% on DREAM-1K, further demonstrating the effectiveness of AutoCaption. The code and data are available at https://github.com/tjunlp-lab/MCTS-VCB.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11155
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search
Yu, Linhao
Ji, Xinguang
Liu, Yahui
Kong, Fanheng
Sun, Chenxi
Zhang, Jingyuan
Zhang, Hongzhi
W., V.
Zhang, Fuzheng
Xiong, Deyi
Computer Vision and Pattern Recognition
Video captioning can be used to assess the video understanding capabilities of Multimodal Large Language Models (MLLMs). However, existing benchmarks and evaluation protocols suffer from crucial issues, such as inadequate or homogeneous creation of key points, exorbitant cost of data creation, and limited evaluation scopes. To address these issues, we propose an automatic framework, named AutoCaption, which leverages Monte Carlo Tree Search (MCTS) to construct numerous and diverse descriptive sentences (\textit{i.e.}, key points) that thoroughly represent video content in an iterative way. This iterative captioning strategy enables the continuous enhancement of video details such as actions, objects' attributes, environment details, etc. We apply AutoCaption to curate MCTS-VCB, a fine-grained video caption benchmark covering video details, thereby enabling a comprehensive evaluation of MLLMs on the video captioning task. We evaluate more than 20 open- and closed-source MLLMs of varying sizes on MCTS-VCB. Results show that MCTS-VCB can effectively and comprehensively evaluate the video captioning capability, with Gemini-1.5-Pro achieving the highest F1 score of 71.2. Interestingly, we fine-tune InternVL2.5-8B with the AutoCaption-generated data, which helps the model achieve an overall improvement of 25.0% on MCTS-VCB and 16.3% on DREAM-1K, further demonstrating the effectiveness of AutoCaption. The code and data are available at https://github.com/tjunlp-lab/MCTS-VCB.
title Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.11155