ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Xuecheng, Liu, Jiaxing, Huang, Danlei, Wang, Yifan, Shi, Yunyun, Chen, Kedi, Xue, Junxiao, Liu, Yang, Chen, Chunlin, Dong, Hairong, Yang, Dingkang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918264836194304
author Wu, Xuecheng
Liu, Jiaxing
Huang, Danlei
Wang, Yifan
Shi, Yunyun
Chen, Kedi
Xue, Junxiao
Liu, Yang
Chen, Chunlin
Dong, Hairong
Yang, Dingkang
author_facet Wu, Xuecheng
Liu, Jiaxing
Huang, Danlei
Wang, Yifan
Shi, Yunyun
Chen, Kedi
Xue, Junxiao
Liu, Yang
Chen, Chunlin
Dong, Hairong
Yang, Dingkang
contents Visual-Interleaved Chain-of-Thought (VI-CoT) enables Multi-modal Large Language Models (MLLMs) to continually update their understanding and decision space based on step-wise intermediate visual states (IVS), much like a human would, which has demonstrated impressive success in various tasks, thereby leading to emerged advancements in related downstream benchmarks. Despite promising progress, current benchmarks provide models with relatively fixed IVS, rather than free-style IVS, whch might forcibly distort the original thinking trajectories, failing to evaluate their intrinsic reasoning capabilities. More importantly, existing benchmarks neglect to systematically explore the impact factors that IVS would impart to the untamed reasoning performance. To tackle above gaps, we introduce a specialized benchmark termed ViC-Bench, consisting of four representive tasks, i.e., maze navigation, jigsaw puzzle, embodied long-horizon planning, as well as complex counting, where each task has dedicated free-style IVS generation pipeline supporting adaptive function calls. To systematically examine VI-CoT capability, we propose a thorough evaluation suite incorporating a progressive three-stage strategy with targeted new metrics. Besides, we establish Incremental Prompting Information Injection strategy to ablatively explore the prompting factors for VI-CoT. We extensively conduct evaluations for 18 advanced MLLMs, revealing key insights into their VI-CoT capability. The introduced ViC-Bench has been made publicly available at Huggingface.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14404
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
Wu, Xuecheng
Liu, Jiaxing
Huang, Danlei
Wang, Yifan
Shi, Yunyun
Chen, Kedi
Xue, Junxiao
Liu, Yang
Chen, Chunlin
Dong, Hairong
Yang, Dingkang
Computer Vision and Pattern Recognition
Visual-Interleaved Chain-of-Thought (VI-CoT) enables Multi-modal Large Language Models (MLLMs) to continually update their understanding and decision space based on step-wise intermediate visual states (IVS), much like a human would, which has demonstrated impressive success in various tasks, thereby leading to emerged advancements in related downstream benchmarks. Despite promising progress, current benchmarks provide models with relatively fixed IVS, rather than free-style IVS, whch might forcibly distort the original thinking trajectories, failing to evaluate their intrinsic reasoning capabilities. More importantly, existing benchmarks neglect to systematically explore the impact factors that IVS would impart to the untamed reasoning performance. To tackle above gaps, we introduce a specialized benchmark termed ViC-Bench, consisting of four representive tasks, i.e., maze navigation, jigsaw puzzle, embodied long-horizon planning, as well as complex counting, where each task has dedicated free-style IVS generation pipeline supporting adaptive function calls. To systematically examine VI-CoT capability, we propose a thorough evaluation suite incorporating a progressive three-stage strategy with targeted new metrics. Besides, we establish Incremental Prompting Information Injection strategy to ablatively explore the prompting factors for VI-CoT. We extensively conduct evaluations for 18 advanced MLLMs, revealing key insights into their VI-CoT capability. The introduced ViC-Bench has been made publicly available at Huggingface.
title ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.14404