QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kao, Kuei-Chun, Tzu-Yin, Hsu, Hong, Yunqi, Wang, Ruochen, Hsieh, Cho-Jui
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909887511920640
author Kao, Kuei-Chun
Tzu-Yin, Hsu
Hong, Yunqi
Wang, Ruochen
Hsieh, Cho-Jui
author_facet Kao, Kuei-Chun
Tzu-Yin, Hsu
Hong, Yunqi
Wang, Ruochen
Hsieh, Cho-Jui
contents Recently, Multimodal Large Language Models (MLLMs) encounter two key issues in multi-image contexts: (1) a lack of fine-grained perception across disparate images, and (2) a diminished capability to effectively reason over and synthesize information from multiple visual inputs. However, while various prompting methods aim to describe visual content, many existing studies focus primarily on single-image settings or specific, constrained scenarios. This leaves a critical gap in understanding and addressing how MLLMs tackle more general and complex multi-image reasoning tasks. Thus, we first extensively investigate how current prompting methods perceive fine-grained visual details and process visual information when dealing with multiple images. Our findings reveal that existing prompting methods fall short in attending to needed clues and seamlessly integrating perception and reasoning. Inspired by the findings, we propose a new zero-shot prompting method, Question-Guided Chain-of-Captions (QG-CoC), a generalized prompting approach that effectively handles problems with an arbitrary number of images. We evaluate our method on various open-source and closed-source MLLMs for multi-image and single-image benchmarks. Experimental results indicate that QG-CoC demonstrates competitive performance across tasks and exhibits robust improvements in the challenging scenarios where existing prompting methods fail.
format Preprint
id arxiv_https___arxiv_org_abs_2511_03206
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models
Kao, Kuei-Chun
Tzu-Yin, Hsu
Hong, Yunqi
Wang, Ruochen
Hsieh, Cho-Jui
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Recently, Multimodal Large Language Models (MLLMs) encounter two key issues in multi-image contexts: (1) a lack of fine-grained perception across disparate images, and (2) a diminished capability to effectively reason over and synthesize information from multiple visual inputs. However, while various prompting methods aim to describe visual content, many existing studies focus primarily on single-image settings or specific, constrained scenarios. This leaves a critical gap in understanding and addressing how MLLMs tackle more general and complex multi-image reasoning tasks. Thus, we first extensively investigate how current prompting methods perceive fine-grained visual details and process visual information when dealing with multiple images. Our findings reveal that existing prompting methods fall short in attending to needed clues and seamlessly integrating perception and reasoning. Inspired by the findings, we propose a new zero-shot prompting method, Question-Guided Chain-of-Captions (QG-CoC), a generalized prompting approach that effectively handles problems with an arbitrary number of images. We evaluate our method on various open-source and closed-source MLLMs for multi-image and single-image benchmarks. Experimental results indicate that QG-CoC demonstrates competitive performance across tasks and exhibits robust improvements in the challenging scenarios where existing prompting methods fail.
title QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.03206