Chain-of-Caption: Training-free improvement of multimodal large language model on referring expression comprehension

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Pang, Yik Lung, Oh, Changjae
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917259345133568
author Pang, Yik Lung
Oh, Changjae
author_facet Pang, Yik Lung
Oh, Changjae
contents Given a textual description, the task of referring expression comprehension (REC) involves the localisation of the referred object in an image. Multimodal large language models (MLLMs) have achieved high accuracy on REC benchmarks through scaling up the model size and training data. Moreover, the performance of MLLMs can be further improved using techniques such as Chain-of-Thought and tool use, which provides additional visual or textual context to the model. In this paper, we analyse the effect of various techniques for providing additional visual and textual context via tool use to the MLLM and its effect on the REC task. Furthermore, we propose a training-free framework named Chain-of-Caption to improve the REC performance of MLLMs. We perform experiments on RefCOCO/RefCOCOg/RefCOCO+ and Ref-L4 datasets and show that individual textual or visual context can improve the REC performance without any fine-tuning. By combining multiple contexts, our training-free framework shows between 5% to 30% performance gain over the baseline model on accuracy at various Intersection over Union (IoU) thresholds.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08211
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Chain-of-Caption: Training-free improvement of multimodal large language model on referring expression comprehension
Pang, Yik Lung
Oh, Changjae
Computer Vision and Pattern Recognition
Given a textual description, the task of referring expression comprehension (REC) involves the localisation of the referred object in an image. Multimodal large language models (MLLMs) have achieved high accuracy on REC benchmarks through scaling up the model size and training data. Moreover, the performance of MLLMs can be further improved using techniques such as Chain-of-Thought and tool use, which provides additional visual or textual context to the model. In this paper, we analyse the effect of various techniques for providing additional visual and textual context via tool use to the MLLM and its effect on the REC task. Furthermore, we propose a training-free framework named Chain-of-Caption to improve the REC performance of MLLMs. We perform experiments on RefCOCO/RefCOCOg/RefCOCO+ and Ref-L4 datasets and show that individual textual or visual context can improve the REC performance without any fine-tuning. By combining multiple contexts, our training-free framework shows between 5% to 30% performance gain over the baseline model on accuracy at various Intersection over Union (IoU) thresholds.
title Chain-of-Caption: Training-free improvement of multimodal large language model on referring expression comprehension
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.08211