CaptionQA: Is Your Caption as Useful as the Image Itself?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Shijia, Liu, Yunong, Zhai, Bohan, Sun, Ximeng, Liu, Zicheng, Barsoum, Emad, Li, Manling, Xu, Chenfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917411085615104
author Yang, Shijia
Liu, Yunong
Zhai, Bohan
Sun, Ximeng
Liu, Zicheng
Barsoum, Emad
Li, Manling
Xu, Chenfeng
author_facet Yang, Shijia
Liu, Yunong
Zhai, Bohan
Sun, Ximeng
Liu, Zicheng
Barsoum, Emad
Li, Manling
Xu, Chenfeng
contents Image captions serve as efficient surrogates for visual content in multimodal systems such as retrieval, recommendation, and multi-step agentic inference pipelines. Yet current evaluation practices miss a fundamental question: Can captions stand-in for images in real downstream tasks? We propose a utility-based benchmark, CaptionQA, to evaluate model-generated captions, where caption quality is measured by how well it supports downstream tasks. CaptionQA is an extensible domain-dependent benchmark covering 4 domains--Natural, Document, E-commerce, and Embodied AI--each with fine-grained taxonomies (25 top-level and 69 subcategories) that identify useful information for domain-specific tasks. CaptionQA builds 33,027 densely annotated multiple-choice questions (50.3 per image on average) that explicitly require visual information to answer, providing a comprehensive probe of caption utility. In our evaluation protocol, an LLM answers these questions using captions alone, directly measuring whether captions preserve image-level utility and are utilizable by a downstream LLM. Evaluating state-of-the-art MLLMs reveals substantial gaps between the image and its caption utility. Notably, models nearly identical on traditional image-QA benchmarks lower by up to 32% in caption utility. We release CaptionQA along with an open-source pipeline for extension to new domains. The code is available at https://github.com/bronyayang/CaptionQA.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21025
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CaptionQA: Is Your Caption as Useful as the Image Itself?
Yang, Shijia
Liu, Yunong
Zhai, Bohan
Sun, Ximeng
Liu, Zicheng
Barsoum, Emad
Li, Manling
Xu, Chenfeng
Computer Vision and Pattern Recognition
Image captions serve as efficient surrogates for visual content in multimodal systems such as retrieval, recommendation, and multi-step agentic inference pipelines. Yet current evaluation practices miss a fundamental question: Can captions stand-in for images in real downstream tasks? We propose a utility-based benchmark, CaptionQA, to evaluate model-generated captions, where caption quality is measured by how well it supports downstream tasks. CaptionQA is an extensible domain-dependent benchmark covering 4 domains--Natural, Document, E-commerce, and Embodied AI--each with fine-grained taxonomies (25 top-level and 69 subcategories) that identify useful information for domain-specific tasks. CaptionQA builds 33,027 densely annotated multiple-choice questions (50.3 per image on average) that explicitly require visual information to answer, providing a comprehensive probe of caption utility. In our evaluation protocol, an LLM answers these questions using captions alone, directly measuring whether captions preserve image-level utility and are utilizable by a downstream LLM. Evaluating state-of-the-art MLLMs reveals substantial gaps between the image and its caption utility. Notably, models nearly identical on traditional image-QA benchmarks lower by up to 32% in caption utility. We release CaptionQA along with an open-source pipeline for extension to new domains. The code is available at https://github.com/bronyayang/CaptionQA.
title CaptionQA: Is Your Caption as Useful as the Image Itself?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.21025