Generating Accurate and Detailed Captions for High-Resolution Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Hankyeol, Seo, Gawon, Lee, Kyounggyu, Kim, Dogun, Song, Kyungwoo, Jung, Jiyoung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909879777624064
author Lee, Hankyeol
Seo, Gawon
Lee, Kyounggyu
Kim, Dogun
Song, Kyungwoo
Jung, Jiyoung
author_facet Lee, Hankyeol
Seo, Gawon
Lee, Kyounggyu
Kim, Dogun
Song, Kyungwoo
Jung, Jiyoung
contents Vision-language models (VLMs) often struggle to generate accurate and detailed captions for high-resolution images since they are typically pre-trained on low-resolution inputs (e.g., 224x224 or 336x336 pixels). Downscaling high-resolution images to these dimensions may result in the loss of visual details and the omission of important objects. To address this limitation, we propose a novel pipeline that integrates vision-language models, large language models (LLMs), and object detection systems to enhance caption quality. Our proposed pipeline refines captions through a novel, multi-stage process. Given a high-resolution image, an initial caption is first generated using a VLM, and key objects in the image are then identified by an LLM. The LLM predicts additional objects likely to co-occur with the identified key objects, and these predictions are verified by object detection systems. Newly detected objects not mentioned in the initial caption undergo focused, region-specific captioning to ensure they are incorporated. This process enriches caption detail while reducing hallucinations by removing references to undetected objects. We evaluate the enhanced captions using pairwise comparison and quantitative scoring from large multimodal models, along with a benchmark for hallucination detection. Experiments on a curated dataset of high-resolution images demonstrate that our pipeline produces more detailed and reliable image captions while effectively minimizing hallucinations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27164
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generating Accurate and Detailed Captions for High-Resolution Images
Lee, Hankyeol
Seo, Gawon
Lee, Kyounggyu
Kim, Dogun
Song, Kyungwoo
Jung, Jiyoung
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision-language models (VLMs) often struggle to generate accurate and detailed captions for high-resolution images since they are typically pre-trained on low-resolution inputs (e.g., 224x224 or 336x336 pixels). Downscaling high-resolution images to these dimensions may result in the loss of visual details and the omission of important objects. To address this limitation, we propose a novel pipeline that integrates vision-language models, large language models (LLMs), and object detection systems to enhance caption quality. Our proposed pipeline refines captions through a novel, multi-stage process. Given a high-resolution image, an initial caption is first generated using a VLM, and key objects in the image are then identified by an LLM. The LLM predicts additional objects likely to co-occur with the identified key objects, and these predictions are verified by object detection systems. Newly detected objects not mentioned in the initial caption undergo focused, region-specific captioning to ensure they are incorporated. This process enriches caption detail while reducing hallucinations by removing references to undetected objects. We evaluate the enhanced captions using pairwise comparison and quantitative scoring from large multimodal models, along with a benchmark for hallucination detection. Experiments on a curated dataset of high-resolution images demonstrate that our pipeline produces more detailed and reliable image captions while effectively minimizing hallucinations.
title Generating Accurate and Detailed Captions for High-Resolution Images
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.27164