DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Xinran, Zhang, Yuxuan, Zhang, Xiao, Yan, Haolong, Diao, Muxi, Xu, Songyu, Yan, Zhonghao, Li, Hongbing, Liang, Kongming, Ma, Zhanyu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915920363913216
author Wang, Xinran
Zhang, Yuxuan
Zhang, Xiao
Yan, Haolong
Diao, Muxi
Xu, Songyu
Yan, Zhonghao
Li, Hongbing
Liang, Kongming
Ma, Zhanyu
author_facet Wang, Xinran
Zhang, Yuxuan
Zhang, Xiao
Yan, Haolong
Diao, Muxi
Xu, Songyu
Yan, Zhonghao
Li, Hongbing
Liang, Kongming
Ma, Zhanyu
contents Accurately detecting and localizing hallucinations is a critical task for ensuring high reliability of image captions. In the era of Multimodal Large Language Models (MLLMs), captions have evolved from brief sentences into comprehensive narratives, often spanning hundreds of words. This shift exponentially increases the challenge: models must now pinpoint specific erroneous spans or words within extensive contexts, rather than merely flag response-level inconsistencies. However, existing benchmarks lack the fine granularity and domain diversity required to evaluate this capability. To bridge this gap, we introduce DetailVerifyBench, a rigorous benchmark comprising 1,000 high-quality images across five distinct domains. With an average caption length of over 200 words and dense, token-level annotations of multiple hallucination types, it stands as the most challenging benchmark for precise hallucination localization in the field of long image captioning to date. Our benchmark is available at https://zyx-hhnkh.github.io/DetailVerifyBench/.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05623
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
Wang, Xinran
Zhang, Yuxuan
Zhang, Xiao
Yan, Haolong
Diao, Muxi
Xu, Songyu
Yan, Zhonghao
Li, Hongbing
Liang, Kongming
Ma, Zhanyu
Computer Vision and Pattern Recognition
Computation and Language
Multimedia
Accurately detecting and localizing hallucinations is a critical task for ensuring high reliability of image captions. In the era of Multimodal Large Language Models (MLLMs), captions have evolved from brief sentences into comprehensive narratives, often spanning hundreds of words. This shift exponentially increases the challenge: models must now pinpoint specific erroneous spans or words within extensive contexts, rather than merely flag response-level inconsistencies. However, existing benchmarks lack the fine granularity and domain diversity required to evaluate this capability. To bridge this gap, we introduce DetailVerifyBench, a rigorous benchmark comprising 1,000 high-quality images across five distinct domains. With an average caption length of over 200 words and dense, token-level annotations of multiple hallucination types, it stands as the most challenging benchmark for precise hallucination localization in the field of long image captioning to date. Our benchmark is available at https://zyx-hhnkh.github.io/DetailVerifyBench/.
title DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
topic Computer Vision and Pattern Recognition
Computation and Language
Multimedia
url https://arxiv.org/abs/2604.05623