DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915920363913216 |
|---|---|
| author | Wang, Xinran Zhang, Yuxuan Zhang, Xiao Yan, Haolong Diao, Muxi Xu, Songyu Yan, Zhonghao Li, Hongbing Liang, Kongming Ma, Zhanyu |
| author_facet | Wang, Xinran Zhang, Yuxuan Zhang, Xiao Yan, Haolong Diao, Muxi Xu, Songyu Yan, Zhonghao Li, Hongbing Liang, Kongming Ma, Zhanyu |
| contents | Accurately detecting and localizing hallucinations is a critical task for ensuring high reliability of image captions. In the era of Multimodal Large Language Models (MLLMs), captions have evolved from brief sentences into comprehensive narratives, often spanning hundreds of words. This shift exponentially increases the challenge: models must now pinpoint specific erroneous spans or words within extensive contexts, rather than merely flag response-level inconsistencies. However, existing benchmarks lack the fine granularity and domain diversity required to evaluate this capability. To bridge this gap, we introduce DetailVerifyBench, a rigorous benchmark comprising 1,000 high-quality images across five distinct domains. With an average caption length of over 200 words and dense, token-level annotations of multiple hallucination types, it stands as the most challenging benchmark for precise hallucination localization in the field of long image captioning to date. Our benchmark is available at https://zyx-hhnkh.github.io/DetailVerifyBench/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_05623 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions Wang, Xinran Zhang, Yuxuan Zhang, Xiao Yan, Haolong Diao, Muxi Xu, Songyu Yan, Zhonghao Li, Hongbing Liang, Kongming Ma, Zhanyu Computer Vision and Pattern Recognition Computation and Language Multimedia Accurately detecting and localizing hallucinations is a critical task for ensuring high reliability of image captions. In the era of Multimodal Large Language Models (MLLMs), captions have evolved from brief sentences into comprehensive narratives, often spanning hundreds of words. This shift exponentially increases the challenge: models must now pinpoint specific erroneous spans or words within extensive contexts, rather than merely flag response-level inconsistencies. However, existing benchmarks lack the fine granularity and domain diversity required to evaluate this capability. To bridge this gap, we introduce DetailVerifyBench, a rigorous benchmark comprising 1,000 high-quality images across five distinct domains. With an average caption length of over 200 words and dense, token-level annotations of multiple hallucination types, it stands as the most challenging benchmark for precise hallucination localization in the field of long image captioning to date. Our benchmark is available at https://zyx-hhnkh.github.io/DetailVerifyBench/. |
| title | DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions |
| topic | Computer Vision and Pattern Recognition Computation and Language Multimedia |
| url | https://arxiv.org/abs/2604.05623 |