MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Keyan, Tang, Zecheng, Ming, Lingfeng, Zhou, Guanghao, Chen, Qiguang, Qiao, Dan, Yang, Zheming, Qin, Libo, Qiu, Minghui, Li, Juntao, Zhang, Min
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917015648731136
author Zhou, Keyan
Tang, Zecheng
Ming, Lingfeng
Zhou, Guanghao
Chen, Qiguang
Qiao, Dan
Yang, Zheming
Qin, Libo
Qiu, Minghui
Li, Juntao
Zhang, Min
author_facet Zhou, Keyan
Tang, Zecheng
Ming, Lingfeng
Zhou, Guanghao
Chen, Qiguang
Qiao, Dan
Yang, Zheming
Qin, Libo
Qiu, Minghui
Li, Juntao
Zhang, Min
contents The rapid advancement of large vision language models (LVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-context faithfulness are predominantly focused on the text-only domain, while multimodal assessments remain limited to short contexts. To bridge this gap, we introduce MMLongCite, a comprehensive benchmark designed to evaluate the fidelity of LVLMs in long-context scenarios. MMLongCite comprises 8 distinct tasks spanning 6 context length intervals and incorporates diverse modalities, including text, images, and videos. Our evaluation of state-of-the-art LVLMs reveals their limited faithfulness in handling long multimodal contexts. Furthermore, we provide an in-depth analysis of how context length and the position of crucial content affect the faithfulness of these models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13276
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models
Zhou, Keyan
Tang, Zecheng
Ming, Lingfeng
Zhou, Guanghao
Chen, Qiguang
Qiao, Dan
Yang, Zheming
Qin, Libo
Qiu, Minghui
Li, Juntao
Zhang, Min
Computer Vision and Pattern Recognition
Computation and Language
The rapid advancement of large vision language models (LVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-context faithfulness are predominantly focused on the text-only domain, while multimodal assessments remain limited to short contexts. To bridge this gap, we introduce MMLongCite, a comprehensive benchmark designed to evaluate the fidelity of LVLMs in long-context scenarios. MMLongCite comprises 8 distinct tasks spanning 6 context length intervals and incorporates diverse modalities, including text, images, and videos. Our evaluation of state-of-the-art LVLMs reveals their limited faithfulness in handling long multimodal contexts. Furthermore, we provide an in-depth analysis of how context length and the position of crucial content affect the faithfulness of these models.
title MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2510.13276