MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917015648731136 |
|---|---|
| author | Zhou, Keyan Tang, Zecheng Ming, Lingfeng Zhou, Guanghao Chen, Qiguang Qiao, Dan Yang, Zheming Qin, Libo Qiu, Minghui Li, Juntao Zhang, Min |
| author_facet | Zhou, Keyan Tang, Zecheng Ming, Lingfeng Zhou, Guanghao Chen, Qiguang Qiao, Dan Yang, Zheming Qin, Libo Qiu, Minghui Li, Juntao Zhang, Min |
| contents | The rapid advancement of large vision language models (LVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-context faithfulness are predominantly focused on the text-only domain, while multimodal assessments remain limited to short contexts. To bridge this gap, we introduce MMLongCite, a comprehensive benchmark designed to evaluate the fidelity of LVLMs in long-context scenarios. MMLongCite comprises 8 distinct tasks spanning 6 context length intervals and incorporates diverse modalities, including text, images, and videos. Our evaluation of state-of-the-art LVLMs reveals their limited faithfulness in handling long multimodal contexts. Furthermore, we provide an in-depth analysis of how context length and the position of crucial content affect the faithfulness of these models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_13276 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models Zhou, Keyan Tang, Zecheng Ming, Lingfeng Zhou, Guanghao Chen, Qiguang Qiao, Dan Yang, Zheming Qin, Libo Qiu, Minghui Li, Juntao Zhang, Min Computer Vision and Pattern Recognition Computation and Language The rapid advancement of large vision language models (LVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-context faithfulness are predominantly focused on the text-only domain, while multimodal assessments remain limited to short contexts. To bridge this gap, we introduce MMLongCite, a comprehensive benchmark designed to evaluate the fidelity of LVLMs in long-context scenarios. MMLongCite comprises 8 distinct tasks spanning 6 context length intervals and incorporates diverse modalities, including text, images, and videos. Our evaluation of state-of-the-art LVLMs reveals their limited faithfulness in handling long multimodal contexts. Furthermore, we provide an in-depth analysis of how context length and the position of crucial content affect the faithfulness of these models. |
| title | MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models |
| topic | Computer Vision and Pattern Recognition Computation and Language |
| url | https://arxiv.org/abs/2510.13276 |