MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Keyan, Tang, Zecheng, Ming, Lingfeng, Zhou, Guanghao, Chen, Qiguang, Qiao, Dan, Yang, Zheming, Qin, Libo, Qiu, Minghui, Li, Juntao, Zhang, Min |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
L-CiteEval: Do Long-Context Models Truly Leverage Context for Responding?
by: Tang, Zecheng, et al.
Published: (2024)
by: Tang, Zecheng, et al.
Published: (2024)
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
by: Wang, Zhaowei, et al.
Published: (2025)
by: Wang, Zhaowei, et al.
Published: (2025)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
by: Ma, Yubo, et al.
Published: (2024)
by: Ma, Yubo, et al.
Published: (2024)
CMD: a framework for Context-aware Model self-Detoxification
by: Tang, Zecheng, et al.
Published: (2023)
by: Tang, Zecheng, et al.
Published: (2023)
Revealing and Mitigating Over-Attention in Knowledge Editing
by: Wang, Pinzheng, et al.
Published: (2025)
by: Wang, Pinzheng, et al.
Published: (2025)
LOOM-Scope: a comprehensive and efficient LOng-cOntext Model evaluation framework
by: Tang, Zecheng, et al.
Published: (2025)
by: Tang, Zecheng, et al.
Published: (2025)
LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling
by: Tang, Zecheng, et al.
Published: (2025)
by: Tang, Zecheng, et al.
Published: (2025)
What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
by: Qin, Libo, et al.
Published: (2024)
by: Qin, Libo, et al.
Published: (2024)
Revisiting Long-context Modeling from Context Denoising Perspective
by: Tang, Zecheng, et al.
Published: (2025)
by: Tang, Zecheng, et al.
Published: (2025)
MemReread: Enhancing Agentic Long-Context Reasoning via Memory-Guided Rereading
by: Ji, Baibei, et al.
Published: (2026)
by: Ji, Baibei, et al.
Published: (2026)
Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought
by: Chen, Qiguang, et al.
Published: (2024)
by: Chen, Qiguang, et al.
Published: (2024)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
MemLong: Memory-Augmented Retrieval for Long Text Modeling
by: Liu, Weijie, et al.
Published: (2024)
by: Liu, Weijie, et al.
Published: (2024)
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
by: Cheng, Zihui, et al.
Published: (2024)
by: Cheng, Zihui, et al.
Published: (2024)
AutoCAP: Towards Automatic Cross-lingual Alignment Planning for Zero-shot Chain-of-Thought
by: Zhang, Yongheng, et al.
Published: (2024)
by: Zhang, Yongheng, et al.
Published: (2024)
CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models
by: Zhang, Yongheng, et al.
Published: (2025)
by: Zhang, Yongheng, et al.
Published: (2025)
LongFlow: Efficient KV Cache Compression for Reasoning Models
by: Su, Yi, et al.
Published: (2026)
by: Su, Yi, et al.
Published: (2026)
When Is Thinking Enough? Early Exit via Sufficiency Assessment for Efficient Reasoning
by: Xiang, Yang, et al.
Published: (2026)
by: Xiang, Yang, et al.
Published: (2026)
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
by: Chen, Qiguang, et al.
Published: (2024)
by: Chen, Qiguang, et al.
Published: (2024)
LOGO -- Long cOntext aliGnment via efficient preference Optimization
by: Tang, Zecheng, et al.
Published: (2024)
by: Tang, Zecheng, et al.
Published: (2024)
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
by: Chen, Qiguang, et al.
Published: (2026)
by: Chen, Qiguang, et al.
Published: (2026)
MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models
by: Ji, Yiyan, et al.
Published: (2025)
by: Ji, Yiyan, et al.
Published: (2025)
DLPO: Towards a Robust, Efficient, and Generalizable Prompt Optimization Framework from a Deep-Learning Perspective
by: Peng, Dengyun, et al.
Published: (2025)
by: Peng, Dengyun, et al.
Published: (2025)
RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning
by: Chen, Qiguang, et al.
Published: (2025)
by: Chen, Qiguang, et al.
Published: (2025)
What are the Essential Factors in Crafting Effective Long Context Multi-Hop Instruction Datasets? Insights and Best Practices
by: Chen, Zhi, et al.
Published: (2024)
by: Chen, Zhi, et al.
Published: (2024)
LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion
by: Ling, Zhan, et al.
Published: (2025)
by: Ling, Zhan, et al.
Published: (2025)
Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers
by: Qin, Libo, et al.
Published: (2024)
by: Qin, Libo, et al.
Published: (2024)
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
by: Chen, Qiguang, et al.
Published: (2025)
by: Chen, Qiguang, et al.
Published: (2025)
CroPrompt: Cross-task Interactive Prompting for Zero-shot Spoken Language Understanding
by: Qin, Libo, et al.
Published: (2024)
by: Qin, Libo, et al.
Published: (2024)
Wrong-of-Thought: An Integrated Reasoning Framework with Multi-Perspective Verification and Wrong Information
by: Zhang, Yongheng, et al.
Published: (2024)
by: Zhang, Yongheng, et al.
Published: (2024)
ContextCite: Attributing Model Generation to Context
by: Cohen-Wang, Benjamin, et al.
Published: (2024)
by: Cohen-Wang, Benjamin, et al.
Published: (2024)
MemoryRewardBench: Benchmarking Reward Models for Long-Term Memory Management in Large Language Models
by: Tang, Zecheng, et al.
Published: (2026)
by: Tang, Zecheng, et al.
Published: (2026)
CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models
by: Zhou, Guanghao, et al.
Published: (2025)
by: Zhou, Guanghao, et al.
Published: (2025)
To Pro-Cite or not to Pro-Cite.
by: Cheney, Debora, et al.
Published: (1988)
by: Cheney, Debora, et al.
Published: (1988)
OpenBA: An Open-sourced 15B Bilingual Asymmetric seq2seq Model Pre-trained from Scratch
by: Li, Juntao, et al.
Published: (2023)
by: Li, Juntao, et al.
Published: (2023)
Revealing and Mitigating the Local Pattern Shortcuts of Mamba
by: You, Wangjie, et al.
Published: (2024)
by: You, Wangjie, et al.
Published: (2024)
Citation Managers and Citing-Cited Data.
by: Simboli, Brian, et al.
Published: (2002)
by: Simboli, Brian, et al.
Published: (2002)
Sequential-NIAH: A Needle-In-A-Haystack Benchmark for Extracting Sequential Needles from Long Contexts
by: Yu, Yifei, et al.
Published: (2025)
by: Yu, Yifei, et al.
Published: (2025)
BenTo: Benchmark Task Reduction with In-Context Transferability
by: Zhao, Hongyu, et al.
Published: (2024)
by: Zhao, Hongyu, et al.
Published: (2024)
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
by: Zhou, Yucheng, et al.
Published: (2024)
by: Zhou, Yucheng, et al.
Published: (2024)
Similar Items
-
L-CiteEval: Do Long-Context Models Truly Leverage Context for Responding?
by: Tang, Zecheng, et al.
Published: (2024) -
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
by: Wang, Zhaowei, et al.
Published: (2025) -
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
by: Ma, Yubo, et al.
Published: (2024) -
CMD: a framework for Context-aware Model self-Detoxification
by: Tang, Zecheng, et al.
Published: (2023) -
Revealing and Mitigating Over-Attention in Knowledge Editing
by: Wang, Pinzheng, et al.
Published: (2025)