MileBench: Benchmarking MLLMs in Long Context
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Dingjie, Chen, Shunian, Chen, Guiming Hardy, Yu, Fei, Wan, Xiang, Wang, Benyou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
von: Chen, Junying, et al.
Veröffentlicht: (2024)
von: Chen, Junying, et al.
Veröffentlicht: (2024)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
von: Cai, Zhenyang, et al.
Veröffentlicht: (2024)
von: Cai, Zhenyang, et al.
Veröffentlicht: (2024)
MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs
von: Gao, Yufei, et al.
Veröffentlicht: (2025)
von: Gao, Yufei, et al.
Veröffentlicht: (2025)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
CMB: A Comprehensive Medical Benchmark in Chinese
von: Wang, Xidong, et al.
Veröffentlicht: (2023)
von: Wang, Xidong, et al.
Veröffentlicht: (2023)
MindBench: A Comprehensive Benchmark for Mind Map Structure Recognition and Analysis
von: Chen, Lei, et al.
Veröffentlicht: (2024)
von: Chen, Lei, et al.
Veröffentlicht: (2024)
MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos
von: Wang, Rongsheng, et al.
Veröffentlicht: (2025)
von: Wang, Rongsheng, et al.
Veröffentlicht: (2025)
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
von: Yu, Chengjun, et al.
Veröffentlicht: (2026)
von: Yu, Chengjun, et al.
Veröffentlicht: (2026)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge
von: Wang, Andong, et al.
Veröffentlicht: (2024)
von: Wang, Andong, et al.
Veröffentlicht: (2024)
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
von: Chen, Guiming Hardy, et al.
Veröffentlicht: (2024)
von: Chen, Guiming Hardy, et al.
Veröffentlicht: (2024)
GRIT: Teaching MLLMs to Think with Images
von: Fan, Yue, et al.
Veröffentlicht: (2025)
von: Fan, Yue, et al.
Veröffentlicht: (2025)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
von: Gao, Xiangbo, et al.
Veröffentlicht: (2026)
von: Gao, Xiangbo, et al.
Veröffentlicht: (2026)
Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
von: Ding, Meidan, et al.
Veröffentlicht: (2025)
von: Ding, Meidan, et al.
Veröffentlicht: (2025)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
von: Chen, Zhuokun, et al.
Veröffentlicht: (2026)
von: Chen, Zhuokun, et al.
Veröffentlicht: (2026)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
von: Lim, Hyeonseok, et al.
Veröffentlicht: (2024)
von: Lim, Hyeonseok, et al.
Veröffentlicht: (2024)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
von: Lin, Junming, et al.
Veröffentlicht: (2024)
von: Lin, Junming, et al.
Veröffentlicht: (2024)
VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models
von: Chen, Kesheng, et al.
Veröffentlicht: (2026)
von: Chen, Kesheng, et al.
Veröffentlicht: (2026)
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2026)
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2026)
VinaBench: Benchmark for Faithful and Consistent Visual Narratives
von: Gao, Silin, et al.
Veröffentlicht: (2025)
von: Gao, Silin, et al.
Veröffentlicht: (2025)
What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
von: Qin, Libo, et al.
Veröffentlicht: (2024)
von: Qin, Libo, et al.
Veröffentlicht: (2024)
APLe: Token-Wise Adaptive for Multi-Modal Prompt Learning
von: Cao, Guiming, et al.
Veröffentlicht: (2024)
von: Cao, Guiming, et al.
Veröffentlicht: (2024)
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
von: Qin, Lixiong, et al.
Veröffentlicht: (2025)
von: Qin, Lixiong, et al.
Veröffentlicht: (2025)
CubeBench: Diagnosing Interactive, Long-Horizon Spatial Reasoning Under Partial Observations
von: Gao, Huan-ang, et al.
Veröffentlicht: (2025)
von: Gao, Huan-ang, et al.
Veröffentlicht: (2025)
II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
von: Kundu, Souvik, et al.
Veröffentlicht: (2025)
von: Kundu, Souvik, et al.
Veröffentlicht: (2025)
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
MedRepBench: A Comprehensive Benchmark for Medical Report Interpretation
von: Shang, Fangxin, et al.
Veröffentlicht: (2025)
von: Shang, Fangxin, et al.
Veröffentlicht: (2025)
SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
von: Zhang, Chenkai, et al.
Veröffentlicht: (2025)
von: Zhang, Chenkai, et al.
Veröffentlicht: (2025)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
von: Yang, Minglai, et al.
Veröffentlicht: (2026)
von: Yang, Minglai, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
von: Song, Dingjie, et al.
Veröffentlicht: (2024) -
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
von: Song, Dingjie, et al.
Veröffentlicht: (2024) -
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture
von: Wang, Xidong, et al.
Veröffentlicht: (2024) -
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
von: Chen, Junying, et al.
Veröffentlicht: (2024) -
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)