CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Fuwen, Chen, Chi, Wan, Zihao, Kang, Zhaolu, Yan, Qidong, Li, Yingjie, Wang, Xiaolong, Wang, Siyu, Wang, Ziyue, Mi, Xiaoyue, Li, Peng, Ma, Ning, Sun, Maosong, Liu, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Browse and Concentrate: Comprehending Multimodal Content via prior-LLM Context Fusion
von: Wang, Ziyue, et al.
Veröffentlicht: (2024)
von: Wang, Ziyue, et al.
Veröffentlicht: (2024)
Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction
von: Liu, Dairu, et al.
Veröffentlicht: (2025)
von: Liu, Dairu, et al.
Veröffentlicht: (2025)
MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025)
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025)
ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
von: Wang, Ziyue, et al.
Veröffentlicht: (2024)
von: Wang, Ziyue, et al.
Veröffentlicht: (2024)
Enabling Stroke-Level Structural Analysis of Hieroglyphic Scripts without Language-Specific Priors
von: Luo, Fuwen, et al.
Veröffentlicht: (2026)
von: Luo, Fuwen, et al.
Veröffentlicht: (2026)
Model Composition for Multimodal Large Language Models
von: Chen, Chi, et al.
Veröffentlicht: (2024)
von: Chen, Chi, et al.
Veröffentlicht: (2024)
EscapeCraft: A 3D Room Escape Environment for Benchmarking Complex Multimodal Reasoning Ability
von: Wang, Ziyue, et al.
Veröffentlicht: (2025)
von: Wang, Ziyue, et al.
Veröffentlicht: (2025)
Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language Models
von: Wang, Xiaolong, et al.
Veröffentlicht: (2024)
von: Wang, Xiaolong, et al.
Veröffentlicht: (2024)
Penny Wise, Pixel Foolish: Bypassing Price Constraints in Multimodal Agents via Visual Adversarial Perturbations
von: Qian, Jiachen, et al.
Veröffentlicht: (2026)
von: Qian, Jiachen, et al.
Veröffentlicht: (2026)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
von: Lin, Junming, et al.
Veröffentlicht: (2024)
von: Lin, Junming, et al.
Veröffentlicht: (2024)
Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency
von: Wang, Zhikai, et al.
Veröffentlicht: (2025)
von: Wang, Zhikai, et al.
Veröffentlicht: (2025)
Perspective Transition of Large Language Models for Solving Subjective Tasks
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025)
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025)
When Helpers Become Hazards: A Benchmark for Analyzing Multimodal LLM-Powered Safety in Daily Life
von: Lou, Xinyue, et al.
Veröffentlicht: (2026)
von: Lou, Xinyue, et al.
Veröffentlicht: (2026)
MER-Bench: A Comprehensive Benchmark for Multimodal Meme Reappraisal
von: Nie, Yiqi, et al.
Veröffentlicht: (2026)
von: Nie, Yiqi, et al.
Veröffentlicht: (2026)
LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
von: Xiao, Yijia, et al.
Veröffentlicht: (2024)
von: Xiao, Yijia, et al.
Veröffentlicht: (2024)
Multimodal Generalized Category Discovery
von: Su, Yuchang, et al.
Veröffentlicht: (2024)
von: Su, Yuchang, et al.
Veröffentlicht: (2024)
Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2023)
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2023)
MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference
von: Wan, Zhongwei, et al.
Veröffentlicht: (2025)
von: Wan, Zhongwei, et al.
Veröffentlicht: (2025)
DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms
von: Bi, Xiaojun, et al.
Veröffentlicht: (2025)
von: Bi, Xiaojun, et al.
Veröffentlicht: (2025)
V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models
von: Wang, Qidong, et al.
Veröffentlicht: (2025)
von: Wang, Qidong, et al.
Veröffentlicht: (2025)
CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
von: Li, Jingyao, et al.
Veröffentlicht: (2025)
von: Li, Jingyao, et al.
Veröffentlicht: (2025)
ReasonAct: Progressive Training for Fine-Grained Video Reasoning in Small Models
von: Liu, Jiaxin, et al.
Veröffentlicht: (2025)
von: Liu, Jiaxin, et al.
Veröffentlicht: (2025)
SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning
von: Zeng, Fanhu, et al.
Veröffentlicht: (2026)
von: Zeng, Fanhu, et al.
Veröffentlicht: (2026)
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
VP-MEL: Visual Prompts Guided Multimodal Entity Linking
von: Mi, Hongze, et al.
Veröffentlicht: (2024)
von: Mi, Hongze, et al.
Veröffentlicht: (2024)
HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
von: Kang, Zhaolu, et al.
Veröffentlicht: (2025)
von: Kang, Zhaolu, et al.
Veröffentlicht: (2025)
Personal Visual Context Learning in Large Multimodal Models
von: Xue, Zihui, et al.
Veröffentlicht: (2026)
von: Xue, Zihui, et al.
Veröffentlicht: (2026)
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
von: Ji, Yifan, et al.
Veröffentlicht: (2026)
von: Ji, Yifan, et al.
Veröffentlicht: (2026)
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
EMODIS: A Benchmark for Context-Dependent Emoji Disambiguation in Large Language Models
von: Huang, Jiacheng, et al.
Veröffentlicht: (2025)
von: Huang, Jiacheng, et al.
Veröffentlicht: (2025)
Filling the Image Information Gap for VQA: Prompting Large Language Models to Proactively Ask Questions
von: Wang, Ziyue, et al.
Veröffentlicht: (2023)
von: Wang, Ziyue, et al.
Veröffentlicht: (2023)
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
von: Luo, Fuwen, et al.
Veröffentlicht: (2025)
von: Luo, Fuwen, et al.
Veröffentlicht: (2025)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
von: Yang, Zhibo, et al.
Veröffentlicht: (2024)
von: Yang, Zhibo, et al.
Veröffentlicht: (2024)
Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models
von: Wang, Hengyi, et al.
Veröffentlicht: (2024)
von: Wang, Hengyi, et al.
Veröffentlicht: (2024)
XL$^2$Bench: A Benchmark for Extremely Long Context Understanding with Long-range Dependencies
von: Ni, Xuanfan, et al.
Veröffentlicht: (2024)
von: Ni, Xuanfan, et al.
Veröffentlicht: (2024)
VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
von: Xu, Mingjie, et al.
Veröffentlicht: (2025)
von: Xu, Mingjie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Browse and Concentrate: Comprehending Multimodal Content via prior-LLM Context Fusion
von: Wang, Ziyue, et al.
Veröffentlicht: (2024) -
Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction
von: Liu, Dairu, et al.
Veröffentlicht: (2025) -
MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025) -
ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
von: Wang, Ziyue, et al.
Veröffentlicht: (2024) -
Enabling Stroke-Level Structural Analysis of Hieroglyphic Scripts without Language-Specific Priors
von: Luo, Fuwen, et al.
Veröffentlicht: (2026)