Visual Room 2.0: Seeing is Not Understanding for MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Haokun, Zhang, Yazhou, Ding, Jizhi, Li, Qiuchi, Zhang, Peng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seeing is Not Understanding: A Benchmark on Perception-Cognition Disparities in Large Language Models
von: Li, Haokun, et al.
Veröffentlicht: (2025)
von: Li, Haokun, et al.
Veröffentlicht: (2025)
Is Sarcasm Detection A Step-by-Step Reasoning Process in Large Language Models?
von: Yao, Ben, et al.
Veröffentlicht: (2024)
von: Yao, Ben, et al.
Veröffentlicht: (2024)
Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
von: Zhang, Yazhou, et al.
Veröffentlicht: (2025)
von: Zhang, Yazhou, et al.
Veröffentlicht: (2025)
Are MLMs Trapped in the Visual Room?
von: Zhang, Yazhou, et al.
Veröffentlicht: (2025)
von: Zhang, Yazhou, et al.
Veröffentlicht: (2025)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
von: Guo, Pinxue, et al.
Veröffentlicht: (2025)
von: Guo, Pinxue, et al.
Veröffentlicht: (2025)
Do MLLMs Really Understand the Charts?
von: Zhang, Xiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xiao, et al.
Veröffentlicht: (2025)
Pushing The Limit of LLM Capacity for Text Classification
von: Zhang, Yazhou, et al.
Veröffentlicht: (2024)
von: Zhang, Yazhou, et al.
Veröffentlicht: (2024)
DialogueLLM: Context and Emotion Knowledge-Tuned Large Language Models for Emotion Recognition in Conversations
von: Zhang, Yazhou, et al.
Veröffentlicht: (2023)
von: Zhang, Yazhou, et al.
Veröffentlicht: (2023)
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2025)
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2025)
Joint Extraction and Classification of Danish Competences for Job Matching
von: Li, Qiuchi, et al.
Veröffentlicht: (2024)
von: Li, Qiuchi, et al.
Veröffentlicht: (2024)
NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
von: Yao, Ben, et al.
Veröffentlicht: (2025)
von: Yao, Ben, et al.
Veröffentlicht: (2025)
Large Language Models for Subjective Language Understanding: A Survey
von: Song, Changhao, et al.
Veröffentlicht: (2025)
von: Song, Changhao, et al.
Veröffentlicht: (2025)
Emotion-o1: Adaptive Long Reasoning for Emotion Understanding in LLMs
von: Song, Changhao, et al.
Veröffentlicht: (2025)
von: Song, Changhao, et al.
Veröffentlicht: (2025)
Bridging External and Parametric Knowledge: Mitigating Hallucination of LLMs with Shared-Private Semantic Synergy in Dual-Stream Knowledge
von: Sui, Yi, et al.
Veröffentlicht: (2025)
von: Sui, Yi, et al.
Veröffentlicht: (2025)
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
von: Zhang, Renrui, et al.
Veröffentlicht: (2024)
von: Zhang, Renrui, et al.
Veröffentlicht: (2024)
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
von: He, Wei, et al.
Veröffentlicht: (2024)
von: He, Wei, et al.
Veröffentlicht: (2024)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
von: Miao, Ziqi, et al.
Veröffentlicht: (2025)
von: Miao, Ziqi, et al.
Veröffentlicht: (2025)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
von: Ding, Xuanwen, et al.
Veröffentlicht: (2025)
von: Ding, Xuanwen, et al.
Veröffentlicht: (2025)
T5Gemma 2: Seeing, Reading, and Understanding Longer
von: Zhang, Biao, et al.
Veröffentlicht: (2025)
von: Zhang, Biao, et al.
Veröffentlicht: (2025)
Roles of MLLMs in Visually Rich Document Retrieval for RAG: A Survey
von: Zhang, Xiantao
Veröffentlicht: (2025)
von: Zhang, Xiantao
Veröffentlicht: (2025)
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
von: Kang, Caixin, et al.
Veröffentlicht: (2025)
von: Kang, Caixin, et al.
Veröffentlicht: (2025)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
Affordance Benchmark for MLLMs
von: Wang, Junying, et al.
Veröffentlicht: (2025)
von: Wang, Junying, et al.
Veröffentlicht: (2025)
Can MLLMs Understand the Deep Implication Behind Chinese Images?
von: Zhang, Chenhao, et al.
Veröffentlicht: (2024)
von: Zhang, Chenhao, et al.
Veröffentlicht: (2024)
GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning
von: Chen, Guizhen, et al.
Veröffentlicht: (2025)
von: Chen, Guizhen, et al.
Veröffentlicht: (2025)
Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs
von: Han, Yujin, et al.
Veröffentlicht: (2025)
von: Han, Yujin, et al.
Veröffentlicht: (2025)
Seeing the Poem: Image-Semantic Detection of AI-Generated Modern Chinese Poetry with MLLMs
von: Wang, Shanshan, et al.
Veröffentlicht: (2026)
von: Wang, Shanshan, et al.
Veröffentlicht: (2026)
The Death of Feature Engineering? BERT with Linguistic Features on SQuAD 2.0
von: Li, Jiawei, et al.
Veröffentlicht: (2024)
von: Li, Jiawei, et al.
Veröffentlicht: (2024)
Large Emotional World Model
von: Song, Changhao, et al.
Veröffentlicht: (2025)
von: Song, Changhao, et al.
Veröffentlicht: (2025)
Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
von: Kang, Caixin, et al.
Veröffentlicht: (2025)
von: Kang, Caixin, et al.
Veröffentlicht: (2025)
See the Text: From Tokenization to Visual Reading
von: Xing, Ling, et al.
Veröffentlicht: (2025)
von: Xing, Ling, et al.
Veröffentlicht: (2025)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
von: Zheng, Naishan, et al.
Veröffentlicht: (2025)
von: Zheng, Naishan, et al.
Veröffentlicht: (2025)
AdaCodec: A Predictive Visual Code for Video MLLMs
von: Hou, Haowen, et al.
Veröffentlicht: (2026)
von: Hou, Haowen, et al.
Veröffentlicht: (2026)
OpenPI2.0: An Improved Dataset for Entity Tracking in Texts
von: Zhang, Li, et al.
Veröffentlicht: (2023)
von: Zhang, Li, et al.
Veröffentlicht: (2023)
MLLMs-Augmented Visual-Language Representation Learning
von: Liu, Yanqing, et al.
Veröffentlicht: (2023)
von: Liu, Yanqing, et al.
Veröffentlicht: (2023)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
von: Wang, Siting, et al.
Veröffentlicht: (2025)
von: Wang, Siting, et al.
Veröffentlicht: (2025)
SarcasmBench: Towards Evaluating Large Language Models on Sarcasm Understanding
von: Zhang, Yazhou, et al.
Veröffentlicht: (2024)
von: Zhang, Yazhou, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Seeing is Not Understanding: A Benchmark on Perception-Cognition Disparities in Large Language Models
von: Li, Haokun, et al.
Veröffentlicht: (2025) -
Is Sarcasm Detection A Step-by-Step Reasoning Process in Large Language Models?
von: Yao, Ben, et al.
Veröffentlicht: (2024) -
Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
von: Zhang, Yazhou, et al.
Veröffentlicht: (2025) -
Are MLMs Trapped in the Visual Room?
von: Zhang, Yazhou, et al.
Veröffentlicht: (2025) -
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
von: Guo, Pinxue, et al.
Veröffentlicht: (2025)