Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Yeh, Chun-Hsiao, Wang, Chenyu, Tong, Shengbang, Cheng, Ta-Ying, Wang, Ruoyu, Chu, Tianzhe, Zhai, Yuexiang, Chen, Yubei, Gao, Shenghua, Ma, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
by: Yu, Yaodong, et al.
Published: (2023)
by: Yu, Yaodong, et al.
Published: (2023)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
by: Chu, Tianzhe, et al.
Published: (2025)
by: Chu, Tianzhe, et al.
Published: (2025)
Recollection from Pensieve: Novel View Synthesis via Learning from Uncalibrated Videos
by: Wang, Ruoyu, et al.
Published: (2025)
by: Wang, Ruoyu, et al.
Published: (2025)
Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection
by: Qi, Dacheng, et al.
Published: (2026)
by: Qi, Dacheng, et al.
Published: (2026)
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models
by: Chu, Tianzhe, et al.
Published: (2023)
by: Chu, Tianzhe, et al.
Published: (2023)
Space-time 2D Gaussian Splatting for Accurate Surface Reconstruction under Complex Dynamic Scenes
by: Wang, Shuo, et al.
Published: (2024)
by: Wang, Shuo, et al.
Published: (2024)
Visual Room 2.0: Seeing is Not Understanding for MLLMs
by: Li, Haokun, et al.
Published: (2025)
by: Li, Haokun, et al.
Published: (2025)
Gen4Gen: Generative Data Pipeline for Generative Multi-Concept Composition
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
3D Spatial Understanding in MLLMs: Disambiguation and Evaluation
by: Chang, Chun-Peng, et al.
Published: (2024)
by: Chang, Chun-Peng, et al.
Published: (2024)
Insight: A Multi-Modal Diagnostic Pipeline using LLMs for Ocular Surface Disease Diagnosis
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
by: Chen, Yiming, et al.
Published: (2025)
by: Chen, Yiming, et al.
Published: (2025)
GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
by: Zhu, Xiaorong, et al.
Published: (2025)
by: Zhu, Xiaorong, et al.
Published: (2025)
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
by: Han, Junlin, et al.
Published: (2025)
by: Han, Junlin, et al.
Published: (2025)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
by: He, Yuping, et al.
Published: (2025)
by: He, Yuping, et al.
Published: (2025)
Ctrl123: Consistent Novel View Synthesis via Closed-Loop Transcription
by: Zhao, Hongxiang, et al.
Published: (2024)
by: Zhao, Hongxiang, et al.
Published: (2024)
Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly
by: Liu, Yexin, et al.
Published: (2024)
by: Liu, Yexin, et al.
Published: (2024)
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
by: Zhai, Yuexiang, et al.
Published: (2024)
by: Zhai, Yuexiang, et al.
Published: (2024)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
by: Yilmaz, Nilay, et al.
Published: (2025)
by: Yilmaz, Nilay, et al.
Published: (2025)
DexHoldem: Playing Texas Hold'em with Dexterous Embodied System
by: Chen, Feng, et al.
Published: (2026)
by: Chen, Feng, et al.
Published: (2026)
Seeing the Poem: Image-Semantic Detection of AI-Generated Modern Chinese Poetry with MLLMs
by: Wang, Shanshan, et al.
Published: (2026)
by: Wang, Shanshan, et al.
Published: (2026)
Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
by: Ou, Siqu, et al.
Published: (2026)
by: Ou, Siqu, et al.
Published: (2026)
LongSplat: Online Generalizable 3D Gaussian Splatting from Long Sequence Images
by: Huang, Guichen, et al.
Published: (2025)
by: Huang, Guichen, et al.
Published: (2025)
Mass-Producing Failures of Multimodal Systems with Language Models
by: Tong, Shengbang, et al.
Published: (2023)
by: Tong, Shengbang, et al.
Published: (2023)
CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM
by: Xu, Jingwei, et al.
Published: (2024)
by: Xu, Jingwei, et al.
Published: (2024)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
by: Lu, Yujie, et al.
Published: (2024)
by: Lu, Yujie, et al.
Published: (2024)
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision
by: Li, Chentao, et al.
Published: (2026)
by: Li, Chentao, et al.
Published: (2026)
FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D Reconstruction
by: Xu, Jiale, et al.
Published: (2024)
by: Xu, Jiale, et al.
Published: (2024)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
by: Lin, Jingli, et al.
Published: (2025)
by: Lin, Jingli, et al.
Published: (2025)
Spatial Preference Rewarding for MLLMs Spatial Understanding
by: Qiu, Han, et al.
Published: (2025)
by: Qiu, Han, et al.
Published: (2025)
Universal Skeleton Understanding via Differentiable Rendering and MLLMs
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
Asymmetric Idiosyncrasies in Multimodal Models
by: Tao, Muzi, et al.
Published: (2026)
by: Tao, Muzi, et al.
Published: (2026)
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?
by: Yang, Liu, et al.
Published: (2025)
by: Yang, Liu, et al.
Published: (2025)
PEACE: Empowering Geologic Map Holistic Understanding with MLLMs
by: Huang, Yangyu, et al.
Published: (2025)
by: Huang, Yangyu, et al.
Published: (2025)
Context Tokens are Anchors: Understanding the Repetition Curse in dMLLMs from an Information Flow Perspective
by: Zhao, Qiyan, et al.
Published: (2026)
by: Zhao, Qiyan, et al.
Published: (2026)
Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs
by: Niu, Jingcheng, et al.
Published: (2025)
by: Niu, Jingcheng, et al.
Published: (2025)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
by: Guo, Pinxue, et al.
Published: (2025)
by: Guo, Pinxue, et al.
Published: (2025)
Similar Items
-
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
by: Yu, Yaodong, et al.
Published: (2023) -
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
by: Chu, Tianzhe, et al.
Published: (2025) -
Recollection from Pensieve: Novel View Synthesis via Learning from Uncalibrated Videos
by: Wang, Ruoyu, et al.
Published: (2025) -
Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection
by: Qi, Dacheng, et al.
Published: (2026) -
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024)