Why MLLMs Struggle to Determine Object Orientations
Fuente:
arXiv
Saved in:
| Main Authors: | Gopinath, Anju, Krishnaswamy, Nikhil, Draper, Bruce |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
by: Zhang, Wanyue, et al.
Published: (2025)
by: Zhang, Wanyue, et al.
Published: (2025)
Are you Struggling? Dataset and Baselines for Struggle Determination in Assembly Videos
by: Feng, Shijia, et al.
Published: (2024)
by: Feng, Shijia, et al.
Published: (2024)
The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
by: Yin, Hao, et al.
Published: (2025)
by: Yin, Hao, et al.
Published: (2025)
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs
by: Tasnim, Nazia, et al.
Published: (2025)
by: Tasnim, Nazia, et al.
Published: (2025)
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs Supplementary
by: Tasnim, Nazia, et al.
Published: (2026)
by: Tasnim, Nazia, et al.
Published: (2026)
Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization
by: Hannan, Darryl, et al.
Published: (2025)
by: Hannan, Darryl, et al.
Published: (2025)
Why Do Vision Language Models Struggle To Recognize Human Emotions?
by: Agarwal, Madhav, et al.
Published: (2026)
by: Agarwal, Madhav, et al.
Published: (2026)
Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment
by: Sarkar, Pritam, et al.
Published: (2024)
by: Sarkar, Pritam, et al.
Published: (2024)
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
by: Yuan, Yuqian, et al.
Published: (2025)
by: Yuan, Yuqian, et al.
Published: (2025)
GOOD: Towards Domain Generalized Orientated Object Detection
by: Bi, Qi, et al.
Published: (2024)
by: Bi, Qi, et al.
Published: (2024)
Determining Fetal Orientations From Blind Sweep Ultrasound Video
by: Wiśniewski, Jakub Maciej, et al.
Published: (2025)
by: Wiśniewski, Jakub Maciej, et al.
Published: (2025)
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding
by: Kim, Jiwan, et al.
Published: (2026)
by: Kim, Jiwan, et al.
Published: (2026)
EvoStruggle: A Dataset Capturing the Evolution of Struggle across Activities and Skill Levels
by: Feng, Shijia, et al.
Published: (2025)
by: Feng, Shijia, et al.
Published: (2025)
R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning
by: Wang, Biao, et al.
Published: (2025)
by: Wang, Biao, et al.
Published: (2025)
Compass Control: Multi Object Orientation Control for Text-to-Image Generation
by: Parihar, Rishubh, et al.
Published: (2025)
by: Parihar, Rishubh, et al.
Published: (2025)
When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs
by: Liu, Jinming, et al.
Published: (2025)
by: Liu, Jinming, et al.
Published: (2025)
From Horizontal to Rotated: Cross-View Object Geo-Localization with Orientation Awareness
by: Fu, Chenlin, et al.
Published: (2026)
by: Fu, Chenlin, et al.
Published: (2026)
On the Role of Domain Experts in Creating Effective Tutoring Systems
by: Sreedharan, Sarath, et al.
Published: (2025)
by: Sreedharan, Sarath, et al.
Published: (2025)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
by: Li, Shuo, et al.
Published: (2025)
by: Li, Shuo, et al.
Published: (2025)
Law of Vision Representation in MLLMs
by: Yang, Shijia, et al.
Published: (2024)
by: Yang, Shijia, et al.
Published: (2024)
Benchmarking Large and Small MLLMs
by: Feng, Xuelu, et al.
Published: (2025)
by: Feng, Xuelu, et al.
Published: (2025)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
OriCon3D: Effective 3D Object Detection using Orientation and Confidence
by: Rajani, Dhyey Manish, et al.
Published: (2023)
by: Rajani, Dhyey Manish, et al.
Published: (2023)
Automated Multi-level Preference for MLLMs
by: Zhang, Mengxi, et al.
Published: (2024)
by: Zhang, Mengxi, et al.
Published: (2024)
Training-Free Reasoning and Reflection in MLLMs
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
Reinforcing Consistency in Video MLLMs with Structured Rewards
by: Quan, Yihao, et al.
Published: (2026)
by: Quan, Yihao, et al.
Published: (2026)
Visual Jigsaw Post-Training Improves MLLMs
by: Wu, Penghao, et al.
Published: (2025)
by: Wu, Penghao, et al.
Published: (2025)
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
by: Tong, Jintao, et al.
Published: (2025)
by: Tong, Jintao, et al.
Published: (2025)
FreeRet: MLLMs as Training-Free Retrievers
by: Zhu, Yuhan, et al.
Published: (2025)
by: Zhu, Yuhan, et al.
Published: (2025)
Aesthetic Image Captioning with Saliency Enhanced MLLMs
by: Tao, Yilin, et al.
Published: (2025)
by: Tao, Yilin, et al.
Published: (2025)
The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
by: Anvekar, Tejas, et al.
Published: (2025)
by: Anvekar, Tejas, et al.
Published: (2025)
Explore How to Inject Beneficial Noise in MLLMs
by: Zhu, Ruishu, et al.
Published: (2025)
by: Zhu, Ruishu, et al.
Published: (2025)
Spatial Preference Rewarding for MLLMs Spatial Understanding
by: Qiu, Han, et al.
Published: (2025)
by: Qiu, Han, et al.
Published: (2025)
Defect Detection in Synthetic Fibre Ropes using Detectron2 Framework
by: Rani, Anju, et al.
Published: (2023)
by: Rani, Anju, et al.
Published: (2023)
Advancements in Point Cloud-Based 3D Defect Detection and Classification for Industrial Systems: A Comprehensive Survey
by: Rani, Anju, et al.
Published: (2024)
by: Rani, Anju, et al.
Published: (2024)
FungalZSL: Zero-Shot Fungal Classification with Image Captioning Using a Synthetic Data Approach
by: Rani, Anju, et al.
Published: (2025)
by: Rani, Anju, et al.
Published: (2025)
Exploiting Unlabeled Data with Multiple Expert Teachers for Open Vocabulary Aerial Object Detection and Its Orientation Adaptation
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
SAM Struggles in Concealed Scenes -- Empirical Study on Segment Anything
by: Ji, Ge-Peng, et al.
Published: (2023)
by: Ji, Ge-Peng, et al.
Published: (2023)
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
by: Jiang, Yankai, et al.
Published: (2026)
by: Jiang, Yankai, et al.
Published: (2026)
Similar Items
-
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
by: Zhang, Wanyue, et al.
Published: (2025) -
Are you Struggling? Dataset and Baselines for Struggle Determination in Assembly Videos
by: Feng, Shijia, et al.
Published: (2024) -
The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
by: Yin, Hao, et al.
Published: (2025) -
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs
by: Tasnim, Nazia, et al.
Published: (2025) -
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs Supplementary
by: Tasnim, Nazia, et al.
Published: (2026)