LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Kong, Fei, Duan, Jinhao, Xu, Kaidi, Guo, Zhenhua, Zhu, Xiaofeng, Shi, Xiaoshuang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
by: Duan, Jinhao, et al.
Published: (2025)
by: Duan, Jinhao, et al.
Published: (2025)
ACT-Diffusion: Efficient Adversarial Consistency Training for One-step Diffusion Models
by: Kong, Fei, et al.
Published: (2023)
by: Kong, Fei, et al.
Published: (2023)
COIN: Uncertainty-Guarding Selective Question Answering for Foundation Models with Provable Risk Guarantees
by: Wang, Zhiyuan, et al.
Published: (2025)
by: Wang, Zhiyuan, et al.
Published: (2025)
ConU: Conformal Uncertainty in Large Language Models with Correctness Coverage Guarantees
by: Wang, Zhiyuan, et al.
Published: (2024)
by: Wang, Zhiyuan, et al.
Published: (2024)
IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model Generation
by: Fan, Haozhi, et al.
Published: (2026)
by: Fan, Haozhi, et al.
Published: (2026)
SConU: Selective Conformal Uncertainty in Large Language Models
by: Wang, Zhiyuan, et al.
Published: (2025)
by: Wang, Zhiyuan, et al.
Published: (2025)
Conformal Lesion Segmentation for 3D Medical Images
by: Tan, Binyu, et al.
Published: (2025)
by: Tan, Binyu, et al.
Published: (2025)
Caterpillar: A Pure-MLP Architecture with Shifted-Pillars-Concatenation
by: Sun, Jin, et al.
Published: (2023)
by: Sun, Jin, et al.
Published: (2023)
AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
by: An, Shengnan, et al.
Published: (2025)
by: An, Shengnan, et al.
Published: (2025)
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
by: Du, Mengfei, et al.
Published: (2024)
by: Du, Mengfei, et al.
Published: (2024)
Frontier LLMs Still Struggle with Simple Reasoning Tasks
by: Malek, Alan, et al.
Published: (2025)
by: Malek, Alan, et al.
Published: (2025)
Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMs
by: Zhang, Zhuoxuan, et al.
Published: (2025)
by: Zhang, Zhuoxuan, et al.
Published: (2025)
Word-Sequence Entropy: Towards Uncertainty Estimation in Free-Form Medical Question Answering Applications and Beyond
by: Wang, Zhiyuan, et al.
Published: (2024)
by: Wang, Zhiyuan, et al.
Published: (2024)
Using Left and Right Brains Together: Towards Vision and Language Planning
by: Cen, Jun, et al.
Published: (2024)
by: Cen, Jun, et al.
Published: (2024)
Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Model
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
D2-LRR: A Dual-Decomposed MDLatLRR Approach for Medical Image Fusion
by: Song, Xu, et al.
Published: (2022)
by: Song, Xu, et al.
Published: (2022)
Design and implementation of tilting‐plate drainage equipment for paddy field
by: Zhenhua Duan, et al.
Published: (2024)
by: Zhenhua Duan, et al.
Published: (2024)
A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly
by: Yao, Yifan, et al.
Published: (2023)
by: Yao, Yifan, et al.
Published: (2023)
GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts
by: Kargaran, Amir Hossein, et al.
Published: (2026)
by: Kargaran, Amir Hossein, et al.
Published: (2026)
DynaCode: A Dynamic Complexity-Aware Code Benchmark for Evaluating Large Language Models in Code Generation
by: Hu, Wenhao, et al.
Published: (2025)
by: Hu, Wenhao, et al.
Published: (2025)
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
by: Kamoi, Ryo, et al.
Published: (2024)
by: Kamoi, Ryo, et al.
Published: (2024)
LRR: Language-Driven Resamplable Continuous Representation against Adversarial Tracking Attacks
by: Chen, Jianlang, et al.
Published: (2024)
by: Chen, Jianlang, et al.
Published: (2024)
The Imitation Game: Using Large Language Models as Chatbots to Combat Chat-Based Cybercrimes
by: Yao, Yifan, et al.
Published: (2025)
by: Yao, Yifan, et al.
Published: (2025)
Unlearnable Examples Give a False Sense of Data Privacy: Understanding and Relearning
by: Dang, Pucheng, et al.
Published: (2023)
by: Dang, Pucheng, et al.
Published: (2023)
The Animal Rights Struggle
by: Traïni, Christophe Traïni
Published: (2017)
by: Traïni, Christophe Traïni
Published: (2017)
Left-Right Symmetry Breaking in CLIP-style Vision-Language Models Trained on Synthetic Spatial-Relation Data
by: Yamamoto, Takaki, et al.
Published: (2026)
by: Yamamoto, Takaki, et al.
Published: (2026)
Language Models Still Struggle to Zero-shot Reason about Time Series
by: Merrill, Mike A., et al.
Published: (2024)
by: Merrill, Mike A., et al.
Published: (2024)
Bi-Level Control of Weaving Sections in Mixed Traffic Environments with Connected and Automated Vehicles
by: Yan, Longhao, et al.
Published: (2024)
by: Yan, Longhao, et al.
Published: (2024)
Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
by: Duan, Jinhao, et al.
Published: (2023)
by: Duan, Jinhao, et al.
Published: (2023)
UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making
by: Duan, Jinhao, et al.
Published: (2025)
by: Duan, Jinhao, et al.
Published: (2025)
Self-supervised Speech Representations Still Struggle with African American Vernacular English
by: Chang, Kalvin, et al.
Published: (2024)
by: Chang, Kalvin, et al.
Published: (2024)
Rotation Equivariant Mamba for Vision Tasks
by: Zhao, Zhongchen, et al.
Published: (2026)
by: Zhao, Zhongchen, et al.
Published: (2026)
SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
by: Guo, Xianda, et al.
Published: (2024)
by: Guo, Xianda, et al.
Published: (2024)
Enhancing the Influence of Labels on Unlabeled Nodes in Graph Convolutional Networks
by: Huang, Jincheng, et al.
Published: (2024)
by: Huang, Jincheng, et al.
Published: (2024)
The Final Layer Holds the Key: A Unified and Efficient GNN Calibration Framework
by: Huang, Jincheng, et al.
Published: (2025)
by: Huang, Jincheng, et al.
Published: (2025)
High nitrogen fertilization suppresses the NtLRR‐RK4 defense pathway to enhance susceptibility to Alternaria alternata in flue‐cured tobacco
by: Meiwei Zhao, et al.
Published: (2026)
by: Meiwei Zhao, et al.
Published: (2026)
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
by: Zhang, Pingrui, et al.
Published: (2025)
by: Zhang, Pingrui, et al.
Published: (2025)
Can Protective Perturbation Safeguard Personal Data from Being Exploited by Stable Diffusion?
by: Zhao, Zhengyue, et al.
Published: (2023)
by: Zhao, Zhengyue, et al.
Published: (2023)
What's Not Said Still Hurts: A Description-Based Evaluation Framework for Measuring Social Bias in LLMs
by: Pan, Jinhao, et al.
Published: (2025)
by: Pan, Jinhao, et al.
Published: (2025)
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
by: Zhang, Wanyue, et al.
Published: (2025)
by: Zhang, Wanyue, et al.
Published: (2025)
Similar Items
-
TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
by: Duan, Jinhao, et al.
Published: (2025) -
ACT-Diffusion: Efficient Adversarial Consistency Training for One-step Diffusion Models
by: Kong, Fei, et al.
Published: (2023) -
COIN: Uncertainty-Guarding Selective Question Answering for Foundation Models with Provable Risk Guarantees
by: Wang, Zhiyuan, et al.
Published: (2025) -
ConU: Conformal Uncertainty in Large Language Models with Correctness Coverage Guarantees
by: Wang, Zhiyuan, et al.
Published: (2024) -
IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model Generation
by: Fan, Haozhi, et al.
Published: (2026)