Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yuangong, Wong, Wai Keung, Li, Jiaxing, Patras, Ioannis, Zheng, Xu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
Leum-VL Technical Report
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
von: Dua, Karan, et al.
Veröffentlicht: (2025)
von: Dua, Karan, et al.
Veröffentlicht: (2025)
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
von: Oliveira, Daniel, et al.
Veröffentlicht: (2026)
von: Oliveira, Daniel, et al.
Veröffentlicht: (2026)
VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors
von: Lyu, Wenbo, et al.
Veröffentlicht: (2025)
von: Lyu, Wenbo, et al.
Veröffentlicht: (2025)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
von: Liu, Bingnan, et al.
Veröffentlicht: (2026)
von: Liu, Bingnan, et al.
Veröffentlicht: (2026)
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
von: Hou, Zhiyi, et al.
Veröffentlicht: (2025)
von: Hou, Zhiyi, et al.
Veröffentlicht: (2025)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025)
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025)
VLM-VPI: A Vision-Language Reasoning Framework for Improving Automated Vehicle-Pedestrian Interactions
von: Pu, Qingwen, et al.
Veröffentlicht: (2026)
von: Pu, Qingwen, et al.
Veröffentlicht: (2026)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
von: Ji, Binbin, et al.
Veröffentlicht: (2025)
von: Ji, Binbin, et al.
Veröffentlicht: (2025)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
von: Rudman, William, et al.
Veröffentlicht: (2026)
von: Rudman, William, et al.
Veröffentlicht: (2026)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
Transfer-learning for video classification: Video Swin Transformer on multiple domains
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2022)
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2022)
GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models
von: Hacheme, Gilles Quentin, et al.
Veröffentlicht: (2025)
von: Hacheme, Gilles Quentin, et al.
Veröffentlicht: (2025)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
Survey Transfer Learning: Recycling Data with Silicon Responses
von: Amini, Ali
Veröffentlicht: (2025)
von: Amini, Ali
Veröffentlicht: (2025)
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
von: Liu, Zhuoyao, et al.
Veröffentlicht: (2026)
von: Liu, Zhuoyao, et al.
Veröffentlicht: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
von: Zhang, Sinin, et al.
Veröffentlicht: (2026)
von: Zhang, Sinin, et al.
Veröffentlicht: (2026)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
von: Dai, Song, et al.
Veröffentlicht: (2025)
von: Dai, Song, et al.
Veröffentlicht: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)
Pedestrian Detection in Low-Light Conditions: A Comprehensive Survey
von: Ghari, Bahareh, et al.
Veröffentlicht: (2024)
von: Ghari, Bahareh, et al.
Veröffentlicht: (2024)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
von: Yang, Shan
Veröffentlicht: (2026)
von: Yang, Shan
Veröffentlicht: (2026)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents
von: Chen, Junting, et al.
Veröffentlicht: (2024)
von: Chen, Junting, et al.
Veröffentlicht: (2024)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
von: Cui, Shaoyang, et al.
Veröffentlicht: (2026)
von: Cui, Shaoyang, et al.
Veröffentlicht: (2026)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
von: Li, Danyang, et al.
Veröffentlicht: (2025)
von: Li, Danyang, et al.
Veröffentlicht: (2025)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
von: Lim, Shoon Kit, et al.
Veröffentlicht: (2025)
von: Lim, Shoon Kit, et al.
Veröffentlicht: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
von: Guo, Yijie, et al.
Veröffentlicht: (2025)
von: Guo, Yijie, et al.
Veröffentlicht: (2025)
YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
von: Bandyopadhyay, Saptarashmi, et al.
Veröffentlicht: (2025)
von: Bandyopadhyay, Saptarashmi, et al.
Veröffentlicht: (2025)
SSD-GS: Scattering and Shadow Decomposition for Relightable 3D Gaussian Splatting
von: Zheng, Iris, et al.
Veröffentlicht: (2026)
von: Zheng, Iris, et al.
Veröffentlicht: (2026)
MSGS: Multispectral 3D Gaussian Splatting
von: Zheng, Iris, et al.
Veröffentlicht: (2026)
von: Zheng, Iris, et al.
Veröffentlicht: (2026)
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering
von: Pandey, Anupam, et al.
Veröffentlicht: (2025)
von: Pandey, Anupam, et al.
Veröffentlicht: (2025)
SERA-H: Beyond Native Sentinel Spatial Limits for High-Resolution Canopy Height Mapping
von: Boudras, Thomas, et al.
Veröffentlicht: (2025)
von: Boudras, Thomas, et al.
Veröffentlicht: (2025)
OmniAcc: Personalized Accessibility Assistant Using Generative AI
von: Karki, Siddhant, et al.
Veröffentlicht: (2025)
von: Karki, Siddhant, et al.
Veröffentlicht: (2025)
Semi supervised GAN for smart microscopy, fast and data efficient cell cycle classification
von: Manick, Rajeev, et al.
Veröffentlicht: (2026)
von: Manick, Rajeev, et al.
Veröffentlicht: (2026)
Universal Adversarial Attack on Aligned Multimodal LLMs
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025)
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025) -
Leum-VL Technical Report
von: He, Yuxuan, et al.
Veröffentlicht: (2026) -
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
von: Dua, Karan, et al.
Veröffentlicht: (2025) -
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
von: Oliveira, Daniel, et al.
Veröffentlicht: (2026) -
VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors
von: Lyu, Wenbo, et al.
Veröffentlicht: (2025)