Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dongfang, Zihao, Zheng, Xu, Weng, Ziqiao, Lyu, Yuanhuiyi, Paudel, Danda Pani, Van Gool, Luc, Yang, Kailun, Hu, Xuming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PANORAMA: The Rise of Omnidirectional Vision in the Embodied AI Era
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
von: Peng, Kunyu, et al.
Veröffentlicht: (2026)
von: Peng, Kunyu, et al.
Veröffentlicht: (2026)
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
von: Motamed, Saman, et al.
Veröffentlicht: (2023)
von: Motamed, Saman, et al.
Veröffentlicht: (2023)
Autonomous Vehicle Controllers From End-to-End Differentiable Simulation
von: Nachkov, Asen, et al.
Veröffentlicht: (2024)
von: Nachkov, Asen, et al.
Veröffentlicht: (2024)
EvenNICER-SLAM: Event-based Neural Implicit Encoding SLAM
von: Chen, Shi, et al.
Veröffentlicht: (2024)
von: Chen, Shi, et al.
Veröffentlicht: (2024)
Implicit-Zoo: A Large-Scale Dataset of Neural Implicit Functions for 2D Images and 3D Scenes
von: Ma, Qi, et al.
Veröffentlicht: (2024)
von: Ma, Qi, et al.
Veröffentlicht: (2024)
Occam's LGS: An Efficient Approach for Language Gaussian Splatting
von: Cheng, Jiahuan, et al.
Veröffentlicht: (2024)
von: Cheng, Jiahuan, et al.
Veröffentlicht: (2024)
Continuous Pose for Monocular Cameras in Neural Implicit Representation
von: Ma, Qi, et al.
Veröffentlicht: (2023)
von: Ma, Qi, et al.
Veröffentlicht: (2023)
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
von: Mahdi, Mohammad, et al.
Veröffentlicht: (2026)
von: Mahdi, Mohammad, et al.
Veröffentlicht: (2026)
Vision encoders should be image size agnostic and task driven
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2025)
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2025)
Self-supervised pretraining for an iterative image size agnostic vision transformer
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2026)
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2026)
A Simple and Generalist Approach for Panoptic Segmentation
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2024)
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2024)
GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond
von: Halacheva, Anna-Maria, et al.
Veröffentlicht: (2025)
von: Halacheva, Anna-Maria, et al.
Veröffentlicht: (2025)
SeasonScapes: Learning Large-scale Re-lightable 3D Landscapes with Seasonal Variation from Sparse Webcams
von: Kleger, Timo, et al.
Veröffentlicht: (2026)
von: Kleger, Timo, et al.
Veröffentlicht: (2026)
Inferring Compositional 4D Scenes without Ever Seeing One
von: Gokmen, Ahmet Berke, et al.
Veröffentlicht: (2025)
von: Gokmen, Ahmet Berke, et al.
Veröffentlicht: (2025)
Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits
von: Balauca, Ada-Astrid, et al.
Veröffentlicht: (2024)
von: Balauca, Ada-Astrid, et al.
Veröffentlicht: (2024)
RICO: Two Realistic Benchmarks and an In-Depth Analysis for Incremental Learning in Object Detection
von: Neuwirth-Trapp, Matthias, et al.
Veröffentlicht: (2025)
von: Neuwirth-Trapp, Matthias, et al.
Veröffentlicht: (2025)
Incremental Object Detection with Prompt-based Methods
von: Neuwirth-Trapp, Matthias, et al.
Veröffentlicht: (2025)
von: Neuwirth-Trapp, Matthias, et al.
Veröffentlicht: (2025)
EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
Partial CLIP is Enough: Chimera-Seg for Zero-shot Semantic Segmentation
von: Chen, Jialei, et al.
Veröffentlicht: (2025)
von: Chen, Jialei, et al.
Veröffentlicht: (2025)
BiXFormer: A Robust Framework for Maximizing Modality Effectiveness in Multi-Modal Semantic Segmentation
von: Chen, Jialei, et al.
Veröffentlicht: (2025)
von: Chen, Jialei, et al.
Veröffentlicht: (2025)
MultiHaystack: Benchmarking Multimodal Retrieval and Reasoning over 40K Images, Videos, and Documents
von: Xu, Dannong, et al.
Veröffentlicht: (2026)
von: Xu, Dannong, et al.
Veröffentlicht: (2026)
Ternary-Type Opacity and Hybrid Odometry for RGB NeRF-SLAM
von: Lin, Junru, et al.
Veröffentlicht: (2023)
von: Lin, Junru, et al.
Veröffentlicht: (2023)
Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
Benchmarking Multi-modal Semantic Segmentation under Sensor Failures: Missing and Noisy Modality Robustness
von: Liao, Chenfei, et al.
Veröffentlicht: (2025)
von: Liao, Chenfei, et al.
Veröffentlicht: (2025)
ProOOD: Prototype-Guided Out-of-Distribution 3D Occupancy Prediction
von: Zhang, Yuheng, et al.
Veröffentlicht: (2026)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2026)
LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation
von: Miao, Yang, et al.
Veröffentlicht: (2025)
von: Miao, Yang, et al.
Veröffentlicht: (2025)
ReVLA: Reverting Visual Domain Limitation of Robotic Foundation Models
von: Dey, Sombit, et al.
Veröffentlicht: (2024)
von: Dey, Sombit, et al.
Veröffentlicht: (2024)
Autonomous Vehicle Path Planning by Searching With Differentiable Simulation
von: Nachkov, Asen, et al.
Veröffentlicht: (2025)
von: Nachkov, Asen, et al.
Veröffentlicht: (2025)
Unlocking Efficient Vehicle Dynamics Modeling via Analytic World Models
von: Nachkov, Asen, et al.
Veröffentlicht: (2025)
von: Nachkov, Asen, et al.
Veröffentlicht: (2025)
Generalist Robot Manipulation beyond Action Labeled Data
von: Spiridonov, Alexander, et al.
Veröffentlicht: (2025)
von: Spiridonov, Alexander, et al.
Veröffentlicht: (2025)
FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle
von: Markov, Mario, et al.
Veröffentlicht: (2025)
von: Markov, Mario, et al.
Veröffentlicht: (2025)
B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation
von: Markov, Mario, et al.
Veröffentlicht: (2026)
von: Markov, Mario, et al.
Veröffentlicht: (2026)
MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models
von: Huo, Jiahao, et al.
Veröffentlicht: (2025)
von: Huo, Jiahao, et al.
Veröffentlicht: (2025)
CityLoc: 6DoF Pose Distributional Localization for Text Descriptions in Large-Scale Scenes with Gaussian Representation
von: Ma, Qi, et al.
Veröffentlicht: (2025)
von: Ma, Qi, et al.
Veröffentlicht: (2025)
Learning Generative Interactive Environments By Trained Agent Exploration
von: Kazemi, Naser, et al.
Veröffentlicht: (2024)
von: Kazemi, Naser, et al.
Veröffentlicht: (2024)
Rethinking Global Context in Crowd Counting
von: Sun, Guolei, et al.
Veröffentlicht: (2021)
von: Sun, Guolei, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
PANORAMA: The Rise of Omnidirectional Vision in the Embodied AI Era
von: Zheng, Xu, et al.
Veröffentlicht: (2025) -
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
von: Zheng, Xu, et al.
Veröffentlicht: (2025) -
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
von: Zheng, Xu, et al.
Veröffentlicht: (2025) -
Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
von: Zheng, Xu, et al.
Veröffentlicht: (2025) -
Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
von: Peng, Kunyu, et al.
Veröffentlicht: (2026)