Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
Fuente:
arXiv
Guardado en:
| Autores principales: | Gholami, Mohsen, Rezaei, Ahmad, Weimin, Zhou, Mao, Sitong, Zhou, Shunbo, Zhang, Yong, Akbari, Mohammad |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CPPO: Contrastive Perception Policy Optimization for VLM Agents
por: Rezaei, Ahmad, et al.
Publicado: (2026)
por: Rezaei, Ahmad, et al.
Publicado: (2026)
CASP: Compression of Large Multimodal Models Based on Attention Sparsity
por: Gholami, Mohsen, et al.
Publicado: (2025)
por: Gholami, Mohsen, et al.
Publicado: (2025)
Omni-Scene: Omni-Gaussian Representation for Ego-Centric Sparse-View Scene Reconstruction
por: Wei, Dongxu, et al.
Publicado: (2024)
por: Wei, Dongxu, et al.
Publicado: (2024)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
por: Feng, Zhiyuan, et al.
Publicado: (2025)
por: Feng, Zhiyuan, et al.
Publicado: (2025)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
por: Tian, Kexin, et al.
Publicado: (2025)
por: Tian, Kexin, et al.
Publicado: (2025)
Dual-Attention Frequency Fusion at Multi-Scale for Joint Segmentation and Deformable Medical Image Registration
por: Zhou, Hongchao, et al.
Publicado: (2024)
por: Zhou, Hongchao, et al.
Publicado: (2024)
LaWa: Using Latent Space for In-Generation Image Watermarking
por: Rezaei, Ahmad, et al.
Publicado: (2024)
por: Rezaei, Ahmad, et al.
Publicado: (2024)
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
por: Li, Chengzu, et al.
Publicado: (2024)
por: Li, Chengzu, et al.
Publicado: (2024)
EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation
por: Hou, Ruibing, et al.
Publicado: (2026)
por: Hou, Ruibing, et al.
Publicado: (2026)
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding
por: Zheng, Henry, et al.
Publicado: (2025)
por: Zheng, Henry, et al.
Publicado: (2025)
Understanding Dynamic Scenes in Ego Centric 4D Point Clouds
por: Huang, Junsheng, et al.
Publicado: (2025)
por: Huang, Junsheng, et al.
Publicado: (2025)
Deformable Image Registration with Multi-scale Feature Fusion from Shared Encoder, Auxiliary and Pyramid Decoders
por: Zhou, Hongchao, et al.
Publicado: (2024)
por: Zhou, Hongchao, et al.
Publicado: (2024)
MultiEgo: A Multi-View Egocentric Video Dataset for 4D Scene Reconstruction
por: Li, Bate, et al.
Publicado: (2025)
por: Li, Bate, et al.
Publicado: (2025)
Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos
por: Ge, Mengmeng, et al.
Publicado: (2026)
por: Ge, Mengmeng, et al.
Publicado: (2026)
View-on-Graph: Zero-shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs
por: Liu, Yuanyuan, et al.
Publicado: (2025)
por: Liu, Yuanyuan, et al.
Publicado: (2025)
Scale Disparity of Instances in Interactive Point Cloud Segmentation
por: Han, Chenrui, et al.
Publicado: (2024)
por: Han, Chenrui, et al.
Publicado: (2024)
Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning
por: Liang, Dayong, et al.
Publicado: (2025)
por: Liang, Dayong, et al.
Publicado: (2025)
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
por: Hassan, Mariam, et al.
Publicado: (2024)
por: Hassan, Mariam, et al.
Publicado: (2024)
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
por: Lee, Youngwan, et al.
Publicado: (2026)
por: Lee, Youngwan, et al.
Publicado: (2026)
PanopticSplatting: End-to-End Panoptic Gaussian Splatting
por: Xie, Yuxuan, et al.
Publicado: (2025)
por: Xie, Yuxuan, et al.
Publicado: (2025)
PanopticRecon: Leverage Open-vocabulary Instance Segmentation for Zero-shot Panoptic Reconstruction
por: Yu, Xuan, et al.
Publicado: (2024)
por: Yu, Xuan, et al.
Publicado: (2024)
CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning
por: Ma, Wenxin, et al.
Publicado: (2026)
por: Ma, Wenxin, et al.
Publicado: (2026)
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
por: Schäfer, Finn Rasmus, et al.
Publicado: (2026)
por: Schäfer, Finn Rasmus, et al.
Publicado: (2026)
EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning
por: Wang, Zeyu, et al.
Publicado: (2026)
por: Wang, Zeyu, et al.
Publicado: (2026)
Ego to World: Collaborative Spatial Reasoning in Embodied Systems via Reinforcement Learning
por: Zhou, Heng, et al.
Publicado: (2026)
por: Zhou, Heng, et al.
Publicado: (2026)
Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
por: Zhou, Shengchao, et al.
Publicado: (2025)
por: Zhou, Shengchao, et al.
Publicado: (2025)
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
por: Liang, Zhixuan, et al.
Publicado: (2025)
por: Liang, Zhixuan, et al.
Publicado: (2025)
ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos
por: Wu, Peiran, et al.
Publicado: (2025)
por: Wu, Peiran, et al.
Publicado: (2025)
3D Scene Change Modeling With Consistent Multi-View Aggregation
por: Zhou, Zirui, et al.
Publicado: (2025)
por: Zhou, Zirui, et al.
Publicado: (2025)
3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
por: Zhan, Shaoxiong, et al.
Publicado: (2026)
por: Zhan, Shaoxiong, et al.
Publicado: (2026)
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
por: Cheng, An-Chieh, et al.
Publicado: (2024)
por: Cheng, An-Chieh, et al.
Publicado: (2024)
Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
por: Bai, Weimin, et al.
Publicado: (2025)
por: Bai, Weimin, et al.
Publicado: (2025)
Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
por: Mirjalili, Vahid, et al.
Publicado: (2025)
por: Mirjalili, Vahid, et al.
Publicado: (2025)
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
por: Fu, Yuqian, et al.
Publicado: (2024)
por: Fu, Yuqian, et al.
Publicado: (2024)
GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language Model
por: Li, Ling, et al.
Publicado: (2024)
por: Li, Ling, et al.
Publicado: (2024)
Pedestrian Intention Prediction via Vision-Language Foundation Models
por: Azarmi, Mohsen, et al.
Publicado: (2025)
por: Azarmi, Mohsen, et al.
Publicado: (2025)
EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
por: Özsoy, Ege, et al.
Publicado: (2025)
por: Özsoy, Ege, et al.
Publicado: (2025)
Learning Multi-View Spatial Reasoning from Cross-View Relations
por: Jeong, Suchae, et al.
Publicado: (2026)
por: Jeong, Suchae, et al.
Publicado: (2026)
SPFFNet: Strip Perception and Feature Fusion Spatial Pyramid Pooling for Fabric Defect Detection
por: Zhao, Peizhe, et al.
Publicado: (2025)
por: Zhao, Peizhe, et al.
Publicado: (2025)
Know Your Neighbors: Improving Single-View Reconstruction via Spatial Vision-Language Reasoning
por: Li, Rui, et al.
Publicado: (2024)
por: Li, Rui, et al.
Publicado: (2024)
Ejemplares similares
-
CPPO: Contrastive Perception Policy Optimization for VLM Agents
por: Rezaei, Ahmad, et al.
Publicado: (2026) -
CASP: Compression of Large Multimodal Models Based on Attention Sparsity
por: Gholami, Mohsen, et al.
Publicado: (2025) -
Omni-Scene: Omni-Gaussian Representation for Ego-Centric Sparse-View Scene Reconstruction
por: Wei, Dongxu, et al.
Publicado: (2024) -
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
por: Feng, Zhiyuan, et al.
Publicado: (2025) -
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
por: Tian, Kexin, et al.
Publicado: (2025)