Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yeh, Chun-Hsiao, Qian, Shengyi, Wang, Manchen, Ma, Yi, Tighe, Joseph, Xiao, Fanyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
von: Ma, Chenyang, et al.
Veröffentlicht: (2024)
von: Ma, Chenyang, et al.
Veröffentlicht: (2024)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
von: Xu, Zhenlin, et al.
Veröffentlicht: (2023)
von: Xu, Zhenlin, et al.
Veröffentlicht: (2023)
PointVLA: Injecting the 3D World into Vision-Language-Action Models
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models
von: Wang, Xiaoyan, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoyan, et al.
Veröffentlicht: (2025)
Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration
von: Cai, Zhongyi, et al.
Veröffentlicht: (2025)
von: Cai, Zhongyi, et al.
Veröffentlicht: (2025)
N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning
von: Liang, Dayong, et al.
Veröffentlicht: (2025)
von: Liang, Dayong, et al.
Veröffentlicht: (2025)
3D-MVP: 3D Multiview Pretraining for Robotic Manipulation
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
Enhancing 3D Lane Detection and Topology Reasoning with 2D Lane Priors
von: Li, Han, et al.
Veröffentlicht: (2024)
von: Li, Han, et al.
Veröffentlicht: (2024)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
von: Ma, Wufei, et al.
Veröffentlicht: (2025)
von: Ma, Wufei, et al.
Veröffentlicht: (2025)
Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations
von: Yuan, Jiangye, et al.
Veröffentlicht: (2026)
von: Yuan, Jiangye, et al.
Veröffentlicht: (2026)
Multi-Object 3D Grounding with Dynamic Modules and Language-Informed Spatial Attention
von: Zhang, Haomeng, et al.
Veröffentlicht: (2024)
von: Zhang, Haomeng, et al.
Veröffentlicht: (2024)
Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT
von: Asfour, Alaa, et al.
Veröffentlicht: (2026)
von: Asfour, Alaa, et al.
Veröffentlicht: (2026)
Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
von: Zheng, Duo, et al.
Veröffentlicht: (2025)
von: Zheng, Duo, et al.
Veröffentlicht: (2025)
Leveraging 3D Geometric Priors in 2D Rotation Symmetry Detection
von: Seo, Ahyun, et al.
Veröffentlicht: (2025)
von: Seo, Ahyun, et al.
Veröffentlicht: (2025)
AnchorSplat: Feed-Forward 3D Gaussian Splatting with 3D Geometric Priors
von: Zhang, Xiaoxue, et al.
Veröffentlicht: (2026)
von: Zhang, Xiaoxue, et al.
Veröffentlicht: (2026)
Unleashing Semantic and Geometric Priors for 3D Scene Completion
von: Chen, Shiyuan, et al.
Veröffentlicht: (2025)
von: Chen, Shiyuan, et al.
Veröffentlicht: (2025)
MGP-KAD: Multimodal Geometric Priors and Kolmogorov-Arnold Decoder for Single-View 3D Reconstruction in Complex Scenes
von: Zhang, Luoxi, et al.
Veröffentlicht: (2026)
von: Zhang, Luoxi, et al.
Veröffentlicht: (2026)
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
von: Xia, Zhongyu, et al.
Veröffentlicht: (2026)
von: Xia, Zhongyu, et al.
Veröffentlicht: (2026)
Grounded 3D-Aware Spatial Vision-Language Modeling
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2026)
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2026)
Resolving Long-Tail Ambiguity in Unsupervised 3D Point Cloud Segmentation with Language Priors
von: Wei, Siqi, et al.
Veröffentlicht: (2026)
von: Wei, Siqi, et al.
Veröffentlicht: (2026)
Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2025)
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2025)
Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation
von: Li, Kailing, et al.
Veröffentlicht: (2026)
von: Li, Kailing, et al.
Veröffentlicht: (2026)
Vision-Language Memory for Spatial Reasoning
von: Liu, Zuntao, et al.
Veröffentlicht: (2025)
von: Liu, Zuntao, et al.
Veröffentlicht: (2025)
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulation
von: Xu, Chongyang, et al.
Veröffentlicht: (2026)
von: Xu, Chongyang, et al.
Veröffentlicht: (2026)
HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
von: Liang, Huizhi, et al.
Veröffentlicht: (2026)
von: Liang, Huizhi, et al.
Veröffentlicht: (2026)
3D Scene Graph Guided Vision-Language Pre-training
von: Liu, Hao, et al.
Veröffentlicht: (2024)
von: Liu, Hao, et al.
Veröffentlicht: (2024)
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
von: Zhang, Jian, et al.
Veröffentlicht: (2026)
von: Zhang, Jian, et al.
Veröffentlicht: (2026)
Advancing Vision Transformer with Enhanced Spatial Priors
von: Fan, Qihang, et al.
Veröffentlicht: (2026)
von: Fan, Qihang, et al.
Veröffentlicht: (2026)
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
von: Liu, Yifan, et al.
Veröffentlicht: (2025)
von: Liu, Yifan, et al.
Veröffentlicht: (2025)
SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes
von: Huang, Jiaxin, et al.
Veröffentlicht: (2025)
von: Huang, Jiaxin, et al.
Veröffentlicht: (2025)
PanAdapter: Two-Stage Fine-Tuning with Spatial-Spectral Priors Injecting for Pansharpening
von: Wu, RuoCheng, et al.
Veröffentlicht: (2024)
von: Wu, RuoCheng, et al.
Veröffentlicht: (2024)
VoxRep: Enhancing 3D Spatial Understanding in 2D Vision-Language Models via Voxel Representation
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
von: Zhou, Shengchao, et al.
Veröffentlicht: (2025)
von: Zhou, Shengchao, et al.
Veröffentlicht: (2025)
Enhancing 2D Representation Learning with a 3D Prior
von: Aygün, Mehmet, et al.
Veröffentlicht: (2024)
von: Aygün, Mehmet, et al.
Veröffentlicht: (2024)
Beyond Flatlands: Unlocking Spatial Intelligence by Decoupling 3D Reasoning from Numerical Regression
von: Guo, Zhongbin, et al.
Veröffentlicht: (2025)
von: Guo, Zhongbin, et al.
Veröffentlicht: (2025)
DUSt3R: Geometric 3D Vision Made Easy
von: Wang, Shuzhe, et al.
Veröffentlicht: (2023)
von: Wang, Shuzhe, et al.
Veröffentlicht: (2023)
On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
von: Bratulić, Jelena, et al.
Veröffentlicht: (2025)
von: Bratulić, Jelena, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
von: Ma, Chenyang, et al.
Veröffentlicht: (2024) -
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
von: Xu, Zhenlin, et al.
Veröffentlicht: (2023) -
PointVLA: Injecting the 3D World into Vision-Language-Action Models
von: Li, Chengmeng, et al.
Veröffentlicht: (2025) -
Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models
von: Wang, Xiaoyan, et al.
Veröffentlicht: (2025) -
Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration
von: Cai, Zhongyi, et al.
Veröffentlicht: (2025)