Do 3D Large Language Models Really Understand 3D Spatial Relationships?
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Xianzheng, Sun, Tao, Chen, Shuai, Bhalgat, Yash, Gu, Jindong, Chang, Angel X, Armeni, Iro, Laina, Iro, Peng, Songyou, Prisacariu, Victor Adrian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reflect3r: Single-View 3D Stereo Reconstruction Aided by Mirror Reflections
by: Wu, Jing, et al.
Published: (2025)
by: Wu, Jing, et al.
Published: (2025)
When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models
by: Ma, Xianzheng, et al.
Published: (2024)
by: Ma, Xianzheng, et al.
Published: (2024)
Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs
by: Smart, Brandon, et al.
Published: (2024)
by: Smart, Brandon, et al.
Published: (2024)
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
by: Bhalgat, Yash, et al.
Published: (2024)
by: Bhalgat, Yash, et al.
Published: (2024)
N2F2: Hierarchical Scene Understanding with Nested Neural Feature Fields
by: Bhalgat, Yash, et al.
Published: (2024)
by: Bhalgat, Yash, et al.
Published: (2024)
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
by: Bucher, Martin JJ., et al.
Published: (2025)
by: Bucher, Martin JJ., et al.
Published: (2025)
DGE: Direct Gaussian 3D Editing by Consistent Multi-view Editing
by: Chen, Minghao, et al.
Published: (2024)
by: Chen, Minghao, et al.
Published: (2024)
SGAligner++: Cross-Modal Language-Aided 3D Scene Graph Alignment
by: Singh, Binod, et al.
Published: (2025)
by: Singh, Binod, et al.
Published: (2025)
Volumetric Semantically Consistent 3D Panoptic Mapping
by: Miao, Yang, et al.
Published: (2023)
by: Miao, Yang, et al.
Published: (2023)
MAP-ADAPT: Real-Time Quality-Adaptive Semantic 3D Maps
by: Zheng, Jianhao, et al.
Published: (2024)
by: Zheng, Jianhao, et al.
Published: (2024)
Living Scenes: Multi-object Relocalization and Reconstruction in Changing 3D Environments
by: Zhu, Liyuan, et al.
Published: (2023)
by: Zhu, Liyuan, et al.
Published: (2023)
Invisible Stitch: Generating Smooth 3D Scenes with Depth Inpainting
by: Engstler, Paul, et al.
Published: (2024)
by: Engstler, Paul, et al.
Published: (2024)
Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos
by: Tschernezki, Vadim, et al.
Published: (2025)
by: Tschernezki, Vadim, et al.
Published: (2025)
GuideFlow3D: Optimization-Guided Rectified Flow For Appearance Transfer
by: Sarkar, Sayan Deb, et al.
Published: (2025)
by: Sarkar, Sayan Deb, et al.
Published: (2025)
WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments
by: Zheng, Jianhao, et al.
Published: (2025)
by: Zheng, Jianhao, et al.
Published: (2025)
Active View Selector: Fast and Accurate Active View Selection with Cross Reference Image Quality Assessment
by: Wang, Zirui, et al.
Published: (2025)
by: Wang, Zirui, et al.
Published: (2025)
GS-CPR: Efficient Camera Pose Refinement via 3D Gaussian Splatting
by: Liu, Changkun, et al.
Published: (2024)
by: Liu, Changkun, et al.
Published: (2024)
SynCity: Training-Free Generation of 3D Worlds
by: Engstler, Paul, et al.
Published: (2025)
by: Engstler, Paul, et al.
Published: (2025)
CrossOver: 3D Scene Cross-Modal Alignment
by: Sarkar, Sayan Deb, et al.
Published: (2025)
by: Sarkar, Sayan Deb, et al.
Published: (2025)
ReScene4D: Temporally Consistent Semantic Instance Segmentation of Evolving Indoor 3D Scenes
by: Steiner, Emily, et al.
Published: (2026)
by: Steiner, Emily, et al.
Published: (2026)
GaussFusion: Improving 3D Reconstruction in the Wild with A Geometry-Informed Video Generator
by: Zhu, Liyuan, et al.
Published: (2026)
by: Zhu, Liyuan, et al.
Published: (2026)
LoopSplat: Loop Closure by Registering 3D Gaussian Splats
by: Zhu, Liyuan, et al.
Published: (2024)
by: Zhu, Liyuan, et al.
Published: (2024)
EPIC Fields: Marrying 3D Geometry and Video Understanding
by: Tschernezki, Vadim, et al.
Published: (2023)
by: Tschernezki, Vadim, et al.
Published: (2023)
Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction
by: Jiang, Zeren, et al.
Published: (2025)
by: Jiang, Zeren, et al.
Published: (2025)
Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video
by: Jiang, Zeren, et al.
Published: (2026)
by: Jiang, Zeren, et al.
Published: (2026)
ReStyle3D: Scene-Level Appearance Transfer with Semantic Correspondences
by: Zhu, Liyuan, et al.
Published: (2025)
by: Zhu, Liyuan, et al.
Published: (2025)
Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
by: Pan, Yue, et al.
Published: (2025)
by: Pan, Yue, et al.
Published: (2025)
IM-3D: Iterative Multiview Diffusion and Reconstruction for High-Quality 3D Generation
by: Melas-Kyriazi, Luke, et al.
Published: (2024)
by: Melas-Kyriazi, Luke, et al.
Published: (2024)
Nothing Stands Still: A Spatiotemporal Benchmark on 3D Point Cloud Registration Under Large Geometric and Temporal Change
by: Sun, Tao, et al.
Published: (2023)
by: Sun, Tao, et al.
Published: (2023)
Deep Sketch-Based 3D Modeling: A Survey
by: Tono, Alberto, et al.
Published: (2026)
by: Tono, Alberto, et al.
Published: (2026)
Deep Sketch‐Based 3D Modeling: A Survey
by: Alberto Tono, et al.
Published: (2026)
by: Alberto Tono, et al.
Published: (2026)
Multiway Point Cloud Mosaicking with Diffusion and Global Optimization
by: Jin, Shengze, et al.
Published: (2024)
by: Jin, Shengze, et al.
Published: (2024)
HouseTour: A Virtual Real Estate A(I)gent
by: Çelen, Ata, et al.
Published: (2025)
by: Çelen, Ata, et al.
Published: (2025)
WildPose: A Unified Framework for Robust Pose Estimation in the Wild
by: Zheng, Jianhao, et al.
Published: (2026)
by: Zheng, Jianhao, et al.
Published: (2026)
Neural Refinement for Absolute Pose Regression with Feature Synthesis
by: Chen, Shuai, et al.
Published: (2023)
by: Chen, Shuai, et al.
Published: (2023)
When Do Diffusion Models learn to Generate Multiple Objects?
by: Jeong, Yujin, et al.
Published: (2026)
by: Jeong, Yujin, et al.
Published: (2026)
Rectified Point Flow: Generic Point Cloud Pose Estimation
by: Sun, Tao, et al.
Published: (2025)
by: Sun, Tao, et al.
Published: (2025)
Diffusion Models for Open-Vocabulary Segmentation
by: Karazija, Laurynas, et al.
Published: (2023)
by: Karazija, Laurynas, et al.
Published: (2023)
Learning segmentation from point trajectories
by: Karazija, Laurynas, et al.
Published: (2025)
by: Karazija, Laurynas, et al.
Published: (2025)
PartGen: Part-level 3D Generation and Reconstruction with Multi-View Diffusion Models
by: Chen, Minghao, et al.
Published: (2024)
by: Chen, Minghao, et al.
Published: (2024)
Similar Items
-
Reflect3r: Single-View 3D Stereo Reconstruction Aided by Mirror Reflections
by: Wu, Jing, et al.
Published: (2025) -
When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models
by: Ma, Xianzheng, et al.
Published: (2024) -
Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs
by: Smart, Brandon, et al.
Published: (2024) -
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
by: Bhalgat, Yash, et al.
Published: (2024) -
N2F2: Hierarchical Scene Understanding with Nested Neural Feature Fields
by: Bhalgat, Yash, et al.
Published: (2024)