When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Xianzheng, Smart, Brandon, Bhalgat, Yash, Chen, Shuai, Li, Xinghui, Ding, Jian, Gu, Jindong, Chen, Dave Zhenyu, Peng, Songyou, Bian, Jia-Wang, Torr, Philip H, Pollefeys, Marc, Nießner, Matthias, Reid, Ian D, Chang, Angel X., Laina, Iro, Prisacariu, Victor Adrian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
by: Ma, Xianzheng, et al.
Published: (2026)
by: Ma, Xianzheng, et al.
Published: (2026)
Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs
by: Smart, Brandon, et al.
Published: (2024)
by: Smart, Brandon, et al.
Published: (2024)
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
by: Bhalgat, Yash, et al.
Published: (2024)
by: Bhalgat, Yash, et al.
Published: (2024)
Reflect3r: Single-View 3D Stereo Reconstruction Aided by Mirror Reflections
by: Wu, Jing, et al.
Published: (2025)
by: Wu, Jing, et al.
Published: (2025)
Neural Refinement for Absolute Pose Regression with Feature Synthesis
by: Chen, Shuai, et al.
Published: (2023)
by: Chen, Shuai, et al.
Published: (2023)
GaussCtrl: Multi-View Consistent Text-Driven 3D Gaussian Splatting Editing
by: Wu, Jing, et al.
Published: (2024)
by: Wu, Jing, et al.
Published: (2024)
N2F2: Hierarchical Scene Understanding with Nested Neural Feature Fields
by: Bhalgat, Yash, et al.
Published: (2024)
by: Bhalgat, Yash, et al.
Published: (2024)
GS-CPR: Efficient Camera Pose Refinement via 3D Gaussian Splatting
by: Liu, Changkun, et al.
Published: (2024)
by: Liu, Changkun, et al.
Published: (2024)
DGE: Direct Gaussian 3D Editing by Consistent Multi-view Editing
by: Chen, Minghao, et al.
Published: (2024)
by: Chen, Minghao, et al.
Published: (2024)
Active View Selector: Fast and Accurate Active View Selection with Cross Reference Image Quality Assessment
by: Wang, Zirui, et al.
Published: (2025)
by: Wang, Zirui, et al.
Published: (2025)
Invisible Stitch: Generating Smooth 3D Scenes with Depth Inpainting
by: Engstler, Paul, et al.
Published: (2024)
by: Engstler, Paul, et al.
Published: (2024)
Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos
by: Tschernezki, Vadim, et al.
Published: (2025)
by: Tschernezki, Vadim, et al.
Published: (2025)
Volumetric Semantically Consistent 3D Panoptic Mapping
by: Miao, Yang, et al.
Published: (2023)
by: Miao, Yang, et al.
Published: (2023)
OpenDAS: Open-Vocabulary Domain Adaptation for 2D and 3D Segmentation
by: Yilmaz, Gonca, et al.
Published: (2024)
by: Yilmaz, Gonca, et al.
Published: (2024)
SynCity: Training-Free Generation of 3D Worlds
by: Engstler, Paul, et al.
Published: (2025)
by: Engstler, Paul, et al.
Published: (2025)
WildGaussians: 3D Gaussian Splatting in the Wild
by: Kulhanek, Jonas, et al.
Published: (2024)
by: Kulhanek, Jonas, et al.
Published: (2024)
Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset
by: Wang, Zirui, et al.
Published: (2025)
by: Wang, Zirui, et al.
Published: (2025)
MAP-ADAPT: Real-Time Quality-Adaptive Semantic 3D Maps
by: Zheng, Jianhao, et al.
Published: (2024)
by: Zheng, Jianhao, et al.
Published: (2024)
Reproducibility Study of CDUL: CLIP-Driven Unsupervised Learning for Multi-Label Image Classification
by: Shah, Manan, et al.
Published: (2024)
by: Shah, Manan, et al.
Published: (2024)
PoRF: Pose Residual Field for Accurate Neural Surface Reconstruction
by: Bian, Jia-Wang, et al.
Published: (2023)
by: Bian, Jia-Wang, et al.
Published: (2023)
WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments
by: Zheng, Jianhao, et al.
Published: (2025)
by: Zheng, Jianhao, et al.
Published: (2025)
IM-3D: Iterative Multiview Diffusion and Reconstruction for High-Quality 3D Generation
by: Melas-Kyriazi, Luke, et al.
Published: (2024)
by: Melas-Kyriazi, Luke, et al.
Published: (2024)
3D Neural Edge Reconstruction
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
PartGen: Part-level 3D Generation and Reconstruction with Multi-View Diffusion Models
by: Chen, Minghao, et al.
Published: (2024)
by: Chen, Minghao, et al.
Published: (2024)
CrossOver: 3D Scene Cross-Modal Alignment
by: Sarkar, Sayan Deb, et al.
Published: (2025)
by: Sarkar, Sayan Deb, et al.
Published: (2025)
Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction
by: Jiang, Zeren, et al.
Published: (2025)
by: Jiang, Zeren, et al.
Published: (2025)
Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video
by: Jiang, Zeren, et al.
Published: (2026)
by: Jiang, Zeren, et al.
Published: (2026)
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
by: Li, Haoyuan, et al.
Published: (2025)
by: Li, Haoyuan, et al.
Published: (2025)
AutoPartGen: Autogressive 3D Part Generation and Discovery
by: Chen, Minghao, et al.
Published: (2025)
by: Chen, Minghao, et al.
Published: (2025)
When Do Diffusion Models learn to Generate Multiple Objects?
by: Jeong, Yujin, et al.
Published: (2026)
by: Jeong, Yujin, et al.
Published: (2026)
EPIC Fields: Marrying 3D Geometry and Video Understanding
by: Tschernezki, Vadim, et al.
Published: (2023)
by: Tschernezki, Vadim, et al.
Published: (2023)
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
by: Motamed, Saman, et al.
Published: (2025)
by: Motamed, Saman, et al.
Published: (2025)
No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images
by: Ye, Botao, et al.
Published: (2024)
by: Ye, Botao, et al.
Published: (2024)
Parametric modal regression with error in covariates
by: Qingyang Liu, et al.
Published: (2024)
by: Qingyang Liu, et al.
Published: (2024)
SD4Match: Learning to Prompt Stable Diffusion Model for Semantic Matching
by: Li, Xinghui, et al.
Published: (2023)
by: Li, Xinghui, et al.
Published: (2023)
DreamEdit3D: Personalization of Multi-View Diffusion Models for 3D Editing
by: Ai, Jinxin, et al.
Published: (2026)
by: Ai, Jinxin, et al.
Published: (2026)
Diffusion Models for Open-Vocabulary Segmentation
by: Karazija, Laurynas, et al.
Published: (2023)
by: Karazija, Laurynas, et al.
Published: (2023)
Learning segmentation from point trajectories
by: Karazija, Laurynas, et al.
Published: (2025)
by: Karazija, Laurynas, et al.
Published: (2025)
REACT3D: Recovering Articulations for Interactive Physical 3D Scenes
by: Huang, Zhao, et al.
Published: (2025)
by: Huang, Zhao, et al.
Published: (2025)
Nothing Stands Still: A Spatiotemporal Benchmark on 3D Point Cloud Registration Under Large Geometric and Temporal Change
by: Sun, Tao, et al.
Published: (2023)
by: Sun, Tao, et al.
Published: (2023)
Similar Items
-
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
by: Ma, Xianzheng, et al.
Published: (2026) -
Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs
by: Smart, Brandon, et al.
Published: (2024) -
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
by: Bhalgat, Yash, et al.
Published: (2024) -
Reflect3r: Single-View 3D Stereo Reconstruction Aided by Mirror Reflections
by: Wu, Jing, et al.
Published: (2025) -
Neural Refinement for Absolute Pose Regression with Feature Synthesis
by: Chen, Shuai, et al.
Published: (2023)