A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Daizong, Liu, Yang, Huang, Wencan, Hu, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding
by: Huang, Wencan, et al.
Published: (2025)
by: Huang, Wencan, et al.
Published: (2025)
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends
by: Liu, Daizong, et al.
Published: (2024)
by: Liu, Daizong, et al.
Published: (2024)
Improving the Transferability of 3D Point Cloud Attack via Spectral-aware Admix and Optimization Designs
by: Hu, Shiyu, et al.
Published: (2024)
by: Hu, Shiyu, et al.
Published: (2024)
From Part to Whole: 3D Generative World Model with an Adaptive Structural Hierarchy
by: Du, Bi'an, et al.
Published: (2026)
by: Du, Bi'an, et al.
Published: (2026)
Revealing Directions for Text-guided 3D Face Editing
by: Chen, Zhuo, et al.
Published: (2024)
by: Chen, Zhuo, et al.
Published: (2024)
Talk2PC: Enhancing 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
by: Luo, Zhi, et al.
Published: (2025)
by: Luo, Zhi, et al.
Published: (2025)
Generative AI for Film Creation: A Survey of Recent Advances
by: Zhang, Ruihan, et al.
Published: (2025)
by: Zhang, Ruihan, et al.
Published: (2025)
TSP3D: Text-guided Sparse Voxel Pruning for Efficient 3D Visual Grounding
by: Guo, Wenxuan, et al.
Published: (2025)
by: Guo, Wenxuan, et al.
Published: (2025)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025)
by: Hu, Miao, et al.
Published: (2025)
A Comprehensive Review of 3D Object Detection in Autonomous Driving: Technological Advances and Future Directions
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
Hard-Label Black-Box Attacks on 3D Point Clouds
by: Liu, Daizong, et al.
Published: (2024)
by: Liu, Daizong, et al.
Published: (2024)
Towards Visual Grounding: A Survey
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
Recent Advances in 3D Object and Scene Generation: A Survey
by: Tang, Xiang, et al.
Published: (2025)
by: Tang, Xiang, et al.
Published: (2025)
Vinedresser3D: Agentic Text-guided 3D Editing
by: Chi, Yankuan, et al.
Published: (2026)
by: Chi, Yankuan, et al.
Published: (2026)
Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation
by: Lei, Yinjie, et al.
Published: (2023)
by: Lei, Yinjie, et al.
Published: (2023)
HandMCM: Multi-modal Point Cloud-based Correspondence State Space Model for 3D Hand Pose Estimation
by: Cheng, Wencan, et al.
Published: (2026)
by: Cheng, Wencan, et al.
Published: (2026)
Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And Detection
by: Wei, Guoting, et al.
Published: (2026)
by: Wei, Guoting, et al.
Published: (2026)
Direct Visual Grounding by Directing Attention of Visual Tokens
by: Esmaeilkhani, Parsa, et al.
Published: (2025)
by: Esmaeilkhani, Parsa, et al.
Published: (2025)
Reasoning Matters for 3D Visual Grounding
by: Huang, Hsiang-Wei, et al.
Published: (2026)
by: Huang, Hsiang-Wei, et al.
Published: (2026)
Text-guided Visual Prompt DINO for Generic Segmentation
by: Guan, Yuchen, et al.
Published: (2025)
by: Guan, Yuchen, et al.
Published: (2025)
OpenGround: Active Cognition-based Reasoning for Open-World 3D Visual Grounding
by: Huang, Wenyuan, et al.
Published: (2025)
by: Huang, Wenyuan, et al.
Published: (2025)
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt
by: Yang, Jiahui, et al.
Published: (2024)
by: Yang, Jiahui, et al.
Published: (2024)
Unified Representation Space for 3D Visual Grounding
by: Zheng, Yinuo, et al.
Published: (2025)
by: Zheng, Yinuo, et al.
Published: (2025)
Advances in 4D Generation: A Survey
by: Miao, Qiaowei, et al.
Published: (2025)
by: Miao, Qiaowei, et al.
Published: (2025)
Recent Advances in Medical Imaging Segmentation: A Survey
by: Bougourzi, Fares, et al.
Published: (2025)
by: Bougourzi, Fares, et al.
Published: (2025)
Recent Advances in 3D Gaussian Splatting
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
Semantic Mapping in Indoor Embodied AI -- A Survey on Advances, Challenges, and Future Directions
by: Raychaudhuri, Sonia, et al.
Published: (2025)
by: Raychaudhuri, Sonia, et al.
Published: (2025)
DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D Data
by: Liu, Qihao, et al.
Published: (2024)
by: Liu, Qihao, et al.
Published: (2024)
AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based Referring
by: Wang, Xinyi, et al.
Published: (2025)
by: Wang, Xinyi, et al.
Published: (2025)
Towards Next-Generation SLAM: A Survey on 3DGS-SLAM Focusing on Performance, Robustness, and Future Directions
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
Hierarchical Local-Global Transformer for Temporal Sentence Grounding
by: Fang, Xiang, et al.
Published: (2022)
by: Fang, Xiang, et al.
Published: (2022)
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding
by: Zheng, Henry, et al.
Published: (2025)
by: Zheng, Henry, et al.
Published: (2025)
A Survey on Video Temporal Grounding with Multimodal Large Language Model
by: Wu, Jianlong, et al.
Published: (2025)
by: Wu, Jianlong, et al.
Published: (2025)
Progressive Language-guided Visual Learning for Multi-Task Visual Grounding
by: Wang, Jingchao, et al.
Published: (2025)
by: Wang, Jingchao, et al.
Published: (2025)
Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer Network
by: Fang, Xiang, et al.
Published: (2024)
by: Fang, Xiang, et al.
Published: (2024)
ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition
by: Huang, Ronggang, et al.
Published: (2025)
by: Huang, Ronggang, et al.
Published: (2025)
Similar Items
-
Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding
by: Huang, Wencan, et al.
Published: (2025) -
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding
by: Liu, Yang, et al.
Published: (2024) -
A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends
by: Liu, Daizong, et al.
Published: (2024) -
Improving the Transferability of 3D Point Cloud Attack via Spectral-aware Admix and Optimization Designs
by: Hu, Shiyu, et al.
Published: (2024) -
From Part to Whole: 3D Generative World Model with an Adaptive Structural Hierarchy
by: Du, Bi'an, et al.
Published: (2026)