Language-to-Space Programming for Training-Free 3D Visual Grounding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mi, Boyu, Wang, Hanqing, Wang, Tai, Chen, Yilun, Pang, Jiangmiao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916922514210816
author Mi, Boyu
Wang, Hanqing
Wang, Tai
Chen, Yilun
Pang, Jiangmiao
author_facet Mi, Boyu
Wang, Hanqing
Wang, Tai
Chen, Yilun
Pang, Jiangmiao
contents 3D visual grounding (3DVG) is challenging due to the need to understand 3D spatial relations. While supervised approaches have achieved superior performance, they are constrained by the scarcity and high annotation costs of 3D vision-language datasets. Training-free approaches based on LLMs/VLMs eliminate the need for large-scale training data, but they either incur prohibitive grounding time and token costs or have unsatisfactory accuracy. To address the challenges, we introduce a novel method for training-free 3D visual grounding, namely Language-to-Space Programming (LaSP). LaSP introduces LLM-generated codes to analyze 3D spatial relations among objects, along with a pipeline that evaluates and optimizes the codes automatically. Experimental results demonstrate that LaSP achieves 52.9% accuracy on the Nr3D benchmark, ranking among the best training-free methods. Moreover, it substantially reduces the grounding time and token costs, offering a balanced trade-off between performance and efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01401
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language-to-Space Programming for Training-Free 3D Visual Grounding
Mi, Boyu
Wang, Hanqing
Wang, Tai
Chen, Yilun
Pang, Jiangmiao
Computer Vision and Pattern Recognition
3D visual grounding (3DVG) is challenging due to the need to understand 3D spatial relations. While supervised approaches have achieved superior performance, they are constrained by the scarcity and high annotation costs of 3D vision-language datasets. Training-free approaches based on LLMs/VLMs eliminate the need for large-scale training data, but they either incur prohibitive grounding time and token costs or have unsatisfactory accuracy. To address the challenges, we introduce a novel method for training-free 3D visual grounding, namely Language-to-Space Programming (LaSP). LaSP introduces LLM-generated codes to analyze 3D spatial relations among objects, along with a pipeline that evaluates and optimizes the codes automatically. Experimental results demonstrate that LaSP achieves 52.9% accuracy on the Nr3D benchmark, ranking among the best training-free methods. Moreover, it substantially reduces the grounding time and token costs, offering a balanced trade-off between performance and efficiency.
title Language-to-Space Programming for Training-Free 3D Visual Grounding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.01401