LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Man, Yunze, Wang, Shihao, Zhang, Guowen, Bjorck, Johan, Li, Zhiqi, Gui, Liang-Yan, Fan, Jim, Kautz, Jan, Wang, Yu-Xiong, Yu, Zhiding
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910029666320384
author Man, Yunze
Wang, Shihao
Zhang, Guowen
Bjorck, Johan
Li, Zhiqi
Gui, Liang-Yan
Fan, Jim
Kautz, Jan
Wang, Yu-Xiong
Yu, Zhiding
author_facet Man, Yunze
Wang, Shihao
Zhang, Guowen
Bjorck, Johan
Li, Zhiqi
Gui, Liang-Yan
Fan, Jim
Kautz, Jan
Wang, Yu-Xiong
Yu, Zhiding
contents To act in the world, a model must name what it sees and know where it is in 3D. Today's vision-language models (VLMs) excel at open-ended 2D description and grounding, yet multi-object 3D detection remains largely missing from the VLM toolbox. We present LocateAnything3D, a VLM-native recipe that casts 3D detection as a next-token prediction problem. The key is a short, explicit Chain-of-Sight (CoS) sequence that mirrors how human reason from images: find an object in 2D, then infer its distance, size, and pose. The decoder first emits 2D detections as a visual chain-of-thought, then predicts 3D boxes under an easy-to-hard curriculum: across objects, a near-to-far order reduces early ambiguity and matches ego-centric utility; within each object, a center-from-camera, dimensions, and rotation factorization ranks information by stability and learnability. This VLM-native interface preserves open-vocabulary and visual-prompting capability without specialized heads. On the challenging Omni3D benchmark, our model achieves state-of-the-art results, with 38.90 AP_3D, surpassing the previous best by +13.98 absolute improvement even when the baseline is given ground-truth 2D boxes. It also generalizes zero-shot to held-out categories with strong robustness. By turning 3D detection into a disciplined next-token problem, LocateAnything3D offers a practical foundation for models to perceive in 3D.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20648
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
Man, Yunze
Wang, Shihao
Zhang, Guowen
Bjorck, Johan
Li, Zhiqi
Gui, Liang-Yan
Fan, Jim
Kautz, Jan
Wang, Yu-Xiong
Yu, Zhiding
Computer Vision and Pattern Recognition
To act in the world, a model must name what it sees and know where it is in 3D. Today's vision-language models (VLMs) excel at open-ended 2D description and grounding, yet multi-object 3D detection remains largely missing from the VLM toolbox. We present LocateAnything3D, a VLM-native recipe that casts 3D detection as a next-token prediction problem. The key is a short, explicit Chain-of-Sight (CoS) sequence that mirrors how human reason from images: find an object in 2D, then infer its distance, size, and pose. The decoder first emits 2D detections as a visual chain-of-thought, then predicts 3D boxes under an easy-to-hard curriculum: across objects, a near-to-far order reduces early ambiguity and matches ego-centric utility; within each object, a center-from-camera, dimensions, and rotation factorization ranks information by stability and learnability. This VLM-native interface preserves open-vocabulary and visual-prompting capability without specialized heads. On the challenging Omni3D benchmark, our model achieves state-of-the-art results, with 38.90 AP_3D, surpassing the previous best by +13.98 absolute improvement even when the baseline is given ground-truth 2D boxes. It also generalizes zero-shot to held-out categories with strong robustness. By turning 3D detection into a disciplined next-token problem, LocateAnything3D offers a practical foundation for models to perceive in 3D.
title LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.20648