LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Shihao, Liu, Shilong, Kuang, Yuanguo, Wei, Xinyu, Liu, Yangzhou, Li, Zhiqi, Man, Yunze, Chen, Guo, Tao, Andrew, Liu, Guilin, Kautz, Jan, Zhang, Lei, Yu, Zhiding
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913165331136512
author Wang, Shihao
Liu, Shilong
Kuang, Yuanguo
Wei, Xinyu
Liu, Yangzhou
Li, Zhiqi
Man, Yunze
Chen, Guo
Tao, Andrew
Liu, Guilin
Kautz, Jan
Zhang, Lei
Yu, Zhiding
author_facet Wang, Shihao
Liu, Shilong
Kuang, Yuanguo
Wei, Xinyu
Liu, Yangzhou
Li, Zhiqi
Man, Yunze
Chen, Guo
Tao, Andrew
Liu, Guilin
Kautz, Jan
Zhang, Lei
Yu, Zhiding
contents Vision-language models (VLMs) commonly formulate visual grounding and detection as a coordinate-token generation problem, serializing each 2D box into multiple 1D tokens that are learned and decoded largely independently. This token-by-token decoding mismatches the coupled structure of box geometry and creates a practical inference bottleneck due to strictly sequential generation. We introduce LocateAnything, a unified generative grounding and detection framework based on Parallel Box Decoding (PBD). By decoding geometric elements such as bounding boxes and points as atomic units in a single step, LocateAnything preserves intra-box geometric coherence and unlocks substantial parallelism. We show that PBD improves both decoding throughput and localization accuracy. We further develop a scalable data engine and curate LocateAnything-Data, a large-scale dataset with more than 138 million training samples, substantially increasing data diversity for high-precision localization. Extensive evaluations show that LocateAnything advances the speed-accuracy frontier, achieving significantly higher decoding throughput while improving high-IoU localization quality across diverse benchmarks. The results highlight the complementary benefits of Parallel Box Decoding and large-scale training data in enabling efficient and precise unified visual grounding and detection.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27365
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
Wang, Shihao
Liu, Shilong
Kuang, Yuanguo
Wei, Xinyu
Liu, Yangzhou
Li, Zhiqi
Man, Yunze
Chen, Guo
Tao, Andrew
Liu, Guilin
Kautz, Jan
Zhang, Lei
Yu, Zhiding
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
Vision-language models (VLMs) commonly formulate visual grounding and detection as a coordinate-token generation problem, serializing each 2D box into multiple 1D tokens that are learned and decoded largely independently. This token-by-token decoding mismatches the coupled structure of box geometry and creates a practical inference bottleneck due to strictly sequential generation. We introduce LocateAnything, a unified generative grounding and detection framework based on Parallel Box Decoding (PBD). By decoding geometric elements such as bounding boxes and points as atomic units in a single step, LocateAnything preserves intra-box geometric coherence and unlocks substantial parallelism. We show that PBD improves both decoding throughput and localization accuracy. We further develop a scalable data engine and curate LocateAnything-Data, a large-scale dataset with more than 138 million training samples, substantially increasing data diversity for high-precision localization. Extensive evaluations show that LocateAnything advances the speed-accuracy frontier, achieving significantly higher decoding throughput while improving high-IoU localization quality across diverse benchmarks. The results highlight the complementary benefits of Parallel Box Decoding and large-scale training data in enabling efficient and precise unified visual grounding and detection.
title LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2605.27365