Point What You Mean: Visually Grounded Instruction Policy
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918405251006464 |
|---|---|
| author | Yu, Hang Zhao, Juntu Liu, Yufeng Li, Kaiyu Ma, Cheng Zhang, Di Hu, Yingdong Chen, Guang Xie, Junyuan Guo, Junliang Zhao, Junqiao Gao, Yang |
| author_facet | Yu, Hang Zhao, Juntu Liu, Yufeng Li, Kaiyu Ma, Cheng Zhang, Di Hu, Yingdong Chen, Guang Xie, Junyuan Guo, Junliang Zhao, Junqiao Gao, Yang |
| contents | Vision-Language-Action (VLA) models align vision and language with embodied control, but their object referring ability remains limited when relying solely on text prompt, especially in cluttered or out-of-distribution (OOD) scenes. In this study, we introduce the Point-VLA, a plug-and-play policy that augments language instructions with explicit visual cues (e.g., bounding boxes) to resolve referential ambiguity and enable precise object-level grounding. To efficiently scale visually grounded datasets, we further develop an automatic data annotation pipeline requiring minimal human effort. We evaluate Point-VLA on diverse real-world referring tasks and observe consistently stronger performance than text-only instruction VLAs, particularly in cluttered or unseen-object scenarios, with robust generalization. These results demonstrate that Point-VLA effectively resolves object referring ambiguity through pixel-level visual grounding, achieving more generalizable embodied control. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_18933 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Point What You Mean: Visually Grounded Instruction Policy Yu, Hang Zhao, Juntu Liu, Yufeng Li, Kaiyu Ma, Cheng Zhang, Di Hu, Yingdong Chen, Guang Xie, Junyuan Guo, Junliang Zhao, Junqiao Gao, Yang Computer Vision and Pattern Recognition Robotics Vision-Language-Action (VLA) models align vision and language with embodied control, but their object referring ability remains limited when relying solely on text prompt, especially in cluttered or out-of-distribution (OOD) scenes. In this study, we introduce the Point-VLA, a plug-and-play policy that augments language instructions with explicit visual cues (e.g., bounding boxes) to resolve referential ambiguity and enable precise object-level grounding. To efficiently scale visually grounded datasets, we further develop an automatic data annotation pipeline requiring minimal human effort. We evaluate Point-VLA on diverse real-world referring tasks and observe consistently stronger performance than text-only instruction VLAs, particularly in cluttered or unseen-object scenarios, with robust generalization. These results demonstrate that Point-VLA effectively resolves object referring ambiguity through pixel-level visual grounding, achieving more generalizable embodied control. |
| title | Point What You Mean: Visually Grounded Instruction Policy |
| topic | Computer Vision and Pattern Recognition Robotics |
| url | https://arxiv.org/abs/2512.18933 |