Point What You Mean: Visually Grounded Instruction Policy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Hang, Zhao, Juntu, Liu, Yufeng, Li, Kaiyu, Ma, Cheng, Zhang, Di, Hu, Yingdong, Chen, Guang, Xie, Junyuan, Guo, Junliang, Zhao, Junqiao, Gao, Yang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918405251006464
author Yu, Hang
Zhao, Juntu
Liu, Yufeng
Li, Kaiyu
Ma, Cheng
Zhang, Di
Hu, Yingdong
Chen, Guang
Xie, Junyuan
Guo, Junliang
Zhao, Junqiao
Gao, Yang
author_facet Yu, Hang
Zhao, Juntu
Liu, Yufeng
Li, Kaiyu
Ma, Cheng
Zhang, Di
Hu, Yingdong
Chen, Guang
Xie, Junyuan
Guo, Junliang
Zhao, Junqiao
Gao, Yang
contents Vision-Language-Action (VLA) models align vision and language with embodied control, but their object referring ability remains limited when relying solely on text prompt, especially in cluttered or out-of-distribution (OOD) scenes. In this study, we introduce the Point-VLA, a plug-and-play policy that augments language instructions with explicit visual cues (e.g., bounding boxes) to resolve referential ambiguity and enable precise object-level grounding. To efficiently scale visually grounded datasets, we further develop an automatic data annotation pipeline requiring minimal human effort. We evaluate Point-VLA on diverse real-world referring tasks and observe consistently stronger performance than text-only instruction VLAs, particularly in cluttered or unseen-object scenarios, with robust generalization. These results demonstrate that Point-VLA effectively resolves object referring ambiguity through pixel-level visual grounding, achieving more generalizable embodied control.
format Preprint
id arxiv_https___arxiv_org_abs_2512_18933
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Point What You Mean: Visually Grounded Instruction Policy
Yu, Hang
Zhao, Juntu
Liu, Yufeng
Li, Kaiyu
Ma, Cheng
Zhang, Di
Hu, Yingdong
Chen, Guang
Xie, Junyuan
Guo, Junliang
Zhao, Junqiao
Gao, Yang
Computer Vision and Pattern Recognition
Robotics
Vision-Language-Action (VLA) models align vision and language with embodied control, but their object referring ability remains limited when relying solely on text prompt, especially in cluttered or out-of-distribution (OOD) scenes. In this study, we introduce the Point-VLA, a plug-and-play policy that augments language instructions with explicit visual cues (e.g., bounding boxes) to resolve referential ambiguity and enable precise object-level grounding. To efficiently scale visually grounded datasets, we further develop an automatic data annotation pipeline requiring minimal human effort. We evaluate Point-VLA on diverse real-world referring tasks and observe consistently stronger performance than text-only instruction VLAs, particularly in cluttered or unseen-object scenarios, with robust generalization. These results demonstrate that Point-VLA effectively resolves object referring ambiguity through pixel-level visual grounding, achieving more generalizable embodied control.
title Point What You Mean: Visually Grounded Instruction Policy
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2512.18933