TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Jiaying, Zhan, Zhihao, Zhai, Ruifeng, Lyu, Qinhan, Liu, Hao, Wang, Keze, Lin, Liang, Wang, Guangrun
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914421632139264
author Zhou, Jiaying
Zhan, Zhihao
Zhai, Ruifeng
Lyu, Qinhan
Liu, Hao
Wang, Keze
Lin, Liang
Wang, Guangrun
author_facet Zhou, Jiaying
Zhan, Zhihao
Zhai, Ruifeng
Lyu, Qinhan
Liu, Hao
Wang, Keze
Lin, Liang
Wang, Guangrun
contents Vision--Language--Action (VLA) policies have shown strong progress in mapping language instructions and visual observations to robotic actions, yet their reliability degrades in cluttered scenes with distractors. By analyzing failure cases, we find that many errors do not arise from infeasible motions, but from instance-level grounding failures: the policy often produces a plausible grasp trajectory that lands slightly off-target or even on the wrong object instance. To address this issue, we propose TAG (Target-Agnostic Guidance), a simple inference-time guidance mechanism that explicitly reduces distractor- and appearance-induced bias in VLA policies. Inspired by classifier-free guidance (CFG), TAG contrasts policy predictions under the original observation and an object-erased observation, and uses their difference as a residual steering signal that strengthens the influence of object evidence in the decision process. TAG does not require modifying the policy architecture and can be integrated with existing VLA policies with minimal training and inference changes. We evaluate TAG on standard manipulation benchmarks, including LIBERO, LIBERO-Plus, and VLABench, where it consistently improves robustness under clutter and reduces near-miss and wrong-object executions.
format Preprint
id arxiv_https___arxiv_org_abs_2603_24584
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
Zhou, Jiaying
Zhan, Zhihao
Zhai, Ruifeng
Lyu, Qinhan
Liu, Hao
Wang, Keze
Lin, Liang
Wang, Guangrun
Computer Vision and Pattern Recognition
Robotics
Vision--Language--Action (VLA) policies have shown strong progress in mapping language instructions and visual observations to robotic actions, yet their reliability degrades in cluttered scenes with distractors. By analyzing failure cases, we find that many errors do not arise from infeasible motions, but from instance-level grounding failures: the policy often produces a plausible grasp trajectory that lands slightly off-target or even on the wrong object instance. To address this issue, we propose TAG (Target-Agnostic Guidance), a simple inference-time guidance mechanism that explicitly reduces distractor- and appearance-induced bias in VLA policies. Inspired by classifier-free guidance (CFG), TAG contrasts policy predictions under the original observation and an object-erased observation, and uses their difference as a residual steering signal that strengthens the influence of object evidence in the decision process. TAG does not require modifying the policy architecture and can be integrated with existing VLA policies with minimal training and inference changes. We evaluate TAG on standard manipulation benchmarks, including LIBERO, LIBERO-Plus, and VLABench, where it consistently improves robustness under clutter and reduces near-miss and wrong-object executions.
title TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2603.24584