AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Wenbo, Wang, Shiyi, Chen, Yiteng, Zhuang, Huiping, Wu, Qingyao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915357185277952
author Li, Wenbo
Wang, Shiyi
Chen, Yiteng
Zhuang, Huiping
Wu, Qingyao
author_facet Li, Wenbo
Wang, Shiyi
Chen, Yiteng
Zhuang, Huiping
Wu, Qingyao
contents Vision-Language Models (VLMs) encode knowledge and reasoning capabilities for robotic manipulation within high-dimensional representation spaces. However, current approaches often project them into compressed intermediate representations, discarding important task-specific information such as fine-grained spatial or semantic details. To address this, we propose AntiGrounding, a new framework that reverses the instruction grounding process. It lifts candidate actions directly into the VLM representation space, renders trajectories from multiple views, and uses structured visual question answering for instruction-based decision making. This enables zero-shot synthesis of optimal closed-loop robot trajectories for new tasks. We also propose an offline policy refinement module that leverages past experience to enhance long-term performance. Experiments in both simulation and real-world environments show that our method outperforms baselines across diverse robotic manipulation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12374
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making
Li, Wenbo
Wang, Shiyi
Chen, Yiteng
Zhuang, Huiping
Wu, Qingyao
Robotics
Artificial Intelligence
I.2.9; I.2.10; I.4.8; H.5.2
Vision-Language Models (VLMs) encode knowledge and reasoning capabilities for robotic manipulation within high-dimensional representation spaces. However, current approaches often project them into compressed intermediate representations, discarding important task-specific information such as fine-grained spatial or semantic details. To address this, we propose AntiGrounding, a new framework that reverses the instruction grounding process. It lifts candidate actions directly into the VLM representation space, renders trajectories from multiple views, and uses structured visual question answering for instruction-based decision making. This enables zero-shot synthesis of optimal closed-loop robot trajectories for new tasks. We also propose an offline policy refinement module that leverages past experience to enhance long-term performance. Experiments in both simulation and real-world environments show that our method outperforms baselines across diverse robotic manipulation tasks.
title AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making
topic Robotics
Artificial Intelligence
I.2.9; I.2.10; I.4.8; H.5.2
url https://arxiv.org/abs/2506.12374