AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915357185277952 |
|---|---|
| author | Li, Wenbo Wang, Shiyi Chen, Yiteng Zhuang, Huiping Wu, Qingyao |
| author_facet | Li, Wenbo Wang, Shiyi Chen, Yiteng Zhuang, Huiping Wu, Qingyao |
| contents | Vision-Language Models (VLMs) encode knowledge and reasoning capabilities for robotic manipulation within high-dimensional representation spaces. However, current approaches often project them into compressed intermediate representations, discarding important task-specific information such as fine-grained spatial or semantic details. To address this, we propose AntiGrounding, a new framework that reverses the instruction grounding process. It lifts candidate actions directly into the VLM representation space, renders trajectories from multiple views, and uses structured visual question answering for instruction-based decision making. This enables zero-shot synthesis of optimal closed-loop robot trajectories for new tasks. We also propose an offline policy refinement module that leverages past experience to enhance long-term performance. Experiments in both simulation and real-world environments show that our method outperforms baselines across diverse robotic manipulation tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_12374 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Li, Wenbo Wang, Shiyi Chen, Yiteng Zhuang, Huiping Wu, Qingyao Robotics Artificial Intelligence I.2.9; I.2.10; I.4.8; H.5.2 Vision-Language Models (VLMs) encode knowledge and reasoning capabilities for robotic manipulation within high-dimensional representation spaces. However, current approaches often project them into compressed intermediate representations, discarding important task-specific information such as fine-grained spatial or semantic details. To address this, we propose AntiGrounding, a new framework that reverses the instruction grounding process. It lifts candidate actions directly into the VLM representation space, renders trajectories from multiple views, and uses structured visual question answering for instruction-based decision making. This enables zero-shot synthesis of optimal closed-loop robot trajectories for new tasks. We also propose an offline policy refinement module that leverages past experience to enhance long-term performance. Experiments in both simulation and real-world environments show that our method outperforms baselines across diverse robotic manipulation tasks. |
| title | AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making |
| topic | Robotics Artificial Intelligence I.2.9; I.2.10; I.4.8; H.5.2 |
| url | https://arxiv.org/abs/2506.12374 |