Rethinking Intermediate Representation for VLM-based Robot Manipulation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915635307479040 |
|---|---|
| author | Tang, Weiliang Gao, Jialin Pan, Jia-Hui Wang, Gang Li, Li Erran Liu, Yunhui Ding, Mingyu Heng, Pheng-Ann Fu, Chi-Wing |
| author_facet | Tang, Weiliang Gao, Jialin Pan, Jia-Hui Wang, Gang Li, Li Erran Liu, Yunhui Ding, Mingyu Heng, Pheng-Ann Fu, Chi-Wing |
| contents | Vision-Language Model (VLM) is an important component to enable robust robot manipulation. Yet, using it to translate human instructions into an action-resolvable intermediate representation often needs a tradeoff between VLM-comprehensibility and generalizability. Inspired by context-free grammar, we design the Semantic Assembly representation named SEAM, by decomposing the intermediate representation into vocabulary and grammar. Doing so leads us to a concise vocabulary of semantically-rich operations and a VLM-friendly grammar for handling diverse unseen tasks. In addition, we design a new open-vocabulary segmentation paradigm with a retrieval-augmented few-shot learning strategy to localize fine-grained object parts for manipulation, effectively with the shortest inference time over all state-of-the-art parallel works. Also, we formulate new metrics for action-generalizability and VLM-comprehensibility, demonstrating the compelling performance of SEAM over mainstream representations on both aspects. Extensive real-world experiments further manifest its SOTA performance under varying settings and tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_19315 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Rethinking Intermediate Representation for VLM-based Robot Manipulation Tang, Weiliang Gao, Jialin Pan, Jia-Hui Wang, Gang Li, Li Erran Liu, Yunhui Ding, Mingyu Heng, Pheng-Ann Fu, Chi-Wing Robotics Vision-Language Model (VLM) is an important component to enable robust robot manipulation. Yet, using it to translate human instructions into an action-resolvable intermediate representation often needs a tradeoff between VLM-comprehensibility and generalizability. Inspired by context-free grammar, we design the Semantic Assembly representation named SEAM, by decomposing the intermediate representation into vocabulary and grammar. Doing so leads us to a concise vocabulary of semantically-rich operations and a VLM-friendly grammar for handling diverse unseen tasks. In addition, we design a new open-vocabulary segmentation paradigm with a retrieval-augmented few-shot learning strategy to localize fine-grained object parts for manipulation, effectively with the shortest inference time over all state-of-the-art parallel works. Also, we formulate new metrics for action-generalizability and VLM-comprehensibility, demonstrating the compelling performance of SEAM over mainstream representations on both aspects. Extensive real-world experiments further manifest its SOTA performance under varying settings and tasks. |
| title | Rethinking Intermediate Representation for VLM-based Robot Manipulation |
| topic | Robotics |
| url | https://arxiv.org/abs/2511.19315 |