Rethinking Intermediate Representation for VLM-based Robot Manipulation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tang, Weiliang, Gao, Jialin, Pan, Jia-Hui, Wang, Gang, Li, Li Erran, Liu, Yunhui, Ding, Mingyu, Heng, Pheng-Ann, Fu, Chi-Wing
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915635307479040
author Tang, Weiliang
Gao, Jialin
Pan, Jia-Hui
Wang, Gang
Li, Li Erran
Liu, Yunhui
Ding, Mingyu
Heng, Pheng-Ann
Fu, Chi-Wing
author_facet Tang, Weiliang
Gao, Jialin
Pan, Jia-Hui
Wang, Gang
Li, Li Erran
Liu, Yunhui
Ding, Mingyu
Heng, Pheng-Ann
Fu, Chi-Wing
contents Vision-Language Model (VLM) is an important component to enable robust robot manipulation. Yet, using it to translate human instructions into an action-resolvable intermediate representation often needs a tradeoff between VLM-comprehensibility and generalizability. Inspired by context-free grammar, we design the Semantic Assembly representation named SEAM, by decomposing the intermediate representation into vocabulary and grammar. Doing so leads us to a concise vocabulary of semantically-rich operations and a VLM-friendly grammar for handling diverse unseen tasks. In addition, we design a new open-vocabulary segmentation paradigm with a retrieval-augmented few-shot learning strategy to localize fine-grained object parts for manipulation, effectively with the shortest inference time over all state-of-the-art parallel works. Also, we formulate new metrics for action-generalizability and VLM-comprehensibility, demonstrating the compelling performance of SEAM over mainstream representations on both aspects. Extensive real-world experiments further manifest its SOTA performance under varying settings and tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19315
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Intermediate Representation for VLM-based Robot Manipulation
Tang, Weiliang
Gao, Jialin
Pan, Jia-Hui
Wang, Gang
Li, Li Erran
Liu, Yunhui
Ding, Mingyu
Heng, Pheng-Ann
Fu, Chi-Wing
Robotics
Vision-Language Model (VLM) is an important component to enable robust robot manipulation. Yet, using it to translate human instructions into an action-resolvable intermediate representation often needs a tradeoff between VLM-comprehensibility and generalizability. Inspired by context-free grammar, we design the Semantic Assembly representation named SEAM, by decomposing the intermediate representation into vocabulary and grammar. Doing so leads us to a concise vocabulary of semantically-rich operations and a VLM-friendly grammar for handling diverse unseen tasks. In addition, we design a new open-vocabulary segmentation paradigm with a retrieval-augmented few-shot learning strategy to localize fine-grained object parts for manipulation, effectively with the shortest inference time over all state-of-the-art parallel works. Also, we formulate new metrics for action-generalizability and VLM-comprehensibility, demonstrating the compelling performance of SEAM over mainstream representations on both aspects. Extensive real-world experiments further manifest its SOTA performance under varying settings and tasks.
title Rethinking Intermediate Representation for VLM-based Robot Manipulation
topic Robotics
url https://arxiv.org/abs/2511.19315