Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhou, Jiaming, Ye, Ke, Liu, Jiayi, Ma, Teli, Wang, Zifan, Qiu, Ronghe, Lin, Kun-Yu, Zhao, Zhilin, Liang, Junwei
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912658698010624
author Zhou, Jiaming
Ye, Ke
Liu, Jiayi
Ma, Teli
Wang, Zifan
Qiu, Ronghe
Lin, Kun-Yu
Zhao, Zhilin
Liang, Junwei
author_facet Zhou, Jiaming
Ye, Ke
Liu, Jiayi
Ma, Teli
Wang, Zifan
Qiu, Ronghe
Lin, Kun-Yu
Zhao, Zhilin
Liang, Junwei
contents The generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings. However, the cross-task generalization capabilities of existing VLA models remain significantly underexplored. To address this gap, we introduce AGNOSTOS, a novel simulation benchmark designed to rigorously evaluate cross-task zero-shot generalization in manipulation. AGNOSTOS comprises 23 unseen manipulation tasks for testing, distinct from common training task distributions, and incorporates two levels of generalization difficulty to assess robustness. Our systematic evaluation reveals that current VLA models, despite being trained on diverse datasets, struggle to generalize effectively to these unseen tasks. To overcome this limitation, we propose Cross-Task In-Context Manipulation (X-ICM), a method that conditions large language models (LLMs) on in-context demonstrations from seen tasks to predict action sequences for unseen tasks. Additionally, we introduce a dynamics-guided sample selection strategy that identifies relevant demonstrations by capturing cross-task dynamics. On AGNOSTOS, X-ICM significantly improves cross-task zero-shot generalization performance over leading VLAs. We believe AGNOSTOS and X-ICM will serve as valuable tools for advancing general-purpose robotic manipulation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15660
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
Zhou, Jiaming
Ye, Ke
Liu, Jiayi
Ma, Teli
Wang, Zifan
Qiu, Ronghe
Lin, Kun-Yu
Zhao, Zhilin
Liang, Junwei
Robotics
Computer Vision and Pattern Recognition
The generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings. However, the cross-task generalization capabilities of existing VLA models remain significantly underexplored. To address this gap, we introduce AGNOSTOS, a novel simulation benchmark designed to rigorously evaluate cross-task zero-shot generalization in manipulation. AGNOSTOS comprises 23 unseen manipulation tasks for testing, distinct from common training task distributions, and incorporates two levels of generalization difficulty to assess robustness. Our systematic evaluation reveals that current VLA models, despite being trained on diverse datasets, struggle to generalize effectively to these unseen tasks. To overcome this limitation, we propose Cross-Task In-Context Manipulation (X-ICM), a method that conditions large language models (LLMs) on in-context demonstrations from seen tasks to predict action sequences for unseen tasks. Additionally, we introduce a dynamics-guided sample selection strategy that identifies relevant demonstrations by capturing cross-task dynamics. On AGNOSTOS, X-ICM significantly improves cross-task zero-shot generalization performance over leading VLAs. We believe AGNOSTOS and X-ICM will serve as valuable tools for advancing general-purpose robotic manipulation.
title Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.15660