Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Fan, Cunxin, Jia, Xiaosong, Sun, Yihang, Wang, Yixiao, Wei, Jianglan, Gong, Ziyang, Zhao, Xiangyu, Tomizuka, Masayoshi, Yang, Xue, Yan, Junchi, Ding, Mingyu
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914081145880576
author Fan, Cunxin
Jia, Xiaosong
Sun, Yihang
Wang, Yixiao
Wei, Jianglan
Gong, Ziyang
Zhao, Xiangyu
Tomizuka, Masayoshi
Yang, Xue
Yan, Junchi
Ding, Mingyu
author_facet Fan, Cunxin
Jia, Xiaosong
Sun, Yihang
Wang, Yixiao
Wei, Jianglan
Gong, Ziyang
Zhao, Xiangyu
Tomizuka, Masayoshi
Yang, Xue
Yan, Junchi
Ding, Mingyu
contents The rise of foundation models paves the way for generalist robot policies in the physical world. Existing methods relying on text-only instructions often struggle to generalize to unseen scenarios. We argue that interleaved image-text inputs offer richer and less biased context and enable robots to better handle unseen tasks with more versatile human-robot interaction. Building on this insight, Interleave-VLA, the first robot learning paradigm capable of comprehending interleaved image-text instructions and directly generating continuous action sequences in the physical world, is introduced. It offers a natural, flexible, and model-agnostic paradigm that extends state-of-the-art vision-language-action (VLA) models with minimal modifications while achieving strong zero-shot generalization. Interleave-VLA also includes an automatic pipeline that converts text instructions from Open X-Embodiment into interleaved image-text instructions, resulting in a large-scale real-world interleaved embodied dataset with 210k episodes. Comprehensive evaluation in simulation and the real world shows that Interleave-VLA offers two major benefits: (1) improves out-of-domain generalization to unseen objects by 2x compared to text input baselines, (2) supports flexible task interfaces and diverse instructions in a zero-shot manner, such as hand-drawn sketches. We attribute Interleave-VLA's strong zero-shot capability to the use of instruction images, which effectively mitigate hallucinations, and the inclusion of heterogeneous multimodal datasets, enriched with Internet-sourced images, offering potential for scalability. More information is available at https://interleave-vla.github.io/Interleave-VLA-Anonymous/
format Preprint
id arxiv_https___arxiv_org_abs_2505_02152
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
Fan, Cunxin
Jia, Xiaosong
Sun, Yihang
Wang, Yixiao
Wei, Jianglan
Gong, Ziyang
Zhao, Xiangyu
Tomizuka, Masayoshi
Yang, Xue
Yan, Junchi
Ding, Mingyu
Robotics
The rise of foundation models paves the way for generalist robot policies in the physical world. Existing methods relying on text-only instructions often struggle to generalize to unseen scenarios. We argue that interleaved image-text inputs offer richer and less biased context and enable robots to better handle unseen tasks with more versatile human-robot interaction. Building on this insight, Interleave-VLA, the first robot learning paradigm capable of comprehending interleaved image-text instructions and directly generating continuous action sequences in the physical world, is introduced. It offers a natural, flexible, and model-agnostic paradigm that extends state-of-the-art vision-language-action (VLA) models with minimal modifications while achieving strong zero-shot generalization. Interleave-VLA also includes an automatic pipeline that converts text instructions from Open X-Embodiment into interleaved image-text instructions, resulting in a large-scale real-world interleaved embodied dataset with 210k episodes. Comprehensive evaluation in simulation and the real world shows that Interleave-VLA offers two major benefits: (1) improves out-of-domain generalization to unseen objects by 2x compared to text input baselines, (2) supports flexible task interfaces and diverse instructions in a zero-shot manner, such as hand-drawn sketches. We attribute Interleave-VLA's strong zero-shot capability to the use of instruction images, which effectively mitigate hallucinations, and the inclusion of heterogeneous multimodal datasets, enriched with Internet-sourced images, offering potential for scalability. More information is available at https://interleave-vla.github.io/Interleave-VLA-Anonymous/
title Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
topic Robotics
url https://arxiv.org/abs/2505.02152