Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Won, John, Lee, Kyungmin, Jang, Huiwon, Kim, Dongyoung, Shin, Jinwoo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910267920613376
author Won, John
Lee, Kyungmin
Jang, Huiwon
Kim, Dongyoung
Shin, Jinwoo
author_facet Won, John
Lee, Kyungmin
Jang, Huiwon
Kim, Dongyoung
Shin, Jinwoo
contents Augmenting vision-language-action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predicting states and actions due to the modality gap. To address this, we propose DUal-STream diffusion (DUST), a world-model augmented VLA framework featuring a multimodal diffusion transformer that maintains separate modality streams while enabling cross-modal knowledge sharing. In addition, DUST utilizes independent noise perturbations and a decoupled flow matching loss to learn cross-modal causal relationships. We further introduce an asynchronous sampling method for action and vision tokens that enhances performance through inference-time scaling. Experimental results on simulated benchmarks like RoboCasa and GR-1 show that DUST achieves up to 6% gains over state-of-the-art VLA and world-modeling baselines, with inference-time scaling providing an additional 2-5% improvement. In real-world tasks using the Franka Research 3, DUST outperforms baselines by 10% in success rate. Finally, we demonstrate that DUST enables effective transfer learning through both pretraining on action-free videos and joint-training with heterogeneous robot and human datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27607
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Won, John
Lee, Kyungmin
Jang, Huiwon
Kim, Dongyoung
Shin, Jinwoo
Computer Vision and Pattern Recognition
Robotics
Augmenting vision-language-action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predicting states and actions due to the modality gap. To address this, we propose DUal-STream diffusion (DUST), a world-model augmented VLA framework featuring a multimodal diffusion transformer that maintains separate modality streams while enabling cross-modal knowledge sharing. In addition, DUST utilizes independent noise perturbations and a decoupled flow matching loss to learn cross-modal causal relationships. We further introduce an asynchronous sampling method for action and vision tokens that enhances performance through inference-time scaling. Experimental results on simulated benchmarks like RoboCasa and GR-1 show that DUST achieves up to 6% gains over state-of-the-art VLA and world-modeling baselines, with inference-time scaling providing an additional 2-5% improvement. In real-world tasks using the Franka Research 3, DUST outperforms baselines by 10% in success rate. Finally, we demonstrate that DUST enables effective transfer learning through both pretraining on action-free videos and joint-training with heterogeneous robot and human datasets.
title Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2510.27607