MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Wenhui, Chen, Changhe, Qi, Han, Lv, Chen, Du, Yilun, Yang, Heng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911227516551168
author Huang, Wenhui
Chen, Changhe
Qi, Han
Lv, Chen
Du, Yilun
Yang, Heng
author_facet Huang, Wenhui
Chen, Changhe
Qi, Han
Lv, Chen
Du, Yilun
Yang, Heng
contents Integrating visual-language instructions into visuomotor policies is gaining momentum in robot learning for enhancing open-world generalization. Despite promising advances, existing approaches face two challenges: limited language steerability when no generated reasoning is used as a condition, or significant inference latency when reasoning is incorporated. In this work, we introduce MoTVLA, a mixture-of-transformers (MoT)-based vision-language-action (VLA) model that integrates fast-slow unified reasoning with behavior policy learning. MoTVLA preserves the general intelligence of pre-trained VLMs (serving as the generalist) for tasks such as perception, scene understanding, and semantic planning, while incorporating a domain expert, a second transformer that shares knowledge with the pretrained VLM, to generate domain-specific fast reasoning (e.g., robot motion decomposition), thereby improving policy execution efficiency. By conditioning the action expert on decomposed motion instructions, MoTVLA can learn diverse behaviors and substantially improve language steerability. Extensive evaluations across natural language processing benchmarks, robotic simulation environments, and real-world experiments confirm the superiority of MoTVLA in both fast-slow reasoning and manipulation task performance.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18337
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning
Huang, Wenhui
Chen, Changhe
Qi, Han
Lv, Chen
Du, Yilun
Yang, Heng
Robotics
Integrating visual-language instructions into visuomotor policies is gaining momentum in robot learning for enhancing open-world generalization. Despite promising advances, existing approaches face two challenges: limited language steerability when no generated reasoning is used as a condition, or significant inference latency when reasoning is incorporated. In this work, we introduce MoTVLA, a mixture-of-transformers (MoT)-based vision-language-action (VLA) model that integrates fast-slow unified reasoning with behavior policy learning. MoTVLA preserves the general intelligence of pre-trained VLMs (serving as the generalist) for tasks such as perception, scene understanding, and semantic planning, while incorporating a domain expert, a second transformer that shares knowledge with the pretrained VLM, to generate domain-specific fast reasoning (e.g., robot motion decomposition), thereby improving policy execution efficiency. By conditioning the action expert on decomposed motion instructions, MoTVLA can learn diverse behaviors and substantially improve language steerability. Extensive evaluations across natural language processing benchmarks, robotic simulation environments, and real-world experiments confirm the superiority of MoTVLA in both fast-slow reasoning and manipulation task performance.
title MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning
topic Robotics
url https://arxiv.org/abs/2510.18337