Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Weijie, Liu, Yitian, Wu, Yuhao, Liang, Zhixuan, Gu, Sijia, Wang, Dehui, Nian, Tian, Xu, Lei, Qin, Yusen, Pang, Jiangmiao, Guan, Xinping, Yang, Xiaokang, Mu, Yao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911214137769984
author Shen, Weijie
Liu, Yitian
Wu, Yuhao
Liang, Zhixuan
Gu, Sijia
Wang, Dehui
Nian, Tian
Xu, Lei
Qin, Yusen
Pang, Jiangmiao
Guan, Xinping
Yang, Xiaokang
Mu, Yao
author_facet Shen, Weijie
Liu, Yitian
Wu, Yuhao
Liang, Zhixuan
Gu, Sijia
Wang, Dehui
Nian, Tian
Xu, Lei
Qin, Yusen
Pang, Jiangmiao
Guan, Xinping
Yang, Xiaokang
Mu, Yao
contents Vision-Language-Action (VLA) models are experiencing rapid development and demonstrating promising capabilities in robotic manipulation tasks. However, scaling up VLA models presents several critical challenges: (1) Training new VLA models from scratch demands substantial computational resources and extensive datasets. Given the current scarcity of robot data, it becomes particularly valuable to fully leverage well-pretrained VLA model weights during the scaling process. (2) Real-time control requires carefully balancing model capacity with computational efficiency. To address these challenges, We propose AdaMoE, a Mixture-of-Experts (MoE) architecture that inherits pretrained weights from dense VLA models, and scales up the action expert by substituting the feedforward layers into sparsely activated MoE layers. AdaMoE employs a decoupling technique that decouples expert selection from expert weighting through an independent scale adapter working alongside the traditional router. This enables experts to be selected based on task relevance while contributing with independently controlled weights, allowing collaborative expert utilization rather than winner-takes-all dynamics. Our approach demonstrates that expertise need not monopolize. Instead, through collaborative expert utilization, we can achieve superior performance while maintaining computational efficiency. AdaMoE consistently outperforms the baseline model across key benchmarks, delivering performance gains of 1.8% on LIBERO and 9.3% on RoboTwin. Most importantly, a substantial 21.5% improvement in real-world experiments validates its practical effectiveness for robotic manipulation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14300
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
Shen, Weijie
Liu, Yitian
Wu, Yuhao
Liang, Zhixuan
Gu, Sijia
Wang, Dehui
Nian, Tian
Xu, Lei
Qin, Yusen
Pang, Jiangmiao
Guan, Xinping
Yang, Xiaokang
Mu, Yao
Robotics
Artificial Intelligence
Vision-Language-Action (VLA) models are experiencing rapid development and demonstrating promising capabilities in robotic manipulation tasks. However, scaling up VLA models presents several critical challenges: (1) Training new VLA models from scratch demands substantial computational resources and extensive datasets. Given the current scarcity of robot data, it becomes particularly valuable to fully leverage well-pretrained VLA model weights during the scaling process. (2) Real-time control requires carefully balancing model capacity with computational efficiency. To address these challenges, We propose AdaMoE, a Mixture-of-Experts (MoE) architecture that inherits pretrained weights from dense VLA models, and scales up the action expert by substituting the feedforward layers into sparsely activated MoE layers. AdaMoE employs a decoupling technique that decouples expert selection from expert weighting through an independent scale adapter working alongside the traditional router. This enables experts to be selected based on task relevance while contributing with independently controlled weights, allowing collaborative expert utilization rather than winner-takes-all dynamics. Our approach demonstrates that expertise need not monopolize. Instead, through collaborative expert utilization, we can achieve superior performance while maintaining computational efficiency. AdaMoE consistently outperforms the baseline model across key benchmarks, delivering performance gains of 1.8% on LIBERO and 9.3% on RoboTwin. Most importantly, a substantial 21.5% improvement in real-world experiments validates its practical effectiveness for robotic manipulation tasks.
title Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2510.14300