$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Siyao, Zhang, Yuhong, Liu, Zhifang, Gao, Zihan, Zhang, Jingye, Choo, Sinwai, Zhong, Dake, Wang, Mengzhe, Lin, Xiao, Zhou, Xianfeng, Jia, Jia, Wang, Haoqian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
by: Zhang, Rongyu, et al.
Published: (2025)
by: Zhang, Rongyu, et al.
Published: (2025)
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
by: Su, Taiyi, et al.
Published: (2026)
by: Su, Taiyi, et al.
Published: (2026)
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
by: Sun, Jianli, et al.
Published: (2026)
by: Sun, Jianli, et al.
Published: (2026)
Uni-Skill: Building Self-Evolving Skill Repository for Generalizable Robotic Manipulation
by: Xie, Senwei, et al.
Published: (2026)
by: Xie, Senwei, et al.
Published: (2026)
SkillEvolver: Skill Learning as a Meta-Skill
by: Zhang, Genrui, et al.
Published: (2026)
by: Zhang, Genrui, et al.
Published: (2026)
SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving
by: You, Zihan, et al.
Published: (2026)
by: You, Zihan, et al.
Published: (2026)
Skill-Aware Diffusion for Generalizable Robotic Manipulation
by: Huang, Aoshen, et al.
Published: (2026)
by: Huang, Aoshen, et al.
Published: (2026)
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
by: Ma, Guoqing, et al.
Published: (2026)
by: Ma, Guoqing, et al.
Published: (2026)
FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
by: Miao, Cui, et al.
Published: (2025)
by: Miao, Cui, et al.
Published: (2025)
AWSPNet: Attention-based Dual-Tree Wavelet Scattering Prototypical Network for MIMO Radar Target Recognition and Jamming Suppression
by: Jia, Yizhen, et al.
Published: (2025)
by: Jia, Yizhen, et al.
Published: (2025)
VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation
by: Zhou, Huayi, et al.
Published: (2025)
by: Zhou, Huayi, et al.
Published: (2025)
FocusVLA: Focused Visual Utilization for Vision-Language-Action Models
by: Zhang, Yichi, et al.
Published: (2026)
by: Zhang, Yichi, et al.
Published: (2026)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
by: Shen, Yichao, et al.
Published: (2025)
by: Shen, Yichao, et al.
Published: (2025)
MoRE-Brain: Routed Mixture of Experts for Interpretable and Generalizable Cross-Subject fMRI Visual Decoding
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation
by: Li, Yutai, et al.
Published: (2026)
by: Li, Yutai, et al.
Published: (2026)
VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
by: Zhou, Hui, et al.
Published: (2025)
by: Zhou, Hui, et al.
Published: (2025)
SELF-VLA: A Skill Enhanced Agentic Vision-Language-Action Framework for Contact-Rich Disassembly
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation
by: Zhou, Hanyu, et al.
Published: (2026)
by: Zhou, Hanyu, et al.
Published: (2026)
MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
by: Fu, Yuxia, et al.
Published: (2025)
by: Fu, Yuxia, et al.
Published: (2025)
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
by: Gu, Chenyang, et al.
Published: (2025)
by: Gu, Chenyang, et al.
Published: (2025)
AnchorVLA: Anchored Diffusion for Efficient End-to-End Mobile Manipulation
by: Lim, Jia Syuen, et al.
Published: (2026)
by: Lim, Jia Syuen, et al.
Published: (2026)
Learning Generalizable Language-Conditioned Cloth Manipulation from Long Demonstrations
by: Zhao, Hanyi, et al.
Published: (2025)
by: Zhao, Hanyi, et al.
Published: (2025)
Boosting Generalizability towards Zero-Shot Cross-Dataset Single-Image Indoor Depth by Meta-Initialization
by: Wu, Cho-Ying, et al.
Published: (2024)
by: Wu, Cho-Ying, et al.
Published: (2024)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
by: Wang, Hongyu, et al.
Published: (2025)
by: Wang, Hongyu, et al.
Published: (2025)
ProgVLA: Progress-Aware Robot Manipulation Skill Learning
by: Kim, Seungsu, et al.
Published: (2026)
by: Kim, Seungsu, et al.
Published: (2026)
BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models
by: Li, Peiyan, et al.
Published: (2025)
by: Li, Peiyan, et al.
Published: (2025)
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
by: Zhao, Han, et al.
Published: (2025)
by: Zhao, Han, et al.
Published: (2025)
TranSplat: Generalizable 3D Gaussian Splatting from Sparse Multi-View Images with Transformers
by: Zhang, Chuanrui, et al.
Published: (2024)
by: Zhang, Chuanrui, et al.
Published: (2024)
SkillVLA: Tackling Combinatorial Diversity in Dual-Arm Manipulation via Skill Reuse
by: Zhai, Xuanran, et al.
Published: (2026)
by: Zhai, Xuanran, et al.
Published: (2026)
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
by: Du, Zhaohui, et al.
Published: (2026)
by: Du, Zhaohui, et al.
Published: (2026)
Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation
by: Li, Yuelei, et al.
Published: (2025)
by: Li, Yuelei, et al.
Published: (2025)
VLA Model-Expert Collaboration for Bi-directional Manipulation Learning
by: Xiang, Tian-Yu, et al.
Published: (2025)
by: Xiang, Tian-Yu, et al.
Published: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
SkillScope: Toward Fine-Grained Least-Privilege Enforcement for Agent Skills
by: Wu, Jiangrong, et al.
Published: (2026)
by: Wu, Jiangrong, et al.
Published: (2026)
iManip: Skill-Incremental Learning for Robotic Manipulation
by: Zheng, Zexin, et al.
Published: (2025)
by: Zheng, Zexin, et al.
Published: (2025)
Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation
by: Wei, Xiangyi, et al.
Published: (2025)
by: Wei, Xiangyi, et al.
Published: (2025)
Vision-Language Semantic Aggregation Leveraging Foundation Model for Generalizable Medical Image Segmentation
by: Yu, Wenjun, et al.
Published: (2025)
by: Yu, Wenjun, et al.
Published: (2025)
ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation
by: Wu, Yuhan, et al.
Published: (2025)
by: Wu, Yuhan, et al.
Published: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
CROSS: A Mixture-of-Experts Reinforcement Learning Framework for Generalizable Large-Scale Traffic Signal Control
by: Chen, Xibei, et al.
Published: (2026)
by: Chen, Xibei, et al.
Published: (2026)
Similar Items
-
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
by: Zhang, Rongyu, et al.
Published: (2025) -
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
by: Su, Taiyi, et al.
Published: (2026) -
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
by: Sun, Jianli, et al.
Published: (2026) -
Uni-Skill: Building Self-Evolving Skill Repository for Generalizable Robotic Manipulation
by: Xie, Senwei, et al.
Published: (2026) -
SkillEvolver: Skill Learning as a Meta-Skill
by: Zhang, Genrui, et al.
Published: (2026)