$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills
Fuente:
arXiv
Salvato in:
| Autori principali: | Xiao, Siyao, Zhang, Yuhong, Liu, Zhifang, Gao, Zihan, Zhang, Jingye, Choo, Sinwai, Zhong, Dake, Wang, Mengzhe, Lin, Xiao, Zhou, Xianfeng, Jia, Jia, Wang, Haoqian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
di: Zhang, Rongyu, et al.
Pubblicazione: (2025)
di: Zhang, Rongyu, et al.
Pubblicazione: (2025)
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
di: Su, Taiyi, et al.
Pubblicazione: (2026)
di: Su, Taiyi, et al.
Pubblicazione: (2026)
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
di: Sun, Jianli, et al.
Pubblicazione: (2026)
di: Sun, Jianli, et al.
Pubblicazione: (2026)
Uni-Skill: Building Self-Evolving Skill Repository for Generalizable Robotic Manipulation
di: Xie, Senwei, et al.
Pubblicazione: (2026)
di: Xie, Senwei, et al.
Pubblicazione: (2026)
SkillEvolver: Skill Learning as a Meta-Skill
di: Zhang, Genrui, et al.
Pubblicazione: (2026)
di: Zhang, Genrui, et al.
Pubblicazione: (2026)
SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving
di: You, Zihan, et al.
Pubblicazione: (2026)
di: You, Zihan, et al.
Pubblicazione: (2026)
Skill-Aware Diffusion for Generalizable Robotic Manipulation
di: Huang, Aoshen, et al.
Pubblicazione: (2026)
di: Huang, Aoshen, et al.
Pubblicazione: (2026)
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
di: Ma, Guoqing, et al.
Pubblicazione: (2026)
di: Ma, Guoqing, et al.
Pubblicazione: (2026)
FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
di: Miao, Cui, et al.
Pubblicazione: (2025)
di: Miao, Cui, et al.
Pubblicazione: (2025)
AWSPNet: Attention-based Dual-Tree Wavelet Scattering Prototypical Network for MIMO Radar Target Recognition and Jamming Suppression
di: Jia, Yizhen, et al.
Pubblicazione: (2025)
di: Jia, Yizhen, et al.
Pubblicazione: (2025)
VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation
di: Zhou, Huayi, et al.
Pubblicazione: (2025)
di: Zhou, Huayi, et al.
Pubblicazione: (2025)
FocusVLA: Focused Visual Utilization for Vision-Language-Action Models
di: Zhang, Yichi, et al.
Pubblicazione: (2026)
di: Zhang, Yichi, et al.
Pubblicazione: (2026)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
di: Shen, Yichao, et al.
Pubblicazione: (2025)
di: Shen, Yichao, et al.
Pubblicazione: (2025)
MoRE-Brain: Routed Mixture of Experts for Interpretable and Generalizable Cross-Subject fMRI Visual Decoding
di: Wei, Yuxiang, et al.
Pubblicazione: (2025)
di: Wei, Yuxiang, et al.
Pubblicazione: (2025)
PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation
di: Li, Yutai, et al.
Pubblicazione: (2026)
di: Li, Yutai, et al.
Pubblicazione: (2026)
VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
di: Zhou, Hui, et al.
Pubblicazione: (2025)
di: Zhou, Hui, et al.
Pubblicazione: (2025)
SELF-VLA: A Skill Enhanced Agentic Vision-Language-Action Framework for Contact-Rich Disassembly
di: Liu, Chang, et al.
Pubblicazione: (2026)
di: Liu, Chang, et al.
Pubblicazione: (2026)
TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation
di: Zhou, Hanyu, et al.
Pubblicazione: (2026)
di: Zhou, Hanyu, et al.
Pubblicazione: (2026)
MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
di: Fu, Yuxia, et al.
Pubblicazione: (2025)
di: Fu, Yuxia, et al.
Pubblicazione: (2025)
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
di: Gu, Chenyang, et al.
Pubblicazione: (2025)
di: Gu, Chenyang, et al.
Pubblicazione: (2025)
AnchorVLA: Anchored Diffusion for Efficient End-to-End Mobile Manipulation
di: Lim, Jia Syuen, et al.
Pubblicazione: (2026)
di: Lim, Jia Syuen, et al.
Pubblicazione: (2026)
Learning Generalizable Language-Conditioned Cloth Manipulation from Long Demonstrations
di: Zhao, Hanyi, et al.
Pubblicazione: (2025)
di: Zhao, Hanyi, et al.
Pubblicazione: (2025)
Boosting Generalizability towards Zero-Shot Cross-Dataset Single-Image Indoor Depth by Meta-Initialization
di: Wu, Cho-Ying, et al.
Pubblicazione: (2024)
di: Wu, Cho-Ying, et al.
Pubblicazione: (2024)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
di: Wang, Hongyu, et al.
Pubblicazione: (2025)
di: Wang, Hongyu, et al.
Pubblicazione: (2025)
ProgVLA: Progress-Aware Robot Manipulation Skill Learning
di: Kim, Seungsu, et al.
Pubblicazione: (2026)
di: Kim, Seungsu, et al.
Pubblicazione: (2026)
BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models
di: Li, Peiyan, et al.
Pubblicazione: (2025)
di: Li, Peiyan, et al.
Pubblicazione: (2025)
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
di: Zhao, Han, et al.
Pubblicazione: (2025)
di: Zhao, Han, et al.
Pubblicazione: (2025)
TranSplat: Generalizable 3D Gaussian Splatting from Sparse Multi-View Images with Transformers
di: Zhang, Chuanrui, et al.
Pubblicazione: (2024)
di: Zhang, Chuanrui, et al.
Pubblicazione: (2024)
SkillVLA: Tackling Combinatorial Diversity in Dual-Arm Manipulation via Skill Reuse
di: Zhai, Xuanran, et al.
Pubblicazione: (2026)
di: Zhai, Xuanran, et al.
Pubblicazione: (2026)
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
di: Du, Zhaohui, et al.
Pubblicazione: (2026)
di: Du, Zhaohui, et al.
Pubblicazione: (2026)
Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation
di: Li, Yuelei, et al.
Pubblicazione: (2025)
di: Li, Yuelei, et al.
Pubblicazione: (2025)
VLA Model-Expert Collaboration for Bi-directional Manipulation Learning
di: Xiang, Tian-Yu, et al.
Pubblicazione: (2025)
di: Xiang, Tian-Yu, et al.
Pubblicazione: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
di: Yang, Shuai, et al.
Pubblicazione: (2025)
di: Yang, Shuai, et al.
Pubblicazione: (2025)
SkillScope: Toward Fine-Grained Least-Privilege Enforcement for Agent Skills
di: Wu, Jiangrong, et al.
Pubblicazione: (2026)
di: Wu, Jiangrong, et al.
Pubblicazione: (2026)
iManip: Skill-Incremental Learning for Robotic Manipulation
di: Zheng, Zexin, et al.
Pubblicazione: (2025)
di: Zheng, Zexin, et al.
Pubblicazione: (2025)
Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation
di: Wei, Xiangyi, et al.
Pubblicazione: (2025)
di: Wei, Xiangyi, et al.
Pubblicazione: (2025)
ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation
di: Wu, Yuhan, et al.
Pubblicazione: (2025)
di: Wu, Yuhan, et al.
Pubblicazione: (2025)
Vision-Language Semantic Aggregation Leveraging Foundation Model for Generalizable Medical Image Segmentation
di: Yu, Wenjun, et al.
Pubblicazione: (2025)
di: Yu, Wenjun, et al.
Pubblicazione: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
di: Shi, Hao, et al.
Pubblicazione: (2025)
di: Shi, Hao, et al.
Pubblicazione: (2025)
CROSS: A Mixture-of-Experts Reinforcement Learning Framework for Generalizable Large-Scale Traffic Signal Control
di: Chen, Xibei, et al.
Pubblicazione: (2026)
di: Chen, Xibei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
di: Zhang, Rongyu, et al.
Pubblicazione: (2025) -
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
di: Su, Taiyi, et al.
Pubblicazione: (2026) -
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
di: Sun, Jianli, et al.
Pubblicazione: (2026) -
Uni-Skill: Building Self-Evolving Skill Repository for Generalizable Robotic Manipulation
di: Xie, Senwei, et al.
Pubblicazione: (2026) -
SkillEvolver: Skill Learning as a Meta-Skill
di: Zhang, Genrui, et al.
Pubblicazione: (2026)