MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Rongyu, Dong, Menghang, Zhang, Yuan, Heng, Liang, Chi, Xiaowei, Dai, Gaole, Du, Li, Du, Yuan, Zhang, Shanghang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
di: Dai, Gaole, et al.
Pubblicazione: (2025)
di: Dai, Gaole, et al.
Pubblicazione: (2025)
Decomposing the Neurons: Activation Sparsity via Mixture of Experts for Continual Test Time Adaptation
di: Zhang, Rongyu, et al.
Pubblicazione: (2024)
di: Zhang, Rongyu, et al.
Pubblicazione: (2024)
FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
di: Zhao, Zhongyu, et al.
Pubblicazione: (2024)
di: Zhao, Zhongyu, et al.
Pubblicazione: (2024)
MoASE++: Mixture of Activation Sparsity Experts with Domain-Adaptive On-policy Distillation for Continual Test Time Adaptation
di: Zhang, Ronyu, et al.
Pubblicazione: (2026)
di: Zhang, Ronyu, et al.
Pubblicazione: (2026)
BEVUDA++: Geometric-aware Unsupervised Domain Adaptation for Multi-View 3D Object Detection
di: Zhang, Rongyu, et al.
Pubblicazione: (2025)
di: Zhang, Rongyu, et al.
Pubblicazione: (2025)
Key-Embedded Privacy for Decentralized AI in Biomedical Omics
di: Zhang, Rongyu, et al.
Pubblicazione: (2026)
di: Zhang, Rongyu, et al.
Pubblicazione: (2026)
T-REX: Mixture-of-Rank-One-Experts with Semantic-aware Intuition for Multi-task Large Language Model Finetuning
di: Zhang, Rongyu, et al.
Pubblicazione: (2024)
di: Zhang, Rongyu, et al.
Pubblicazione: (2024)
Orochi: Versatile Biomedical Image Processor
di: Dai, Gaole, et al.
Pubblicazione: (2025)
di: Dai, Gaole, et al.
Pubblicazione: (2025)
HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
di: Du, Zhiying, et al.
Pubblicazione: (2025)
di: Du, Zhiying, et al.
Pubblicazione: (2025)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
di: Li, Xiaoqi, et al.
Pubblicazione: (2025)
di: Li, Xiaoqi, et al.
Pubblicazione: (2025)
$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills
di: Xiao, Siyao, et al.
Pubblicazione: (2026)
di: Xiao, Siyao, et al.
Pubblicazione: (2026)
DySL-VLA: Efficient Vision-Language-Action Model Inference via Dynamic-Static Layer-Skipping for Robot Manipulation
di: Yang, Zebin, et al.
Pubblicazione: (2026)
di: Yang, Zebin, et al.
Pubblicazione: (2026)
Coupled Cluster con MōLe: Molecular Orbital Learning for Neural Wavefunctions
di: Thiede, Luca, et al.
Pubblicazione: (2026)
di: Thiede, Luca, et al.
Pubblicazione: (2026)
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
di: Liu, Jiaming, et al.
Pubblicazione: (2024)
di: Liu, Jiaming, et al.
Pubblicazione: (2024)
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
di: Chi, Xiaowei, et al.
Pubblicazione: (2023)
di: Chi, Xiaowei, et al.
Pubblicazione: (2023)
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
di: Gu, Chenyang, et al.
Pubblicazione: (2025)
di: Gu, Chenyang, et al.
Pubblicazione: (2025)
MōLe-Λ: Learning the Coupled-Cluster Response State for Energies, Gradients, and Properties
di: Burger, Andreas, et al.
Pubblicazione: (2026)
di: Burger, Andreas, et al.
Pubblicazione: (2026)
Multimodal Large Language Models for Bioimage Analysis
di: Zhang, Shanghang, et al.
Pubblicazione: (2024)
di: Zhang, Shanghang, et al.
Pubblicazione: (2024)
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
di: Du, Fan, et al.
Pubblicazione: (2026)
di: Du, Fan, et al.
Pubblicazione: (2026)
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
di: Zhang, Qizhe, et al.
Pubblicazione: (2023)
di: Zhang, Qizhe, et al.
Pubblicazione: (2023)
RepCaM++: Exploring Transparent Visual Prompt With Inference-Time Re-Parameterization for Neural Video Delivery
di: Zhang, Rongyu, et al.
Pubblicazione: (2025)
di: Zhang, Rongyu, et al.
Pubblicazione: (2025)
ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance
di: Li, Ying, et al.
Pubblicazione: (2025)
di: Li, Ying, et al.
Pubblicazione: (2025)
VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation
di: Bian, Jinyue, et al.
Pubblicazione: (2025)
di: Bian, Jinyue, et al.
Pubblicazione: (2025)
FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
di: Miao, Cui, et al.
Pubblicazione: (2025)
di: Miao, Cui, et al.
Pubblicazione: (2025)
TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation
di: Zhang, Kaidi, et al.
Pubblicazione: (2026)
di: Zhang, Kaidi, et al.
Pubblicazione: (2026)
SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models
di: Fang, Hengyu, et al.
Pubblicazione: (2025)
di: Fang, Hengyu, et al.
Pubblicazione: (2025)
FocusVLA: Focused Visual Utilization for Vision-Language-Action Models
di: Zhang, Yichi, et al.
Pubblicazione: (2026)
di: Zhang, Yichi, et al.
Pubblicazione: (2026)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
di: Shi, Hao, et al.
Pubblicazione: (2025)
di: Shi, Hao, et al.
Pubblicazione: (2025)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
di: Peng, Xiongfeng, et al.
Pubblicazione: (2026)
di: Peng, Xiongfeng, et al.
Pubblicazione: (2026)
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
di: Cao, Jiajun, et al.
Pubblicazione: (2025)
di: Cao, Jiajun, et al.
Pubblicazione: (2025)
Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation
di: Wei, Xiangyi, et al.
Pubblicazione: (2025)
di: Wei, Xiangyi, et al.
Pubblicazione: (2025)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
di: Wen, Junjie, et al.
Pubblicazione: (2024)
di: Wen, Junjie, et al.
Pubblicazione: (2024)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
di: Wen, Junjie, et al.
Pubblicazione: (2025)
di: Wen, Junjie, et al.
Pubblicazione: (2025)
VeCAF: Vision-language Collaborative Active Finetuning with Training Objective Awareness
di: Zhang, Rongyu, et al.
Pubblicazione: (2024)
di: Zhang, Rongyu, et al.
Pubblicazione: (2024)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
di: Sun, Jianli, et al.
Pubblicazione: (2026)
di: Sun, Jianli, et al.
Pubblicazione: (2026)
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
di: Li, Chenxuan, et al.
Pubblicazione: (2024)
di: Li, Chenxuan, et al.
Pubblicazione: (2024)
Adaptive Layer-skipping in Pre-trained LLMs
di: Luo, Xuan, et al.
Pubblicazione: (2025)
di: Luo, Xuan, et al.
Pubblicazione: (2025)
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation
di: Li, Yajie, et al.
Pubblicazione: (2026)
di: Li, Yajie, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
di: Dai, Gaole, et al.
Pubblicazione: (2025) -
Decomposing the Neurons: Activation Sparsity via Mixture of Experts for Continual Test Time Adaptation
di: Zhang, Rongyu, et al.
Pubblicazione: (2024) -
FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
di: Zhao, Zhongyu, et al.
Pubblicazione: (2024) -
MoASE++: Mixture of Activation Sparsity Experts with Domain-Adaptive On-policy Distillation for Continual Test Time Adaptation
di: Zhang, Ronyu, et al.
Pubblicazione: (2026) -
BEVUDA++: Geometric-aware Unsupervised Domain Adaptation for Multi-View 3D Object Detection
di: Zhang, Rongyu, et al.
Pubblicazione: (2025)