OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Can, Ding, Pengxiang, Song, Wenxuan, Bai, Shuanghao, Tong, Xinyang, Ge, Zirui, Suo, Runze, Zhou, Wanqi, Liu, Yang, Jia, Bofang, Zhao, Han, Huang, Siteng, Wang, Donglin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation
by: Bai, Shuanghao, et al.
Published: (2025)
by: Bai, Shuanghao, et al.
Published: (2025)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
by: Fan, Yiguo, et al.
Published: (2025)
by: Fan, Yiguo, et al.
Published: (2025)
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
by: Lin, Minghui, et al.
Published: (2025)
by: Lin, Minghui, et al.
Published: (2025)
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
by: Wang, Yihao, et al.
Published: (2025)
by: Wang, Yihao, et al.
Published: (2025)
Prompt-based Distribution Alignment for Unsupervised Domain Adaptation
by: Bai, Shuanghao, et al.
Published: (2023)
by: Bai, Shuanghao, et al.
Published: (2023)
Score and Distribution Matching Policy: Advanced Accelerated Visuomotor Policies via Matched Distillation
by: Jia, Bofang, et al.
Published: (2024)
by: Jia, Bofang, et al.
Published: (2024)
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
by: Zhao, Han, et al.
Published: (2025)
by: Zhao, Han, et al.
Published: (2025)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
by: Ding, Pengxiang, et al.
Published: (2023)
by: Ding, Pengxiang, et al.
Published: (2023)
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
by: Cui, Can, et al.
Published: (2024)
by: Cui, Can, et al.
Published: (2024)
VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
by: Li, Hengtao, et al.
Published: (2025)
by: Li, Hengtao, et al.
Published: (2025)
VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models
by: Zhao, Han, et al.
Published: (2025)
by: Zhao, Han, et al.
Published: (2025)
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey
by: Bai, Shuanghao, et al.
Published: (2025)
by: Bai, Shuanghao, et al.
Published: (2025)
PromptTA: Prompt-driven Text Adapter for Source-free Domain Generalization
by: Zhang, Haoran, et al.
Published: (2024)
by: Zhang, Haoran, et al.
Published: (2024)
Dual-Path Stable Soft Prompt Generation for Domain Generalization
by: Zhang, Yuedi, et al.
Published: (2025)
by: Zhang, Yuedi, et al.
Published: (2025)
VCoT-Grasp: Grasp Foundation Models with Visual Chain-of-Thought Reasoning for Language-driven Grasp Generation
by: Zhang, Haoran, et al.
Published: (2025)
by: Zhang, Haoran, et al.
Published: (2025)
RationalVLA: A Rational Vision-Language-Action Model with Dual System
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
by: Bai, Shuanghao, et al.
Published: (2025)
by: Bai, Shuanghao, et al.
Published: (2025)
Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
by: Ding, Pengxiang, et al.
Published: (2025)
by: Ding, Pengxiang, et al.
Published: (2025)
BlockVLA: Accelerating Autoregressive VLA via Block Diffusion Finetuning
by: Wang, Ruiheng, et al.
Published: (2026)
by: Wang, Ruiheng, et al.
Published: (2026)
CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
OpenVLA: An Open-Source Vision-Language-Action Model
by: Kim, Moo Jin, et al.
Published: (2024)
by: Kim, Moo Jin, et al.
Published: (2024)
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
by: Song, Wenxuan, et al.
Published: (2026)
by: Song, Wenxuan, et al.
Published: (2026)
QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
by: Tong, Xinyang, et al.
Published: (2024)
by: Tong, Xinyang, et al.
Published: (2024)
Jacobian Regularizer-based Neural Granger Causality
by: Zhou, Wanqi, et al.
Published: (2024)
by: Zhou, Wanqi, et al.
Published: (2024)
Soft Prompt Generation for Domain Generalization
by: Bai, Shuanghao, et al.
Published: (2024)
by: Bai, Shuanghao, et al.
Published: (2024)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
by: Chen, Jiayi, et al.
Published: (2025)
by: Chen, Jiayi, et al.
Published: (2025)
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference
by: Zhao, Han, et al.
Published: (2024)
by: Zhao, Han, et al.
Published: (2024)
Revisiting the Adversarial Robustness of Vision Language Models: a Multimodal Perspective
by: Zhou, Wanqi, et al.
Published: (2024)
by: Zhou, Wanqi, et al.
Published: (2024)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
by: Bai, Shuanghao, et al.
Published: (2026)
by: Bai, Shuanghao, et al.
Published: (2026)
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot
by: Song, Wenxuan, et al.
Published: (2024)
by: Song, Wenxuan, et al.
Published: (2024)
Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
by: Zhao, Pu, et al.
Published: (2025)
by: Zhao, Pu, et al.
Published: (2025)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025)
by: Zhu, Minjie, et al.
Published: (2025)
Why Does the LLM Stop Computing: An Empirical Study of User-Reported Failures in Open-Source LLMs
by: Yu, Guangba, et al.
Published: (2026)
by: Yu, Guangba, et al.
Published: (2026)
RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
Similar Items
-
Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation
by: Bai, Shuanghao, et al.
Published: (2025) -
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
by: Fan, Yiguo, et al.
Published: (2025) -
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
by: Lin, Minghui, et al.
Published: (2025) -
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
by: Wang, Yihao, et al.
Published: (2025) -
Prompt-based Distribution Alignment for Unsupervised Domain Adaptation
by: Bai, Shuanghao, et al.
Published: (2023)