Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wen, Junjie, Zhu, Minjie, Zhu, Yichen, Tang, Zhibin, Li, Jinming, Zhou, Zhongyi, Li, Chengmeng, Liu, Xiaoyu, Peng, Yaxin, Shen, Chaomin, Feng, Feifei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
von: Zhu, Minjie, et al.
Veröffentlicht: (2025)
von: Zhu, Minjie, et al.
Veröffentlicht: (2025)
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
von: Li, Jinming, et al.
Veröffentlicht: (2024)
von: Li, Jinming, et al.
Veröffentlicht: (2024)
PointVLA: Injecting the 3D World into Vision-Language-Action Models
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
Object-Centric Instruction Augmentation for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
Language-Conditioned Robotic Manipulation with Fast and Slow Thinking
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual Learning
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
Visual Robotic Manipulation with Depth-Aware Pretraining
von: Wang, Wanying, et al.
Veröffentlicht: (2024)
von: Wang, Wanying, et al.
Veröffentlicht: (2024)
WorldEval: World Model as Real-World Robot Policies Evaluator
von: Li, Yaxuan, et al.
Veröffentlicht: (2025)
von: Li, Yaxuan, et al.
Veröffentlicht: (2025)
MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?
von: Li, Jinming, et al.
Veröffentlicht: (2024)
von: Li, Jinming, et al.
Veröffentlicht: (2024)
Embodiment Transfer Learning for Vision-Language-Action Models
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
von: Zeng, Qiyuan, et al.
Veröffentlicht: (2025)
von: Zeng, Qiyuan, et al.
Veröffentlicht: (2025)
Mipha: A Comprehensive Overhaul of Multimodal Assistant with Small Language Models
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
Efficient Feature Fusion for UAV Object Detection
von: Wang, Xudong, et al.
Veröffentlicht: (2025)
von: Wang, Xudong, et al.
Veröffentlicht: (2025)
Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation
von: Wu, Kun, et al.
Veröffentlicht: (2024)
von: Wu, Kun, et al.
Veröffentlicht: (2024)
Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation
von: Zhu, Yichen, et al.
Veröffentlicht: (2025)
von: Zhu, Yichen, et al.
Veröffentlicht: (2025)
dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model
von: Li, Yaxuan, et al.
Veröffentlicht: (2026)
von: Li, Yaxuan, et al.
Veröffentlicht: (2026)
VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts
von: Jiang, Yuhua, et al.
Veröffentlicht: (2026)
von: Jiang, Yuhua, et al.
Veröffentlicht: (2026)
A Survey on Robotics with Foundation Models: toward Embodied AI
von: Xu, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Xu, Zhiyuan, et al.
Veröffentlicht: (2024)
IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
von: Chen, Yandu, et al.
Veröffentlicht: (2025)
von: Chen, Yandu, et al.
Veröffentlicht: (2025)
Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt
von: Zhu, Xiang, et al.
Veröffentlicht: (2025)
von: Zhu, Xiang, et al.
Veröffentlicht: (2025)
EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like Sketching
von: Chen, Xinwang, et al.
Veröffentlicht: (2024)
von: Chen, Xinwang, et al.
Veröffentlicht: (2024)
HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model
von: Zhu, Xiang, et al.
Veröffentlicht: (2026)
von: Zhu, Xiang, et al.
Veröffentlicht: (2026)
PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation
von: Li, Yutai, et al.
Veröffentlicht: (2026)
von: Li, Yutai, et al.
Veröffentlicht: (2026)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
von: Shen, Yichao, et al.
Veröffentlicht: (2025)
von: Shen, Yichao, et al.
Veröffentlicht: (2025)
TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches
von: Huang, Zhengxian, et al.
Veröffentlicht: (2026)
von: Huang, Zhengxian, et al.
Veröffentlicht: (2026)
Whether We Care, How We Reason: The Dual Role of Anthropomorphism and Moral Foundations in Robot Abuse
von: Yang, Fan, et al.
Veröffentlicht: (2026)
von: Yang, Fan, et al.
Veröffentlicht: (2026)
Distilling Reinforcement Learning Policies for Interpretable Robot Locomotion: Gradient Boosting Machines and Symbolic Regression
von: Acero, Fernando, et al.
Veröffentlicht: (2024)
von: Acero, Fernando, et al.
Veröffentlicht: (2024)
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
von: Su, Taiyi, et al.
Veröffentlicht: (2026)
von: Su, Taiyi, et al.
Veröffentlicht: (2026)
Can Emotional Quotient Boost Intelligence Quotient? The Influence of Service Robots' Communication Styles on Tourists' Memorable Experiences
von: Zhangxiang Zhu, et al.
Veröffentlicht: (2025)
von: Zhangxiang Zhu, et al.
Veröffentlicht: (2025)
LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
von: Zhu, Yichen, et al.
Veröffentlicht: (2024)
von: Zhu, Yichen, et al.
Veröffentlicht: (2024)
A Pragmatic VLA Foundation Model
von: Wu, Wei, et al.
Veröffentlicht: (2026)
von: Wu, Wei, et al.
Veröffentlicht: (2026)
Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
von: Liu, Chenyv, et al.
Veröffentlicht: (2026)
von: Liu, Chenyv, et al.
Veröffentlicht: (2026)
Inversion of the Windowed Special Affine Fourier Transform and Numerical Results
von: Yaoyao Han, et al.
Veröffentlicht: (2026)
von: Yaoyao Han, et al.
Veröffentlicht: (2026)
OrthoDiffusion: A Generalizable Multi-Task Diffusion Foundation Model for Musculoskeletal MRI Interpretation
von: Lan, Tian, et al.
Veröffentlicht: (2026)
von: Lan, Tian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
von: Wen, Junjie, et al.
Veröffentlicht: (2025) -
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
von: Zhu, Minjie, et al.
Veröffentlicht: (2025) -
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
von: Li, Jinming, et al.
Veröffentlicht: (2024) -
PointVLA: Injecting the 3D World into Vision-Language-Action Models
von: Li, Chengmeng, et al.
Veröffentlicht: (2025) -
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)