Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Minjie, Zhu, Yichen, Li, Jinming, Wen, Junjie, Xu, Zhiyuan, Liu, Ning, Cheng, Ran, Shen, Chaomin, Peng, Yaxin, Feng, Feifei, Tang, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Object-Centric Instruction Augmentation for Robotic Manipulation
by: Wen, Junjie, et al.
Published: (2024)
by: Wen, Junjie, et al.
Published: (2024)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
by: Wen, Junjie, et al.
Published: (2024)
by: Wen, Junjie, et al.
Published: (2024)
Language-Conditioned Robotic Manipulation with Fast and Slow Thinking
by: Zhu, Minjie, et al.
Published: (2024)
by: Zhu, Minjie, et al.
Published: (2024)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025)
by: Zhu, Minjie, et al.
Published: (2025)
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
by: Wen, Junjie, et al.
Published: (2025)
by: Wen, Junjie, et al.
Published: (2025)
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
by: Wen, Junjie, et al.
Published: (2024)
by: Wen, Junjie, et al.
Published: (2024)
Visual Robotic Manipulation with Depth-Aware Pretraining
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
by: Zhou, Zhongyi, et al.
Published: (2025)
by: Zhou, Zhongyi, et al.
Published: (2025)
Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation
by: Wu, Kun, et al.
Published: (2024)
by: Wu, Kun, et al.
Published: (2024)
Mipha: A Comprehensive Overhaul of Multimodal Assistant with Small Language Models
by: Zhu, Minjie, et al.
Published: (2024)
by: Zhu, Minjie, et al.
Published: (2024)
MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?
by: Li, Jinming, et al.
Published: (2024)
by: Li, Jinming, et al.
Published: (2024)
WorldEval: World Model as Real-World Robot Policies Evaluator
by: Li, Yaxuan, et al.
Published: (2025)
by: Li, Yaxuan, et al.
Published: (2025)
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
by: Li, Jinming, et al.
Published: (2024)
by: Li, Jinming, et al.
Published: (2024)
Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation
by: Zhu, Yichen, et al.
Published: (2025)
by: Zhu, Yichen, et al.
Published: (2025)
Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual Learning
by: Zhou, Zhongyi, et al.
Published: (2025)
by: Zhou, Zhongyi, et al.
Published: (2025)
EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like Sketching
by: Chen, Xinwang, et al.
Published: (2024)
by: Chen, Xinwang, et al.
Published: (2024)
PointVLA: Injecting the 3D World into Vision-Language-Action Models
by: Li, Chengmeng, et al.
Published: (2025)
by: Li, Chengmeng, et al.
Published: (2025)
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
by: Zhou, Zhongyi, et al.
Published: (2025)
by: Zhou, Zhongyi, et al.
Published: (2025)
Scaling Diffusion Transformers to 16 Billion Parameters
by: Fei, Zhengcong, et al.
Published: (2024)
by: Fei, Zhengcong, et al.
Published: (2024)
Efficient Feature Fusion for UAV Object Detection
by: Wang, Xudong, et al.
Published: (2025)
by: Wang, Xudong, et al.
Published: (2025)
A Survey on Robotics with Foundation Models: toward Embodied AI
by: Xu, Zhiyuan, et al.
Published: (2024)
by: Xu, Zhiyuan, et al.
Published: (2024)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
by: Wen, Junjie, et al.
Published: (2025)
by: Wen, Junjie, et al.
Published: (2025)
LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
by: Zhu, Yichen, et al.
Published: (2024)
by: Zhu, Yichen, et al.
Published: (2024)
Scaling Recommender Transformers to One Billion Parameters
by: Khrylchenko, Kirill, et al.
Published: (2025)
by: Khrylchenko, Kirill, et al.
Published: (2025)
Learning from Imperfect Demonstrations with Self-Supervision for Robotic Manipulation
by: Wu, Kun, et al.
Published: (2024)
by: Wu, Kun, et al.
Published: (2024)
SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation
by: Gui, Youqiang, et al.
Published: (2026)
by: Gui, Youqiang, et al.
Published: (2026)
U-DiT Policy: U-shaped Diffusion Transformers for Robotic Manipulation
by: Wu, Linzhi, et al.
Published: (2025)
by: Wu, Linzhi, et al.
Published: (2025)
ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
by: Zeng, Qiyuan, et al.
Published: (2025)
by: Zeng, Qiyuan, et al.
Published: (2025)
On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
by: Wu, Yiming, et al.
Published: (2025)
by: Wu, Yiming, et al.
Published: (2025)
AnchorDP3: 3D Affordance Guided Sparse Diffusion Policy for Robotic Manipulation
by: Zhao, Ziyan, et al.
Published: (2025)
by: Zhao, Ziyan, et al.
Published: (2025)
HumanoidExo: Scalable Whole-Body Humanoid Manipulation via Wearable Exoskeleton
by: Zhong, Rui, et al.
Published: (2025)
by: Zhong, Rui, et al.
Published: (2025)
dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model
by: Li, Yaxuan, et al.
Published: (2026)
by: Li, Yaxuan, et al.
Published: (2026)
Can Emotional Quotient Boost Intelligence Quotient? The Influence of Service Robots' Communication Styles on Tourists' Memorable Experiences
by: Zhangxiang Zhu, et al.
Published: (2025)
by: Zhangxiang Zhu, et al.
Published: (2025)
GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
by: Xiong, Tianwei, et al.
Published: (2025)
by: Xiong, Tianwei, et al.
Published: (2025)
VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation
by: Wen, Youpeng, et al.
Published: (2024)
by: Wen, Youpeng, et al.
Published: (2024)
Adaptive Diffusion Policy Optimization for Robotic Manipulation
by: Jiang, Huiyun, et al.
Published: (2025)
by: Jiang, Huiyun, et al.
Published: (2025)
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
by: Sun, Quan, et al.
Published: (2024)
by: Sun, Quan, et al.
Published: (2024)
Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt
by: Zhu, Xiang, et al.
Published: (2025)
by: Zhu, Xiang, et al.
Published: (2025)
Cross-Scenario Unified Modeling of User Interests at Billion Scale
by: Xu, Manjie, et al.
Published: (2025)
by: Xu, Manjie, et al.
Published: (2025)
Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning
by: Wang, Yixiao, et al.
Published: (2024)
by: Wang, Yixiao, et al.
Published: (2024)
Similar Items
-
Object-Centric Instruction Augmentation for Robotic Manipulation
by: Wen, Junjie, et al.
Published: (2024) -
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
by: Wen, Junjie, et al.
Published: (2024) -
Language-Conditioned Robotic Manipulation with Fast and Slow Thinking
by: Zhu, Minjie, et al.
Published: (2024) -
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025) -
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
by: Wen, Junjie, et al.
Published: (2025)