MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Haoyun, Zhang, Ivan, Ouyang, Runqi, Wang, Xiaofeng, Zhu, Zheng, Yang, Zhiqin, Zhang, Zhentao, Wang, Boyuan, Ni, Chaojun, Qin, Wenkang, Chen, Xinze, Ye, Yun, Huang, Guan, Song, Zhenbo, Wang, Xingang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation
by: Wang, Boyuan, et al.
Published: (2025)
by: Wang, Boyuan, et al.
Published: (2025)
ReconDreamer++: Harmonizing Generative and Reconstructive Models for Driving Scene Representation
by: Zhao, Guosheng, et al.
Published: (2025)
by: Zhao, Guosheng, et al.
Published: (2025)
HumanDreamer-X: Photorealistic Single-image Human Avatars Reconstruction via Gaussian Restoration
by: Wang, Boyuan, et al.
Published: (2025)
by: Wang, Boyuan, et al.
Published: (2025)
EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling
by: Wang, Boyuan, et al.
Published: (2025)
by: Wang, Boyuan, et al.
Published: (2025)
DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation
by: Zhao, Guosheng, et al.
Published: (2024)
by: Zhao, Guosheng, et al.
Published: (2024)
ReconDreamer-RL: Enhancing Reinforcement Learning via Diffusion-based Scene Reconstruction
by: Ni, Chaojun, et al.
Published: (2025)
by: Ni, Chaojun, et al.
Published: (2025)
EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transfer
by: Dong, Zhehao, et al.
Published: (2025)
by: Dong, Zhehao, et al.
Published: (2025)
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
by: Ni, Chaojun, et al.
Published: (2025)
by: Ni, Chaojun, et al.
Published: (2025)
DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation
by: Zhao, Guosheng, et al.
Published: (2024)
by: Zhao, Guosheng, et al.
Published: (2024)
Motion-R1: Enhancing Motion Generation with Decomposed Chain-of-Thought and RL Binding
by: Ouyang, Runqi, et al.
Published: (2025)
by: Ouyang, Runqi, et al.
Published: (2025)
ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration
by: Ni, Chaojun, et al.
Published: (2024)
by: Ni, Chaojun, et al.
Published: (2024)
WonderTurbo: Generating Interactive 3D World in 0.72 Seconds
by: Ni, Chaojun, et al.
Published: (2025)
by: Ni, Chaojun, et al.
Published: (2025)
WonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene Exploration
by: Ni, Chaojun, et al.
Published: (2025)
by: Ni, Chaojun, et al.
Published: (2025)
VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
by: Ye, Angen, et al.
Published: (2025)
by: Ye, Angen, et al.
Published: (2025)
GigaBrain-0: A World Model-Powered Vision-Language-Action Model
by: GigaBrain Team, et al.
Published: (2025)
by: GigaBrain Team, et al.
Published: (2025)
DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion
by: Wang, Weijie, et al.
Published: (2025)
by: Wang, Weijie, et al.
Published: (2025)
ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video
by: Wang, Boyuan, et al.
Published: (2026)
by: Wang, Boyuan, et al.
Published: (2026)
UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving
by: Zhao, Guosheng, et al.
Published: (2026)
by: Zhao, Guosheng, et al.
Published: (2026)
GigaWorld-0: World Models as Data Engine to Empower Embodied AI
by: GigaWorld Team, et al.
Published: (2025)
by: GigaWorld Team, et al.
Published: (2025)
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
by: Luo, Hao, et al.
Published: (2026)
by: Luo, Hao, et al.
Published: (2026)
EgoDemoGen: Egocentric Demonstration Generation for Viewpoint Generalization in Robotic Manipulation
by: Xu, Yuan, et al.
Published: (2025)
by: Xu, Yuan, et al.
Published: (2025)
GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning
by: GigaBrain Team, et al.
Published: (2026)
by: GigaBrain Team, et al.
Published: (2026)
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
by: Jiang, Yuming, et al.
Published: (2025)
by: Jiang, Yuming, et al.
Published: (2025)
GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning
by: Bao, Xiaoyi, et al.
Published: (2025)
by: Bao, Xiaoyi, et al.
Published: (2025)
SkillMimic: Learning Basketball Interaction Skills from Demonstrations
by: Wang, Yinhuai, et al.
Published: (2024)
by: Wang, Yinhuai, et al.
Published: (2024)
VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
by: Lu, Guanxing, et al.
Published: (2025)
by: Lu, Guanxing, et al.
Published: (2025)
Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
by: Chang, Yifan, et al.
Published: (2025)
by: Chang, Yifan, et al.
Published: (2025)
SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Mimic Grasping: A Modular and Flexible Programming‐by‐Demonstration Robotic Grasping Solution
by: João P. C. de Souza, et al.
Published: (2026)
by: João P. C. de Souza, et al.
Published: (2026)
TeethDreamer: 3D Teeth Reconstruction from Five Intra-oral Photographs
by: Xu, Chenfan, et al.
Published: (2024)
by: Xu, Chenfan, et al.
Published: (2024)
GED-Consistent Disentanglement of Aligned and Unaligned Substructures for Graph Similarity Learning
by: Zhan, Zhentao, et al.
Published: (2025)
by: Zhan, Zhentao, et al.
Published: (2025)
Hydroxylamine Hydrochloride as Bifunctional Reagent for Aminochlorination of Alkenes via Iron Catalysis
by: Guan‐Wang Huang, et al.
Published: (2026)
by: Guan‐Wang Huang, et al.
Published: (2026)
Hydroxylamine Hydrochloride as Bifunctional Reagent for Aminochlorination of Alkenes via Iron Catalysis
by: Guan‐Wang Huang, et al.
Published: (2026)
by: Guan‐Wang Huang, et al.
Published: (2026)
OmniVTLA: Vision-Tactile-Language-Action Model with Semantic-Aligned Tactile Sensing
by: Cheng, Zhengxue, et al.
Published: (2025)
by: Cheng, Zhengxue, et al.
Published: (2025)
ViVa: A Video-Generative Value Model for Robot Reinforcement Learning
by: Lv, Jindi, et al.
Published: (2026)
by: Lv, Jindi, et al.
Published: (2026)
Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning
by: Zahid, Azizul, et al.
Published: (2025)
by: Zahid, Azizul, et al.
Published: (2025)
Learning from Demonstration with Failure Awareness for Safe Robot Navigation
by: Wang, Xianghui, et al.
Published: (2026)
by: Wang, Xianghui, et al.
Published: (2026)
A Survey on Convex Optimization for Guidance and Control of Vehicular Systems
by: Wang, Zhenbo
Published: (2023)
by: Wang, Zhenbo
Published: (2023)
SimVLA: A Simple VLA Baseline for Robotic Manipulation
by: Luo, Yuankai, et al.
Published: (2026)
by: Luo, Yuankai, et al.
Published: (2026)
Similar Items
-
HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation
by: Wang, Boyuan, et al.
Published: (2025) -
ReconDreamer++: Harmonizing Generative and Reconstructive Models for Driving Scene Representation
by: Zhao, Guosheng, et al.
Published: (2025) -
HumanDreamer-X: Photorealistic Single-image Human Avatars Reconstruction via Gaussian Restoration
by: Wang, Boyuan, et al.
Published: (2025) -
EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling
by: Wang, Boyuan, et al.
Published: (2025) -
DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation
by: Zhao, Guosheng, et al.
Published: (2024)