EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transfer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dong, Zhehao, Wang, Xiaofeng, Zhu, Zheng, Wang, Yirui, Wang, Yang, Zhou, Yukun, Wang, Boyuan, Ni, Chaojun, Ouyang, Runqi, Qin, Wenkang, Chen, Xinze, Ye, Yun, Huang, Guan, Lu, Zhen, Yang, Yue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training
von: Li, Haoyun, et al.
Veröffentlicht: (2025)
von: Li, Haoyun, et al.
Veröffentlicht: (2025)
EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling
von: Wang, Boyuan, et al.
Veröffentlicht: (2025)
von: Wang, Boyuan, et al.
Veröffentlicht: (2025)
HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation
von: Wang, Boyuan, et al.
Veröffentlicht: (2025)
von: Wang, Boyuan, et al.
Veröffentlicht: (2025)
ReconDreamer++: Harmonizing Generative and Reconstructive Models for Driving Scene Representation
von: Zhao, Guosheng, et al.
Veröffentlicht: (2025)
von: Zhao, Guosheng, et al.
Veröffentlicht: (2025)
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
von: Wang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Wang, Xiaofeng, et al.
Veröffentlicht: (2024)
WonderTurbo: Generating Interactive 3D World in 0.72 Seconds
von: Ni, Chaojun, et al.
Veröffentlicht: (2025)
von: Ni, Chaojun, et al.
Veröffentlicht: (2025)
GigaBrain-0: A World Model-Powered Vision-Language-Action Model
von: GigaBrain Team, et al.
Veröffentlicht: (2025)
von: GigaBrain Team, et al.
Veröffentlicht: (2025)
GigaWorld-0: World Models as Data Engine to Empower Embodied AI
von: GigaWorld Team, et al.
Veröffentlicht: (2025)
von: GigaWorld Team, et al.
Veröffentlicht: (2025)
ReconDreamer-RL: Enhancing Reinforcement Learning via Diffusion-based Scene Reconstruction
von: Ni, Chaojun, et al.
Veröffentlicht: (2025)
von: Ni, Chaojun, et al.
Veröffentlicht: (2025)
HumanDreamer-X: Photorealistic Single-image Human Avatars Reconstruction via Gaussian Restoration
von: Wang, Boyuan, et al.
Veröffentlicht: (2025)
von: Wang, Boyuan, et al.
Veröffentlicht: (2025)
DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation
von: Zhao, Guosheng, et al.
Veröffentlicht: (2024)
von: Zhao, Guosheng, et al.
Veröffentlicht: (2024)
DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion
von: Wang, Weijie, et al.
Veröffentlicht: (2025)
von: Wang, Weijie, et al.
Veröffentlicht: (2025)
EgoDemoGen: Egocentric Demonstration Generation for Viewpoint Generalization in Robotic Manipulation
von: Xu, Yuan, et al.
Veröffentlicht: (2025)
von: Xu, Yuan, et al.
Veröffentlicht: (2025)
VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis
von: Lang, Xiaolei, et al.
Veröffentlicht: (2026)
von: Lang, Xiaolei, et al.
Veröffentlicht: (2026)
RoboTransfer: Controllable Geometry-Consistent Video Diffusion for Manipulation Policy Transfer
von: Liu, Liu, et al.
Veröffentlicht: (2025)
von: Liu, Liu, et al.
Veröffentlicht: (2025)
DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation
von: Zhao, Guosheng, et al.
Veröffentlicht: (2024)
von: Zhao, Guosheng, et al.
Veröffentlicht: (2024)
SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
von: Ni, Chaojun, et al.
Veröffentlicht: (2025)
von: Ni, Chaojun, et al.
Veröffentlicht: (2025)
GigaWorld-Policy: An Efficient Action-Centered World--Action Model
von: Ye, Angen, et al.
Veröffentlicht: (2026)
von: Ye, Angen, et al.
Veröffentlicht: (2026)
ShapeGen: Robotic Data Generation for Category-Level Manipulation
von: Wang, Yirui, et al.
Veröffentlicht: (2026)
von: Wang, Yirui, et al.
Veröffentlicht: (2026)
ViVa: A Video-Generative Value Model for Robot Reinforcement Learning
von: Lv, Jindi, et al.
Veröffentlicht: (2026)
von: Lv, Jindi, et al.
Veröffentlicht: (2026)
Motion-R1: Enhancing Motion Generation with Decomposed Chain-of-Thought and RL Binding
von: Ouyang, Runqi, et al.
Veröffentlicht: (2025)
von: Ouyang, Runqi, et al.
Veröffentlicht: (2025)
ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration
von: Ni, Chaojun, et al.
Veröffentlicht: (2024)
von: Ni, Chaojun, et al.
Veröffentlicht: (2024)
EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
von: He, Xin, et al.
Veröffentlicht: (2025)
von: He, Xin, et al.
Veröffentlicht: (2025)
FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model
von: Gao, Chongkai, et al.
Veröffentlicht: (2024)
von: Gao, Chongkai, et al.
Veröffentlicht: (2024)
STORM: Search-Guided Generative World Models for Robotic Manipulation
von: Lin, Wenjun, et al.
Veröffentlicht: (2025)
von: Lin, Wenjun, et al.
Veröffentlicht: (2025)
GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning
von: GigaBrain Team, et al.
Veröffentlicht: (2026)
von: GigaBrain Team, et al.
Veröffentlicht: (2026)
Real‐Time Motion Generation for Robot Manipulators in Complex Dynamic Environments
von: Tianyu Zhang, et al.
Veröffentlicht: (2025)
von: Tianyu Zhang, et al.
Veröffentlicht: (2025)
WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation
von: Qian, Zezhong, et al.
Veröffentlicht: (2025)
von: Qian, Zezhong, et al.
Veröffentlicht: (2025)
Real Garment Benchmark (RGBench): A Comprehensive Benchmark for Robotic Garment Manipulation featuring a High-Fidelity Scalable Simulator
von: Hu, Wenkang, et al.
Veröffentlicht: (2025)
von: Hu, Wenkang, et al.
Veröffentlicht: (2025)
WonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene Exploration
von: Ni, Chaojun, et al.
Veröffentlicht: (2025)
von: Ni, Chaojun, et al.
Veröffentlicht: (2025)
Beyond Visual Safety: Jailbreaking Multimodal Large Language Models for Harmful Image Generation via Semantic-Agnostic Inputs
von: Yu, Mingyu, et al.
Veröffentlicht: (2026)
von: Yu, Mingyu, et al.
Veröffentlicht: (2026)
EmbodiedGen: Towards a Generative 3D World Engine for Embodied Intelligence
von: Wang, Xinjie, et al.
Veröffentlicht: (2025)
von: Wang, Xinjie, et al.
Veröffentlicht: (2025)
ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video
von: Wang, Boyuan, et al.
Veröffentlicht: (2026)
von: Wang, Boyuan, et al.
Veröffentlicht: (2026)
ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model
von: Yu, Tengbo, et al.
Veröffentlicht: (2025)
von: Yu, Tengbo, et al.
Veröffentlicht: (2025)
Stable-Makeup: When Real-World Makeup Transfer Meets Diffusion Model
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2024)
FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception
von: Zhang, Zhen, et al.
Veröffentlicht: (2026)
von: Zhang, Zhen, et al.
Veröffentlicht: (2026)
VoltSchemer: Use Voltage Noise to Manipulate Your Wireless Charger
von: Zhan, Zihao, et al.
Veröffentlicht: (2024)
von: Zhan, Zihao, et al.
Veröffentlicht: (2024)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025)
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025)
Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation
von: Xie, Senwei, et al.
Veröffentlicht: (2025)
von: Xie, Senwei, et al.
Veröffentlicht: (2025)
ManiAgent: An Agentic Framework for General Robotic Manipulation
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training
von: Li, Haoyun, et al.
Veröffentlicht: (2025) -
EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling
von: Wang, Boyuan, et al.
Veröffentlicht: (2025) -
HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation
von: Wang, Boyuan, et al.
Veröffentlicht: (2025) -
ReconDreamer++: Harmonizing Generative and Reconstructive Models for Driving Scene Representation
von: Zhao, Guosheng, et al.
Veröffentlicht: (2025) -
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
von: Wang, Xiaofeng, et al.
Veröffentlicht: (2024)