Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917058941288448 |
|---|---|
| author | Liao, Yue Zhou, Pengfei Huang, Siyuan Yang, Donglin Chen, Shengcong Jiang, Yuxin Hu, Yue Cai, Jingbin Liu, Si Luo, Jianlan Chen, Liliang Yan, Shuicheng Yao, Maoqing Ren, Guanghui |
| author_facet | Liao, Yue Zhou, Pengfei Huang, Siyuan Yang, Donglin Chen, Shengcong Jiang, Yuxin Hu, Yue Cai, Jingbin Liu, Si Luo, Jianlan Chen, Liliang Yan, Shuicheng Yao, Maoqing Ren, Guanghui |
| contents | We introduce Genie Envisioner (GE), a unified world foundation platform for robotic manipulation that integrates policy learning, evaluation, and simulation within a single video-generative framework. At its core, GE-Base is a large-scale, instruction-conditioned video diffusion model that captures the spatial, temporal, and semantic dynamics of real-world robotic interactions in a structured latent space. Built upon this foundation, GE-Act maps latent representations to executable action trajectories through a lightweight, flow-matching decoder, enabling precise and generalizable policy inference across diverse embodiments with minimal supervision. To support scalable evaluation and training, GE-Sim serves as an action-conditioned neural simulator, producing high-fidelity rollouts for closed-loop policy development. The platform is further equipped with EWMBench, a standardized benchmark suite measuring visual fidelity, physical consistency, and instruction-action alignment. Together, these components establish Genie Envisioner as a scalable and practical foundation for instruction-driven, general-purpose embodied intelligence. All code, models, and benchmarks will be released publicly. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_05635 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation Liao, Yue Zhou, Pengfei Huang, Siyuan Yang, Donglin Chen, Shengcong Jiang, Yuxin Hu, Yue Cai, Jingbin Liu, Si Luo, Jianlan Chen, Liliang Yan, Shuicheng Yao, Maoqing Ren, Guanghui Robotics Computer Vision and Pattern Recognition We introduce Genie Envisioner (GE), a unified world foundation platform for robotic manipulation that integrates policy learning, evaluation, and simulation within a single video-generative framework. At its core, GE-Base is a large-scale, instruction-conditioned video diffusion model that captures the spatial, temporal, and semantic dynamics of real-world robotic interactions in a structured latent space. Built upon this foundation, GE-Act maps latent representations to executable action trajectories through a lightweight, flow-matching decoder, enabling precise and generalizable policy inference across diverse embodiments with minimal supervision. To support scalable evaluation and training, GE-Sim serves as an action-conditioned neural simulator, producing high-fidelity rollouts for closed-loop policy development. The platform is further equipped with EWMBench, a standardized benchmark suite measuring visual fidelity, physical consistency, and instruction-action alignment. Together, these components establish Genie Envisioner as a scalable and practical foundation for instruction-driven, general-purpose embodied intelligence. All code, models, and benchmarks will be released publicly. |
| title | Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation |
| topic | Robotics Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2508.05635 |