Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liao, Yue, Zhou, Pengfei, Huang, Siyuan, Yang, Donglin, Chen, Shengcong, Jiang, Yuxin, Hu, Yue, Cai, Jingbin, Liu, Si, Luo, Jianlan, Chen, Liliang, Yan, Shuicheng, Yao, Maoqing, Ren, Guanghui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917058941288448
author Liao, Yue
Zhou, Pengfei
Huang, Siyuan
Yang, Donglin
Chen, Shengcong
Jiang, Yuxin
Hu, Yue
Cai, Jingbin
Liu, Si
Luo, Jianlan
Chen, Liliang
Yan, Shuicheng
Yao, Maoqing
Ren, Guanghui
author_facet Liao, Yue
Zhou, Pengfei
Huang, Siyuan
Yang, Donglin
Chen, Shengcong
Jiang, Yuxin
Hu, Yue
Cai, Jingbin
Liu, Si
Luo, Jianlan
Chen, Liliang
Yan, Shuicheng
Yao, Maoqing
Ren, Guanghui
contents We introduce Genie Envisioner (GE), a unified world foundation platform for robotic manipulation that integrates policy learning, evaluation, and simulation within a single video-generative framework. At its core, GE-Base is a large-scale, instruction-conditioned video diffusion model that captures the spatial, temporal, and semantic dynamics of real-world robotic interactions in a structured latent space. Built upon this foundation, GE-Act maps latent representations to executable action trajectories through a lightweight, flow-matching decoder, enabling precise and generalizable policy inference across diverse embodiments with minimal supervision. To support scalable evaluation and training, GE-Sim serves as an action-conditioned neural simulator, producing high-fidelity rollouts for closed-loop policy development. The platform is further equipped with EWMBench, a standardized benchmark suite measuring visual fidelity, physical consistency, and instruction-action alignment. Together, these components establish Genie Envisioner as a scalable and practical foundation for instruction-driven, general-purpose embodied intelligence. All code, models, and benchmarks will be released publicly.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05635
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
Liao, Yue
Zhou, Pengfei
Huang, Siyuan
Yang, Donglin
Chen, Shengcong
Jiang, Yuxin
Hu, Yue
Cai, Jingbin
Liu, Si
Luo, Jianlan
Chen, Liliang
Yan, Shuicheng
Yao, Maoqing
Ren, Guanghui
Robotics
Computer Vision and Pattern Recognition
We introduce Genie Envisioner (GE), a unified world foundation platform for robotic manipulation that integrates policy learning, evaluation, and simulation within a single video-generative framework. At its core, GE-Base is a large-scale, instruction-conditioned video diffusion model that captures the spatial, temporal, and semantic dynamics of real-world robotic interactions in a structured latent space. Built upon this foundation, GE-Act maps latent representations to executable action trajectories through a lightweight, flow-matching decoder, enabling precise and generalizable policy inference across diverse embodiments with minimal supervision. To support scalable evaluation and training, GE-Sim serves as an action-conditioned neural simulator, producing high-fidelity rollouts for closed-loop policy development. The platform is further equipped with EWMBench, a standardized benchmark suite measuring visual fidelity, physical consistency, and instruction-action alignment. Together, these components establish Genie Envisioner as a scalable and practical foundation for instruction-driven, general-purpose embodied intelligence. All code, models, and benchmarks will be released publicly.
title Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.05635