Yume: An Interactive World Generation Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Xiaofeng, Lin, Shaoheng, Li, Zhen, Li, Chuanhao, Peng, Wenshuo, He, Tong, Pang, Jiangmiao, Chi, Mingmin, Qiao, Yu, Zhang, Kaipeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913957000773632
author Mao, Xiaofeng
Lin, Shaoheng
Li, Zhen
Li, Chuanhao
Peng, Wenshuo
He, Tong
Pang, Jiangmiao
Chi, Mingmin
Qiao, Yu
Zhang, Kaipeng
author_facet Mao, Xiaofeng
Lin, Shaoheng
Li, Zhen
Li, Chuanhao
Peng, Wenshuo
He, Tong
Pang, Jiangmiao
Chi, Mingmin
Qiao, Yu
Zhang, Kaipeng
contents Yume aims to use images, text, or videos to create an interactive, realistic, and dynamic world, which allows exploration and control using peripheral devices or neural signals. In this report, we present a preview version of \method, which creates a dynamic world from an input image and allows exploration of the world using keyboard actions. To achieve this high-fidelity and interactive video world generation, we introduce a well-designed framework, which consists of four main components, including camera motion quantization, video generation architecture, advanced sampler, and model acceleration. First, we quantize camera motions for stable training and user-friendly interaction using keyboard inputs. Then, we introduce the Masked Video Diffusion Transformer~(MVDT) with a memory module for infinite video generation in an autoregressive manner. After that, training-free Anti-Artifact Mechanism (AAM) and Time Travel Sampling based on Stochastic Differential Equations (TTS-SDE) are introduced to the sampler for better visual quality and more precise control. Moreover, we investigate model acceleration by synergistic optimization of adversarial distillation and caching mechanisms. We use the high-quality world exploration dataset \sekai to train \method, and it achieves remarkable results in diverse scenes and applications. All data, codebase, and model weights are available on https://github.com/stdstu12/YUME. Yume will update monthly to achieve its original goal. Project page: https://stdstu12.github.io/YUME-Project/.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17744
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Yume: An Interactive World Generation Model
Mao, Xiaofeng
Lin, Shaoheng
Li, Zhen
Li, Chuanhao
Peng, Wenshuo
He, Tong
Pang, Jiangmiao
Chi, Mingmin
Qiao, Yu
Zhang, Kaipeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
Yume aims to use images, text, or videos to create an interactive, realistic, and dynamic world, which allows exploration and control using peripheral devices or neural signals. In this report, we present a preview version of \method, which creates a dynamic world from an input image and allows exploration of the world using keyboard actions. To achieve this high-fidelity and interactive video world generation, we introduce a well-designed framework, which consists of four main components, including camera motion quantization, video generation architecture, advanced sampler, and model acceleration. First, we quantize camera motions for stable training and user-friendly interaction using keyboard inputs. Then, we introduce the Masked Video Diffusion Transformer~(MVDT) with a memory module for infinite video generation in an autoregressive manner. After that, training-free Anti-Artifact Mechanism (AAM) and Time Travel Sampling based on Stochastic Differential Equations (TTS-SDE) are introduced to the sampler for better visual quality and more precise control. Moreover, we investigate model acceleration by synergistic optimization of adversarial distillation and caching mechanisms. We use the high-quality world exploration dataset \sekai to train \method, and it achieves remarkable results in diverse scenes and applications. All data, codebase, and model weights are available on https://github.com/stdstu12/YUME. Yume will update monthly to achieve its original goal. Project page: https://stdstu12.github.io/YUME-Project/.
title Yume: An Interactive World Generation Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2507.17744