Yume-1.5: A Text-Controlled Interactive World Generation Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mao, Xiaofeng, Li, Zhen, Li, Chuanhao, Xu, Xiaojie, Ying, Kaining, He, Tong, Pang, Jiangmiao, Qiao, Yu, Zhang, Kaipeng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911340117884928
author Mao, Xiaofeng
Li, Zhen
Li, Chuanhao
Xu, Xiaojie
Ying, Kaining
He, Tong
Pang, Jiangmiao
Qiao, Yu
Zhang, Kaipeng
author_facet Mao, Xiaofeng
Li, Zhen
Li, Chuanhao
Xu, Xiaojie
Ying, Kaining
He, Tong
Pang, Jiangmiao
Qiao, Yu
Zhang, Kaipeng
contents Recent approaches have demonstrated the promise of using diffusion models to generate interactive and explorable worlds. However, most of these methods face critical challenges such as excessively large parameter sizes, reliance on lengthy inference steps, and rapidly growing historical context, which severely limit real-time performance and lack text-controlled generation capabilities. To address these challenges, we propose \method, a novel framework designed to generate realistic, interactive, and continuous worlds from a single image or text prompt. \method achieves this through a carefully designed framework that supports keyboard-based exploration of the generated worlds. The framework comprises three core components: (1) a long-video generation framework integrating unified context compression with linear attention; (2) a real-time streaming acceleration strategy powered by bidirectional attention distillation and an enhanced text embedding scheme; (3) a text-controlled method for generating world events. We have provided the codebase in the supplementary material.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22096
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Yume-1.5: A Text-Controlled Interactive World Generation Model
Mao, Xiaofeng
Li, Zhen
Li, Chuanhao
Xu, Xiaojie
Ying, Kaining
He, Tong
Pang, Jiangmiao
Qiao, Yu
Zhang, Kaipeng
Computer Vision and Pattern Recognition
Recent approaches have demonstrated the promise of using diffusion models to generate interactive and explorable worlds. However, most of these methods face critical challenges such as excessively large parameter sizes, reliance on lengthy inference steps, and rapidly growing historical context, which severely limit real-time performance and lack text-controlled generation capabilities. To address these challenges, we propose \method, a novel framework designed to generate realistic, interactive, and continuous worlds from a single image or text prompt. \method achieves this through a carefully designed framework that supports keyboard-based exploration of the generated worlds. The framework comprises three core components: (1) a long-video generation framework integrating unified context compression with linear attention; (2) a real-time streaming acceleration strategy powered by bidirectional attention distillation and an enhanced text embedding scheme; (3) a text-controlled method for generating world events. We have provided the codebase in the supplementary material.
title Yume-1.5: A Text-Controlled Interactive World Generation Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.22096