Yan: Foundational Interactive Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Deheng, Zhou, Fangyun, Lv, Jiacheng, Ma, Jianqi, Zhang, Jun, Lv, Junyan, Li, Junyou, Deng, Minwen, Yang, Mingyu, Fu, Qiang, Yang, Wei, Lv, Wenkai, Yu, Yangbin, Wang, Yewen, Guan, Yonghang, Hu, Zhihao, Fang, Zhongbin, Sun, Zhongqian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913990517456896
author Ye, Deheng
Zhou, Fangyun
Lv, Jiacheng
Ma, Jianqi
Zhang, Jun
Lv, Junyan
Li, Junyou
Deng, Minwen
Yang, Mingyu
Fu, Qiang
Yang, Wei
Lv, Wenkai
Yu, Yangbin
Wang, Yewen
Guan, Yonghang
Hu, Zhihao
Fang, Zhongbin
Sun, Zhongqian
author_facet Ye, Deheng
Zhou, Fangyun
Lv, Jiacheng
Ma, Jianqi
Zhang, Jun
Lv, Junyan
Li, Junyou
Deng, Minwen
Yang, Mingyu
Fu, Qiang
Yang, Wei
Lv, Wenkai
Yu, Yangbin
Wang, Yewen
Guan, Yonghang
Hu, Zhihao
Fang, Zhongbin
Sun, Zhongqian
contents We present Yan, a foundational framework for interactive video generation, covering the entire pipeline from simulation and generation to editing. Specifically, Yan comprises three core modules. AAA-level Simulation: We design a highly-compressed, low-latency 3D-VAE coupled with a KV-cache-based shift-window denoising inference process, achieving real-time 1080P/60FPS interactive simulation. Multi-Modal Generation: We introduce a hierarchical autoregressive caption method that injects game-specific knowledge into open-domain multi-modal video diffusion models (VDMs), then transforming the VDM into a frame-wise, action-controllable, real-time infinite interactive video generator. Notably, when the textual and visual prompts are sourced from different domains, the model demonstrates strong generalization, allowing it to blend and compose the style and mechanics across domains flexibly according to user prompts. Multi-Granularity Editing: We propose a hybrid model that explicitly disentangles interactive mechanics simulation from visual rendering, enabling multi-granularity video content editing during interaction through text. Collectively, Yan offers an integration of these modules, pushing interactive video generation beyond isolated capabilities toward a comprehensive AI-driven interactive creation paradigm, paving the way for the next generation of creative tools, media, and entertainment. The project page is: https://greatx3.github.io/Yan/.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08601
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Yan: Foundational Interactive Video Generation
Ye, Deheng
Zhou, Fangyun
Lv, Jiacheng
Ma, Jianqi
Zhang, Jun
Lv, Junyan
Li, Junyou
Deng, Minwen
Yang, Mingyu
Fu, Qiang
Yang, Wei
Lv, Wenkai
Yu, Yangbin
Wang, Yewen
Guan, Yonghang
Hu, Zhihao
Fang, Zhongbin
Sun, Zhongqian
Computer Vision and Pattern Recognition
Artificial Intelligence
We present Yan, a foundational framework for interactive video generation, covering the entire pipeline from simulation and generation to editing. Specifically, Yan comprises three core modules. AAA-level Simulation: We design a highly-compressed, low-latency 3D-VAE coupled with a KV-cache-based shift-window denoising inference process, achieving real-time 1080P/60FPS interactive simulation. Multi-Modal Generation: We introduce a hierarchical autoregressive caption method that injects game-specific knowledge into open-domain multi-modal video diffusion models (VDMs), then transforming the VDM into a frame-wise, action-controllable, real-time infinite interactive video generator. Notably, when the textual and visual prompts are sourced from different domains, the model demonstrates strong generalization, allowing it to blend and compose the style and mechanics across domains flexibly according to user prompts. Multi-Granularity Editing: We propose a hybrid model that explicitly disentangles interactive mechanics simulation from visual rendering, enabling multi-granularity video content editing during interaction through text. Collectively, Yan offers an integration of these modules, pushing interactive video generation beyond isolated capabilities toward a comprehensive AI-driven interactive creation paradigm, paving the way for the next generation of creative tools, media, and entertainment. The project page is: https://greatx3.github.io/Yan/.
title Yan: Foundational Interactive Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2508.08601