Generative World Renderer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Zheng-Hui, Wang, Zhixiang, Tan, Jiaming, Yu, Ruihan, Zhang, Yidan, Zheng, Bo, Liu, Yu-Lun, Chuang, Yung-Yu, Zhang, Kaipeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910098228510720
author Huang, Zheng-Hui
Wang, Zhixiang
Tan, Jiaming
Yu, Ruihan
Zhang, Yidan
Zheng, Bo
Liu, Yu-Lun
Chuang, Yung-Yu
Zhang, Kaipeng
author_facet Huang, Zheng-Hui
Wang, Zhixiang
Tan, Jiaming
Yu, Ruihan
Zhang, Yidan
Zheng, Bo
Liu, Yu-Lun
Chuang, Yung-Yu
Zhang, Kaipeng
contents Scaling generative inverse and forward rendering to real-world scenarios is bottlenecked by the limited realism and temporal coherence of existing synthetic datasets. To bridge this persistent domain gap, we introduce a large-scale, dynamic dataset curated from visually complex AAA games. Using a novel dual-screen stitched capture method, we extracted 4M continuous frames (720p/30 FPS) of synchronized RGB and five G-buffer channels across diverse scenes, visual effects, and environments, including adverse weather and motion-blur variants. This dataset uniquely advances bidirectional rendering: enabling robust in-the-wild geometry and material decomposition, and facilitating high-fidelity G-buffer-guided video generation. Furthermore, to evaluate the real-world performance of inverse rendering without ground truth, we propose a novel VLM-based assessment protocol measuring semantic, spatial, and temporal consistency. Experiments demonstrate that inverse renderers fine-tuned on our data achieve superior cross-dataset generalization and controllable generation, while our VLM evaluation strongly correlates with human judgment. Combined with our toolkit, our forward renderer enables users to edit styles of AAA games from G-buffers using text prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02329
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Generative World Renderer
Huang, Zheng-Hui
Wang, Zhixiang
Tan, Jiaming
Yu, Ruihan
Zhang, Yidan
Zheng, Bo
Liu, Yu-Lun
Chuang, Yung-Yu
Zhang, Kaipeng
Computer Vision and Pattern Recognition
Scaling generative inverse and forward rendering to real-world scenarios is bottlenecked by the limited realism and temporal coherence of existing synthetic datasets. To bridge this persistent domain gap, we introduce a large-scale, dynamic dataset curated from visually complex AAA games. Using a novel dual-screen stitched capture method, we extracted 4M continuous frames (720p/30 FPS) of synchronized RGB and five G-buffer channels across diverse scenes, visual effects, and environments, including adverse weather and motion-blur variants. This dataset uniquely advances bidirectional rendering: enabling robust in-the-wild geometry and material decomposition, and facilitating high-fidelity G-buffer-guided video generation. Furthermore, to evaluate the real-world performance of inverse rendering without ground truth, we propose a novel VLM-based assessment protocol measuring semantic, spatial, and temporal consistency. Experiments demonstrate that inverse renderers fine-tuned on our data achieve superior cross-dataset generalization and controllable generation, while our VLM evaluation strongly correlates with human judgment. Combined with our toolkit, our forward renderer enables users to edit styles of AAA games from G-buffers using text prompts.
title Generative World Renderer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.02329