_version_ 1866914476783042560
author HY-World, Team
Cao, Chenjie
Zuo, Xuhui
Wang, Zhenwei
Zhang, Yisu
Wu, Junta
Liu, Zhenyang
Gong, Yuning
Liu, Yang
Yuan, Bo
Zhang, Chao
Li, Coopers
Guo, Dongyuan
Yang, Fan
Zhang, Haiyu
Cao, Hang
Zhu, Jianchen
Lin, Jiaxin
Xiao, Jie
Zhang, Jihong
Yu, Junlin
Wang, Lei
Wang, Lifu
Wang, Lilin
Linus
Chen, Minghui
He, Peng
Zhao, Penghao
Chen, Qi
Chen, Rui
Shao, Rui
Liu, Sicong
Qin, Wangchen
Niu, Xiaochuan
Yuan, Xiang
Sun, Yi
Tang, Yifei
Sun, Yifu
Lian, Yihang
Tan, Yonghao
Liu, Yuhong
Yin, Yuyang
Min, Zhiyuan
Wang, Tengfei
Guo, Chunchao
author_facet HY-World, Team
Cao, Chenjie
Zuo, Xuhui
Wang, Zhenwei
Zhang, Yisu
Wu, Junta
Liu, Zhenyang
Gong, Yuning
Liu, Yang
Yuan, Bo
Zhang, Chao
Li, Coopers
Guo, Dongyuan
Yang, Fan
Zhang, Haiyu
Cao, Hang
Zhu, Jianchen
Lin, Jiaxin
Xiao, Jie
Zhang, Jihong
Yu, Junlin
Wang, Lei
Wang, Lifu
Wang, Lilin
Linus
Chen, Minghui
He, Peng
Zhao, Penghao
Chen, Qi
Chen, Rui
Shao, Rui
Liu, Sicong
Qin, Wangchen
Niu, Xiaochuan
Yuan, Xiang
Sun, Yi
Tang, Yifei
Sun, Yifu
Lian, Yihang
Tan, Yonghao
Liu, Yuhong
Yin, Yuyang
Min, Zhiyuan
Wang, Tengfei
Guo, Chunchao
contents We introduce HY-World 2.0, a multi-modal world model framework that advances our prior project HY-World 1.0. HY-World 2.0 accommodates diverse input modalities, including text prompts, single-view images, multi-view images, and videos, and produces 3D world representations. With text or single-view image inputs, the model performs world generation, synthesizing high-fidelity, navigable 3D Gaussian Splatting (3DGS) scenes. This is achieved through a four-stage method: a) Panorama Generation with HY-Pano 2.0, b) Trajectory Planning with WorldNav, c) World Expansion with WorldStereo 2.0, and d) World Composition with WorldMirror 2.0. Specifically, we introduce key innovations to enhance panorama fidelity, enable 3D scene understanding and planning, and upgrade WorldStereo, our keyframe-based view generation model with consistent memory. We also upgrade WorldMirror, a feed-forward model for universal 3D prediction, by refining model architecture and learning strategy, enabling world reconstruction from multi-view images or videos. Also, we introduce WorldLens, a high-performance 3DGS rendering platform featuring a flexible engine-agnostic architecture, automatic IBL lighting, efficient collision detection, and training-rendering co-design, enabling interactive exploration of 3D worlds with character support. Extensive experiments demonstrate that HY-World 2.0 achieves state-of-the-art performance on several benchmarks among open-source approaches, delivering results comparable to the closed-source model Marble. We release all model weights, code, and technical details to facilitate reproducibility and support further research on 3D world models.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14268
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
HY-World, Team
Cao, Chenjie
Zuo, Xuhui
Wang, Zhenwei
Zhang, Yisu
Wu, Junta
Liu, Zhenyang
Gong, Yuning
Liu, Yang
Yuan, Bo
Zhang, Chao
Li, Coopers
Guo, Dongyuan
Yang, Fan
Zhang, Haiyu
Cao, Hang
Zhu, Jianchen
Lin, Jiaxin
Xiao, Jie
Zhang, Jihong
Yu, Junlin
Wang, Lei
Wang, Lifu
Wang, Lilin
Linus
Chen, Minghui
He, Peng
Zhao, Penghao
Chen, Qi
Chen, Rui
Shao, Rui
Liu, Sicong
Qin, Wangchen
Niu, Xiaochuan
Yuan, Xiang
Sun, Yi
Tang, Yifei
Sun, Yifu
Lian, Yihang
Tan, Yonghao
Liu, Yuhong
Yin, Yuyang
Min, Zhiyuan
Wang, Tengfei
Guo, Chunchao
Computer Vision and Pattern Recognition
We introduce HY-World 2.0, a multi-modal world model framework that advances our prior project HY-World 1.0. HY-World 2.0 accommodates diverse input modalities, including text prompts, single-view images, multi-view images, and videos, and produces 3D world representations. With text or single-view image inputs, the model performs world generation, synthesizing high-fidelity, navigable 3D Gaussian Splatting (3DGS) scenes. This is achieved through a four-stage method: a) Panorama Generation with HY-Pano 2.0, b) Trajectory Planning with WorldNav, c) World Expansion with WorldStereo 2.0, and d) World Composition with WorldMirror 2.0. Specifically, we introduce key innovations to enhance panorama fidelity, enable 3D scene understanding and planning, and upgrade WorldStereo, our keyframe-based view generation model with consistent memory. We also upgrade WorldMirror, a feed-forward model for universal 3D prediction, by refining model architecture and learning strategy, enabling world reconstruction from multi-view images or videos. Also, we introduce WorldLens, a high-performance 3DGS rendering platform featuring a flexible engine-agnostic architecture, automatic IBL lighting, efficient collision detection, and training-rendering co-design, enabling interactive exploration of 3D worlds with character support. Extensive experiments demonstrate that HY-World 2.0 achieves state-of-the-art performance on several benchmarks among open-source approaches, delivering results comparable to the closed-source model Marble. We release all model weights, code, and technical details to facilitate reproducibility and support further research on 3D world models.
title HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.14268