Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.14268 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914476783042560 |
|---|---|
| author | HY-World, Team Cao, Chenjie Zuo, Xuhui Wang, Zhenwei Zhang, Yisu Wu, Junta Liu, Zhenyang Gong, Yuning Liu, Yang Yuan, Bo Zhang, Chao Li, Coopers Guo, Dongyuan Yang, Fan Zhang, Haiyu Cao, Hang Zhu, Jianchen Lin, Jiaxin Xiao, Jie Zhang, Jihong Yu, Junlin Wang, Lei Wang, Lifu Wang, Lilin Linus Chen, Minghui He, Peng Zhao, Penghao Chen, Qi Chen, Rui Shao, Rui Liu, Sicong Qin, Wangchen Niu, Xiaochuan Yuan, Xiang Sun, Yi Tang, Yifei Sun, Yifu Lian, Yihang Tan, Yonghao Liu, Yuhong Yin, Yuyang Min, Zhiyuan Wang, Tengfei Guo, Chunchao |
| author_facet | HY-World, Team Cao, Chenjie Zuo, Xuhui Wang, Zhenwei Zhang, Yisu Wu, Junta Liu, Zhenyang Gong, Yuning Liu, Yang Yuan, Bo Zhang, Chao Li, Coopers Guo, Dongyuan Yang, Fan Zhang, Haiyu Cao, Hang Zhu, Jianchen Lin, Jiaxin Xiao, Jie Zhang, Jihong Yu, Junlin Wang, Lei Wang, Lifu Wang, Lilin Linus Chen, Minghui He, Peng Zhao, Penghao Chen, Qi Chen, Rui Shao, Rui Liu, Sicong Qin, Wangchen Niu, Xiaochuan Yuan, Xiang Sun, Yi Tang, Yifei Sun, Yifu Lian, Yihang Tan, Yonghao Liu, Yuhong Yin, Yuyang Min, Zhiyuan Wang, Tengfei Guo, Chunchao |
| contents | We introduce HY-World 2.0, a multi-modal world model framework that advances our prior project HY-World 1.0. HY-World 2.0 accommodates diverse input modalities, including text prompts, single-view images, multi-view images, and videos, and produces 3D world representations. With text or single-view image inputs, the model performs world generation, synthesizing high-fidelity, navigable 3D Gaussian Splatting (3DGS) scenes. This is achieved through a four-stage method: a) Panorama Generation with HY-Pano 2.0, b) Trajectory Planning with WorldNav, c) World Expansion with WorldStereo 2.0, and d) World Composition with WorldMirror 2.0. Specifically, we introduce key innovations to enhance panorama fidelity, enable 3D scene understanding and planning, and upgrade WorldStereo, our keyframe-based view generation model with consistent memory. We also upgrade WorldMirror, a feed-forward model for universal 3D prediction, by refining model architecture and learning strategy, enabling world reconstruction from multi-view images or videos. Also, we introduce WorldLens, a high-performance 3DGS rendering platform featuring a flexible engine-agnostic architecture, automatic IBL lighting, efficient collision detection, and training-rendering co-design, enabling interactive exploration of 3D worlds with character support. Extensive experiments demonstrate that HY-World 2.0 achieves state-of-the-art performance on several benchmarks among open-source approaches, delivering results comparable to the closed-source model Marble. We release all model weights, code, and technical details to facilitate reproducibility and support further research on 3D world models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_14268 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds HY-World, Team Cao, Chenjie Zuo, Xuhui Wang, Zhenwei Zhang, Yisu Wu, Junta Liu, Zhenyang Gong, Yuning Liu, Yang Yuan, Bo Zhang, Chao Li, Coopers Guo, Dongyuan Yang, Fan Zhang, Haiyu Cao, Hang Zhu, Jianchen Lin, Jiaxin Xiao, Jie Zhang, Jihong Yu, Junlin Wang, Lei Wang, Lifu Wang, Lilin Linus Chen, Minghui He, Peng Zhao, Penghao Chen, Qi Chen, Rui Shao, Rui Liu, Sicong Qin, Wangchen Niu, Xiaochuan Yuan, Xiang Sun, Yi Tang, Yifei Sun, Yifu Lian, Yihang Tan, Yonghao Liu, Yuhong Yin, Yuyang Min, Zhiyuan Wang, Tengfei Guo, Chunchao Computer Vision and Pattern Recognition We introduce HY-World 2.0, a multi-modal world model framework that advances our prior project HY-World 1.0. HY-World 2.0 accommodates diverse input modalities, including text prompts, single-view images, multi-view images, and videos, and produces 3D world representations. With text or single-view image inputs, the model performs world generation, synthesizing high-fidelity, navigable 3D Gaussian Splatting (3DGS) scenes. This is achieved through a four-stage method: a) Panorama Generation with HY-Pano 2.0, b) Trajectory Planning with WorldNav, c) World Expansion with WorldStereo 2.0, and d) World Composition with WorldMirror 2.0. Specifically, we introduce key innovations to enhance panorama fidelity, enable 3D scene understanding and planning, and upgrade WorldStereo, our keyframe-based view generation model with consistent memory. We also upgrade WorldMirror, a feed-forward model for universal 3D prediction, by refining model architecture and learning strategy, enabling world reconstruction from multi-view images or videos. Also, we introduce WorldLens, a high-performance 3DGS rendering platform featuring a flexible engine-agnostic architecture, automatic IBL lighting, efficient collision detection, and training-rendering co-design, enabling interactive exploration of 3D worlds with character support. Extensive experiments demonstrate that HY-World 2.0 achieves state-of-the-art performance on several benchmarks among open-source approaches, delivering results comparable to the closed-source model Marble. We release all model weights, code, and technical details to facilitate reproducibility and support further research on 3D world models. |
| title | HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2604.14268 |