Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Tianyu, Zheng, Wangguandong, Wang, Tengfei, Liu, Yuhao, Wang, Zhenwei, Wu, Junta, Jiang, Jie, Li, Hui, Lau, Rynson W. H., Zuo, Wangmeng, Guo, Chunchao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy
by: Liu, Yuhao, et al.
Published: (2025)
by: Liu, Yuhao, et al.
Published: (2025)
WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
by: Sun, Wenqiang, et al.
Published: (2025)
by: Sun, Wenqiang, et al.
Published: (2025)
WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories
by: Zhang, Yisu, et al.
Published: (2026)
by: Zhang, Yisu, et al.
Published: (2026)
DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion Priors
by: Huang, Tianyu, et al.
Published: (2024)
by: Huang, Tianyu, et al.
Published: (2024)
WorldCompass: Reinforcement Learning for Long-Horizon World Models
by: Wang, Zehan, et al.
Published: (2026)
by: Wang, Zehan, et al.
Published: (2026)
WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Phidias: A Generative Model for Creating 3D Content from Text, Image, and 3D Conditions with Reference-Augmented Diffusion
by: Wang, Zhenwei, et al.
Published: (2024)
by: Wang, Zhenwei, et al.
Published: (2024)
ThemeStation: Generating Theme-Aware 3D Assets from Few Exemplars
by: Wang, Zhenwei, et al.
Published: (2024)
by: Wang, Zhenwei, et al.
Published: (2024)
HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels
by: HunyuanWorld Team, et al.
Published: (2025)
by: HunyuanWorld Team, et al.
Published: (2025)
TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields
by: Huang, Tianyu, et al.
Published: (2023)
by: Huang, Tianyu, et al.
Published: (2023)
PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion
by: Yin, Yuyang, et al.
Published: (2025)
by: Yin, Yuyang, et al.
Published: (2025)
Pathwise Test-Time Correction for Autoregressive Long Video Generation
by: Xiang, Xunzhi, et al.
Published: (2026)
by: Xiang, Xunzhi, et al.
Published: (2026)
Recasting Regional Lighting for Shadow Removal
by: Liu, Yuhao, et al.
Published: (2024)
by: Liu, Yuhao, et al.
Published: (2024)
FlashWorld: High-quality 3D Scene Generation within Seconds
by: Li, Xinyang, et al.
Published: (2025)
by: Li, Xinyang, et al.
Published: (2025)
Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
by: Yang, Zaiquan, et al.
Published: (2025)
by: Yang, Zaiquan, et al.
Published: (2025)
One2Scene: Geometric Consistent Explorable 3D Scene Generation from a Single Image
by: Wang, Pengfei, et al.
Published: (2026)
by: Wang, Pengfei, et al.
Published: (2026)
DreamControl: Control-Based Text-to-3D Generation with 3D Self-Prior
by: Huang, Tianyu, et al.
Published: (2023)
by: Huang, Tianyu, et al.
Published: (2023)
Self-Supervised High Dynamic Range Imaging with Multi-Exposure Images in Dynamic Scenes
by: Zhang, Zhilu, et al.
Published: (2023)
by: Zhang, Zhilu, et al.
Published: (2023)
World-Shaper: A Unified Framework for 360° Panoramic Editing
by: Liang, Dong, et al.
Published: (2026)
by: Liang, Dong, et al.
Published: (2026)
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
GenEx: Generating an Explorable World
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
Diff-Plugin: Revitalizing Details for Diffusion-based Low-level Tasks
by: Liu, Yuhao, et al.
Published: (2024)
by: Liu, Yuhao, et al.
Published: (2024)
MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation
by: Li, Zhiqi, et al.
Published: (2025)
by: Li, Zhiqi, et al.
Published: (2025)
Lie Flow: Video Dynamic Fields Modeling and Predicting with Lie Algebra as Geometric Physics Principle
by: Qiao, Weidong, et al.
Published: (2026)
by: Qiao, Weidong, et al.
Published: (2026)
Terra: Explorable Native 3D World Model with Point Latents
by: Huang, Yuanhui, et al.
Published: (2025)
by: Huang, Yuanhui, et al.
Published: (2025)
StyleSculptor: Zero-Shot Style-Controllable 3D Asset Generation with Texture-Geometry Dual Guidance
by: Qu, Zefan, et al.
Published: (2025)
by: Qu, Zefan, et al.
Published: (2025)
RelayAttention for Efficient Large Language Model Serving with Long System Prompts
by: Zhu, Lei, et al.
Published: (2024)
by: Zhu, Lei, et al.
Published: (2024)
MirrorMamba: Towards Scalable and Robust Mirror Detection in Videos
by: Song, Rui, et al.
Published: (2025)
by: Song, Rui, et al.
Published: (2025)
PhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration Synthesis
by: Yang, Yu, et al.
Published: (2025)
by: Yang, Yu, et al.
Published: (2025)
HOComp: Interaction-Aware Human-Object Composition
by: Liang, Dong, et al.
Published: (2025)
by: Liang, Dong, et al.
Published: (2025)
VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models
by: Zhang, Yabo, et al.
Published: (2024)
by: Zhang, Yabo, et al.
Published: (2024)
FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors
by: Zhang, Yabo, et al.
Published: (2025)
by: Zhang, Yabo, et al.
Published: (2025)
Explorable Ideas: Externalizing Ideas as Explorable Environments
by: Jung, Euijun, et al.
Published: (2025)
by: Jung, Euijun, et al.
Published: (2025)
RefSTAR: Blind Facial Image Restoration with Reference Selection, Transfer, and Reconstruction
by: Yin, Zhicun, et al.
Published: (2025)
by: Yin, Zhicun, et al.
Published: (2025)
ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Embedded Representation Learning Network for Animating Styled Video Portrait
by: Wang, Tianyong, et al.
Published: (2024)
by: Wang, Tianyong, et al.
Published: (2024)
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
by: Wu, Haoyu, et al.
Published: (2025)
by: Wu, Haoyu, et al.
Published: (2025)
PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
by: Zhang, Haoze, et al.
Published: (2025)
by: Zhang, Haoze, et al.
Published: (2025)
LuSh-NeRF: Lighting up and Sharpening NeRFs for Low-light Scenes
by: Qu, Zefan, et al.
Published: (2024)
by: Qu, Zefan, et al.
Published: (2024)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
by: Zhou, Yupeng, et al.
Published: (2024)
by: Zhou, Yupeng, et al.
Published: (2024)
Similar Items
-
Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy
by: Liu, Yuhao, et al.
Published: (2025) -
WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
by: Sun, Wenqiang, et al.
Published: (2025) -
WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories
by: Zhang, Yisu, et al.
Published: (2026) -
DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion Priors
by: Huang, Tianyu, et al.
Published: (2024) -
WorldCompass: Reinforcement Learning for Long-Horizon World Models
by: Wang, Zehan, et al.
Published: (2026)