DreamWorld: Unified World Modeling in Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Boming, Zhang, Xiangdong, Liao, Ning, Zhang, Yuqing, Zhang, Shaofeng, Yang, Xue, Fan, Qi, Zhang, Yanyong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ConsisDrive: Identity-Preserving Driving World Models for Video Generation by Instance Mask
by: Yang, Zhuoran, et al.
Published: (2026)
by: Yang, Zhuoran, et al.
Published: (2026)
SWIFT: Prompt-Adaptive Memory for Efficient Interactive Long Video Generation
by: Tan, Shanwen, et al.
Published: (2026)
by: Tan, Shanwen, et al.
Published: (2026)
VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
by: Zhang, Xiangdong, et al.
Published: (2025)
by: Zhang, Xiangdong, et al.
Published: (2025)
InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation
by: Yang, Zhuoran, et al.
Published: (2026)
by: Yang, Zhuoran, et al.
Published: (2026)
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
by: Zhang, Wenyao, et al.
Published: (2025)
by: Zhang, Wenyao, et al.
Published: (2025)
PCP-MAE: Learning to Predict Centers for Point Masked Autoencoders
by: Zhang, Xiangdong, et al.
Published: (2024)
by: Zhang, Xiangdong, et al.
Published: (2024)
Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
by: Zhang, Xiangdong, et al.
Published: (2025)
by: Zhang, Xiangdong, et al.
Published: (2025)
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
RISE-Video: Can Video Generators Decode Implicit World Rules?
by: Liu, Mingxin, et al.
Published: (2026)
by: Liu, Mingxin, et al.
Published: (2026)
FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction
by: Dai, Yixiang, et al.
Published: (2025)
by: Dai, Yixiang, et al.
Published: (2025)
LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
by: Duan, Zicheng, et al.
Published: (2026)
by: Duan, Zicheng, et al.
Published: (2026)
WorldMark: A Unified Benchmark Suite for Interactive Video World Models
by: Xu, Xiaojie, et al.
Published: (2026)
by: Xu, Xiaojie, et al.
Published: (2026)
EVA: An Embodied World Model for Future Video Anticipation
by: Chi, Xiaowei, et al.
Published: (2024)
by: Chi, Xiaowei, et al.
Published: (2024)
Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
by: Zhang, Shaofeng, et al.
Published: (2025)
by: Zhang, Shaofeng, et al.
Published: (2025)
Dream to Drive with Predictive Individual World Model
by: Gao, Yinfeng, et al.
Published: (2025)
by: Gao, Yinfeng, et al.
Published: (2025)
DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
by: Hu, Xiaotao, et al.
Published: (2024)
by: Hu, Xiaotao, et al.
Published: (2024)
UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
by: Ni, Shuo, et al.
Published: (2025)
by: Ni, Shuo, et al.
Published: (2025)
World Models as Group Actions
by: Wang, Zijie, et al.
Published: (2026)
by: Wang, Zijie, et al.
Published: (2026)
Dual-Branch Center-Surrounding Contrast: Rethinking Contrastive Learning for 3D Point Clouds
by: Zhang, Shaofeng, et al.
Published: (2025)
by: Zhang, Shaofeng, et al.
Published: (2025)
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
by: Zhou, Xin, et al.
Published: (2026)
by: Zhou, Xin, et al.
Published: (2026)
WorldSimBench: Towards Video Generation Models as World Simulators
by: Qin, Yiran, et al.
Published: (2024)
by: Qin, Yiran, et al.
Published: (2024)
Pathwise Test-Time Correction for Autoregressive Long Video Generation
by: Xiang, Xunzhi, et al.
Published: (2026)
by: Xiang, Xunzhi, et al.
Published: (2026)
PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language Models
by: Su, Yuanhao, et al.
Published: (2026)
by: Su, Yuanhao, et al.
Published: (2026)
UVLM: Benchmarking Video Language Model for Underwater World Understanding
by: Xue, Xizhe, et al.
Published: (2025)
by: Xue, Xizhe, et al.
Published: (2025)
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
DreamOmni: Unified Image Generation and Editing
by: Xia, Bin, et al.
Published: (2024)
by: Xia, Bin, et al.
Published: (2024)
Pre-Trained Video Generative Models as World Simulators
by: He, Haoran, et al.
Published: (2025)
by: He, Haoran, et al.
Published: (2025)
World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
by: Wang, Weijie, et al.
Published: (2026)
by: Wang, Weijie, et al.
Published: (2026)
DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
by: Guo, Xu, et al.
Published: (2026)
by: Guo, Xu, et al.
Published: (2026)
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
by: Xie, Wulin, et al.
Published: (2025)
by: Xie, Wulin, et al.
Published: (2025)
DreamVE: Unified Instruction-based Image and Video Editing
by: Xia, Bin, et al.
Published: (2025)
by: Xia, Bin, et al.
Published: (2025)
NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
by: Yang, Yuxue, et al.
Published: (2026)
by: Yang, Yuxue, et al.
Published: (2026)
DreamPoster: A Unified Framework for Image-Conditioned Generative Poster Design
by: Hu, Xiwei, et al.
Published: (2025)
by: Hu, Xiwei, et al.
Published: (2025)
WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning
by: Zhang, Yuanhan, et al.
Published: (2024)
by: Zhang, Yuanhan, et al.
Published: (2024)
Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow
by: Dharmarajan, Karthik, et al.
Published: (2025)
by: Dharmarajan, Karthik, et al.
Published: (2025)
UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
by: Huang, Jiehui, et al.
Published: (2025)
by: Huang, Jiehui, et al.
Published: (2025)
Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation
by: Gui, Xingtai, et al.
Published: (2026)
by: Gui, Xingtai, et al.
Published: (2026)
Motion Control for Enhanced Complex Action Video Generation
by: Zhou, Qiang, et al.
Published: (2024)
by: Zhou, Qiang, et al.
Published: (2024)
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
DreamRelation: Relation-Centric Video Customization
by: Wei, Yujie, et al.
Published: (2025)
by: Wei, Yujie, et al.
Published: (2025)
Similar Items
-
ConsisDrive: Identity-Preserving Driving World Models for Video Generation by Instance Mask
by: Yang, Zhuoran, et al.
Published: (2026) -
SWIFT: Prompt-Adaptive Memory for Efficient Interactive Long Video Generation
by: Tan, Shanwen, et al.
Published: (2026) -
VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
by: Zhang, Xiangdong, et al.
Published: (2025) -
InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation
by: Yang, Zhuoran, et al.
Published: (2026) -
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
by: Zhang, Wenyao, et al.
Published: (2025)