Vid2World: Crafting Video Diffusion Models to Interactive World Models
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Siqiao, Wu, Jialong, Zhou, Qixing, Miao, Shangchen, Long, Mingsheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
iVideoGPT: Interactive VideoGPTs are Scalable World Models
by: Wu, Jialong, et al.
Published: (2024)
by: Wu, Jialong, et al.
Published: (2024)
JEPA-VLA: Video Predictive Embedding is Needed for VLA Models
by: Miao, Shangchen, et al.
Published: (2026)
by: Miao, Shangchen, et al.
Published: (2026)
Nano World Models: A Minimalist Implementation of Future Video Prediction
by: Huang, Siqiao, et al.
Published: (2026)
by: Huang, Siqiao, et al.
Published: (2026)
AVID: Adapting Video Diffusion Models to World Models
by: Rigter, Marc, et al.
Published: (2024)
by: Rigter, Marc, et al.
Published: (2024)
Diffusion Tuning: Transferring Diffusion Models via Chain of Forgetting
by: Zhong, Jincheng, et al.
Published: (2024)
by: Zhong, Jincheng, et al.
Published: (2024)
Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models
by: Lu, Jiacheng, et al.
Published: (2026)
by: Lu, Jiacheng, et al.
Published: (2026)
Dynamical Diffusion: Learning Temporal Dynamics with Diffusion Models
by: Guo, Xingzhuo, et al.
Published: (2025)
by: Guo, Xingzhuo, et al.
Published: (2025)
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models
by: Gu, Bohai, et al.
Published: (2026)
by: Gu, Bohai, et al.
Published: (2026)
WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling
by: Oshima, Yuta, et al.
Published: (2025)
by: Oshima, Yuta, et al.
Published: (2025)
WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
by: Zhou, Runjie, et al.
Published: (2026)
by: Zhou, Runjie, et al.
Published: (2026)
The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation
by: Yin, Aoxiong, et al.
Published: (2025)
by: Yin, Aoxiong, et al.
Published: (2025)
DiWA: Diffusion Policy Adaptation with World Models
by: Chandra, Akshay L, et al.
Published: (2025)
by: Chandra, Akshay L, et al.
Published: (2025)
Foresight Diffusion: Improving Sampling Consistency in Predictive Diffusion Models
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Sora Detector: A Unified Hallucination Detection for Large Text-to-Video Models
by: Chu, Zhixuan, et al.
Published: (2024)
by: Chu, Zhixuan, et al.
Published: (2024)
Olaf-World: Orienting Latent Actions for Video World Modeling
by: Jiang, Yuxin, et al.
Published: (2026)
by: Jiang, Yuxin, et al.
Published: (2026)
Domain Guidance: A Simple Transfer Approach for a Pre-trained Diffusion Model
by: Zhong, Jincheng, et al.
Published: (2025)
by: Zhong, Jincheng, et al.
Published: (2025)
SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications
by: Hasson, Yana, et al.
Published: (2025)
by: Hasson, Yana, et al.
Published: (2025)
VidLA: Video-Language Alignment at Scale
by: Rizve, Mamshad Nayeem, et al.
Published: (2024)
by: Rizve, Mamshad Nayeem, et al.
Published: (2024)
ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration
by: Ni, Chaojun, et al.
Published: (2024)
by: Ni, Chaojun, et al.
Published: (2024)
Owl-1: Omni World Model for Consistent Long Video Generation
by: Huang, Yuanhui, et al.
Published: (2024)
by: Huang, Yuanhui, et al.
Published: (2024)
An Empirical Study of World Model Quantization
by: Fu, Zhongqian, et al.
Published: (2026)
by: Fu, Zhongqian, et al.
Published: (2026)
PAN: A World Model for General, Interactable, and Long-Horizon World Simulation
by: PAN Team, et al.
Published: (2025)
by: PAN Team, et al.
Published: (2025)
Astra: General Interactive World Model with Autoregressive Denoising
by: Zhu, Yixuan, et al.
Published: (2025)
by: Zhu, Yixuan, et al.
Published: (2025)
GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction
by: Zuo, Sicheng, et al.
Published: (2024)
by: Zuo, Sicheng, et al.
Published: (2024)
Cross-View World Models
by: Sharma, Rishabh, et al.
Published: (2026)
by: Sharma, Rishabh, et al.
Published: (2026)
Slot Structured World Models
by: Collu, Jonathan, et al.
Published: (2024)
by: Collu, Jonathan, et al.
Published: (2024)
Progressive Autoregressive Video Diffusion Models
by: Xie, Desai, et al.
Published: (2024)
by: Xie, Desai, et al.
Published: (2024)
WorldEval: World Model as Real-World Robot Policies Evaluator
by: Li, Yaxuan, et al.
Published: (2025)
by: Li, Yaxuan, et al.
Published: (2025)
WorldCache: Content-Aware Caching for Accelerated Video World Models
by: Nawaz, Umair, et al.
Published: (2026)
by: Nawaz, Umair, et al.
Published: (2026)
Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models
by: Fan, Weichen, et al.
Published: (2025)
by: Fan, Weichen, et al.
Published: (2025)
VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
by: Tian, Zeyue, et al.
Published: (2024)
by: Tian, Zeyue, et al.
Published: (2024)
Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model
by: Tang, Junshu, et al.
Published: (2025)
by: Tang, Junshu, et al.
Published: (2025)
Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Diffusion for World Modeling: Visual Details Matter in Atari
by: Alonso, Eloi, et al.
Published: (2024)
by: Alonso, Eloi, et al.
Published: (2024)
VidTwin: Video VAE with Decoupled Structure and Dynamics
by: Wang, Yuchi, et al.
Published: (2024)
by: Wang, Yuchi, et al.
Published: (2024)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
by: Zheng, Sixiao, et al.
Published: (2025)
by: Zheng, Sixiao, et al.
Published: (2025)
TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos
by: Lee, Seungjae, et al.
Published: (2025)
by: Lee, Seungjae, et al.
Published: (2025)
LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
by: Duan, Yinglin, et al.
Published: (2025)
by: Duan, Yinglin, et al.
Published: (2025)
Deterministic World Models for Verification of Closed-loop Vision-based Systems
by: Geng, Yuang, et al.
Published: (2025)
by: Geng, Yuang, et al.
Published: (2025)
VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL
by: Park, Kyoungjun, et al.
Published: (2025)
by: Park, Kyoungjun, et al.
Published: (2025)
Similar Items
-
iVideoGPT: Interactive VideoGPTs are Scalable World Models
by: Wu, Jialong, et al.
Published: (2024) -
JEPA-VLA: Video Predictive Embedding is Needed for VLA Models
by: Miao, Shangchen, et al.
Published: (2026) -
Nano World Models: A Minimalist Implementation of Future Video Prediction
by: Huang, Siqiao, et al.
Published: (2026) -
AVID: Adapting Video Diffusion Models to World Models
by: Rigter, Marc, et al.
Published: (2024) -
Diffusion Tuning: Transferring Diffusion Models via Chain of Forgetting
by: Zhong, Jincheng, et al.
Published: (2024)