AVID: Adapting Video Diffusion Models to World Models
Fuente:
arXiv
Saved in:
| Main Authors: | Rigter, Marc, Gupta, Tarun, Hilmkil, Agrin, Ma, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AVID: Any-Length Video Inpainting with Diffusion Model
by: Zhang, Zhixing, et al.
Published: (2023)
by: Zhang, Zhixing, et al.
Published: (2023)
Vid2World: Crafting Video Diffusion Models to Interactive World Models
by: Huang, Siqiao, et al.
Published: (2025)
by: Huang, Siqiao, et al.
Published: (2025)
Adapting a World Model for Trajectory Following in a 3D Game
by: Tot, Marko, et al.
Published: (2025)
by: Tot, Marko, et al.
Published: (2025)
The Essential Role of Causality in Foundation World Models for Embodied AI
by: Gupta, Tarun, et al.
Published: (2024)
by: Gupta, Tarun, et al.
Published: (2024)
Adapting Vision-Language Models for Evaluating World Models
by: Hendriksen, Mariya, et al.
Published: (2025)
by: Hendriksen, Mariya, et al.
Published: (2025)
Progressive Autoregressive Video Diffusion Models
by: Xie, Desai, et al.
Published: (2024)
by: Xie, Desai, et al.
Published: (2024)
Video Diffusion Models: A Survey
by: Melnik, Andrew, et al.
Published: (2024)
by: Melnik, Andrew, et al.
Published: (2024)
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction
by: Xing, Zhen, et al.
Published: (2024)
by: Xing, Zhen, et al.
Published: (2024)
Measuring Style Similarity in Diffusion Models
by: Somepalli, Gowthami, et al.
Published: (2024)
by: Somepalli, Gowthami, et al.
Published: (2024)
PAVE: Patching and Adapting Video Large Language Models
by: Liu, Zhuoming, et al.
Published: (2025)
by: Liu, Zhuoming, et al.
Published: (2025)
WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling
by: Oshima, Yuta, et al.
Published: (2025)
by: Oshima, Yuta, et al.
Published: (2025)
Semantically Consistent Video Inpainting with Conditional Diffusion Models
by: Green, Dylan, et al.
Published: (2024)
by: Green, Dylan, et al.
Published: (2024)
Unlearning Concepts from Text-to-Video Diffusion Models
by: Liu, Shiqi, et al.
Published: (2024)
by: Liu, Shiqi, et al.
Published: (2024)
Lifelong Learning of Video Diffusion Models From a Single Video Stream
by: Yoo, Jason, et al.
Published: (2024)
by: Yoo, Jason, et al.
Published: (2024)
The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation
by: Yin, Aoxiong, et al.
Published: (2025)
by: Yin, Aoxiong, et al.
Published: (2025)
Corruption-Aware Training of Latent Video Diffusion Models for Robust Text-to-Video Generation
by: Maduabuchi, Chika, et al.
Published: (2025)
by: Maduabuchi, Chika, et al.
Published: (2025)
Redefining Temporal Modeling in Video Diffusion: The Vectorized Timestep Approach
by: Liu, Yaofang, et al.
Published: (2024)
by: Liu, Yaofang, et al.
Published: (2024)
In-Context Symmetries: Self-Supervised Learning through Contextual World Models
by: Gupta, Sharut, et al.
Published: (2024)
by: Gupta, Sharut, et al.
Published: (2024)
iVideoGPT: Interactive VideoGPTs are Scalable World Models
by: Wu, Jialong, et al.
Published: (2024)
by: Wu, Jialong, et al.
Published: (2024)
Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models
by: Lu, Jiacheng, et al.
Published: (2026)
by: Lu, Jiacheng, et al.
Published: (2026)
LD-ViCE: Latent Diffusion Model for Video Counterfactual Explanations
by: Varshney, Payal, et al.
Published: (2025)
by: Varshney, Payal, et al.
Published: (2025)
BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models
by: Po, Ryan, et al.
Published: (2025)
by: Po, Ryan, et al.
Published: (2025)
PathoTune: Adapting Visual Foundation Model to Pathological Specialists
by: Lu, Jiaxuan, et al.
Published: (2024)
by: Lu, Jiaxuan, et al.
Published: (2024)
SelfAdapt: Unsupervised Domain Adaptation of Cell Segmentation Models
by: Reith, Fabian H., et al.
Published: (2025)
by: Reith, Fabian H., et al.
Published: (2025)
DP-RDM: Adapting Diffusion Models to Private Domains Without Fine-Tuning
by: Lebensold, Jonathan, et al.
Published: (2024)
by: Lebensold, Jonathan, et al.
Published: (2024)
Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states
by: Gupta, Samarth, et al.
Published: (2025)
by: Gupta, Samarth, et al.
Published: (2025)
DiWA: Diffusion Policy Adaptation with World Models
by: Chandra, Akshay L, et al.
Published: (2025)
by: Chandra, Akshay L, et al.
Published: (2025)
Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition
by: Yu, Sihyun, et al.
Published: (2024)
by: Yu, Sihyun, et al.
Published: (2024)
Diffusion Model-based Activity Completion for AI Motion Capture from Videos
by: Huayu, Gao, et al.
Published: (2025)
by: Huayu, Gao, et al.
Published: (2025)
Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models
by: Fan, Weichen, et al.
Published: (2025)
by: Fan, Weichen, et al.
Published: (2025)
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
by: Wang, Yuduo, et al.
Published: (2023)
by: Wang, Yuduo, et al.
Published: (2023)
Olaf-World: Orienting Latent Actions for Video World Modeling
by: Jiang, Yuxin, et al.
Published: (2026)
by: Jiang, Yuxin, et al.
Published: (2026)
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
by: Lin, Han, et al.
Published: (2024)
by: Lin, Han, et al.
Published: (2024)
Promptable Foundation Models for SAR Remote Sensing: Adapting the Segment Anything Model for Snow Avalanche Segmentation
by: Gelato, Riccardo, et al.
Published: (2026)
by: Gelato, Riccardo, et al.
Published: (2026)
Adapting by Analogy: OOD Generalization of Visuomotor Policies via Functional Correspondence
by: Gupta, Pranay, et al.
Published: (2025)
by: Gupta, Pranay, et al.
Published: (2025)
VMDT: Decoding the Trustworthiness of Video Foundation Models
by: Potter, Yujin, et al.
Published: (2025)
by: Potter, Yujin, et al.
Published: (2025)
Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Zero-Shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model
by: Cao, Cong, et al.
Published: (2024)
by: Cao, Cong, et al.
Published: (2024)
Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
by: Bao, Fan, et al.
Published: (2024)
by: Bao, Fan, et al.
Published: (2024)
A Survey on Video Diffusion Models
by: Xing, Zhen, et al.
Published: (2023)
by: Xing, Zhen, et al.
Published: (2023)
Similar Items
-
AVID: Any-Length Video Inpainting with Diffusion Model
by: Zhang, Zhixing, et al.
Published: (2023) -
Vid2World: Crafting Video Diffusion Models to Interactive World Models
by: Huang, Siqiao, et al.
Published: (2025) -
Adapting a World Model for Trajectory Following in a 3D Game
by: Tot, Marko, et al.
Published: (2025) -
The Essential Role of Causality in Foundation World Models for Embodied AI
by: Gupta, Tarun, et al.
Published: (2024) -
Adapting Vision-Language Models for Evaluating World Models
by: Hendriksen, Mariya, et al.
Published: (2025)