Image-to-Video Diffusion: From Foundations to Open Frontiers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xianlong, Pan, Wenbo, Zhou, Shijia, Li, Ke, Wang, Yuqi, Ye, Zeyu, Zhang, Hangtao, Zhang, Leo Yu, Jia, Xiaohua
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911691763089408
author Wang, Xianlong
Pan, Wenbo
Zhou, Shijia
Li, Ke
Wang, Yuqi
Ye, Zeyu
Zhang, Hangtao
Zhang, Leo Yu
Jia, Xiaohua
author_facet Wang, Xianlong
Pan, Wenbo
Zhou, Shijia
Li, Ke
Wang, Yuqi
Ye, Zeyu
Zhang, Hangtao
Zhang, Leo Yu
Jia, Xiaohua
contents Diffusion-based \textit{image-to-video} (I2V) generation has become a central direction in generative models by turning a reference image, with optional conditions, into a temporally coherent video. Compared with broader video generation settings, this task places stricter demands on content consistency, identity preservation, and motion coherence. Although the literature grows rapidly, existing works mostly discuss I2V generation within broader topics and still lack a dedicated taxonomy together with a systematic analysis centered on this field. This work addresses that gap by treating diffusion I2V generation as a standalone subject. It first reviews the task formulation, model architectures, datasets, and evaluation metrics, and then organizes existing methods through a taxonomy based on architecture and training paradigm. It further distills four core designs, namely condition encoding, temporal modeling, noise prior design, and spatial-temporal upsampling, and discusses representative application scenarios together with major open challenges.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17248
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Image-to-Video Diffusion: From Foundations to Open Frontiers
Wang, Xianlong
Pan, Wenbo
Zhou, Shijia
Li, Ke
Wang, Yuqi
Ye, Zeyu
Zhang, Hangtao
Zhang, Leo Yu
Jia, Xiaohua
Computer Vision and Pattern Recognition
Diffusion-based \textit{image-to-video} (I2V) generation has become a central direction in generative models by turning a reference image, with optional conditions, into a temporally coherent video. Compared with broader video generation settings, this task places stricter demands on content consistency, identity preservation, and motion coherence. Although the literature grows rapidly, existing works mostly discuss I2V generation within broader topics and still lack a dedicated taxonomy together with a systematic analysis centered on this field. This work addresses that gap by treating diffusion I2V generation as a standalone subject. It first reviews the task formulation, model architectures, datasets, and evaluation metrics, and then organizes existing methods through a taxonomy based on architecture and training paradigm. It further distills four core designs, namely condition encoding, temporal modeling, noise prior design, and spatial-temporal upsampling, and discusses representative application scenarios together with major open challenges.
title Image-to-Video Diffusion: From Foundations to Open Frontiers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.17248