FashionFlow: Leveraging Diffusion Models for Dynamic Fashion Video Synthesis from Static Imagery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Islam, Tasin, Miron, Alina, Liu, XiaoHui, Li, Yongmin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910304024133632
author Islam, Tasin
Miron, Alina
Liu, XiaoHui
Li, Yongmin
author_facet Islam, Tasin
Miron, Alina
Liu, XiaoHui
Li, Yongmin
contents Our study introduces a new image-to-video generator called FashionFlow to generate fashion videos. By utilising a diffusion model, we are able to create short videos from still fashion images. Our approach involves developing and connecting relevant components with the diffusion model, which results in the creation of high-fidelity videos that are aligned with the conditional image. The components include the use of pseudo-3D convolutional layers to generate videos efficiently. VAE and CLIP encoders capture vital characteristics from still images to condition the diffusion model at a global level. Our research demonstrates a successful synthesis of fashion videos featuring models posing from various angles, showcasing the fit and appearance of the garment. Our findings hold great promise for improving and enhancing the shopping experience for the online fashion industry.
format Preprint
id arxiv_https___arxiv_org_abs_2310_00106
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle FashionFlow: Leveraging Diffusion Models for Dynamic Fashion Video Synthesis from Static Imagery
Islam, Tasin
Miron, Alina
Liu, XiaoHui
Li, Yongmin
Computer Vision and Pattern Recognition
Artificial Intelligence
Our study introduces a new image-to-video generator called FashionFlow to generate fashion videos. By utilising a diffusion model, we are able to create short videos from still fashion images. Our approach involves developing and connecting relevant components with the diffusion model, which results in the creation of high-fidelity videos that are aligned with the conditional image. The components include the use of pseudo-3D convolutional layers to generate videos efficiently. VAE and CLIP encoders capture vital characteristics from still images to condition the diffusion model at a global level. Our research demonstrates a successful synthesis of fashion videos featuring models posing from various angles, showcasing the fit and appearance of the garment. Our findings hold great promise for improving and enhancing the shopping experience for the online fashion industry.
title FashionFlow: Leveraging Diffusion Models for Dynamic Fashion Video Synthesis from Static Imagery
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2310.00106