Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhai, Yuanhao, Lin, Kevin, Yang, Zhengyuan, Li, Linjie, Wang, Jianfeng, Lin, Chung-Ching, Doermann, David, Yuan, Junsong, Wang, Lijuan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910669990789120
author Zhai, Yuanhao
Lin, Kevin
Yang, Zhengyuan
Li, Linjie
Wang, Jianfeng
Lin, Chung-Ching
Doermann, David
Yuan, Junsong
Wang, Lijuan
author_facet Zhai, Yuanhao
Lin, Kevin
Yang, Zhengyuan
Li, Linjie
Wang, Jianfeng
Lin, Chung-Ching
Doermann, David
Yuan, Junsong
Wang, Lijuan
contents Image diffusion distillation achieves high-fidelity generation with very few sampling steps. However, applying these techniques directly to video diffusion often results in unsatisfactory frame quality due to the limited visual quality in public video datasets. This affects the performance of both teacher and student video diffusion models. Our study aims to improve video diffusion distillation while improving frame appearance using abundant high-quality image data. We propose motion consistency model (MCM), a single-stage video diffusion distillation method that disentangles motion and appearance learning. Specifically, MCM includes a video consistency model that distills motion from the video teacher model, and an image discriminator that enhances frame appearance to match high-quality image data. This combination presents two challenges: (1) conflicting frame learning objectives, as video distillation learns from low-quality video frames while the image discriminator targets high-quality images; and (2) training-inference discrepancies due to the differing quality of video samples used during training and inference. To address these challenges, we introduce disentangled motion distillation and mixed trajectory distillation. The former applies the distillation objective solely to the motion representation, while the latter mitigates training-inference discrepancies by mixing distillation trajectories from both the low- and high-quality video domains. Extensive experiments show that our MCM achieves the state-of-the-art video diffusion distillation performance. Additionally, our method can enhance frame quality in video diffusion models, producing frames with high aesthetic scores or specific styles without corresponding video data.
format Preprint
id arxiv_https___arxiv_org_abs_2406_06890
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
Zhai, Yuanhao
Lin, Kevin
Yang, Zhengyuan
Li, Linjie
Wang, Jianfeng
Lin, Chung-Ching
Doermann, David
Yuan, Junsong
Wang, Lijuan
Computer Vision and Pattern Recognition
Image diffusion distillation achieves high-fidelity generation with very few sampling steps. However, applying these techniques directly to video diffusion often results in unsatisfactory frame quality due to the limited visual quality in public video datasets. This affects the performance of both teacher and student video diffusion models. Our study aims to improve video diffusion distillation while improving frame appearance using abundant high-quality image data. We propose motion consistency model (MCM), a single-stage video diffusion distillation method that disentangles motion and appearance learning. Specifically, MCM includes a video consistency model that distills motion from the video teacher model, and an image discriminator that enhances frame appearance to match high-quality image data. This combination presents two challenges: (1) conflicting frame learning objectives, as video distillation learns from low-quality video frames while the image discriminator targets high-quality images; and (2) training-inference discrepancies due to the differing quality of video samples used during training and inference. To address these challenges, we introduce disentangled motion distillation and mixed trajectory distillation. The former applies the distillation objective solely to the motion representation, while the latter mitigates training-inference discrepancies by mixing distillation trajectories from both the low- and high-quality video domains. Extensive experiments show that our MCM achieves the state-of-the-art video diffusion distillation performance. Additionally, our method can enhance frame quality in video diffusion models, producing frames with high aesthetic scores or specific styles without corresponding video data.
title Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.06890