Dual-Expert Consistency Model for Efficient and High-Quality Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lv, Zhengyao, Si, Chenyang, Pan, Tianlin, Chen, Zhaoxi, Wong, Kwan-Yee K., Qiao, Yu, Liu, Ziwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911093513781248
author Lv, Zhengyao
Si, Chenyang
Pan, Tianlin
Chen, Zhaoxi
Wong, Kwan-Yee K.
Qiao, Yu
Liu, Ziwei
author_facet Lv, Zhengyao
Si, Chenyang
Pan, Tianlin
Chen, Zhaoxi
Wong, Kwan-Yee K.
Qiao, Yu
Liu, Ziwei
contents Diffusion Models have achieved remarkable results in video synthesis but require iterative denoising steps, leading to substantial computational overhead. Consistency Models have made significant progress in accelerating diffusion models. However, directly applying them to video diffusion models often results in severe degradation of temporal consistency and appearance details. In this paper, by analyzing the training dynamics of Consistency Models, we identify a key conflicting learning dynamics during the distillation process: there is a significant discrepancy in the optimization gradients and loss contributions across different timesteps. This discrepancy prevents the distilled student model from achieving an optimal state, leading to compromised temporal consistency and degraded appearance details. To address this issue, we propose a parameter-efficient \textbf{Dual-Expert Consistency Model~(DCM)}, where a semantic expert focuses on learning semantic layout and motion, while a detail expert specializes in fine detail refinement. Furthermore, we introduce Temporal Coherence Loss to improve motion consistency for the semantic expert and apply GAN and Feature Matching Loss to enhance the synthesis quality of the detail expert.Our approach achieves state-of-the-art visual quality with significantly reduced sampling steps, demonstrating the effectiveness of expert specialization in video diffusion model distillation. Our code and models are available at \href{https://github.com/Vchitect/DCM}{https://github.com/Vchitect/DCM}.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03123
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
Lv, Zhengyao
Si, Chenyang
Pan, Tianlin
Chen, Zhaoxi
Wong, Kwan-Yee K.
Qiao, Yu
Liu, Ziwei
Computer Vision and Pattern Recognition
Diffusion Models have achieved remarkable results in video synthesis but require iterative denoising steps, leading to substantial computational overhead. Consistency Models have made significant progress in accelerating diffusion models. However, directly applying them to video diffusion models often results in severe degradation of temporal consistency and appearance details. In this paper, by analyzing the training dynamics of Consistency Models, we identify a key conflicting learning dynamics during the distillation process: there is a significant discrepancy in the optimization gradients and loss contributions across different timesteps. This discrepancy prevents the distilled student model from achieving an optimal state, leading to compromised temporal consistency and degraded appearance details. To address this issue, we propose a parameter-efficient \textbf{Dual-Expert Consistency Model~(DCM)}, where a semantic expert focuses on learning semantic layout and motion, while a detail expert specializes in fine detail refinement. Furthermore, we introduce Temporal Coherence Loss to improve motion consistency for the semantic expert and apply GAN and Feature Matching Loss to enhance the synthesis quality of the detail expert.Our approach achieves state-of-the-art visual quality with significantly reduced sampling steps, demonstrating the effectiveness of expert specialization in video diffusion model distillation. Our code and models are available at \href{https://github.com/Vchitect/DCM}{https://github.com/Vchitect/DCM}.
title Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.03123