GPD: Guided Progressive Distillation for Fast and High-Quality Video Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liang, Xiao, Zhang, Yunzhu, Zhu, Linchao
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917241485787136
author Liang, Xiao
Zhang, Yunzhu
Zhu, Linchao
author_facet Liang, Xiao
Zhang, Yunzhu
Zhu, Linchao
contents Diffusion models have achieved remarkable success in video generation; however, the high computational cost of the denoising process remains a major bottleneck. Existing approaches have shown promise in reducing the number of diffusion steps, but they often suffer from significant quality degradation when applied to video generation. We propose Guided Progressive Distillation (GPD), a framework that accelerates the diffusion process for fast and high-quality video generation. GPD introduces a novel training strategy in which a teacher model progressively guides a student model to operate with larger step sizes. The framework consists of two key components: (1) an online-generated training target that reduces optimization difficulty while improving computational efficiency, and (2) frequency-domain constraints in the latent space that promote the preservation of fine-grained details and temporal dynamics. Applied to the Wan2.1 model, GPD reduces the number of sampling steps from 48 to 6 while maintaining competitive visual quality on VBench. Compared with existing distillation methods, GPD demonstrates clear advantages in both pipeline simplicity and quality preservation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01814
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GPD: Guided Progressive Distillation for Fast and High-Quality Video Generation
Liang, Xiao
Zhang, Yunzhu
Zhu, Linchao
Computer Vision and Pattern Recognition
Diffusion models have achieved remarkable success in video generation; however, the high computational cost of the denoising process remains a major bottleneck. Existing approaches have shown promise in reducing the number of diffusion steps, but they often suffer from significant quality degradation when applied to video generation. We propose Guided Progressive Distillation (GPD), a framework that accelerates the diffusion process for fast and high-quality video generation. GPD introduces a novel training strategy in which a teacher model progressively guides a student model to operate with larger step sizes. The framework consists of two key components: (1) an online-generated training target that reduces optimization difficulty while improving computational efficiency, and (2) frequency-domain constraints in the latent space that promote the preservation of fine-grained details and temporal dynamics. Applied to the Wan2.1 model, GPD reduces the number of sampling steps from 48 to 6 while maintaining competitive visual quality on VBench. Compared with existing distillation methods, GPD demonstrates clear advantages in both pipeline simplicity and quality preservation.
title GPD: Guided Progressive Distillation for Fast and High-Quality Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.01814