SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Zhuguanyu, Gong, Ruihao, Yong, Yang, Huang, Yushi, Fan, Xiangyu, Yang, Lei, Lin, Dahua, Liu, Xianglong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910270329192448
author Wu, Zhuguanyu
Gong, Ruihao
Yong, Yang
Huang, Yushi
Fan, Xiangyu
Yang, Lei
Lin, Dahua
Liu, Xianglong
author_facet Wu, Zhuguanyu
Gong, Ruihao
Yong, Yang
Huang, Yushi
Fan, Xiangyu
Yang, Lei
Lin, Dahua
Liu, Xianglong
contents Distribution Matching Distillation (DMD) is a widely used paradigm for accelerating inference in few-step video diffusion models. However, DMD-style video distillation faces two coupled challenges: the fake score must track a continuously evolving generator, making training costly when frequent updates are required, while reverse-KL-style matching can be mode-seeking and conservative for preserving strong motion dynamics. To address these issues, we propose \textbf{Score Gradient Matching Distillation (SGMD)}. SGMD adopts a fake-score perspective by directly optimizing the fake score toward the teacher, while using teacher stop-gradient Fisher as a stable distribution-matching objective. We provide a gradient analysis that motivates this objective choice under ideal tracking. Building on this, SGMD introduces a pair of dual potentials: negative-residual (NR) for outer-loop correction and residual-contraction (RC) for inner-loop tracking. Empirically, compared to DMD2, SGMD achieves an approximately $\sim 3\times$ training speedup and substantially improves motion dynamics for 4-step distilled models while preserving temporal consistency. A human study confirms that SGMD is preferred in motion quality and overall preference, while visual quality and text alignment remain comparable. Code is available at https://github.com/ModelTC/LightX2V.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30116
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation
Wu, Zhuguanyu
Gong, Ruihao
Yong, Yang
Huang, Yushi
Fan, Xiangyu
Yang, Lei
Lin, Dahua
Liu, Xianglong
Computer Vision and Pattern Recognition
Machine Learning
Distribution Matching Distillation (DMD) is a widely used paradigm for accelerating inference in few-step video diffusion models. However, DMD-style video distillation faces two coupled challenges: the fake score must track a continuously evolving generator, making training costly when frequent updates are required, while reverse-KL-style matching can be mode-seeking and conservative for preserving strong motion dynamics. To address these issues, we propose \textbf{Score Gradient Matching Distillation (SGMD)}. SGMD adopts a fake-score perspective by directly optimizing the fake score toward the teacher, while using teacher stop-gradient Fisher as a stable distribution-matching objective. We provide a gradient analysis that motivates this objective choice under ideal tracking. Building on this, SGMD introduces a pair of dual potentials: negative-residual (NR) for outer-loop correction and residual-contraction (RC) for inner-loop tracking. Empirically, compared to DMD2, SGMD achieves an approximately $\sim 3\times$ training speedup and substantially improves motion dynamics for 4-step distilled models while preserving temporal consistency. A human study confirms that SGMD is preferred in motion quality and overall preference, while visual quality and text alignment remain comparable. Code is available at https://github.com/ModelTC/LightX2V.
title SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2605.30116