Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ge, Xingtong, Zhang, Yi, Huang, Yushi, He, Dailan, Wang, Xiahong, Ma, Bingqi, Song, Guanglu, Liu, Yu, Zhang, Jun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913002109796352
author Ge, Xingtong
Zhang, Yi
Huang, Yushi
He, Dailan
Wang, Xiahong
Ma, Bingqi
Song, Guanglu
Liu, Yu
Zhang, Jun
author_facet Ge, Xingtong
Zhang, Yi
Huang, Yushi
He, Dailan
Wang, Xiahong
Ma, Bingqi
Song, Guanglu
Liu, Yu
Zhang, Jun
contents Distilling video generation models to extremely low inference budgets (e.g., 2--4 NFEs) is crucial for real-time deployment, yet remains challenging. Trajectory-style consistency distillation often becomes conservative under complex video dynamics, yielding an over-smoothed appearance and weak motion. Distribution matching distillation (DMD) can recover sharp, mode-seeking samples, but its local training signals do not explicitly regularize how denoising updates compose across timesteps, making composed rollouts prone to drift. To overcome this challenge, we propose Self-Consistent Distribution Matching Distillation (SC-DMD), which explicitly regularizes the endpoint-consistent composition of consecutive denoising updates. For real-time autoregressive video generation, we further treat the KV cache as a quality parameterized condition and propose Cache-Distribution-Aware training. This training scheme applies SC-DMD over multi-step rollouts and introduces a cache-conditioned feature alignment objective that steers low-quality outputs toward high-quality references. Across extensive experiments on both non-autoregressive backbones (e.g., Wan~2.1) and autoregressive real-time paradigms (e.g., Self Forcing), our method, dubbed \textbf{Salt}, consistently improves low-NFE video generation quality while remaining compatible with diverse KV-cache memory mechanisms. Source code will be released at \href{https://github.com/XingtongGe/Salt}{https://github.com/XingtongGe/Salt}.
format Preprint
id arxiv_https___arxiv_org_abs_2604_03118
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation
Ge, Xingtong
Zhang, Yi
Huang, Yushi
He, Dailan
Wang, Xiahong
Ma, Bingqi
Song, Guanglu
Liu, Yu
Zhang, Jun
Computer Vision and Pattern Recognition
Image and Video Processing
Distilling video generation models to extremely low inference budgets (e.g., 2--4 NFEs) is crucial for real-time deployment, yet remains challenging. Trajectory-style consistency distillation often becomes conservative under complex video dynamics, yielding an over-smoothed appearance and weak motion. Distribution matching distillation (DMD) can recover sharp, mode-seeking samples, but its local training signals do not explicitly regularize how denoising updates compose across timesteps, making composed rollouts prone to drift. To overcome this challenge, we propose Self-Consistent Distribution Matching Distillation (SC-DMD), which explicitly regularizes the endpoint-consistent composition of consecutive denoising updates. For real-time autoregressive video generation, we further treat the KV cache as a quality parameterized condition and propose Cache-Distribution-Aware training. This training scheme applies SC-DMD over multi-step rollouts and introduces a cache-conditioned feature alignment objective that steers low-quality outputs toward high-quality references. Across extensive experiments on both non-autoregressive backbones (e.g., Wan~2.1) and autoregressive real-time paradigms (e.g., Self Forcing), our method, dubbed \textbf{Salt}, consistently improves low-NFE video generation quality while remaining compatible with diverse KV-cache memory mechanisms. Source code will be released at \href{https://github.com/XingtongGe/Salt}{https://github.com/XingtongGe/Salt}.
title Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2604.03118