Taming Consistency Distillation for Accelerated Human Image Animation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Xiang, Zhang, Shiwei, Yuan, Hangjie, Wei, Yujie, Zhang, Yingya, Gao, Changxin, Wang, Yuehuan, Sang, Nong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909580153323520
author Wang, Xiang
Zhang, Shiwei
Yuan, Hangjie
Wei, Yujie
Zhang, Yingya
Gao, Changxin
Wang, Yuehuan
Sang, Nong
author_facet Wang, Xiang
Zhang, Shiwei
Yuan, Hangjie
Wei, Yujie
Zhang, Yingya
Gao, Changxin
Wang, Yuehuan
Sang, Nong
contents Recent advancements in human image animation have been propelled by video diffusion models, yet their reliance on numerous iterative denoising steps results in high inference costs and slow speeds. An intuitive solution involves adopting consistency models, which serve as an effective acceleration paradigm through consistency distillation. However, simply employing this strategy in human image animation often leads to quality decline, including visual blurring, motion degradation, and facial distortion, particularly in dynamic regions. In this paper, we propose the DanceLCM approach complemented by several enhancements to improve visual quality and motion continuity at low-step regime: (1) segmented consistency distillation with an auxiliary light-weight head to incorporate supervision from real video latents, mitigating cumulative errors resulting from single full-trajectory generation; (2) a motion-focused loss to centre on motion regions, and explicit injection of facial fidelity features to improve face authenticity. Extensive qualitative and quantitative experiments demonstrate that DanceLCM achieves results comparable to state-of-the-art video diffusion models with a mere 2-4 inference steps, significantly reducing the inference burden without compromising video quality. The code and models will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11143
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Taming Consistency Distillation for Accelerated Human Image Animation
Wang, Xiang
Zhang, Shiwei
Yuan, Hangjie
Wei, Yujie
Zhang, Yingya
Gao, Changxin
Wang, Yuehuan
Sang, Nong
Computer Vision and Pattern Recognition
Recent advancements in human image animation have been propelled by video diffusion models, yet their reliance on numerous iterative denoising steps results in high inference costs and slow speeds. An intuitive solution involves adopting consistency models, which serve as an effective acceleration paradigm through consistency distillation. However, simply employing this strategy in human image animation often leads to quality decline, including visual blurring, motion degradation, and facial distortion, particularly in dynamic regions. In this paper, we propose the DanceLCM approach complemented by several enhancements to improve visual quality and motion continuity at low-step regime: (1) segmented consistency distillation with an auxiliary light-weight head to incorporate supervision from real video latents, mitigating cumulative errors resulting from single full-trajectory generation; (2) a motion-focused loss to centre on motion regions, and explicit injection of facial fidelity features to improve face authenticity. Extensive qualitative and quantitative experiments demonstrate that DanceLCM achieves results comparable to state-of-the-art video diffusion models with a mere 2-4 inference steps, significantly reducing the inference burden without compromising video quality. The code and models will be made publicly available.
title Taming Consistency Distillation for Accelerated Human Image Animation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.11143