FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhong, Tianyun, Liang, Chao, Jiang, Jianwen, Lin, Gaojie, Yang, Jiaqi, Zhao, Zhou
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915227220574208
author Zhong, Tianyun
Liang, Chao
Jiang, Jianwen
Lin, Gaojie
Yang, Jiaqi
Zhao, Zhou
author_facet Zhong, Tianyun
Liang, Chao
Jiang, Jianwen
Lin, Gaojie
Yang, Jiaqi
Zhao, Zhou
contents Diffusion-based audio-driven talking avatar methods have recently gained attention for their high-fidelity, vivid, and expressive results. However, their slow inference speed limits practical applications. Despite the development of various distillation techniques for diffusion models, we found that naive diffusion distillation methods do not yield satisfactory results. Distilled models exhibit reduced robustness with open-set input images and a decreased correlation between audio and video compared to teacher models, undermining the advantages of diffusion models. To address this, we propose FADA (Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation). We first designed a mixed-supervised loss to leverage data of varying quality and enhance the overall model capability as well as robustness. Additionally, we propose a multi-CFG distillation with learnable tokens to utilize the correlation between audio and reference image conditions, reducing the threefold inference runs caused by multi-CFG with acceptable quality degradation. Extensive experiments across multiple datasets show that FADA generates vivid videos comparable to recent diffusion model-based methods while achieving an NFE speedup of 4.17-12.5 times. Demos are available at our webpage http://fadavatar.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16915
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation
Zhong, Tianyun
Liang, Chao
Jiang, Jianwen
Lin, Gaojie
Yang, Jiaqi
Zhao, Zhou
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Sound
Audio and Speech Processing
Diffusion-based audio-driven talking avatar methods have recently gained attention for their high-fidelity, vivid, and expressive results. However, their slow inference speed limits practical applications. Despite the development of various distillation techniques for diffusion models, we found that naive diffusion distillation methods do not yield satisfactory results. Distilled models exhibit reduced robustness with open-set input images and a decreased correlation between audio and video compared to teacher models, undermining the advantages of diffusion models. To address this, we propose FADA (Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation). We first designed a mixed-supervised loss to leverage data of varying quality and enhance the overall model capability as well as robustness. Additionally, we propose a multi-CFG distillation with learnable tokens to utilize the correlation between audio and reference image conditions, reducing the threefold inference runs caused by multi-CFG with acceptable quality degradation. Extensive experiments across multiple datasets show that FADA generates vivid videos comparable to recent diffusion model-based methods while achieving an NFE speedup of 4.17-12.5 times. Demos are available at our webpage http://fadavatar.github.io.
title FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.16915