MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Seyeon, Jin, Siyoon, Park, Jihye, Kim, Kihong, Kim, Jiyoung, Nam, Jisu, Kim, Seungryong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910388619051008
author Kim, Seyeon
Jin, Siyoon
Park, Jihye
Kim, Kihong
Kim, Jiyoung
Nam, Jisu
Kim, Seungryong
author_facet Kim, Seyeon
Jin, Siyoon
Park, Jihye
Kim, Kihong
Kim, Jiyoung
Nam, Jisu
Kim, Seungryong
contents Conventional GAN-based models for talking head generation often suffer from limited quality and unstable training. Recent approaches based on diffusion models aimed to address these limitations and improve fidelity. However, they still face challenges, including extensive sampling times and difficulties in maintaining temporal consistency due to the high stochasticity of diffusion models. To overcome these challenges, we propose a novel motion-disentangled diffusion model for high-quality talking head generation, dubbed MoDiTalker. We introduce the two modules: audio-to-motion (AToM), designed to generate a synchronized lip motion from audio, and motion-to-video (MToV), designed to produce high-quality head video following the generated motion. AToM excels in capturing subtle lip movements by leveraging an audio attention mechanism. In addition, MToV enhances temporal consistency by leveraging an efficient tri-plane representation. Our experiments conducted on standard benchmarks demonstrate that our model achieves superior performance compared to existing models. We also provide comprehensive ablation studies and user study results.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19144
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation
Kim, Seyeon
Jin, Siyoon
Park, Jihye
Kim, Kihong
Kim, Jiyoung
Nam, Jisu
Kim, Seungryong
Computer Vision and Pattern Recognition
Conventional GAN-based models for talking head generation often suffer from limited quality and unstable training. Recent approaches based on diffusion models aimed to address these limitations and improve fidelity. However, they still face challenges, including extensive sampling times and difficulties in maintaining temporal consistency due to the high stochasticity of diffusion models. To overcome these challenges, we propose a novel motion-disentangled diffusion model for high-quality talking head generation, dubbed MoDiTalker. We introduce the two modules: audio-to-motion (AToM), designed to generate a synchronized lip motion from audio, and motion-to-video (MToV), designed to produce high-quality head video following the generated motion. AToM excels in capturing subtle lip movements by leveraging an audio attention mechanism. In addition, MToV enhances temporal consistency by leveraging an efficient tri-plane representation. Our experiments conducted on standard benchmarks demonstrate that our model achieves superior performance compared to existing models. We also provide comprehensive ablation studies and user study results.
title MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.19144