ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vo, Hoang-Son, Nguyen, Quang-Vinh, Kim, Seungwon, Yang, Hyung-Jeong, Yeom, Soonja, Kim, Soo-Hyung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911061240709120
author Vo, Hoang-Son
Nguyen, Quang-Vinh
Kim, Seungwon
Yang, Hyung-Jeong
Yeom, Soonja
Kim, Soo-Hyung
author_facet Vo, Hoang-Son
Nguyen, Quang-Vinh
Kim, Seungwon
Yang, Hyung-Jeong
Yeom, Soonja
Kim, Soo-Hyung
contents Audio-driven talking head generation requires precise synchronization between facial animations and audio signals. This paper introduces ATL-Diff, a novel approach addressing synchronization limitations while reducing noise and computational costs. Our framework features three key components: a Landmark Generation Module converting audio to facial landmarks, a Landmarks-Guide Noise approach that decouples audio by distributing noise according to landmarks, and a 3D Identity Diffusion network preserving identity characteristics. Experiments on MEAD and CREMA-D datasets demonstrate that ATL-Diff outperforms state-of-the-art methods across all metrics. Our approach achieves near real-time processing with high-quality animations, computational efficiency, and exceptional preservation of facial nuances. This advancement offers promising applications for virtual assistants, education, medical communication, and digital platforms. The source code is available at: \href{https://github.com/sonvth/ATL-Diff}{https://github.com/sonvth/ATL-Diff}
format Preprint
id arxiv_https___arxiv_org_abs_2507_12804
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion
Vo, Hoang-Son
Nguyen, Quang-Vinh
Kim, Seungwon
Yang, Hyung-Jeong
Yeom, Soonja
Kim, Soo-Hyung
Computer Vision and Pattern Recognition
Audio-driven talking head generation requires precise synchronization between facial animations and audio signals. This paper introduces ATL-Diff, a novel approach addressing synchronization limitations while reducing noise and computational costs. Our framework features three key components: a Landmark Generation Module converting audio to facial landmarks, a Landmarks-Guide Noise approach that decouples audio by distributing noise according to landmarks, and a 3D Identity Diffusion network preserving identity characteristics. Experiments on MEAD and CREMA-D datasets demonstrate that ATL-Diff outperforms state-of-the-art methods across all metrics. Our approach achieves near real-time processing with high-quality animations, computational efficiency, and exceptional preservation of facial nuances. This advancement offers promising applications for virtual assistants, education, medical communication, and digital platforms. The source code is available at: \href{https://github.com/sonvth/ATL-Diff}{https://github.com/sonvth/ATL-Diff}
title ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.12804