ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911061240709120 |
|---|---|
| author | Vo, Hoang-Son Nguyen, Quang-Vinh Kim, Seungwon Yang, Hyung-Jeong Yeom, Soonja Kim, Soo-Hyung |
| author_facet | Vo, Hoang-Son Nguyen, Quang-Vinh Kim, Seungwon Yang, Hyung-Jeong Yeom, Soonja Kim, Soo-Hyung |
| contents | Audio-driven talking head generation requires precise synchronization between facial animations and audio signals. This paper introduces ATL-Diff, a novel approach addressing synchronization limitations while reducing noise and computational costs. Our framework features three key components: a Landmark Generation Module converting audio to facial landmarks, a Landmarks-Guide Noise approach that decouples audio by distributing noise according to landmarks, and a 3D Identity Diffusion network preserving identity characteristics. Experiments on MEAD and CREMA-D datasets demonstrate that ATL-Diff outperforms state-of-the-art methods across all metrics. Our approach achieves near real-time processing with high-quality animations, computational efficiency, and exceptional preservation of facial nuances. This advancement offers promising applications for virtual assistants, education, medical communication, and digital platforms. The source code is available at: \href{https://github.com/sonvth/ATL-Diff}{https://github.com/sonvth/ATL-Diff} |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_12804 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion Vo, Hoang-Son Nguyen, Quang-Vinh Kim, Seungwon Yang, Hyung-Jeong Yeom, Soonja Kim, Soo-Hyung Computer Vision and Pattern Recognition Audio-driven talking head generation requires precise synchronization between facial animations and audio signals. This paper introduces ATL-Diff, a novel approach addressing synchronization limitations while reducing noise and computational costs. Our framework features three key components: a Landmark Generation Module converting audio to facial landmarks, a Landmarks-Guide Noise approach that decouples audio by distributing noise according to landmarks, and a 3D Identity Diffusion network preserving identity characteristics. Experiments on MEAD and CREMA-D datasets demonstrate that ATL-Diff outperforms state-of-the-art methods across all metrics. Our approach achieves near real-time processing with high-quality animations, computational efficiency, and exceptional preservation of facial nuances. This advancement offers promising applications for virtual assistants, education, medical communication, and digital platforms. The source code is available at: \href{https://github.com/sonvth/ATL-Diff}{https://github.com/sonvth/ATL-Diff} |
| title | ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2507.12804 |