Unlock Pose Diversity: Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking Portrait

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Chaolong, Yao, Kai, Yan, Yuyao, Jiang, Chenru, Zhao, Weiguang, Sun, Jie, Cheng, Guangliang, Zhang, Yifei, Dong, Bin, Huang, Kaizhu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916654400667648
author Yang, Chaolong
Yao, Kai
Yan, Yuyao
Jiang, Chenru
Zhao, Weiguang
Sun, Jie
Cheng, Guangliang
Zhang, Yifei
Dong, Bin
Huang, Kaizhu
author_facet Yang, Chaolong
Yao, Kai
Yan, Yuyao
Jiang, Chenru
Zhao, Weiguang
Sun, Jie
Cheng, Guangliang
Zhang, Yifei
Dong, Bin
Huang, Kaizhu
contents Audio-driven single-image talking portrait generation plays a crucial role in virtual reality, digital human creation, and filmmaking. Existing approaches are generally categorized into keypoint-based and image-based methods. Keypoint-based methods effectively preserve character identity but struggle to capture fine facial details due to the fixed points limitation of the 3D Morphable Model. Moreover, traditional generative networks face challenges in establishing causality between audio and keypoints on limited datasets, resulting in low pose diversity. In contrast, image-based approaches produce high-quality portraits with diverse details using the diffusion network but incur identity distortion and expensive computational costs. In this work, we propose KDTalker, the first framework to combine unsupervised implicit 3D keypoint with a spatiotemporal diffusion model. Leveraging unsupervised implicit 3D keypoints, KDTalker adapts facial information densities, allowing the diffusion process to model diverse head poses and capture fine facial details flexibly. The custom-designed spatiotemporal attention mechanism ensures accurate lip synchronization, producing temporally consistent, high-quality animations while enhancing computational efficiency. Experimental results demonstrate that KDTalker achieves state-of-the-art performance regarding lip synchronization accuracy, head pose diversity, and execution efficiency.Our codes are available at https://github.com/chaolongy/KDTalker.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12963
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unlock Pose Diversity: Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking Portrait
Yang, Chaolong
Yao, Kai
Yan, Yuyao
Jiang, Chenru
Zhao, Weiguang
Sun, Jie
Cheng, Guangliang
Zhang, Yifei
Dong, Bin
Huang, Kaizhu
Computer Vision and Pattern Recognition
Audio-driven single-image talking portrait generation plays a crucial role in virtual reality, digital human creation, and filmmaking. Existing approaches are generally categorized into keypoint-based and image-based methods. Keypoint-based methods effectively preserve character identity but struggle to capture fine facial details due to the fixed points limitation of the 3D Morphable Model. Moreover, traditional generative networks face challenges in establishing causality between audio and keypoints on limited datasets, resulting in low pose diversity. In contrast, image-based approaches produce high-quality portraits with diverse details using the diffusion network but incur identity distortion and expensive computational costs. In this work, we propose KDTalker, the first framework to combine unsupervised implicit 3D keypoint with a spatiotemporal diffusion model. Leveraging unsupervised implicit 3D keypoints, KDTalker adapts facial information densities, allowing the diffusion process to model diverse head poses and capture fine facial details flexibly. The custom-designed spatiotemporal attention mechanism ensures accurate lip synchronization, producing temporally consistent, high-quality animations while enhancing computational efficiency. Experimental results demonstrate that KDTalker achieves state-of-the-art performance regarding lip synchronization accuracy, head pose diversity, and execution efficiency.Our codes are available at https://github.com/chaolongy/KDTalker.
title Unlock Pose Diversity: Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking Portrait
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.12963