AI killed the video star. Audio-driven diffusion model for expressive talking head generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chopin, Baptiste, Dhamija, Tashvik, Balaji, Pranav, Wang, Yaohui, Dantcheva, Antitza
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918221334970368
author Chopin, Baptiste
Dhamija, Tashvik
Balaji, Pranav
Wang, Yaohui
Dantcheva, Antitza
author_facet Chopin, Baptiste
Dhamija, Tashvik
Balaji, Pranav
Wang, Yaohui
Dantcheva, Antitza
contents We propose Dimitra++, a novel framework for audio-driven talking head generation, streamlined to learn lip motion, facial expression, as well as head pose motion. Specifically, we propose a conditional Motion Diffusion Transformer (cMDT) to model facial motion sequences, employing a 3D representation. The cMDT is conditioned on two inputs: a reference facial image, which determines appearance, as well as an audio sequence, which drives the motion. Quantitative and qualitative experiments, as well as a user study on two widely employed datasets, i.e., VoxCeleb2 and CelebV-HQ, suggest that Dimitra++ is able to outperform existing approaches in generating realistic talking heads imparting lip motion, facial expression, and head pose.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22488
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AI killed the video star. Audio-driven diffusion model for expressive talking head generation
Chopin, Baptiste
Dhamija, Tashvik
Balaji, Pranav
Wang, Yaohui
Dantcheva, Antitza
Computer Vision and Pattern Recognition
We propose Dimitra++, a novel framework for audio-driven talking head generation, streamlined to learn lip motion, facial expression, as well as head pose motion. Specifically, we propose a conditional Motion Diffusion Transformer (cMDT) to model facial motion sequences, employing a 3D representation. The cMDT is conditioned on two inputs: a reference facial image, which determines appearance, as well as an audio sequence, which drives the motion. Quantitative and qualitative experiments, as well as a user study on two widely employed datasets, i.e., VoxCeleb2 and CelebV-HQ, suggest that Dimitra++ is able to outperform existing approaches in generating realistic talking heads imparting lip motion, facial expression, and head pose.
title AI killed the video star. Audio-driven diffusion model for expressive talking head generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.22488