AI killed the video star. Audio-driven diffusion model for expressive talking head generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918221334970368 |
|---|---|
| author | Chopin, Baptiste Dhamija, Tashvik Balaji, Pranav Wang, Yaohui Dantcheva, Antitza |
| author_facet | Chopin, Baptiste Dhamija, Tashvik Balaji, Pranav Wang, Yaohui Dantcheva, Antitza |
| contents | We propose Dimitra++, a novel framework for audio-driven talking head generation, streamlined to learn lip motion, facial expression, as well as head pose motion. Specifically, we propose a conditional Motion Diffusion Transformer (cMDT) to model facial motion sequences, employing a 3D representation. The cMDT is conditioned on two inputs: a reference facial image, which determines appearance, as well as an audio sequence, which drives the motion. Quantitative and qualitative experiments, as well as a user study on two widely employed datasets, i.e., VoxCeleb2 and CelebV-HQ, suggest that Dimitra++ is able to outperform existing approaches in generating realistic talking heads imparting lip motion, facial expression, and head pose. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_22488 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | AI killed the video star. Audio-driven diffusion model for expressive talking head generation Chopin, Baptiste Dhamija, Tashvik Balaji, Pranav Wang, Yaohui Dantcheva, Antitza Computer Vision and Pattern Recognition We propose Dimitra++, a novel framework for audio-driven talking head generation, streamlined to learn lip motion, facial expression, as well as head pose motion. Specifically, we propose a conditional Motion Diffusion Transformer (cMDT) to model facial motion sequences, employing a 3D representation. The cMDT is conditioned on two inputs: a reference facial image, which determines appearance, as well as an audio sequence, which drives the motion. Quantitative and qualitative experiments, as well as a user study on two widely employed datasets, i.e., VoxCeleb2 and CelebV-HQ, suggest that Dimitra++ is able to outperform existing approaches in generating realistic talking heads imparting lip motion, facial expression, and head pose. |
| title | AI killed the video star. Audio-driven diffusion model for expressive talking head generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.22488 |