GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Agarwal, Madhav, Zhang, Mingtian, Sevilla-Lara, Laura, McDonagh, Steven
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915669897904128
author Agarwal, Madhav
Zhang, Mingtian
Sevilla-Lara, Laura
McDonagh, Steven
author_facet Agarwal, Madhav
Zhang, Mingtian
Sevilla-Lara, Laura
McDonagh, Steven
contents Speech-driven talking heads have recently emerged and enable interactive avatars. However, real-world applications are limited, as current methods achieve high visual fidelity but slow or fast yet temporally unstable. Diffusion methods provide realistic image generation, yet struggle with oneshot settings. Gaussian Splatting approaches are real-time, yet inaccuracies in facial tracking, or inconsistent Gaussian mappings, lead to unstable outputs and video artifacts that are detrimental to realistic use cases. We address this problem by mapping Gaussian Splatting using 3D Morphable Models to generate person-specific avatars. We introduce transformer-based prediction of model parameters, directly from audio, to drive temporal consistency. From monocular video and independent audio speech inputs, our method enables generation of real-time talking head videos where we report competitive quantitative and qualitative performance.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10939
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting
Agarwal, Madhav
Zhang, Mingtian
Sevilla-Lara, Laura
McDonagh, Steven
Computer Vision and Pattern Recognition
Speech-driven talking heads have recently emerged and enable interactive avatars. However, real-world applications are limited, as current methods achieve high visual fidelity but slow or fast yet temporally unstable. Diffusion methods provide realistic image generation, yet struggle with oneshot settings. Gaussian Splatting approaches are real-time, yet inaccuracies in facial tracking, or inconsistent Gaussian mappings, lead to unstable outputs and video artifacts that are detrimental to realistic use cases. We address this problem by mapping Gaussian Splatting using 3D Morphable Models to generate person-specific avatars. We introduce transformer-based prediction of model parameters, directly from audio, to drive temporal consistency. From monocular video and independent audio speech inputs, our method enables generation of real-time talking head videos where we report competitive quantitative and qualitative performance.
title GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.10939