Talking Head Generation via AU-Guided Landmark Prediction

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chang, Shao-Yu, Xu, Jingyi, Le, Hieu, Samaras, Dimitris
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916967032553472
author Chang, Shao-Yu
Xu, Jingyi
Le, Hieu
Samaras, Dimitris
author_facet Chang, Shao-Yu
Xu, Jingyi
Le, Hieu
Samaras, Dimitris
contents We propose a two-stage framework for audio-driven talking head generation with fine-grained expression control via facial Action Units (AUs). Unlike prior methods relying on emotion labels or implicit AU conditioning, our model explicitly maps AUs to 2D facial landmarks, enabling physically grounded, per-frame expression control. In the first stage, a variational motion generator predicts temporally coherent landmark sequences from audio and AU intensities. In the second stage, a diffusion-based synthesizer generates realistic, lip-synced videos conditioned on these landmarks and a reference image. This separation of motion and appearance improves expression accuracy, temporal stability, and visual realism. Experiments on the MEAD dataset show that our method outperforms state-of-the-art baselines across multiple metrics, demonstrating the effectiveness of explicit AU-to-landmark modeling for expressive talking head generation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19749
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Talking Head Generation via AU-Guided Landmark Prediction
Chang, Shao-Yu
Xu, Jingyi
Le, Hieu
Samaras, Dimitris
Computer Vision and Pattern Recognition
We propose a two-stage framework for audio-driven talking head generation with fine-grained expression control via facial Action Units (AUs). Unlike prior methods relying on emotion labels or implicit AU conditioning, our model explicitly maps AUs to 2D facial landmarks, enabling physically grounded, per-frame expression control. In the first stage, a variational motion generator predicts temporally coherent landmark sequences from audio and AU intensities. In the second stage, a diffusion-based synthesizer generates realistic, lip-synced videos conditioned on these landmarks and a reference image. This separation of motion and appearance improves expression accuracy, temporal stability, and visual realism. Experiments on the MEAD dataset show that our method outperforms state-of-the-art baselines across multiple metrics, demonstrating the effectiveness of explicit AU-to-landmark modeling for expressive talking head generation.
title Talking Head Generation via AU-Guided Landmark Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.19749