FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion model

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Mu, Lingzhou, Liu, Baiji, Zhang, Ruonan, Mo, Guiming, Jin, Jiawei, Zhang, Kai, Huang, Haozhi
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916703355535360
author Mu, Lingzhou
Liu, Baiji
Zhang, Ruonan
Mo, Guiming
Jin, Jiawei
Zhang, Kai
Huang, Haozhi
author_facet Mu, Lingzhou
Liu, Baiji
Zhang, Ruonan
Mo, Guiming
Jin, Jiawei
Zhang, Kai
Huang, Haozhi
contents Diffusion-based video generation techniques have significantly improved zero-shot talking-head avatar generation, enhancing the naturalness of both head motion and facial expressions. However, existing methods suffer from poor controllability, making them less applicable to real-world scenarios such as filmmaking and live streaming for e-commerce. To address this limitation, we propose FLAP, a novel approach that integrates explicit 3D intermediate parameters (head poses and facial expressions) into the diffusion model for end-to-end generation of realistic portrait videos. The proposed architecture allows the model to generate vivid portrait videos from audio while simultaneously incorporating additional control signals, such as head rotation angles and eye-blinking frequency. Furthermore, the decoupling of head pose and facial expression allows for independent control of each, offering precise manipulation of both the avatar's pose and facial expressions. We also demonstrate its flexibility in integrating with existing 3D head generation methods, bridging the gap between 3D model-based approaches and end-to-end diffusion techniques. Extensive experiments show that our method outperforms recent audio-driven portrait video models in both naturalness and controllability.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19455
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion model
Mu, Lingzhou
Liu, Baiji
Zhang, Ruonan
Mo, Guiming
Jin, Jiawei
Zhang, Kai
Huang, Haozhi
Graphics
Diffusion-based video generation techniques have significantly improved zero-shot talking-head avatar generation, enhancing the naturalness of both head motion and facial expressions. However, existing methods suffer from poor controllability, making them less applicable to real-world scenarios such as filmmaking and live streaming for e-commerce. To address this limitation, we propose FLAP, a novel approach that integrates explicit 3D intermediate parameters (head poses and facial expressions) into the diffusion model for end-to-end generation of realistic portrait videos. The proposed architecture allows the model to generate vivid portrait videos from audio while simultaneously incorporating additional control signals, such as head rotation angles and eye-blinking frequency. Furthermore, the decoupling of head pose and facial expression allows for independent control of each, offering precise manipulation of both the avatar's pose and facial expressions. We also demonstrate its flexibility in integrating with existing 3D head generation methods, bridging the gap between 3D model-based approaches and end-to-end diffusion techniques. Extensive experiments show that our method outperforms recent audio-driven portrait video models in both naturalness and controllability.
title FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion model
topic Graphics
url https://arxiv.org/abs/2502.19455