MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nazarieh, Fatemeh, Feng, Zhenhua, Kanojia, Diptesh, Awais, Muhammad, Kittler, Josef
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915579395309568
author Nazarieh, Fatemeh
Feng, Zhenhua
Kanojia, Diptesh
Awais, Muhammad
Kittler, Josef
author_facet Nazarieh, Fatemeh
Feng, Zhenhua
Kanojia, Diptesh
Awais, Muhammad
Kittler, Josef
contents Audio-driven talking face generation has gained significant attention for applications in digital media and virtual avatars. While recent methods improve audio-lip synchronization, they often struggle with temporal consistency, identity preservation, and customization, especially in long video generation. To address these issues, we propose MAGIC-Talk, a one-shot diffusion-based framework for customizable and temporally stable talking face generation. MAGIC-Talk consists of ReferenceNet, which preserves identity and enables fine-grained facial editing via text prompts, and AnimateNet, which enhances motion coherence using structured motion priors. Unlike previous methods requiring multiple reference images or fine-tuning, MAGIC-Talk maintains identity from a single image while ensuring smooth transitions across frames. Additionally, a progressive latent fusion strategy is introduced to improve long-form video quality by reducing motion inconsistencies and flickering. Extensive experiments demonstrate that MAGIC-Talk outperforms state-of-the-art methods in visual quality, identity preservation, and synchronization accuracy, offering a robust solution for talking face generation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22810
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
Nazarieh, Fatemeh
Feng, Zhenhua
Kanojia, Diptesh
Awais, Muhammad
Kittler, Josef
Computer Vision and Pattern Recognition
Audio-driven talking face generation has gained significant attention for applications in digital media and virtual avatars. While recent methods improve audio-lip synchronization, they often struggle with temporal consistency, identity preservation, and customization, especially in long video generation. To address these issues, we propose MAGIC-Talk, a one-shot diffusion-based framework for customizable and temporally stable talking face generation. MAGIC-Talk consists of ReferenceNet, which preserves identity and enables fine-grained facial editing via text prompts, and AnimateNet, which enhances motion coherence using structured motion priors. Unlike previous methods requiring multiple reference images or fine-tuning, MAGIC-Talk maintains identity from a single image while ensuring smooth transitions across frames. Additionally, a progressive latent fusion strategy is introduced to improve long-form video quality by reducing motion inconsistencies and flickering. Extensive experiments demonstrate that MAGIC-Talk outperforms state-of-the-art methods in visual quality, identity preservation, and synchronization accuracy, offering a robust solution for talking face generation.
title MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.22810