MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915579395309568 |
|---|---|
| author | Nazarieh, Fatemeh Feng, Zhenhua Kanojia, Diptesh Awais, Muhammad Kittler, Josef |
| author_facet | Nazarieh, Fatemeh Feng, Zhenhua Kanojia, Diptesh Awais, Muhammad Kittler, Josef |
| contents | Audio-driven talking face generation has gained significant attention for applications in digital media and virtual avatars. While recent methods improve audio-lip synchronization, they often struggle with temporal consistency, identity preservation, and customization, especially in long video generation. To address these issues, we propose MAGIC-Talk, a one-shot diffusion-based framework for customizable and temporally stable talking face generation. MAGIC-Talk consists of ReferenceNet, which preserves identity and enables fine-grained facial editing via text prompts, and AnimateNet, which enhances motion coherence using structured motion priors. Unlike previous methods requiring multiple reference images or fine-tuning, MAGIC-Talk maintains identity from a single image while ensuring smooth transitions across frames. Additionally, a progressive latent fusion strategy is introduced to improve long-form video quality by reducing motion inconsistencies and flickering. Extensive experiments demonstrate that MAGIC-Talk outperforms state-of-the-art methods in visual quality, identity preservation, and synchronization accuracy, offering a robust solution for talking face generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_22810 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control Nazarieh, Fatemeh Feng, Zhenhua Kanojia, Diptesh Awais, Muhammad Kittler, Josef Computer Vision and Pattern Recognition Audio-driven talking face generation has gained significant attention for applications in digital media and virtual avatars. While recent methods improve audio-lip synchronization, they often struggle with temporal consistency, identity preservation, and customization, especially in long video generation. To address these issues, we propose MAGIC-Talk, a one-shot diffusion-based framework for customizable and temporally stable talking face generation. MAGIC-Talk consists of ReferenceNet, which preserves identity and enables fine-grained facial editing via text prompts, and AnimateNet, which enhances motion coherence using structured motion priors. Unlike previous methods requiring multiple reference images or fine-tuning, MAGIC-Talk maintains identity from a single image while ensuring smooth transitions across frames. Additionally, a progressive latent fusion strategy is introduced to improve long-form video quality by reducing motion inconsistencies and flickering. Extensive experiments demonstrate that MAGIC-Talk outperforms state-of-the-art methods in visual quality, identity preservation, and synchronization accuracy, offering a robust solution for talking face generation. |
| title | MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2510.22810 |