Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation
Fuente:
arXiv
Saved in:
| Main Authors: | Yaman, Dogucan, Eyiokur, Fevziye Irem, Bärmann, Leonard, Ekenel, Hazım Kemal, Waibel, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Audio-driven Talking Face Generation with Stabilized Synchronization Loss
by: Yaman, Dogucan, et al.
Published: (2023)
by: Yaman, Dogucan, et al.
Published: (2023)
Audio-Visual Speech Representation Expert for Enhanced Talking Face Video Generation and Evaluation
by: Yaman, Dogucan, et al.
Published: (2024)
by: Yaman, Dogucan, et al.
Published: (2024)
Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework
by: Yaman, Dogucan, et al.
Published: (2025)
by: Yaman, Dogucan, et al.
Published: (2025)
CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
A Multimodal Depth-Aware Method For Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis
by: Yaman, Dogucan, et al.
Published: (2025)
by: Yaman, Dogucan, et al.
Published: (2025)
Impact of Face Alignment on Face Image Quality
by: Onaran, Eren, et al.
Published: (2024)
by: Onaran, Eren, et al.
Published: (2024)
Analyzing the Effect of Combined Degradations on Face Recognition
by: Sarıtaş, Erdi, et al.
Published: (2024)
by: Sarıtaş, Erdi, et al.
Published: (2024)
Analyzing the Feature Extractor Networks for Face Image Synthesis
by: Sarıtaş, Erdi, et al.
Published: (2024)
by: Sarıtaş, Erdi, et al.
Published: (2024)
Assessing the Use of Face Swapping Methods as Face Anonymizers in Videos
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
Bias-Aware Face Mask Detection Dataset
by: Kantarcı, Alperen, et al.
Published: (2022)
by: Kantarcı, Alperen, et al.
Published: (2022)
Employing Vision-Language Models for Face Image Quality Assessment
by: Sarıtaş, Erdi, et al.
Published: (2026)
by: Sarıtaş, Erdi, et al.
Published: (2026)
Facial Attribute Based Text Guided Face Anonymization
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
Titanic Calling: Low Bandwidth Video Conference from the Titanic Wreck
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
Yolo-Key-6D: Single Stage Monocular 6D Pose Estimation with Keypoint Enhancements
by: Çetiner, Kemal Alperen, et al.
Published: (2026)
by: Çetiner, Kemal Alperen, et al.
Published: (2026)
Impact of Surface Reflections in Maritime Obstacle Detection
by: Yalçın, Samed, et al.
Published: (2024)
by: Yalçın, Samed, et al.
Published: (2024)
Improved MambdaBDA Framework for Robust Building Damage Assessment Across Disaster Domains
by: Gençoğlu, Alp Eren, et al.
Published: (2026)
by: Gençoğlu, Alp Eren, et al.
Published: (2026)
Attention-Enhanced Hybrid Feature Aggregation Network for 3D Brain Tumor Segmentation
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
In-Bed Pose Estimation: A Review
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
GLIMS: Attention-Guided Lightweight Multi-Scale Hybrid Network for Volumetric Semantic Segmentation
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
On Applicability of Synthetic Datasets for Facial Expression Recognition
by: Azmoudeh, Ali, et al.
Published: (2026)
by: Azmoudeh, Ali, et al.
Published: (2026)
MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
by: Nazarieh, Fatemeh, et al.
Published: (2025)
by: Nazarieh, Fatemeh, et al.
Published: (2025)
AVI-Talking: Learning Audio-Visual Instructions for Expressive 3D Talking Face Generation
by: Sun, Yasheng, et al.
Published: (2024)
by: Sun, Yasheng, et al.
Published: (2024)
IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion Transfer
by: Chen, Bo, et al.
Published: (2025)
by: Chen, Bo, et al.
Published: (2025)
RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network
by: Ji, Xiaozhong, et al.
Published: (2024)
by: Ji, Xiaozhong, et al.
Published: (2024)
Text-driven Talking Face Synthesis by Reprogramming Audio-driven Models
by: Choi, Jeongsoo, et al.
Published: (2023)
by: Choi, Jeongsoo, et al.
Published: (2023)
Learning to Forget -- Hierarchical Episodic Memory for Lifelong Robot Deployment
by: Bärmann, Leonard, et al.
Published: (2026)
by: Bärmann, Leonard, et al.
Published: (2026)
PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation
by: Nazarieh, Fatemeh, et al.
Published: (2024)
by: Nazarieh, Fatemeh, et al.
Published: (2024)
PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation
by: Wang, Baiqin, et al.
Published: (2025)
by: Wang, Baiqin, et al.
Published: (2025)
SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space
by: Zhang, Zeren, et al.
Published: (2024)
by: Zhang, Zeren, et al.
Published: (2024)
Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning
by: Xie, Yifan, et al.
Published: (2025)
by: Xie, Yifan, et al.
Published: (2025)
VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time
by: Xu, Sicheng, et al.
Published: (2024)
by: Xu, Sicheng, et al.
Published: (2024)
A Survey on Class-Agnostic Counting: Advancements from Reference-Based to Open-World Text-Guided Approaches
by: Ciampi, Luca, et al.
Published: (2025)
by: Ciampi, Luca, et al.
Published: (2025)
A Comparative Study of Perceptual Quality Metrics for Audio-driven Talking Head Videos
by: Zhang, Weixia, et al.
Published: (2024)
by: Zhang, Weixia, et al.
Published: (2024)
FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases
by: Tan, Shuai, et al.
Published: (2025)
by: Tan, Shuai, et al.
Published: (2025)
Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation
by: Chopin, Baptiste, et al.
Published: (2025)
by: Chopin, Baptiste, et al.
Published: (2025)
SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local Editing
by: Xiong, Lingyu, et al.
Published: (2024)
by: Xiong, Lingyu, et al.
Published: (2024)
Audio-Driven Talking Face Generation with Blink Embedding and Hash Grid Landmarks Encoding
by: Zhang, Yuhui, et al.
Published: (2026)
by: Zhang, Yuhui, et al.
Published: (2026)
JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation
by: Chakkera, Sai Tanmay Reddy, et al.
Published: (2024)
by: Chakkera, Sai Tanmay Reddy, et al.
Published: (2024)
PIER: A Novel Metric for Evaluating What Matters in Code-Switching
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Similar Items
-
Audio-driven Talking Face Generation with Stabilized Synchronization Loss
by: Yaman, Dogucan, et al.
Published: (2023) -
Audio-Visual Speech Representation Expert for Enhanced Talking Face Video Generation and Evaluation
by: Yaman, Dogucan, et al.
Published: (2024) -
Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework
by: Yaman, Dogucan, et al.
Published: (2025) -
CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025) -
A Multimodal Depth-Aware Method For Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)