Audio-driven Talking Face Generation with Stabilized Synchronization Loss
Fuente:
arXiv
Salvato in:
| Autori principali: | Yaman, Dogucan, Eyiokur, Fevziye Irem, Bärmann, Leonard, Ekenel, Hazim Kemal, Waibel, Alexander |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation
di: Yaman, Dogucan, et al.
Pubblicazione: (2025)
di: Yaman, Dogucan, et al.
Pubblicazione: (2025)
Audio-Visual Speech Representation Expert for Enhanced Talking Face Video Generation and Evaluation
di: Yaman, Dogucan, et al.
Pubblicazione: (2024)
di: Yaman, Dogucan, et al.
Pubblicazione: (2024)
Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework
di: Yaman, Dogucan, et al.
Pubblicazione: (2025)
di: Yaman, Dogucan, et al.
Pubblicazione: (2025)
CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
di: Eyiokur, Fevziye Irem, et al.
Pubblicazione: (2025)
di: Eyiokur, Fevziye Irem, et al.
Pubblicazione: (2025)
A Multimodal Depth-Aware Method For Embodied Reference Understanding
di: Eyiokur, Fevziye Irem, et al.
Pubblicazione: (2025)
di: Eyiokur, Fevziye Irem, et al.
Pubblicazione: (2025)
Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis
di: Yaman, Dogucan, et al.
Pubblicazione: (2025)
di: Yaman, Dogucan, et al.
Pubblicazione: (2025)
Analyzing the Effect of Combined Degradations on Face Recognition
di: Sarıtaş, Erdi, et al.
Pubblicazione: (2024)
di: Sarıtaş, Erdi, et al.
Pubblicazione: (2024)
Analyzing the Feature Extractor Networks for Face Image Synthesis
di: Sarıtaş, Erdi, et al.
Pubblicazione: (2024)
di: Sarıtaş, Erdi, et al.
Pubblicazione: (2024)
Assessing the Use of Face Swapping Methods as Face Anonymizers in Videos
di: Muştu, Mustafa İzzet, et al.
Pubblicazione: (2025)
di: Muştu, Mustafa İzzet, et al.
Pubblicazione: (2025)
Impact of Face Alignment on Face Image Quality
di: Onaran, Eren, et al.
Pubblicazione: (2024)
di: Onaran, Eren, et al.
Pubblicazione: (2024)
Employing Vision-Language Models for Face Image Quality Assessment
di: Sarıtaş, Erdi, et al.
Pubblicazione: (2026)
di: Sarıtaş, Erdi, et al.
Pubblicazione: (2026)
Improved MambdaBDA Framework for Robust Building Damage Assessment Across Disaster Domains
di: Gençoğlu, Alp Eren, et al.
Pubblicazione: (2026)
di: Gençoğlu, Alp Eren, et al.
Pubblicazione: (2026)
Impact of Surface Reflections in Maritime Obstacle Detection
di: Yalçın, Samed, et al.
Pubblicazione: (2024)
di: Yalçın, Samed, et al.
Pubblicazione: (2024)
Facial Attribute Based Text Guided Face Anonymization
di: Muştu, Mustafa İzzet, et al.
Pubblicazione: (2025)
di: Muştu, Mustafa İzzet, et al.
Pubblicazione: (2025)
Yolo-Key-6D: Single Stage Monocular 6D Pose Estimation with Keypoint Enhancements
di: Çetiner, Kemal Alperen, et al.
Pubblicazione: (2026)
di: Çetiner, Kemal Alperen, et al.
Pubblicazione: (2026)
Bias-Aware Face Mask Detection Dataset
di: Kantarcı, Alperen, et al.
Pubblicazione: (2022)
di: Kantarcı, Alperen, et al.
Pubblicazione: (2022)
In-Bed Pose Estimation: A Review
di: Yazıcı, Ziya Ata, et al.
Pubblicazione: (2024)
di: Yazıcı, Ziya Ata, et al.
Pubblicazione: (2024)
On Applicability of Synthetic Datasets for Facial Expression Recognition
di: Azmoudeh, Ali, et al.
Pubblicazione: (2026)
di: Azmoudeh, Ali, et al.
Pubblicazione: (2026)
GLIMS: Attention-Guided Lightweight Multi-Scale Hybrid Network for Volumetric Semantic Segmentation
di: Yazıcı, Ziya Ata, et al.
Pubblicazione: (2024)
di: Yazıcı, Ziya Ata, et al.
Pubblicazione: (2024)
Attention-Enhanced Hybrid Feature Aggregation Network for 3D Brain Tumor Segmentation
di: Yazıcı, Ziya Ata, et al.
Pubblicazione: (2024)
di: Yazıcı, Ziya Ata, et al.
Pubblicazione: (2024)
IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion Transfer
di: Chen, Bo, et al.
Pubblicazione: (2025)
di: Chen, Bo, et al.
Pubblicazione: (2025)
MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
di: Nazarieh, Fatemeh, et al.
Pubblicazione: (2025)
di: Nazarieh, Fatemeh, et al.
Pubblicazione: (2025)
AVI-Talking: Learning Audio-Visual Instructions for Expressive 3D Talking Face Generation
di: Sun, Yasheng, et al.
Pubblicazione: (2024)
di: Sun, Yasheng, et al.
Pubblicazione: (2024)
Titanic Calling: Low Bandwidth Video Conference from the Titanic Wreck
di: Eyiokur, Fevziye Irem, et al.
Pubblicazione: (2024)
di: Eyiokur, Fevziye Irem, et al.
Pubblicazione: (2024)
VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time
di: Xu, Sicheng, et al.
Pubblicazione: (2024)
di: Xu, Sicheng, et al.
Pubblicazione: (2024)
A Survey on Class-Agnostic Counting: Advancements from Reference-Based to Open-World Text-Guided Approaches
di: Ciampi, Luca, et al.
Pubblicazione: (2025)
di: Ciampi, Luca, et al.
Pubblicazione: (2025)
Text-driven Talking Face Synthesis by Reprogramming Audio-driven Models
di: Choi, Jeongsoo, et al.
Pubblicazione: (2023)
di: Choi, Jeongsoo, et al.
Pubblicazione: (2023)
PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation
di: Wang, Baiqin, et al.
Pubblicazione: (2025)
di: Wang, Baiqin, et al.
Pubblicazione: (2025)
Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation
di: Chopin, Baptiste, et al.
Pubblicazione: (2025)
di: Chopin, Baptiste, et al.
Pubblicazione: (2025)
Audio-Driven Talking Face Generation with Blink Embedding and Hash Grid Landmarks Encoding
di: Zhang, Yuhui, et al.
Pubblicazione: (2026)
di: Zhang, Yuhui, et al.
Pubblicazione: (2026)
JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation
di: Chakkera, Sai Tanmay Reddy, et al.
Pubblicazione: (2024)
di: Chakkera, Sai Tanmay Reddy, et al.
Pubblicazione: (2024)
RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network
di: Ji, Xiaozhong, et al.
Pubblicazione: (2024)
di: Ji, Xiaozhong, et al.
Pubblicazione: (2024)
SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space
di: Zhang, Zeren, et al.
Pubblicazione: (2024)
di: Zhang, Zeren, et al.
Pubblicazione: (2024)
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
di: Yang, Shaoshu, et al.
Pubblicazione: (2025)
di: Yang, Shaoshu, et al.
Pubblicazione: (2025)
Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning
di: Xie, Yifan, et al.
Pubblicazione: (2025)
di: Xie, Yifan, et al.
Pubblicazione: (2025)
GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian Splatting
di: Chen, Bo, et al.
Pubblicazione: (2024)
di: Chen, Bo, et al.
Pubblicazione: (2024)
PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation
di: Nazarieh, Fatemeh, et al.
Pubblicazione: (2024)
di: Nazarieh, Fatemeh, et al.
Pubblicazione: (2024)
SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
di: Yee, Phyo Thet, et al.
Pubblicazione: (2025)
di: Yee, Phyo Thet, et al.
Pubblicazione: (2025)
FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled Audio
di: Xu, Chao, et al.
Pubblicazione: (2024)
di: Xu, Chao, et al.
Pubblicazione: (2024)
SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis
di: Peng, Ziqiao, et al.
Pubblicazione: (2023)
di: Peng, Ziqiao, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation
di: Yaman, Dogucan, et al.
Pubblicazione: (2025) -
Audio-Visual Speech Representation Expert for Enhanced Talking Face Video Generation and Evaluation
di: Yaman, Dogucan, et al.
Pubblicazione: (2024) -
Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework
di: Yaman, Dogucan, et al.
Pubblicazione: (2025) -
CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
di: Eyiokur, Fevziye Irem, et al.
Pubblicazione: (2025) -
A Multimodal Depth-Aware Method For Embodied Reference Understanding
di: Eyiokur, Fevziye Irem, et al.
Pubblicazione: (2025)