Audio-Visual Speech Representation Expert for Enhanced Talking Face Video Generation and Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Yaman, Dogucan, Eyiokur, Fevziye Irem, Bärmann, Leonard, Aktı, Seymanur, Ekenel, Hazım Kemal, Waibel, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation
by: Yaman, Dogucan, et al.
Published: (2025)
by: Yaman, Dogucan, et al.
Published: (2025)
Audio-driven Talking Face Generation with Stabilized Synchronization Loss
by: Yaman, Dogucan, et al.
Published: (2023)
by: Yaman, Dogucan, et al.
Published: (2023)
Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework
by: Yaman, Dogucan, et al.
Published: (2025)
by: Yaman, Dogucan, et al.
Published: (2025)
Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis
by: Yaman, Dogucan, et al.
Published: (2025)
by: Yaman, Dogucan, et al.
Published: (2025)
CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
A Multimodal Depth-Aware Method For Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
by: Akti, Seymanur, et al.
Published: (2026)
by: Akti, Seymanur, et al.
Published: (2026)
Titanic Calling: Low Bandwidth Video Conference from the Titanic Wreck
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
Assessing the Use of Face Swapping Methods as Face Anonymizers in Videos
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
by: Akti, Seymanur, et al.
Published: (2025)
by: Akti, Seymanur, et al.
Published: (2025)
Analyzing the Effect of Combined Degradations on Face Recognition
by: Sarıtaş, Erdi, et al.
Published: (2024)
by: Sarıtaş, Erdi, et al.
Published: (2024)
Analyzing the Feature Extractor Networks for Face Image Synthesis
by: Sarıtaş, Erdi, et al.
Published: (2024)
by: Sarıtaş, Erdi, et al.
Published: (2024)
Impact of Face Alignment on Face Image Quality
by: Onaran, Eren, et al.
Published: (2024)
by: Onaran, Eren, et al.
Published: (2024)
Facial Attribute Based Text Guided Face Anonymization
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS
by: Nguyen, Tuan Nam, et al.
Published: (2024)
by: Nguyen, Tuan Nam, et al.
Published: (2024)
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025
by: Koneru, Sai, et al.
Published: (2025)
by: Koneru, Sai, et al.
Published: (2025)
Yolo-Key-6D: Single Stage Monocular 6D Pose Estimation with Keypoint Enhancements
by: Çetiner, Kemal Alperen, et al.
Published: (2026)
by: Çetiner, Kemal Alperen, et al.
Published: (2026)
Impact of Surface Reflections in Maritime Obstacle Detection
by: Yalçın, Samed, et al.
Published: (2024)
by: Yalçın, Samed, et al.
Published: (2024)
Bias-Aware Face Mask Detection Dataset
by: Kantarcı, Alperen, et al.
Published: (2022)
by: Kantarcı, Alperen, et al.
Published: (2022)
Employing Vision-Language Models for Face Image Quality Assessment
by: Sarıtaş, Erdi, et al.
Published: (2026)
by: Sarıtaş, Erdi, et al.
Published: (2026)
Improved MambdaBDA Framework for Robust Building Damage Assessment Across Disaster Domains
by: Gençoğlu, Alp Eren, et al.
Published: (2026)
by: Gençoğlu, Alp Eren, et al.
Published: (2026)
Attention-Enhanced Hybrid Feature Aggregation Network for 3D Brain Tumor Segmentation
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
GLIMS: Attention-Guided Lightweight Multi-Scale Hybrid Network for Volumetric Semantic Segmentation
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
In-Bed Pose Estimation: A Review
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
by: Yazıcı, Ziya Ata, et al.
Published: (2024)
Learning to Forget -- Hierarchical Episodic Memory for Lifelong Robot Deployment
by: Bärmann, Leonard, et al.
Published: (2026)
by: Bärmann, Leonard, et al.
Published: (2026)
PIER: A Novel Metric for Evaluating What Matters in Code-Switching
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
On Applicability of Synthetic Datasets for Facial Expression Recognition
by: Azmoudeh, Ali, et al.
Published: (2026)
by: Azmoudeh, Ali, et al.
Published: (2026)
Cocktail-Party Audio-Visual Speech Recognition
by: Nguyen, Thai-Binh, et al.
Published: (2025)
by: Nguyen, Thai-Binh, et al.
Published: (2025)
BOOM: Beyond Only One Modality KIT's Multimodal Multilingual Lecture Companion
by: Koneru, Sai, et al.
Published: (2025)
by: Koneru, Sai, et al.
Published: (2025)
Episodic Memory Verbalization using Hierarchical Representations of Life-Long Robot Experience
by: Bärmann, Leonard, et al.
Published: (2024)
by: Bärmann, Leonard, et al.
Published: (2024)
Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning
by: Xie, Yifan, et al.
Published: (2025)
by: Xie, Yifan, et al.
Published: (2025)
AVI-Talking: Learning Audio-Visual Instructions for Expressive 3D Talking Face Generation
by: Sun, Yasheng, et al.
Published: (2024)
by: Sun, Yasheng, et al.
Published: (2024)
SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection
by: Liang, Yachao, et al.
Published: (2025)
by: Liang, Yachao, et al.
Published: (2025)
SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation
by: Ling, Zeyu, et al.
Published: (2025)
by: Ling, Zeyu, et al.
Published: (2025)
Incremental Learning of Humanoid Robot Behavior from Natural Interaction and Large Language Models
by: Bärmann, Leonard, et al.
Published: (2023)
by: Bärmann, Leonard, et al.
Published: (2023)
Handling Numeric Expressions in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation
by: Nazarieh, Fatemeh, et al.
Published: (2024)
by: Nazarieh, Fatemeh, et al.
Published: (2024)
Continuously Learning New Words in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
Context Biasing for Pronunciation-Orthography Mismatch in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2025)
by: Huber, Christian, et al.
Published: (2025)
Similar Items
-
Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation
by: Yaman, Dogucan, et al.
Published: (2025) -
Audio-driven Talking Face Generation with Stabilized Synchronization Loss
by: Yaman, Dogucan, et al.
Published: (2023) -
Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework
by: Yaman, Dogucan, et al.
Published: (2025) -
Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis
by: Yaman, Dogucan, et al.
Published: (2025) -
CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)