Shared Multi-modal Embedding Space for Face-Voice Association
Fuente:
arXiv
Saved in:
| Main Authors: | Simic, Christopher, Riedhammer, Korbinian, Bocklet, Tobias |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR Models
by: Simic, Christopher, et al.
Published: (2025)
by: Simic, Christopher, et al.
Published: (2025)
XM-ALIGN: Unified Cross-Modal Embedding Alignment for Face-Voice Association
by: Fang, Zhihua, et al.
Published: (2025)
by: Fang, Zhihua, et al.
Published: (2025)
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
by: Baumann, Ilja, et al.
Published: (2025)
by: Baumann, Ilja, et al.
Published: (2025)
Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0
by: Engert, Natalie, et al.
Published: (2026)
by: Engert, Natalie, et al.
Published: (2026)
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
by: Kang, Fang, et al.
Published: (2025)
by: Kang, Fang, et al.
Published: (2025)
Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection
by: Chuchra, Akanksha, et al.
Published: (2026)
by: Chuchra, Akanksha, et al.
Published: (2026)
Large Language Models for Dysfluency Detection in Stuttered Speech
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
by: Tang, Changli, et al.
Published: (2025)
by: Tang, Changli, et al.
Published: (2025)
Face-voice Association in Multilingual Environments (FAME) Challenge 2024 Evaluation Plan
by: Saeed, Muhammad Saad, et al.
Published: (2024)
by: Saeed, Muhammad Saad, et al.
Published: (2024)
From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
by: Kim, Ji-Hoon, et al.
Published: (2025)
by: Kim, Ji-Hoon, et al.
Published: (2025)
WavFlow: Audio Generation in Waveform Space
by: Zhou, Feiyan, et al.
Published: (2026)
by: Zhou, Feiyan, et al.
Published: (2026)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
by: Tian, Zeyue, et al.
Published: (2026)
by: Tian, Zeyue, et al.
Published: (2026)
Emotional Face-to-Speech
by: Ye, Jiaxin, et al.
Published: (2025)
by: Ye, Jiaxin, et al.
Published: (2025)
Hear Your Face: Face-based voice conversion with F0 estimation
by: Lee, Jaejun, et al.
Published: (2024)
by: Lee, Jaejun, et al.
Published: (2024)
Voice Pathology Detection Using Phonation
by: Siva, Sri Raksha, et al.
Published: (2025)
by: Siva, Sri Raksha, et al.
Published: (2025)
Infusing Acoustic Pause Context into Text-Based Dementia Assessment
by: Braun, Franziska, et al.
Published: (2024)
by: Braun, Franziska, et al.
Published: (2024)
Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion
by: Rong, Yan, et al.
Published: (2024)
by: Rong, Yan, et al.
Published: (2024)
Differentiable Room Acoustic Rendering with Multi-View Vision Priors
by: Jin, Derong, et al.
Published: (2025)
by: Jin, Derong, et al.
Published: (2025)
ReelWave: Multi-Agentic Movie Sound Generation through Multimodal LLM Conversation
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape Mapping
by: Khanal, Subash, et al.
Published: (2024)
by: Khanal, Subash, et al.
Published: (2024)
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
by: Li, Kai, et al.
Published: (2025)
by: Li, Kai, et al.
Published: (2025)
Toward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization
by: Vu, Tung, et al.
Published: (2026)
by: Vu, Tung, et al.
Published: (2026)
Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control
by: Li, Bingliang, et al.
Published: (2024)
by: Li, Bingliang, et al.
Published: (2024)
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls
by: Chen, Zihao, et al.
Published: (2024)
by: Chen, Zihao, et al.
Published: (2024)
VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized Diffusion-based Voice Cloning
by: Hu, Qianyue, et al.
Published: (2025)
by: Hu, Qianyue, et al.
Published: (2025)
Training-Free Deepfake Voice Recognition by Leveraging Large-Scale Pre-Trained Models
by: Pianese, Alessandro, et al.
Published: (2024)
by: Pianese, Alessandro, et al.
Published: (2024)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
by: Yang, Jianxuan, et al.
Published: (2025)
by: Yang, Jianxuan, et al.
Published: (2025)
EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model
by: Zhang, Bingyuan, et al.
Published: (2024)
by: Zhang, Bingyuan, et al.
Published: (2024)
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
by: Li, Junjie, et al.
Published: (2024)
by: Li, Junjie, et al.
Published: (2024)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
TRACE: Training-Free Partial Audio Deepfake Detection via Embedding Trajectory Analysis of Speech Foundation Models
by: Khan, Awais, et al.
Published: (2026)
by: Khan, Awais, et al.
Published: (2026)
VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models
by: Hu, Rui, et al.
Published: (2025)
by: Hu, Rui, et al.
Published: (2025)
Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?
by: Guan, Yiwen, et al.
Published: (2024)
by: Guan, Yiwen, et al.
Published: (2024)
MAG: Multi-Modal Aligned Autoregressive Co-Speech Gesture Generation without Vector Quantization
by: Liu, Binjie, et al.
Published: (2025)
by: Liu, Binjie, et al.
Published: (2025)
Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment
by: Hong, Jiaying, et al.
Published: (2025)
by: Hong, Jiaying, et al.
Published: (2025)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
by: Wagner, Dominik, et al.
Published: (2025)
by: Wagner, Dominik, et al.
Published: (2025)
A Low-rank Matching Attention based Cross-modal Feature Fusion Method for Conversational Emotion Recognition
by: Shou, Yuntao, et al.
Published: (2023)
by: Shou, Yuntao, et al.
Published: (2023)
A-JEPA: Joint-Embedding Predictive Architecture Can Listen
by: Fei, Zhengcong, et al.
Published: (2023)
by: Fei, Zhengcong, et al.
Published: (2023)
VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection
by: Cheng, Hao, et al.
Published: (2025)
by: Cheng, Hao, et al.
Published: (2025)
Similar Items
-
Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR Models
by: Simic, Christopher, et al.
Published: (2025) -
XM-ALIGN: Unified Cross-Modal Embedding Alignment for Face-Voice Association
by: Fang, Zhihua, et al.
Published: (2025) -
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
by: Baumann, Ilja, et al.
Published: (2025) -
Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0
by: Engert, Natalie, et al.
Published: (2026) -
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
by: Wagner, Dominik, et al.
Published: (2024)