Towards Language-Independent Face-Voice Association with Multimodal Foundation Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Farhadipour, Aref, Vukovic, Teodora, Dellwo, Volker |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
por: Farhadipour, Aref, et al.
Publicado: (2025)
por: Farhadipour, Aref, et al.
Publicado: (2025)
Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
por: Farhadipour, Aref, et al.
Publicado: (2024)
por: Farhadipour, Aref, et al.
Publicado: (2024)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
por: Farhadipour, Aref, et al.
Publicado: (2024)
por: Farhadipour, Aref, et al.
Publicado: (2024)
CL-UZH submission to the NIST SRE 2024 Speaker Recognition Evaluation
por: Farhadipour, Aref, et al.
Publicado: (2025)
por: Farhadipour, Aref, et al.
Publicado: (2025)
TidyVoice 2026 Challenge Evaluation Plan
por: Farhadipour, Aref, et al.
Publicado: (2026)
por: Farhadipour, Aref, et al.
Publicado: (2026)
Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts
por: Farhadipour, Aref, et al.
Publicado: (2025)
por: Farhadipour, Aref, et al.
Publicado: (2025)
State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
por: Farhadipour, Aref, et al.
Publicado: (2025)
por: Farhadipour, Aref, et al.
Publicado: (2025)
Efficient Face Detection with Audio-Based Region Proposals for Human-Robot Interactions
por: Aris, William, et al.
Publicado: (2023)
por: Aris, William, et al.
Publicado: (2023)
TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
por: Farhadipour, Aref, et al.
Publicado: (2026)
por: Farhadipour, Aref, et al.
Publicado: (2026)
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
por: Park, Jay, et al.
Publicado: (2025)
por: Park, Jay, et al.
Publicado: (2025)
CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing
por: Yue, Xianghu, et al.
Publicado: (2024)
por: Yue, Xianghu, et al.
Publicado: (2024)
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
por: Wang, Yi, et al.
Publicado: (2025)
por: Wang, Yi, et al.
Publicado: (2025)
KunquDB: An Attempt for Speaker Verification in the Chinese Opera Scenario
por: Zhou, Huali, et al.
Publicado: (2024)
por: Zhou, Huali, et al.
Publicado: (2024)
Improvement Of Audiovisual Quality Estimation Using A Nonlinear Autoregressive Exogenous Neural Network And Bitstream Parameters
por: Kossi, Koffi, et al.
Publicado: (2024)
por: Kossi, Koffi, et al.
Publicado: (2024)
Beyond Correlation: Evaluating Multimedia Quality Models with the Constrained Concordance Index
por: Ragano, Alessandro, et al.
Publicado: (2024)
por: Ragano, Alessandro, et al.
Publicado: (2024)
Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis
por: Zhang, Miao, et al.
Publicado: (2025)
por: Zhang, Miao, et al.
Publicado: (2025)
Multimodal Marvels of Deep Learning in Medical Diagnosis: A Comprehensive Review of COVID-19 Detection
por: Islam, Md Shofiqul, et al.
Publicado: (2025)
por: Islam, Md Shofiqul, et al.
Publicado: (2025)
UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
por: Cao, Yuqin, et al.
Publicado: (2024)
por: Cao, Yuqin, et al.
Publicado: (2024)
Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
por: Mingote, Victoria, et al.
Publicado: (2024)
por: Mingote, Victoria, et al.
Publicado: (2024)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
por: Farhadipour, Aref, et al.
Publicado: (2023)
por: Farhadipour, Aref, et al.
Publicado: (2023)
End-to-end audio-visual learning for cochlear implant sound coding simulations in noisy environments
por: Lin, Meng-Ping, et al.
Publicado: (2025)
por: Lin, Meng-Ping, et al.
Publicado: (2025)
SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with Silhouettes
por: Tanigawa, Risako, et al.
Publicado: (2024)
por: Tanigawa, Risako, et al.
Publicado: (2024)
A multi-modal approach for identifying schizophrenia using cross-modal attention
por: Premananth, Gowtham, et al.
Publicado: (2023)
por: Premananth, Gowtham, et al.
Publicado: (2023)
Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles
por: Farhadipour, Aref, et al.
Publicado: (2026)
por: Farhadipour, Aref, et al.
Publicado: (2026)
Faces that Speak: Jointly Synthesising Talking Face and Speech from Text
por: Jang, Youngjoon, et al.
Publicado: (2024)
por: Jang, Youngjoon, et al.
Publicado: (2024)
ToS: A Team of Specialists ensemble framework for Stereo Sound Event Localization and Detection with distance estimation in Video
por: Berghi, Davide, et al.
Publicado: (2026)
por: Berghi, Davide, et al.
Publicado: (2026)
Livestock feeding behaviour: A review on automated systems for ruminant monitoring
por: Chelotti, José, et al.
Publicado: (2023)
por: Chelotti, José, et al.
Publicado: (2023)
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
por: Zhong, Zhi, et al.
Publicado: (2025)
por: Zhong, Zhi, et al.
Publicado: (2025)
Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling
por: Chen, Xuanjun, et al.
Publicado: (2025)
por: Chen, Xuanjun, et al.
Publicado: (2025)
Korean aegyo speech shows systematic F1 increase to signal childlike qualities
por: Kim, Ji-eun, et al.
Publicado: (2026)
por: Kim, Ji-eun, et al.
Publicado: (2026)
Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
por: Nguyen, Hong, et al.
Publicado: (2024)
por: Nguyen, Hong, et al.
Publicado: (2024)
Two Web Toolkits for Multimodal Piano Performance Dataset Acquisition and Fingering Annotation
por: Park, Junhyung, et al.
Publicado: (2025)
por: Park, Junhyung, et al.
Publicado: (2025)
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
por: Phukan, Orchid Chetia, et al.
Publicado: (2025)
por: Phukan, Orchid Chetia, et al.
Publicado: (2025)
Exploring Phonetic Context-Aware Lip-Sync For Talking Face Generation
por: Park, Se Jin, et al.
Publicado: (2023)
por: Park, Se Jin, et al.
Publicado: (2023)
Audio-Visual Approach For Multimodal Concurrent Speaker Detection
por: Eliav, Amit, et al.
Publicado: (2024)
por: Eliav, Amit, et al.
Publicado: (2024)
Sound Source Localization for Spatial Mapping of Surgical Actions in Dynamic Scenes
por: Hein, Jonas, et al.
Publicado: (2025)
por: Hein, Jonas, et al.
Publicado: (2025)
PrismAudio: Decomposed Chain-of-Thoughts and Multi-dimensional Rewards for Video-to-Audio Generation
por: Liu, Huadai, et al.
Publicado: (2025)
por: Liu, Huadai, et al.
Publicado: (2025)
Cross Attentional Audio-Visual Fusion for Dimensional Emotion Recognition
por: Praveen, R. Gnana, et al.
Publicado: (2021)
por: Praveen, R. Gnana, et al.
Publicado: (2021)
Lip Reading for Low-resource Languages by Learning and Combining General Speech Knowledge and Language-specific Knowledge
por: Kim, Minsu, et al.
Publicado: (2023)
por: Kim, Minsu, et al.
Publicado: (2023)
Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification
por: Shimada, Kazuki, et al.
Publicado: (2025)
por: Shimada, Kazuki, et al.
Publicado: (2025)
Ejemplares similares
-
Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
por: Farhadipour, Aref, et al.
Publicado: (2025) -
Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
por: Farhadipour, Aref, et al.
Publicado: (2024) -
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
por: Farhadipour, Aref, et al.
Publicado: (2024) -
CL-UZH submission to the NIST SRE 2024 Speaker Recognition Evaluation
por: Farhadipour, Aref, et al.
Publicado: (2025) -
TidyVoice 2026 Challenge Evaluation Plan
por: Farhadipour, Aref, et al.
Publicado: (2026)