Magnitude and Phase-based Feature Fusion Using Co-attention Mechanism for Speaker recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Su, Rongfeng, Du, Mengjie, Liu, Xiaokang, Wang, Lan, Yan, Nan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Audio-textual Diffusion Model For Converting Speech Signals Into Ultrasound Tongue Imaging Data
von: Yang, Yudong, et al.
Veröffentlicht: (2024)
von: Yang, Yudong, et al.
Veröffentlicht: (2024)
Automatic Assessment of Dysarthria Using Audio-visual Vowel Graph Attention Network
von: Liu, Xiaokang, et al.
Veröffentlicht: (2024)
von: Liu, Xiaokang, et al.
Veröffentlicht: (2024)
An End-To-End Stuttering Detection Method Based On Conformer And BILSTM
von: Liu, Xiaokang, et al.
Veröffentlicht: (2024)
von: Liu, Xiaokang, et al.
Veröffentlicht: (2024)
Speaker Contrastive Learning for Source Speaker Tracing
von: Wang, Qing, et al.
Veröffentlicht: (2024)
von: Wang, Qing, et al.
Veröffentlicht: (2024)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
Graph-based multi-Feature fusion method for speech emotion recognition
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
Rhythm Features for Speaker Identification
von: Mehlman, Nick, et al.
Veröffentlicht: (2025)
von: Mehlman, Nick, et al.
Veröffentlicht: (2025)
Exploring Frequency-Domain Feature Modeling for HRTF Magnitude Upsampling
von: Chen, Xingyu, et al.
Veröffentlicht: (2026)
von: Chen, Xingyu, et al.
Veröffentlicht: (2026)
MC-LExt: Multi-Channel Target Speaker Extraction with Onset-Prompted Speaker Conditioning Mechanism
von: Ling, Tongtao, et al.
Veröffentlicht: (2025)
von: Ling, Tongtao, et al.
Veröffentlicht: (2025)
SCDNet: Self-supervised Learning Feature-based Speaker Change Detection
von: Li, Yue, et al.
Veröffentlicht: (2024)
von: Li, Yue, et al.
Veröffentlicht: (2024)
Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
von: Jin, Zhan, et al.
Veröffentlicht: (2025)
von: Jin, Zhan, et al.
Veröffentlicht: (2025)
Revisiting and Improving Scoring Fusion for Spoofing-aware Speaker Verification Using Compositional Data Analysis
von: Wang, Xin, et al.
Veröffentlicht: (2024)
von: Wang, Xin, et al.
Veröffentlicht: (2024)
Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
von: Shao, Yiwen, et al.
Veröffentlicht: (2024)
von: Shao, Yiwen, et al.
Veröffentlicht: (2024)
Speech-Based Estimation of Schizophrenia Severity Using Feature Fusion
von: Premananth, Gowtham, et al.
Veröffentlicht: (2024)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2024)
MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
Comparison of Frequency-Fusion Mechanisms for Binaural Direction-of-Arrival Estimation for Multiple Speakers
von: Fejgin, Daniel, et al.
Veröffentlicht: (2024)
von: Fejgin, Daniel, et al.
Veröffentlicht: (2024)
On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization
von: Deegen, Marc, et al.
Veröffentlicht: (2026)
von: Deegen, Marc, et al.
Veröffentlicht: (2026)
SpeakerRPL v2: Robust Open-set Speaker Identification through Enhanced Few-shot Foundation Tuning and Model Fusion
von: Chen, Zhiyong, et al.
Veröffentlicht: (2026)
von: Chen, Zhiyong, et al.
Veröffentlicht: (2026)
Investigating the Potential of Multi-Stage Score Fusion in Spoofing-Aware Speaker Verification
von: Kurnaz, Oguzhan, et al.
Veröffentlicht: (2025)
von: Kurnaz, Oguzhan, et al.
Veröffentlicht: (2025)
Heterogeneous bimodal attention fusion for speech emotion recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion
von: Li, Xiao, et al.
Veröffentlicht: (2025)
von: Li, Xiao, et al.
Veröffentlicht: (2025)
UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition
von: Gan, Chong-Xin, et al.
Veröffentlicht: (2026)
von: Gan, Chong-Xin, et al.
Veröffentlicht: (2026)
MGFF-TDNN: A Multi-Granularity Feature Fusion TDNN Model with Depth-Wise Separable Module for Speaker Verification
von: Li, Ya, et al.
Veröffentlicht: (2025)
von: Li, Ya, et al.
Veröffentlicht: (2025)
Effective Modeling of Critical Contextual Information for TDNN-based Speaker Verification
von: Weng, Shilong, et al.
Veröffentlicht: (2025)
von: Weng, Shilong, et al.
Veröffentlicht: (2025)
Multi-Speaker DOA Estimation in Binaural Hearing Aids using Deep Learning and Speaker Count Fusion
von: Jazaeri, Farnaz, et al.
Veröffentlicht: (2025)
von: Jazaeri, Farnaz, et al.
Veröffentlicht: (2025)
Investigating Acoustic-Textual Emotional Inconsistency Information for Automatic Depression Detection
von: Su, Rongfeng, et al.
Veröffentlicht: (2024)
von: Su, Rongfeng, et al.
Veröffentlicht: (2024)
Phase Aware Ear-Conditioned Learning for Multi-Channel Binaural Speaker Separation
von: Jeremiah, Ruben Johnson Robert, et al.
Veröffentlicht: (2025)
von: Jeremiah, Ruben Johnson Robert, et al.
Veröffentlicht: (2025)
Vclip: Face-based Speaker Generation by Face-voice Association Learning
von: Shi, Yao, et al.
Veröffentlicht: (2026)
von: Shi, Yao, et al.
Veröffentlicht: (2026)
A Probabilistic Fusion Framework for Spoofing Aware Speaker Verification
von: Zhang, You, et al.
Veröffentlicht: (2022)
von: Zhang, You, et al.
Veröffentlicht: (2022)
Neighborhood Attention Transformer with Progressive Channel Fusion for Speaker Verification
von: Li, Nian, et al.
Veröffentlicht: (2024)
von: Li, Nian, et al.
Veröffentlicht: (2024)
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
Neural Codec-based Adversarial Sample Detection for Speaker Verification
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024)
Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
von: Li, Guinan, et al.
Veröffentlicht: (2024)
von: Li, Guinan, et al.
Veröffentlicht: (2024)
Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
Token-based Attractors and Cross-attention in Spoof Diarization
von: Koo, Kyo-Won, et al.
Veröffentlicht: (2025)
von: Koo, Kyo-Won, et al.
Veröffentlicht: (2025)
IDMap: A Pseudo-Speaker Generator Framework Based on Speaker Identity Index to Vector Mapping
von: Liu, Zeyan, et al.
Veröffentlicht: (2025)
von: Liu, Zeyan, et al.
Veröffentlicht: (2025)
Study on Inter and Intra Speaker Variability in Speaker Recognition
von: Okhotnikov, Anton, et al.
Veröffentlicht: (2024)
von: Okhotnikov, Anton, et al.
Veröffentlicht: (2024)
Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization
von: Han, Jiangyu, et al.
Veröffentlicht: (2025)
von: Han, Jiangyu, et al.
Veröffentlicht: (2025)
EvoTSE: Evolving Enrollment for Target Speaker Extraction
von: Liu, Zikai, et al.
Veröffentlicht: (2026)
von: Liu, Zikai, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
An Audio-textual Diffusion Model For Converting Speech Signals Into Ultrasound Tongue Imaging Data
von: Yang, Yudong, et al.
Veröffentlicht: (2024) -
Automatic Assessment of Dysarthria Using Audio-visual Vowel Graph Attention Network
von: Liu, Xiaokang, et al.
Veröffentlicht: (2024) -
An End-To-End Stuttering Detection Method Based On Conformer And BILSTM
von: Liu, Xiaokang, et al.
Veröffentlicht: (2024) -
Speaker Contrastive Learning for Source Speaker Tracing
von: Wang, Qing, et al.
Veröffentlicht: (2024) -
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)