Disentangled Representation Learning for Environment-agnostic Speaker Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nam, KiHyun, Heo, Hee-Soo, Jung, Jee-weon, Chung, Joon Son |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
von: Nam, KiHyun, et al.
Veröffentlicht: (2026)
von: Nam, KiHyun, et al.
Veröffentlicht: (2026)
SEED: Speaker Embedding Enhancement Diffusion Model
von: Nam, KiHyun, et al.
Veröffentlicht: (2025)
von: Nam, KiHyun, et al.
Veröffentlicht: (2025)
The VoxCeleb Speaker Recognition Challenge: A Retrospective
von: Huh, Jaesung, et al.
Veröffentlicht: (2024)
von: Huh, Jaesung, et al.
Veröffentlicht: (2024)
Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap
von: Nam, KiHyun, et al.
Veröffentlicht: (2025)
von: Nam, KiHyun, et al.
Veröffentlicht: (2025)
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
Improving Design of Input Condition Invariant Speech Enhancement
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
von: Jung, Jihoo, et al.
Veröffentlicht: (2026)
von: Jung, Jihoo, et al.
Veröffentlicht: (2026)
Curriculum learning for self-supervised speaker verification
von: Heo, Hee-Soo, et al.
Veröffentlicht: (2022)
von: Heo, Hee-Soo, et al.
Veröffentlicht: (2022)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
von: Yang, Da-Hee, et al.
Veröffentlicht: (2026)
von: Yang, Da-Hee, et al.
Veröffentlicht: (2026)
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
von: Murata, Masato, et al.
Veröffentlicht: (2025)
von: Murata, Masato, et al.
Veröffentlicht: (2025)
Learning Emotion-Invariant Speaker Representations for Speaker Verification
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification
von: Heo, Hyun-Jun, et al.
Veröffentlicht: (2023)
von: Heo, Hyun-Jun, et al.
Veröffentlicht: (2023)
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
A Comprehensive Investigation on Speaker Augmentation for Speaker Recognition
von: Zhou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenyu, et al.
Veröffentlicht: (2024)
EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training
von: Kwak, Doyeop, et al.
Veröffentlicht: (2025)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2025)
Beyond Silence: Bias Analysis through Loss and Asymmetric Approach in Audio Anti-Spoofing
von: Shim, Hye-jin, et al.
Veröffentlicht: (2024)
von: Shim, Hye-jin, et al.
Veröffentlicht: (2024)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
von: Chung, Soo-Whan, et al.
Veröffentlicht: (2025)
von: Chung, Soo-Whan, et al.
Veröffentlicht: (2025)
Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
Multi-Level Speaker Representation for Target Speaker Extraction
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
von: Zhu, Xiaoxu, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoxu, et al.
Veröffentlicht: (2025)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
Speaker Recognition Using Isomorphic Graph Attention Network Based Pooling on Self-Supervised Representation
von: Ge, Zirui, et al.
Veröffentlicht: (2023)
von: Ge, Zirui, et al.
Veröffentlicht: (2023)
Adapting General Disentanglement-Based Speaker Anonymization for Enhanced Emotion Preservation
von: Miao, Xiaoxiao, et al.
Veröffentlicht: (2024)
von: Miao, Xiaoxiao, et al.
Veröffentlicht: (2024)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
von: Park, Nohil, et al.
Veröffentlicht: (2024)
von: Park, Nohil, et al.
Veröffentlicht: (2024)
A Joint Noise Disentanglement and Adversarial Training Framework for Robust Speaker Verification
von: Xing, Xujiang, et al.
Veröffentlicht: (2024)
von: Xing, Xujiang, et al.
Veröffentlicht: (2024)
Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
von: Li, Guinan, et al.
Veröffentlicht: (2024)
von: Li, Guinan, et al.
Veröffentlicht: (2024)
Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion
von: Li, Xiao, et al.
Veröffentlicht: (2025)
von: Li, Xiao, et al.
Veröffentlicht: (2025)
Speaker Contrastive Learning for Source Speaker Tracing
von: Wang, Qing, et al.
Veröffentlicht: (2024)
von: Wang, Qing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
von: Nam, KiHyun, et al.
Veröffentlicht: (2026) -
SEED: Speaker Embedding Enhancement Diffusion Model
von: Nam, KiHyun, et al.
Veröffentlicht: (2025) -
The VoxCeleb Speaker Recognition Challenge: A Retrospective
von: Huh, Jaesung, et al.
Veröffentlicht: (2024) -
Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap
von: Nam, KiHyun, et al.
Veröffentlicht: (2025) -
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)