Text adaptation for speaker verification with speaker-text factorized embeddings
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Yexin, Wang, Shuai, Gong, Xun, Qian, Yanmin, Yu, Kai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving fairness in speaker verification via Group-adapted Fusion Network
von: Shen, Hua, et al.
Veröffentlicht: (2022)
von: Shen, Hua, et al.
Veröffentlicht: (2022)
A framework of text-dependent speaker verification for chinese numerical string corpus
von: Zheng, Litong, et al.
Veröffentlicht: (2024)
von: Zheng, Litong, et al.
Veröffentlicht: (2024)
Hierarchical speaker representation for target speaker extraction
von: He, Shulin, et al.
Veröffentlicht: (2022)
von: He, Shulin, et al.
Veröffentlicht: (2022)
On the influence of language similarity in non-target speaker verification trials
von: Reuter, Paul M., et al.
Veröffentlicht: (2025)
von: Reuter, Paul M., et al.
Veröffentlicht: (2025)
Gradient weighting for speaker verification in extremely low Signal-to-Noise Ratio
von: Ma, Yi, et al.
Veröffentlicht: (2024)
von: Ma, Yi, et al.
Veröffentlicht: (2024)
Improving curriculum learning for target speaker extraction with synthetic speakers
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
Improving speaker verification robustness with synthetic emotional utterances
von: Koditala, Nikhil Kumar, et al.
Veröffentlicht: (2024)
von: Koditala, Nikhil Kumar, et al.
Veröffentlicht: (2024)
X-CrossNet: A complex spectral mapping approach to target speaker extraction with cross attention speaker embedding fusion
von: Sun, Chang, et al.
Veröffentlicht: (2024)
von: Sun, Chang, et al.
Veröffentlicht: (2024)
How phonemes contribute to deep speaker models?
von: Li, Pengqi, et al.
Veröffentlicht: (2024)
von: Li, Pengqi, et al.
Veröffentlicht: (2024)
Clustering-based hard negative sampling for supervised contrastive speaker verification
von: Masztalski, Piotr, et al.
Veröffentlicht: (2025)
von: Masztalski, Piotr, et al.
Veröffentlicht: (2025)
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
von: Cord-Landwehr, Tobias, et al.
Veröffentlicht: (2025)
von: Cord-Landwehr, Tobias, et al.
Veröffentlicht: (2025)
Quantifying the effect of speech pathology on automatic and human speaker verification
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2024)
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2024)
The importance of spatial and spectral information in multiple speaker tracking
von: Beit-On, Hanan, et al.
Veröffentlicht: (2024)
von: Beit-On, Hanan, et al.
Veröffentlicht: (2024)
SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
von: Grossman, Raymond, et al.
Veröffentlicht: (2025)
von: Grossman, Raymond, et al.
Veröffentlicht: (2025)
Audio-visual child-adult speaker classification in dyadic interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2023)
von: Xu, Anfeng, et al.
Veröffentlicht: (2023)
Spoken language change detection inspired by speaker change detection
von: Mishra, Jagabandhu, et al.
Veröffentlicht: (2023)
von: Mishra, Jagabandhu, et al.
Veröffentlicht: (2023)
Multi-speaker Text-to-speech Training with Speaker Anonymized Data
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
A Benchmark for Multi-speaker Anonymization
von: Miao, Xiaoxiao, et al.
Veröffentlicht: (2024)
von: Miao, Xiaoxiao, et al.
Veröffentlicht: (2024)
Spectral or spatial? Leveraging both for speaker extraction in challenging data conditions
von: Eisenberg, Aviad, et al.
Veröffentlicht: (2025)
von: Eisenberg, Aviad, et al.
Veröffentlicht: (2025)
Why disentanglement-based speaker anonymization systems fail at preserving emotions?
von: Gaznepoglu, Ünal Ege, et al.
Veröffentlicht: (2025)
von: Gaznepoglu, Ünal Ege, et al.
Veröffentlicht: (2025)
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
von: Murata, Masato, et al.
Veröffentlicht: (2025)
von: Murata, Masato, et al.
Veröffentlicht: (2025)
HeightCeleb - an enrichment of VoxCeleb dataset with speaker height information
von: Kacprzak, Stanisław, et al.
Veröffentlicht: (2024)
von: Kacprzak, Stanisław, et al.
Veröffentlicht: (2024)
EEND-M2F: Masked-attention mask transformers for speaker diarization
von: Härkönen, Marc, et al.
Veröffentlicht: (2024)
von: Härkönen, Marc, et al.
Veröffentlicht: (2024)
On the calibration of powerset speaker diarization models
von: Plaquet, Alexis, et al.
Veröffentlicht: (2024)
von: Plaquet, Alexis, et al.
Veröffentlicht: (2024)
Visual-based spatial audio generation system for multi-speaker environments
von: Liu, Xiaojing, et al.
Veröffentlicht: (2025)
von: Liu, Xiaojing, et al.
Veröffentlicht: (2025)
An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
von: Han, Runduo, et al.
Veröffentlicht: (2024)
von: Han, Runduo, et al.
Veröffentlicht: (2024)
Exploring synthetic data for cross-speaker style transfer in style representation based TTS
von: Ueda, Lucas H., et al.
Veröffentlicht: (2024)
von: Ueda, Lucas H., et al.
Veröffentlicht: (2024)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
Triage knowledge distillation for speaker verification
von: Kim, Ju-ho, et al.
Veröffentlicht: (2026)
von: Kim, Ju-ho, et al.
Veröffentlicht: (2026)
Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment
von: Zhu, Haolin, et al.
Veröffentlicht: (2024)
von: Zhu, Haolin, et al.
Veröffentlicht: (2024)
Extending Whisper with prompt tuning to target-speaker ASR
von: Ma, Hao, et al.
Veröffentlicht: (2023)
von: Ma, Hao, et al.
Veröffentlicht: (2023)
An approach to optimize inference of the DIART speaker diarization pipeline
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
Malacopula: adversarial automatic speaker verification attacks using a neural-based generalised Hammerstein model
von: Todisco, Massimiliano, et al.
Veröffentlicht: (2024)
von: Todisco, Massimiliano, et al.
Veröffentlicht: (2024)
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
von: Gong, Xun, et al.
Veröffentlicht: (2025)
von: Gong, Xun, et al.
Veröffentlicht: (2025)
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
von: Gong, Xun, et al.
Veröffentlicht: (2024)
von: Gong, Xun, et al.
Veröffentlicht: (2024)
Tandem spoofing-robust automatic speaker verification based on time-domain embeddings
von: Weizman, Avishai, et al.
Veröffentlicht: (2024)
von: Weizman, Avishai, et al.
Veröffentlicht: (2024)
ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
End-to-end multi-channel speaker extraction and binaural speech synthesis
von: Chi, Cheng, et al.
Veröffentlicht: (2024)
von: Chi, Cheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Improving fairness in speaker verification via Group-adapted Fusion Network
von: Shen, Hua, et al.
Veröffentlicht: (2022) -
A framework of text-dependent speaker verification for chinese numerical string corpus
von: Zheng, Litong, et al.
Veröffentlicht: (2024) -
Hierarchical speaker representation for target speaker extraction
von: He, Shulin, et al.
Veröffentlicht: (2022) -
On the influence of language similarity in non-target speaker verification trials
von: Reuter, Paul M., et al.
Veröffentlicht: (2025) -
Gradient weighting for speaker verification in extremely low Signal-to-Noise Ratio
von: Ma, Yi, et al.
Veröffentlicht: (2024)