A framework of text-dependent speaker verification for chinese numerical string corpus
Fuente:
arXiv
Salvato in:
| Autori principali: | Zheng, Litong, Hong, Feng, Xu, Weijie, Zheng, Wan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Text adaptation for speaker verification with speaker-text factorized embeddings
di: Yang, Yexin, et al.
Pubblicazione: (2025)
di: Yang, Yexin, et al.
Pubblicazione: (2025)
On the influence of language similarity in non-target speaker verification trials
di: Reuter, Paul M., et al.
Pubblicazione: (2025)
di: Reuter, Paul M., et al.
Pubblicazione: (2025)
Gradient weighting for speaker verification in extremely low Signal-to-Noise Ratio
di: Ma, Yi, et al.
Pubblicazione: (2024)
di: Ma, Yi, et al.
Pubblicazione: (2024)
Improving fairness in speaker verification via Group-adapted Fusion Network
di: Shen, Hua, et al.
Pubblicazione: (2022)
di: Shen, Hua, et al.
Pubblicazione: (2022)
Improving speaker verification robustness with synthetic emotional utterances
di: Koditala, Nikhil Kumar, et al.
Pubblicazione: (2024)
di: Koditala, Nikhil Kumar, et al.
Pubblicazione: (2024)
Hierarchical speaker representation for target speaker extraction
di: He, Shulin, et al.
Pubblicazione: (2022)
di: He, Shulin, et al.
Pubblicazione: (2022)
Clustering-based hard negative sampling for supervised contrastive speaker verification
di: Masztalski, Piotr, et al.
Pubblicazione: (2025)
di: Masztalski, Piotr, et al.
Pubblicazione: (2025)
Audio-visual child-adult speaker classification in dyadic interactions
di: Xu, Anfeng, et al.
Pubblicazione: (2023)
di: Xu, Anfeng, et al.
Pubblicazione: (2023)
Improving curriculum learning for target speaker extraction with synthetic speakers
di: Liu, Yun, et al.
Pubblicazione: (2024)
di: Liu, Yun, et al.
Pubblicazione: (2024)
Quantifying the effect of speech pathology on automatic and human speaker verification
di: Halpern, Bence Mark, et al.
Pubblicazione: (2024)
di: Halpern, Bence Mark, et al.
Pubblicazione: (2024)
How phonemes contribute to deep speaker models?
di: Li, Pengqi, et al.
Pubblicazione: (2024)
di: Li, Pengqi, et al.
Pubblicazione: (2024)
The importance of spatial and spectral information in multiple speaker tracking
di: Beit-On, Hanan, et al.
Pubblicazione: (2024)
di: Beit-On, Hanan, et al.
Pubblicazione: (2024)
SCDiar: a streaming diarization system based on speaker change detection and speech recognition
di: Zheng, Naijun, et al.
Pubblicazione: (2025)
di: Zheng, Naijun, et al.
Pubblicazione: (2025)
Spoken language change detection inspired by speaker change detection
di: Mishra, Jagabandhu, et al.
Pubblicazione: (2023)
di: Mishra, Jagabandhu, et al.
Pubblicazione: (2023)
An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
di: Han, Runduo, et al.
Pubblicazione: (2024)
di: Han, Runduo, et al.
Pubblicazione: (2024)
HeightCeleb - an enrichment of VoxCeleb dataset with speaker height information
di: Kacprzak, Stanisław, et al.
Pubblicazione: (2024)
di: Kacprzak, Stanisław, et al.
Pubblicazione: (2024)
Spectral or spatial? Leveraging both for speaker extraction in challenging data conditions
di: Eisenberg, Aviad, et al.
Pubblicazione: (2025)
di: Eisenberg, Aviad, et al.
Pubblicazione: (2025)
Why disentanglement-based speaker anonymization systems fail at preserving emotions?
di: Gaznepoglu, Ünal Ege, et al.
Pubblicazione: (2025)
di: Gaznepoglu, Ünal Ege, et al.
Pubblicazione: (2025)
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
di: Murata, Masato, et al.
Pubblicazione: (2025)
di: Murata, Masato, et al.
Pubblicazione: (2025)
An efficient text augmentation approach for contextualized Mandarin speech recognition
di: Zheng, Naijun, et al.
Pubblicazione: (2024)
di: Zheng, Naijun, et al.
Pubblicazione: (2024)
EEND-M2F: Masked-attention mask transformers for speaker diarization
di: Härkönen, Marc, et al.
Pubblicazione: (2024)
di: Härkönen, Marc, et al.
Pubblicazione: (2024)
X-CrossNet: A complex spectral mapping approach to target speaker extraction with cross attention speaker embedding fusion
di: Sun, Chang, et al.
Pubblicazione: (2024)
di: Sun, Chang, et al.
Pubblicazione: (2024)
Exploring synthetic data for cross-speaker style transfer in style representation based TTS
di: Ueda, Lucas H., et al.
Pubblicazione: (2024)
di: Ueda, Lucas H., et al.
Pubblicazione: (2024)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
di: Zhang, Yiru, et al.
Pubblicazione: (2025)
di: Zhang, Yiru, et al.
Pubblicazione: (2025)
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
di: Kunešová, Marie, et al.
Pubblicazione: (2025)
di: Kunešová, Marie, et al.
Pubblicazione: (2025)
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2025)
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2025)
Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment
di: Zhu, Haolin, et al.
Pubblicazione: (2024)
di: Zhu, Haolin, et al.
Pubblicazione: (2024)
End-to-end multi-channel speaker extraction and binaural speech synthesis
di: Chi, Cheng, et al.
Pubblicazione: (2024)
di: Chi, Cheng, et al.
Pubblicazione: (2024)
$\text{M}^3\text{PDB}$: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation
di: Zhu, Boyu, et al.
Pubblicazione: (2025)
di: Zhu, Boyu, et al.
Pubblicazione: (2025)
THAI Speech Emotion Recognition (THAI-SER) corpus
di: Wongpithayadisai, Jilamika, et al.
Pubblicazione: (2025)
di: Wongpithayadisai, Jilamika, et al.
Pubblicazione: (2025)
SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation
di: Ueda, Lucas H., et al.
Pubblicazione: (2026)
di: Ueda, Lucas H., et al.
Pubblicazione: (2026)
A Benchmark for Multi-speaker Anonymization
di: Miao, Xiaoxiao, et al.
Pubblicazione: (2024)
di: Miao, Xiaoxiao, et al.
Pubblicazione: (2024)
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
di: Grossman, Raymond, et al.
Pubblicazione: (2025)
di: Grossman, Raymond, et al.
Pubblicazione: (2025)
Building speech corpus with diverse voice characteristics for its prompt-based representation
di: Watanabe, Aya, et al.
Pubblicazione: (2024)
di: Watanabe, Aya, et al.
Pubblicazione: (2024)
Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
di: Kang, Wei, et al.
Pubblicazione: (2023)
di: Kang, Wei, et al.
Pubblicazione: (2023)
Visual-based spatial audio generation system for multi-speaker environments
di: Liu, Xiaojing, et al.
Pubblicazione: (2025)
di: Liu, Xiaojing, et al.
Pubblicazione: (2025)
A Joint Noise Disentanglement and Adversarial Training Framework for Robust Speaker Verification
di: Xing, Xujiang, et al.
Pubblicazione: (2024)
di: Xing, Xujiang, et al.
Pubblicazione: (2024)
STASE: A spatialized text-to-audio synthesis engine for music generation
di: Chi, Tutti, et al.
Pubblicazione: (2025)
di: Chi, Tutti, et al.
Pubblicazione: (2025)
Malacopula: adversarial automatic speaker verification attacks using a neural-based generalised Hammerstein model
di: Todisco, Massimiliano, et al.
Pubblicazione: (2024)
di: Todisco, Massimiliano, et al.
Pubblicazione: (2024)
SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Text adaptation for speaker verification with speaker-text factorized embeddings
di: Yang, Yexin, et al.
Pubblicazione: (2025) -
On the influence of language similarity in non-target speaker verification trials
di: Reuter, Paul M., et al.
Pubblicazione: (2025) -
Gradient weighting for speaker verification in extremely low Signal-to-Noise Ratio
di: Ma, Yi, et al.
Pubblicazione: (2024) -
Improving fairness in speaker verification via Group-adapted Fusion Network
di: Shen, Hua, et al.
Pubblicazione: (2022) -
Improving speaker verification robustness with synthetic emotional utterances
di: Koditala, Nikhil Kumar, et al.
Pubblicazione: (2024)