Gespeichert in:
| Hauptverfasser: | Olijslager, Mariëtte, Ziabari, Seyed Sahand Mohammadi, Alsahag, Ali Mohammed Mansoor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.01363 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Probing the Feasibility of Multilingual Speaker Anonymization
von: Meyer, Sarina, et al.
Veröffentlicht: (2024)
von: Meyer, Sarina, et al.
Veröffentlicht: (2024)
Disentangling Speaker Traits for Deepfake Source Verification via Chebyshev Polynomial and Riemannian Metric Learning
von: Xuan, Xi, et al.
Veröffentlicht: (2026)
von: Xuan, Xi, et al.
Veröffentlicht: (2026)
Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
von: Komatsu, Ryota, et al.
Veröffentlicht: (2024)
von: Komatsu, Ryota, et al.
Veröffentlicht: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
System Description for the Displace Speaker Diarization Challenge 2023
von: Aliyev, Ali
Veröffentlicht: (2024)
von: Aliyev, Ali
Veröffentlicht: (2024)
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
Advancing Topic Segmentation of Broadcasted Speech with Multilingual Semantic Embeddings
von: Shukla, Sakshi Deo, et al.
Veröffentlicht: (2024)
von: Shukla, Sakshi Deo, et al.
Veröffentlicht: (2024)
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
von: Li, Junjie, et al.
Veröffentlicht: (2023)
von: Li, Junjie, et al.
Veröffentlicht: (2023)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
Investigation of Speaker Representation for Target-Speaker Speech Processing
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition
von: Yang, Zhengdong, et al.
Veröffentlicht: (2025)
von: Yang, Zhengdong, et al.
Veröffentlicht: (2025)
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026)
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026)
Exploring Multilingual Unseen Speaker Emotion Recognition: Leveraging Co-Attention Cues in Multitask Learning
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR
von: Lin, Zhennan, et al.
Veröffentlicht: (2026)
von: Lin, Zhennan, et al.
Veröffentlicht: (2026)
R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2023)
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2023)
Multilingual Prosody Transfer: Comparing Supervised & Transfer Learning
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning
von: He, Haorui, et al.
Veröffentlicht: (2024)
von: He, Haorui, et al.
Veröffentlicht: (2024)
Speaker Contrastive Learning for Source Speaker Tracing
von: Wang, Qing, et al.
Veröffentlicht: (2024)
von: Wang, Qing, et al.
Veröffentlicht: (2024)
Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning
von: Yang, Qian, et al.
Veröffentlicht: (2025)
von: Yang, Qian, et al.
Veröffentlicht: (2025)
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
von: Roll, Nathan, et al.
Veröffentlicht: (2025)
von: Roll, Nathan, et al.
Veröffentlicht: (2025)
Speaker-Aware Simulation Improves Conversational Speech Recognition
von: Gedeon, Máté, et al.
Veröffentlicht: (2026)
von: Gedeon, Máté, et al.
Veröffentlicht: (2026)
CoLMbo: Speaker Language Model for Descriptive Profiling
von: Baali, Massa, et al.
Veröffentlicht: (2025)
von: Baali, Massa, et al.
Veröffentlicht: (2025)
A Review of Common Online Speaker Diarization Methods
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
DiariST: Streaming Speech Translation with Speaker Diarization
von: Yang, Mu, et al.
Veröffentlicht: (2023)
von: Yang, Mu, et al.
Veröffentlicht: (2023)
Continual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
kNN For Whisper And Its Effect On Bias And Speaker Adaptation
von: Nachesa, Maya K., et al.
Veröffentlicht: (2024)
von: Nachesa, Maya K., et al.
Veröffentlicht: (2024)
Systematic Evaluation of Online Speaker Diarization Systems Regarding their Latency
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
Romanization Encoding For Multilingual ASR
von: Ding, Wen, et al.
Veröffentlicht: (2024)
von: Ding, Wen, et al.
Veröffentlicht: (2024)
Speech Separation based on Contrastive Learning and Deep Modularization
von: Ochieng, Peter
Veröffentlicht: (2023)
von: Ochieng, Peter
Veröffentlicht: (2023)
Factorized RVQ-GAN For Disentangled Speech Tokenization
von: Khurana, Sameer, et al.
Veröffentlicht: (2025)
von: Khurana, Sameer, et al.
Veröffentlicht: (2025)
Disentanglement in a GAN for Unconditional Speech Synthesis
von: Baas, Matthew, et al.
Veröffentlicht: (2023)
von: Baas, Matthew, et al.
Veröffentlicht: (2023)
Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training
von: Denisov, Pavel, et al.
Veröffentlicht: (2024)
von: Denisov, Pavel, et al.
Veröffentlicht: (2024)
Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2024)
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2024)
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
von: Kong, Jungil, et al.
Veröffentlicht: (2023)
von: Kong, Jungil, et al.
Veröffentlicht: (2023)
Emotion-Anchored Contrastive Learning Framework for Emotion Recognition in Conversation
von: Yu, Fangxu, et al.
Veröffentlicht: (2024)
von: Yu, Fangxu, et al.
Veröffentlicht: (2024)
LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservation
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2026)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2026)
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation
von: Cheng, Luyao, et al.
Veröffentlicht: (2023)
von: Cheng, Luyao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Probing the Feasibility of Multilingual Speaker Anonymization
von: Meyer, Sarina, et al.
Veröffentlicht: (2024) -
Disentangling Speaker Traits for Deepfake Source Verification via Chebyshev Polynomial and Riemannian Metric Learning
von: Xuan, Xi, et al.
Veröffentlicht: (2026) -
Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
von: Komatsu, Ryota, et al.
Veröffentlicht: (2024) -
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024) -
System Description for the Displace Speaker Diarization Challenge 2023
von: Aliyev, Ali
Veröffentlicht: (2024)