Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Horiguchi, Shota, Ando, Atsushi, Moriya, Takafumi, Ashihara, Takanori, Sato, Hiroshi, Tawara, Naohiro, Delcroix, Marc |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Guided Speaker Embedding
von: Horiguchi, Shota, et al.
Veröffentlicht: (2024)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2024)
Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
von: Sato, Hiroshi, et al.
Veröffentlicht: (2024)
von: Sato, Hiroshi, et al.
Veröffentlicht: (2024)
Investigation of Speaker Representation for Target-Speaker Speech Processing
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
Mamba-based Segmentation Model for Speaker Diarization
von: Plaquet, Alexis, et al.
Veröffentlicht: (2024)
von: Plaquet, Alexis, et al.
Veröffentlicht: (2024)
Alignment-Free Training for Transducer-based Multi-Talker ASR
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
von: Plaquet, Alexis, et al.
Veröffentlicht: (2025)
von: Plaquet, Alexis, et al.
Veröffentlicht: (2025)
What Do Self-Supervised Speech and Speaker Models Learn? New Findings From a Cross Model Layer-Wise Analysis
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
Frontend Token Enhancement for Token-Based Speech Recognition
von: Ashihara, Takanori, et al.
Veröffentlicht: (2026)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2026)
NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
von: Kamo, Naoyuki, et al.
Veröffentlicht: (2024)
von: Kamo, Naoyuki, et al.
Veröffentlicht: (2024)
Factor-Conditioned Speaking-Style Captioning
von: Ando, Atsushi, et al.
Veröffentlicht: (2024)
von: Ando, Atsushi, et al.
Veröffentlicht: (2024)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
von: Pálka, Petr, et al.
Veröffentlicht: (2024)
von: Pálka, Petr, et al.
Veröffentlicht: (2024)
CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over
von: Ogawa, Atsunori, et al.
Veröffentlicht: (2024)
von: Ogawa, Atsunori, et al.
Veröffentlicht: (2024)
Interaural time difference loss for binaural target sound extraction
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)
Speaker Characterization by means of Attention Pooling
von: Costa, Federico, et al.
Veröffentlicht: (2024)
von: Costa, Federico, et al.
Veröffentlicht: (2024)
Lightweight Zero-shot Text-to-Speech with Mixture of Adapters
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers
von: Tammen, Marvin, et al.
Veröffentlicht: (2024)
von: Tammen, Marvin, et al.
Veröffentlicht: (2024)
Probing Self-supervised Learning Models with Target Speech Extraction
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
von: Kamo, Naoyuki, et al.
Veröffentlicht: (2025)
von: Kamo, Naoyuki, et al.
Veröffentlicht: (2025)
USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction
von: Zeng, Bang, et al.
Veröffentlicht: (2024)
von: Zeng, Bang, et al.
Veröffentlicht: (2024)
Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2024)
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2024)
Multi-Level Speaker Representation for Target Speaker Extraction
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
Interpreting the Dimensions of Speaker Embedding Space
von: Huckvale, Mark
Veröffentlicht: (2025)
von: Huckvale, Mark
Veröffentlicht: (2025)
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
Listen to Extract: Onset-Prompted Target Speaker Extraction
von: Shen, Pengjie, et al.
Veröffentlicht: (2025)
von: Shen, Pengjie, et al.
Veröffentlicht: (2025)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2025)
von: Wang, Weiqing, et al.
Veröffentlicht: (2025)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
The Reasonable Effectiveness of Speaker Embeddings for Violence Detection
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
Explaining Speaker and Spoof Embeddings via Probing
von: Liu, Xuechen, et al.
Veröffentlicht: (2024)
von: Liu, Xuechen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Guided Speaker Embedding
von: Horiguchi, Shota, et al.
Veröffentlicht: (2024) -
Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025) -
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025) -
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025) -
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
von: Sato, Hiroshi, et al.
Veröffentlicht: (2024)