Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization
Fuente:
arXiv
Guardado en:
| Autores principales: | Polok, Alexander, Medennikov, Ivan, Černocký, Jan, Watanabe, Shinji, Burget, Lukáš, Cornell, Samuele |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
por: Kocour, Martin, et al.
Publicado: (2025)
por: Kocour, Martin, et al.
Publicado: (2025)
Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics
por: Tawara, Naohiro, et al.
Publicado: (2026)
por: Tawara, Naohiro, et al.
Publicado: (2026)
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
por: Polok, Alexander, et al.
Publicado: (2026)
por: Polok, Alexander, et al.
Publicado: (2026)
Target Speaker ASR with Whisper
por: Polok, Alexander, et al.
Publicado: (2024)
por: Polok, Alexander, et al.
Publicado: (2024)
BUT System for the MLC-SLM Challenge
por: Polok, Alexander, et al.
Publicado: (2025)
por: Polok, Alexander, et al.
Publicado: (2025)
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
por: Polok, Alexander, et al.
Publicado: (2024)
por: Polok, Alexander, et al.
Publicado: (2024)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
por: He, Xiluo, et al.
Publicado: (2025)
por: He, Xiluo, et al.
Publicado: (2025)
Improving Automatic Speech Recognition with Decoder-Centric Regularisation in Encoder-Decoder Models
por: Polok, Alexander, et al.
Publicado: (2024)
por: Polok, Alexander, et al.
Publicado: (2024)
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
por: Han, Jiangyu, et al.
Publicado: (2025)
por: Han, Jiangyu, et al.
Publicado: (2025)
Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
por: Han, Jiangyu, et al.
Publicado: (2025)
por: Han, Jiangyu, et al.
Publicado: (2025)
DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
por: Polok, Alexander, et al.
Publicado: (2025)
por: Polok, Alexander, et al.
Publicado: (2025)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
por: Shakeel, Muhammad, et al.
Publicado: (2025)
por: Shakeel, Muhammad, et al.
Publicado: (2025)
Leveraging Self-Supervised Learning for Speaker Diarization
por: Han, Jiangyu, et al.
Publicado: (2024)
por: Han, Jiangyu, et al.
Publicado: (2024)
Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization
por: Han, Jiangyu, et al.
Publicado: (2025)
por: Han, Jiangyu, et al.
Publicado: (2025)
Modeling Overlapped Speech with Shuffles
por: Wiesner, Matthew, et al.
Publicado: (2026)
por: Wiesner, Matthew, et al.
Publicado: (2026)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
por: Cornell, Samuele, et al.
Publicado: (2024)
por: Cornell, Samuele, et al.
Publicado: (2024)
Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
por: Medennikov, Ivan, et al.
Publicado: (2025)
por: Medennikov, Ivan, et al.
Publicado: (2025)
CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification
por: Peng, Junyi, et al.
Publicado: (2024)
por: Peng, Junyi, et al.
Publicado: (2024)
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
por: Wang, Weiqing, et al.
Publicado: (2025)
por: Wang, Weiqing, et al.
Publicado: (2025)
MAPSS: Manifold-based Assessment of Perceptual Source Separation
por: Ivry, Amir, et al.
Publicado: (2025)
por: Ivry, Amir, et al.
Publicado: (2025)
On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization
por: Deegen, Marc, et al.
Publicado: (2026)
por: Deegen, Marc, et al.
Publicado: (2026)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
por: Pálka, Petr, et al.
Publicado: (2024)
por: Pálka, Petr, et al.
Publicado: (2024)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
por: Horiguchi, Shota, et al.
Publicado: (2025)
por: Horiguchi, Shota, et al.
Publicado: (2025)
Do End-to-End Neural Diarization Attractors Need to Encode Speaker Characteristic Information?
por: Zhang, Lin, et al.
Publicado: (2024)
por: Zhang, Lin, et al.
Publicado: (2024)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
por: Wang, Weiqing, et al.
Publicado: (2024)
por: Wang, Weiqing, et al.
Publicado: (2024)
SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
por: Fan, Zhiyun, et al.
Publicado: (2024)
por: Fan, Zhiyun, et al.
Publicado: (2024)
Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and Speech Recognition
por: Yang, Yufeng, et al.
Publicado: (2025)
por: Yang, Yufeng, et al.
Publicado: (2025)
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
Unsupervised Speech Enhancement using Data-defined Priors
por: Klement, Dominik, et al.
Publicado: (2025)
por: Klement, Dominik, et al.
Publicado: (2025)
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
por: Jin, Zengrui, et al.
Publicado: (2024)
por: Jin, Zengrui, et al.
Publicado: (2024)
Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
por: Xu, Anfeng, et al.
Publicado: (2024)
por: Xu, Anfeng, et al.
Publicado: (2024)
Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement
por: Li, Chenda, et al.
Publicado: (2024)
por: Li, Chenda, et al.
Publicado: (2024)
State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data
por: Barahona, Sara, et al.
Publicado: (2024)
por: Barahona, Sara, et al.
Publicado: (2024)
VBx for End-to-End Neural and Clustering-based Diarization
por: Pálka, Petr, et al.
Publicado: (2025)
por: Pálka, Petr, et al.
Publicado: (2025)
The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization
por: Cornell, Samuele, et al.
Publicado: (2024)
por: Cornell, Samuele, et al.
Publicado: (2024)
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR
por: Morrone, Giovanni, et al.
Publicado: (2024)
por: Morrone, Giovanni, et al.
Publicado: (2024)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
por: Guo, Pengcheng, et al.
Publicado: (2024)
por: Guo, Pengcheng, et al.
Publicado: (2024)
ASR-Synchronized Speaker-Role Diarization
por: Ghosh, Arindam, et al.
Publicado: (2025)
por: Ghosh, Arindam, et al.
Publicado: (2025)
Beyond the Labels: Unveiling Text-Dependency in Paralinguistic Speech Recognition Datasets
por: Pešán, Jan, et al.
Publicado: (2024)
por: Pešán, Jan, et al.
Publicado: (2024)
DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors
por: Landini, Federico, et al.
Publicado: (2023)
por: Landini, Federico, et al.
Publicado: (2023)
Ejemplares similares
-
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
por: Kocour, Martin, et al.
Publicado: (2025) -
Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics
por: Tawara, Naohiro, et al.
Publicado: (2026) -
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
por: Polok, Alexander, et al.
Publicado: (2026) -
Target Speaker ASR with Whisper
por: Polok, Alexander, et al.
Publicado: (2024) -
BUT System for the MLC-SLM Challenge
por: Polok, Alexander, et al.
Publicado: (2025)