DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Polok, Alexander, Klement, Dominik, Kocour, Martin, Han, Jiangyu, Landini, Federico, Yusuf, Bolaji, Wiesner, Matthew, Khudanpur, Sanjeev, Černocký, Jan, Burget, Lukáš |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
by: Polok, Alexander, et al.
Published: (2026)
by: Polok, Alexander, et al.
Published: (2026)
Target Speaker ASR with Whisper
by: Polok, Alexander, et al.
Published: (2024)
by: Polok, Alexander, et al.
Published: (2024)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
by: Kocour, Martin, et al.
Published: (2025)
by: Kocour, Martin, et al.
Published: (2025)
Unsupervised Speech Enhancement using Data-defined Priors
by: Klement, Dominik, et al.
Published: (2025)
by: Klement, Dominik, et al.
Published: (2025)
DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
by: Polok, Alexander, et al.
Published: (2025)
by: Polok, Alexander, et al.
Published: (2025)
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
by: Han, Jiangyu, et al.
Published: (2025)
by: Han, Jiangyu, et al.
Published: (2025)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
by: Pálka, Petr, et al.
Published: (2024)
by: Pálka, Petr, et al.
Published: (2024)
BUT System for the MLC-SLM Challenge
by: Polok, Alexander, et al.
Published: (2025)
by: Polok, Alexander, et al.
Published: (2025)
Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
by: Han, Jiangyu, et al.
Published: (2025)
by: Han, Jiangyu, et al.
Published: (2025)
Leveraging Self-Supervised Learning for Speaker Diarization
by: Han, Jiangyu, et al.
Published: (2024)
by: Han, Jiangyu, et al.
Published: (2024)
Improving Automatic Speech Recognition with Decoder-Centric Regularisation in Encoder-Decoder Models
by: Polok, Alexander, et al.
Published: (2024)
by: Polok, Alexander, et al.
Published: (2024)
Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization
by: Polok, Alexander, et al.
Published: (2026)
by: Polok, Alexander, et al.
Published: (2026)
Modeling Overlapped Speech with Shuffles
by: Wiesner, Matthew, et al.
Published: (2026)
by: Wiesner, Matthew, et al.
Published: (2026)
Do End-to-End Neural Diarization Attractors Need to Encode Speaker Characteristic Information?
by: Zhang, Lin, et al.
Published: (2024)
by: Zhang, Lin, et al.
Published: (2024)
DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors
by: Landini, Federico, et al.
Published: (2023)
by: Landini, Federico, et al.
Published: (2023)
From Modular to End-to-End Speaker Diarization
by: Landini, Federico
Published: (2024)
by: Landini, Federico
Published: (2024)
BUT System Description for CHiME-9 MCoRec Challenge
by: Klement, Dominik, et al.
Published: (2026)
by: Klement, Dominik, et al.
Published: (2026)
Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units
by: Yusuf, Bolaji, et al.
Published: (2024)
by: Yusuf, Bolaji, et al.
Published: (2024)
Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization
by: Han, Jiangyu, et al.
Published: (2025)
by: Han, Jiangyu, et al.
Published: (2025)
On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization
by: Deegen, Marc, et al.
Published: (2026)
by: Deegen, Marc, et al.
Published: (2026)
Beyond the Labels: Unveiling Text-Dependency in Paralinguistic Speech Recognition Datasets
by: Pešán, Jan, et al.
Published: (2024)
by: Pešán, Jan, et al.
Published: (2024)
VBx for End-to-End Neural and Clustering-based Diarization
by: Pálka, Petr, et al.
Published: (2025)
by: Pálka, Petr, et al.
Published: (2025)
WST: Weakly Supervised Transducer for Automatic Speech Recognition
by: Gao, Dongji, et al.
Published: (2025)
by: Gao, Dongji, et al.
Published: (2025)
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
by: Bhuiyan, Mohammed Aman, et al.
Published: (2026)
by: Bhuiyan, Mohammed Aman, et al.
Published: (2026)
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
by: Chowdhury, MD. Sagor, et al.
Published: (2026)
by: Chowdhury, MD. Sagor, et al.
Published: (2026)
Whisper-UT: A Unified Translation Framework for Speech and Text
by: Xiao, Cihan, et al.
Published: (2025)
by: Xiao, Cihan, et al.
Published: (2025)
On Speaker Attribution with SURT
by: Raj, Desh, et al.
Published: (2024)
by: Raj, Desh, et al.
Published: (2024)
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
by: Lin, Yuke, et al.
Published: (2025)
by: Lin, Yuke, et al.
Published: (2025)
HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation
by: Hussein, Amir, et al.
Published: (2025)
by: Hussein, Amir, et al.
Published: (2025)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
by: Vendrame, Katia, et al.
Published: (2025)
by: Vendrame, Katia, et al.
Published: (2025)
Prompting Whisper for Joint Speech Transcription and Diarization
by: Zamyrova, Mariia, et al.
Published: (2026)
by: Zamyrova, Mariia, et al.
Published: (2026)
Integrated Spoofing-Robust Automatic Speaker Verification via a Three-Class Formulation and LLR
by: Tan, Kai, et al.
Published: (2026)
by: Tan, Kai, et al.
Published: (2026)
Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
by: Sedláček, Šimon, et al.
Published: (2025)
by: Sedláček, Šimon, et al.
Published: (2025)
Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation
by: Huang, Ruizhe, et al.
Published: (2024)
by: Huang, Ruizhe, et al.
Published: (2024)
CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification
by: Peng, Junyi, et al.
Published: (2024)
by: Peng, Junyi, et al.
Published: (2024)
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
by: Shao, Yiwen, et al.
Published: (2024)
by: Shao, Yiwen, et al.
Published: (2024)
The Impact of Automatic Speech Transcription on Speaker Attribution
by: Aggazzotti, Cristina, et al.
Published: (2025)
by: Aggazzotti, Cristina, et al.
Published: (2025)
PhoWhisper: Automatic Speech Recognition for Vietnamese
by: Le, Thanh-Thien, et al.
Published: (2024)
by: Le, Thanh-Thien, et al.
Published: (2024)
The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization
by: Cornell, Samuele, et al.
Published: (2024)
by: Cornell, Samuele, et al.
Published: (2024)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
by: Dai, Yuhang, et al.
Published: (2025)
by: Dai, Yuhang, et al.
Published: (2025)
Similar Items
-
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
by: Polok, Alexander, et al.
Published: (2026) -
Target Speaker ASR with Whisper
by: Polok, Alexander, et al.
Published: (2024) -
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
by: Kocour, Martin, et al.
Published: (2025) -
Unsupervised Speech Enhancement using Data-defined Priors
by: Klement, Dominik, et al.
Published: (2025) -
DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
by: Polok, Alexander, et al.
Published: (2025)