DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Li, Cheng, Ming, Zhu, Weixin, Wang, Yannan, Liu, Juan, Li, Ming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
by: Lin, Yuke, et al.
Published: (2025)
by: Lin, Yuke, et al.
Published: (2025)
Spatially-Augmented Sequence-to-Sequence Neural Diarization for Meetings
by: Li, Li, et al.
Published: (2025)
by: Li, Li, et al.
Published: (2025)
Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge
by: Cheng, Ming, et al.
Published: (2025)
by: Cheng, Ming, et al.
Published: (2025)
Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization
by: Cheng, Ming, et al.
Published: (2024)
by: Cheng, Ming, et al.
Published: (2024)
Speech Diarization and ASR with GMM
by: Sharma, Aayush Kumar, et al.
Published: (2023)
by: Sharma, Aayush Kumar, et al.
Published: (2023)
Extending Whisper with prompt tuning to target-speaker ASR
by: Ma, Hao, et al.
Published: (2023)
by: Ma, Hao, et al.
Published: (2023)
Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization
by: Polok, Alexander, et al.
Published: (2026)
by: Polok, Alexander, et al.
Published: (2026)
Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation
by: Cheng, Ming, et al.
Published: (2024)
by: Cheng, Ming, et al.
Published: (2024)
Leveraging ASR Pretrained Conformers for Speaker Verification through Transfer Learning and Knowledge Distillation
by: Cai, Danwei, et al.
Published: (2023)
by: Cai, Danwei, et al.
Published: (2023)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
by: Shakeel, Muhammad, et al.
Published: (2025)
by: Shakeel, Muhammad, et al.
Published: (2025)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
by: Horiguchi, Shota, et al.
Published: (2025)
by: Horiguchi, Shota, et al.
Published: (2025)
Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction
by: Baroudi, Séverin, et al.
Published: (2026)
by: Baroudi, Séverin, et al.
Published: (2026)
Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio
by: Shi, Mohan, et al.
Published: (2025)
by: Shi, Mohan, et al.
Published: (2025)
ASR-Synchronized Speaker-Role Diarization
by: Ghosh, Arindam, et al.
Published: (2025)
by: Ghosh, Arindam, et al.
Published: (2025)
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
by: Xu, Anfeng, et al.
Published: (2026)
by: Xu, Anfeng, et al.
Published: (2026)
Bengali-Loop: Community Benchmarks for Long-Form Bangla ASR and Speaker Diarization
by: Tabib, H. M. Shadman, et al.
Published: (2026)
by: Tabib, H. M. Shadman, et al.
Published: (2026)
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR
by: Morrone, Giovanni, et al.
Published: (2024)
by: Morrone, Giovanni, et al.
Published: (2024)
Reverb: Open-Source ASR and Diarization from Rev
by: Bhandari, Nishchal, et al.
Published: (2024)
by: Bhandari, Nishchal, et al.
Published: (2024)
Advancing Multi-talker ASR Performance with Large Language Models
by: Shi, Mohan, et al.
Published: (2024)
by: Shi, Mohan, et al.
Published: (2024)
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
by: Wei, Victor Junqiu, et al.
Published: (2024)
by: Wei, Victor Junqiu, et al.
Published: (2024)
TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding
by: Huo, Mingyue, et al.
Published: (2026)
by: Huo, Mingyue, et al.
Published: (2026)
Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses
by: Yang, Yufeng, et al.
Published: (2025)
by: Yang, Yufeng, et al.
Published: (2025)
Speech Emotion Recognition with ASR Integration
by: Li, Yuanchao
Published: (2026)
by: Li, Yuanchao
Published: (2026)
Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge
by: Huang, Shangkun, et al.
Published: (2025)
by: Huang, Shangkun, et al.
Published: (2025)
Semi-supervised Learning for Code-Switching ASR with Large Language Model Filter
by: Xi, Yu, et al.
Published: (2024)
by: Xi, Yu, et al.
Published: (2024)
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
by: Guan, Wenhao, et al.
Published: (2025)
by: Guan, Wenhao, et al.
Published: (2025)
LibriConvo: Simulating Conversations from Read Literature for ASR and Diarization
by: Gedeon, Máté, et al.
Published: (2025)
by: Gedeon, Máté, et al.
Published: (2025)
Phonemes vs. Projectors: An Investigation of Speech-Language Interfaces for LLM-based ASR
by: Li, Ziwei, et al.
Published: (2026)
by: Li, Ziwei, et al.
Published: (2026)
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
by: Moriya, Takafumi, et al.
Published: (2025)
by: Moriya, Takafumi, et al.
Published: (2025)
PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings
by: Kalda, Joonas, et al.
Published: (2024)
by: Kalda, Joonas, et al.
Published: (2024)
MDM-ASR: Bridging Accuracy and Efficiency in ASR with Diffusion-Based Non-Autoregressive Decoding
by: Yen, Hao, et al.
Published: (2026)
by: Yen, Hao, et al.
Published: (2026)
Transferable Adversarial Attacks against ASR
by: Gao, Xiaoxue, et al.
Published: (2024)
by: Gao, Xiaoxue, et al.
Published: (2024)
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
by: Saon, George, et al.
Published: (2025)
by: Saon, George, et al.
Published: (2025)
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
by: Gong, Xun, et al.
Published: (2025)
by: Gong, Xun, et al.
Published: (2025)
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
by: Kumar, Shashi, et al.
Published: (2024)
by: Kumar, Shashi, et al.
Published: (2024)
ASR Error Correction using Large Language Models
by: Ma, Rao, et al.
Published: (2024)
by: Ma, Rao, et al.
Published: (2024)
Complexity boosted adaptive training for better low resource ASR performance
by: Lu, Hongxuan, et al.
Published: (2024)
by: Lu, Hongxuan, et al.
Published: (2024)
Large Language Model Should Understand Pinyin for Chinese ASR Error Correction
by: Li, Yuang, et al.
Published: (2024)
by: Li, Yuang, et al.
Published: (2024)
Mel-FullSubNet: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
by: Zhou, Rui, et al.
Published: (2024)
by: Zhou, Rui, et al.
Published: (2024)
Large Language Models based ASR Error Correction for Child Conversations
by: Xu, Anfeng, et al.
Published: (2025)
by: Xu, Anfeng, et al.
Published: (2025)
Similar Items
-
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
by: Lin, Yuke, et al.
Published: (2025) -
Spatially-Augmented Sequence-to-Sequence Neural Diarization for Meetings
by: Li, Li, et al.
Published: (2025) -
Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge
by: Cheng, Ming, et al.
Published: (2025) -
Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization
by: Cheng, Ming, et al.
Published: (2024) -
Speech Diarization and ASR with GMM
by: Sharma, Aayush Kumar, et al.
Published: (2023)