Improving Speaker Assignment in Speaker-Attributed ASR for Real Meeting Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Can, Sheikh, Imran Ahamad, Sadeghi, Mostafa, Vincent, Emmanuel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Joint Beamforming and Speaker-Attributed ASR for Real Distant-Microphone Meeting Transcription
by: Cui, Can, et al.
Published: (2024)
by: Cui, Can, et al.
Published: (2024)
End-to-end Joint Punctuated and Normalized ASR with a Limited Amount of Punctuated Training Data
by: Cui, Can, et al.
Published: (2023)
by: Cui, Can, et al.
Published: (2023)
Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR
by: Lin, Zhennan, et al.
Published: (2026)
by: Lin, Zhennan, et al.
Published: (2026)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
by: Nguyen, Thai-Binh, et al.
Published: (2024)
by: Nguyen, Thai-Binh, et al.
Published: (2024)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
by: Shakeel, Muhammad, et al.
Published: (2025)
by: Shakeel, Muhammad, et al.
Published: (2025)
The Impact of Automatic Speech Transcription on Speaker Attribution
by: Aggazzotti, Cristina, et al.
Published: (2025)
by: Aggazzotti, Cristina, et al.
Published: (2025)
Towards Fair ASR For Second Language Speakers Using Fairness Prompted Finetuning
by: Swain, Monorama, et al.
Published: (2025)
by: Swain, Monorama, et al.
Published: (2025)
Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
by: Tomashenko, Natalia, et al.
Published: (2024)
by: Tomashenko, Natalia, et al.
Published: (2024)
Data-independent Beamforming for End-to-end Multichannel Multi-speaker ASR
by: Cui, Can, et al.
Published: (2025)
by: Cui, Can, et al.
Published: (2025)
Can Authorship Attribution Models Distinguish Speakers in Speech Transcripts?
by: Aggazzotti, Cristina, et al.
Published: (2023)
by: Aggazzotti, Cristina, et al.
Published: (2023)
CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
by: Shakeel, Muhammad, et al.
Published: (2026)
by: Shakeel, Muhammad, et al.
Published: (2026)
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR
by: Morrone, Giovanni, et al.
Published: (2024)
by: Morrone, Giovanni, et al.
Published: (2024)
Who Gets the Mic? Investigating Gender Bias in the Speaker Assignment of a Speech-LLM
by: Puhach, Dariia, et al.
Published: (2025)
by: Puhach, Dariia, et al.
Published: (2025)
TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding
by: Huo, Mingyue, et al.
Published: (2026)
by: Huo, Mingyue, et al.
Published: (2026)
Investigation of Speaker Representation for Target-Speaker Speech Processing
by: Ashihara, Takanori, et al.
Published: (2024)
by: Ashihara, Takanori, et al.
Published: (2024)
WhoSaidIt: Human-LLM Collaborative Annotation for Text-Based Multilingual Speaker-Attribute Classification
by: Gao, Lingyu, et al.
Published: (2026)
by: Gao, Lingyu, et al.
Published: (2026)
M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models
by: Kwon, Yejin, et al.
Published: (2025)
by: Kwon, Yejin, et al.
Published: (2025)
Speaker-Aware Simulation Improves Conversational Speech Recognition
by: Gedeon, Máté, et al.
Published: (2026)
by: Gedeon, Máté, et al.
Published: (2026)
LaERC-S: Improving LLM-based Emotion Recognition in Conversation with Speaker Characteristics
by: Fu, Yumeng, et al.
Published: (2024)
by: Fu, Yumeng, et al.
Published: (2024)
An Exhaustive Evaluation of TTS- and VC-based Data Augmentation for ASR
by: Ogun, Sewade, et al.
Published: (2025)
by: Ogun, Sewade, et al.
Published: (2025)
Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
by: Yong, Zheng-Xin, et al.
Published: (2025)
by: Yong, Zheng-Xin, et al.
Published: (2025)
Speaker Verification in Agent-Generated Conversations
by: Yang, Yizhe, et al.
Published: (2024)
by: Yang, Yizhe, et al.
Published: (2024)
SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?
by: Lee, Jonggeun, et al.
Published: (2026)
by: Lee, Jonggeun, et al.
Published: (2026)
Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
by: Wu, Junkai, et al.
Published: (2024)
by: Wu, Junkai, et al.
Published: (2024)
Speaker Embeddings to Improve Tracking of Intermittent and Moving Speakers
by: Iatariene, Taous, et al.
Published: (2025)
by: Iatariene, Taous, et al.
Published: (2025)
Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
by: Carbonneau, Marc-André, et al.
Published: (2025)
by: Carbonneau, Marc-André, et al.
Published: (2025)
Speakers Fill Lexical Semantic Gaps with Context
by: Pimentel, Tiago, et al.
Published: (2020)
by: Pimentel, Tiago, et al.
Published: (2020)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
by: Wang, Hsuan-Yu, et al.
Published: (2025)
by: Wang, Hsuan-Yu, et al.
Published: (2025)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
by: Sakuma, Asahi, et al.
Published: (2025)
by: Sakuma, Asahi, et al.
Published: (2025)
Identifying Speakers and Addressees of Quotations in Novels with Prompt Learning
by: Yan, Yuchen, et al.
Published: (2024)
by: Yan, Yuchen, et al.
Published: (2024)
Designing Explainable Conversational Agentic Systems for Guaraní Speakers
by: Adorno, Samantha, et al.
Published: (2026)
by: Adorno, Samantha, et al.
Published: (2026)
Toward Responsible ASR for African American English Speakers: A Scoping Review of Bias and Equity in Speech Technology
by: Cunningham, Jay L., et al.
Published: (2025)
by: Cunningham, Jay L., et al.
Published: (2025)
Speaker Style-Aware Phoneme Anchoring for Improved Cross-Lingual Speech Emotion Recognition
by: Upadhyay, Shreya G., et al.
Published: (2025)
by: Upadhyay, Shreya G., et al.
Published: (2025)
Learning Speaker-Invariant Visual Features for Lipreading
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
Communicating with Speakers and Listeners of Different Pragmatic Levels
by: Naszadi, Kata, et al.
Published: (2024)
by: Naszadi, Kata, et al.
Published: (2024)
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation
by: Cheng, Luyao, et al.
Published: (2023)
by: Cheng, Luyao, et al.
Published: (2023)
SPECTRUM: Speaker-Enhanced Pre-Training for Long Dialogue Summarization
by: Cho, Sangwoo, et al.
Published: (2024)
by: Cho, Sangwoo, et al.
Published: (2024)
Large Language Models Discriminate Against Speakers of German Dialects
by: Bui, Minh Duc, et al.
Published: (2025)
by: Bui, Minh Duc, et al.
Published: (2025)
Probing the Feasibility of Multilingual Speaker Anonymization
by: Meyer, Sarina, et al.
Published: (2024)
by: Meyer, Sarina, et al.
Published: (2024)
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
by: Jin, Zengrui, et al.
Published: (2024)
by: Jin, Zengrui, et al.
Published: (2024)
Similar Items
-
Joint Beamforming and Speaker-Attributed ASR for Real Distant-Microphone Meeting Transcription
by: Cui, Can, et al.
Published: (2024) -
End-to-end Joint Punctuated and Normalized ASR with a Limited Amount of Punctuated Training Data
by: Cui, Can, et al.
Published: (2023) -
Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR
by: Lin, Zhennan, et al.
Published: (2026) -
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
by: Nguyen, Thai-Binh, et al.
Published: (2024) -
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
by: Shakeel, Muhammad, et al.
Published: (2025)