Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lin, Yuke, Cheng, Ming, Li, Ze, Tang, Beilong, Li, Ming
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909640995897344
author Lin, Yuke
Cheng, Ming
Li, Ze
Tang, Beilong
Li, Ming
author_facet Lin, Yuke
Cheng, Ming
Li, Ze
Tang, Beilong
Li, Ming
contents Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and conversational analysis. While serialized output training (SOT)-style methods serve as common solutions, they often discard absolute timing information, limiting their utility in time-sensitive scenarios. Leveraging recent advances in large language models (LLMs) for conversational audio processing, we propose a novel diarization-aware multi-speaker ASR system that integrates speaker diarization with LLM-based transcription. Our framework processes structured diarization inputs alongside frame-level speaker and semantic embeddings, enabling the LLM to generate segment-level transcriptions. Experiments demonstrate that the system achieves robust performance in multilingual dyadic conversations and excels in complex, high-overlap multi-speaker meeting scenarios. This work highlights the potential of LLMs as unified back-ends for joint speaker-aware segmentation and transcription.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05796
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
Lin, Yuke
Cheng, Ming
Li, Ze
Tang, Beilong
Li, Ming
Audio and Speech Processing
Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and conversational analysis. While serialized output training (SOT)-style methods serve as common solutions, they often discard absolute timing information, limiting their utility in time-sensitive scenarios. Leveraging recent advances in large language models (LLMs) for conversational audio processing, we propose a novel diarization-aware multi-speaker ASR system that integrates speaker diarization with LLM-based transcription. Our framework processes structured diarization inputs alongside frame-level speaker and semantic embeddings, enabling the LLM to generate segment-level transcriptions. Experiments demonstrate that the system achieves robust performance in multilingual dyadic conversations and excels in complex, high-overlap multi-speaker meeting scenarios. This work highlights the potential of LLMs as unified back-ends for joint speaker-aware segmentation and transcription.
title Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
topic Audio and Speech Processing
url https://arxiv.org/abs/2506.05796