Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Peng, Yizhou, Liu, Hexin, Chng, Eng Siong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909675454201856
author Peng, Yizhou
Liu, Hexin
Chng, Eng Siong
author_facet Peng, Yizhou
Liu, Hexin
Chng, Eng Siong
contents This paper introduces the integration of language-specific bi-directional context into a speech large language model (SLLM) to improve multilingual continuous conversational automatic speech recognition (ASR). We propose a character-level contextual masking strategy during training, which randomly removes portions of the context to enhance robustness and better emulate the flawed transcriptions that may occur during inference. For decoding, a two-stage pipeline is utilized: initial isolated segment decoding followed by context-aware re-decoding using neighboring hypotheses. Evaluated on the 1500-hour Multilingual Conversational Speech and Language Model (MLC-SLM) corpus covering eleven languages, our method achieves an 18% relative improvement compared to a strong baseline, outperforming even the model trained on 6000 hours of data for the MLC-SLM competition. These results underscore the significant benefit of incorporating contextual information in multilingual continuous conversational ASR.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13396
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR
Peng, Yizhou
Liu, Hexin
Chng, Eng Siong
Computation and Language
Audio and Speech Processing
This paper introduces the integration of language-specific bi-directional context into a speech large language model (SLLM) to improve multilingual continuous conversational automatic speech recognition (ASR). We propose a character-level contextual masking strategy during training, which randomly removes portions of the context to enhance robustness and better emulate the flawed transcriptions that may occur during inference. For decoding, a two-stage pipeline is utilized: initial isolated segment decoding followed by context-aware re-decoding using neighboring hypotheses. Evaluated on the 1500-hour Multilingual Conversational Speech and Language Model (MLC-SLM) corpus covering eleven languages, our method achieves an 18% relative improvement compared to a strong baseline, outperforming even the model trained on 6000 hours of data for the MLC-SLM competition. These results underscore the significant benefit of incorporating contextual information in multilingual continuous conversational ASR.
title Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2506.13396