Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cui, Mingyu, Geng, Mengzhe, Deng, Jiajun, Deng, Chengxi, Kang, Jiawen, Hu, Shujie, Li, Guinan, Wang, Tianzi, Li, Zhaoqing, Chen, Xie, Liu, Xunying
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915446494593024
author Cui, Mingyu
Geng, Mengzhe
Deng, Jiajun
Deng, Chengxi
Kang, Jiawen
Hu, Shujie
Li, Guinan
Wang, Tianzi
Li, Zhaoqing
Chen, Xie
Liu, Xunying
author_facet Cui, Mingyu
Geng, Mengzhe
Deng, Jiajun
Deng, Chengxi
Kang, Jiawen
Hu, Shujie
Li, Guinan
Wang, Tianzi
Li, Zhaoqing
Chen, Xie
Liu, Xunying
contents This paper investigates four types of cross-utterance speech contexts modeling approaches for streaming and non-streaming Conformer-Transformer (C-T) ASR systems: i) input audio feature concatenation; ii) cross-utterance Encoder embedding concatenation; iii) cross-utterance Encoder embedding pooling projection; or iv) a novel chunk-based approach applied to C-T models for the first time. An efficient batch-training scheme is proposed for contextual C-Ts that uses spliced speech utterances within each minibatch to minimize the synchronization overhead while preserving the sequential order of cross-utterance speech contexts. Experiments are conducted on four benchmark speech datasets across three languages: the English GigaSpeech and Mandarin Wenetspeech corpora used in contextual C-T models pre-training; and the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets used in domain fine-tuning. The best performing contextual C-T systems consistently outperform their respective baselines using no cross-utterance speech contexts in pre-training and fine-tuning stages with statistically significant average word error rate (WER) or character error rate (CER) reductions up to 0.9%, 1.1%, 0.51%, and 0.98% absolute (6.0%, 5.4%, 2.0%, and 3.4% relative) on the four tasks respectively. Their performance competitiveness against Wav2vec2.0-Conformer, XLSR-128, and Whisper models highlights the potential benefit of incorporating cross-utterance speech contexts into current speech foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10456
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
Cui, Mingyu
Geng, Mengzhe
Deng, Jiajun
Deng, Chengxi
Kang, Jiawen
Hu, Shujie
Li, Guinan
Wang, Tianzi
Li, Zhaoqing
Chen, Xie
Liu, Xunying
Audio and Speech Processing
This paper investigates four types of cross-utterance speech contexts modeling approaches for streaming and non-streaming Conformer-Transformer (C-T) ASR systems: i) input audio feature concatenation; ii) cross-utterance Encoder embedding concatenation; iii) cross-utterance Encoder embedding pooling projection; or iv) a novel chunk-based approach applied to C-T models for the first time. An efficient batch-training scheme is proposed for contextual C-Ts that uses spliced speech utterances within each minibatch to minimize the synchronization overhead while preserving the sequential order of cross-utterance speech contexts. Experiments are conducted on four benchmark speech datasets across three languages: the English GigaSpeech and Mandarin Wenetspeech corpora used in contextual C-T models pre-training; and the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets used in domain fine-tuning. The best performing contextual C-T systems consistently outperform their respective baselines using no cross-utterance speech contexts in pre-training and fine-tuning stages with statistically significant average word error rate (WER) or character error rate (CER) reductions up to 0.9%, 1.1%, 0.51%, and 0.98% absolute (6.0%, 5.4%, 2.0%, and 3.4% relative) on the four tasks respectively. Their performance competitiveness against Wav2vec2.0-Conformer, XLSR-128, and Whisper models highlights the potential benefit of incorporating cross-utterance speech contexts into current speech foundation models.
title Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
topic Audio and Speech Processing
url https://arxiv.org/abs/2508.10456