Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915446494593024 |
|---|---|
| author | Cui, Mingyu Geng, Mengzhe Deng, Jiajun Deng, Chengxi Kang, Jiawen Hu, Shujie Li, Guinan Wang, Tianzi Li, Zhaoqing Chen, Xie Liu, Xunying |
| author_facet | Cui, Mingyu Geng, Mengzhe Deng, Jiajun Deng, Chengxi Kang, Jiawen Hu, Shujie Li, Guinan Wang, Tianzi Li, Zhaoqing Chen, Xie Liu, Xunying |
| contents | This paper investigates four types of cross-utterance speech contexts modeling approaches for streaming and non-streaming Conformer-Transformer (C-T) ASR systems: i) input audio feature concatenation; ii) cross-utterance Encoder embedding concatenation; iii) cross-utterance Encoder embedding pooling projection; or iv) a novel chunk-based approach applied to C-T models for the first time. An efficient batch-training scheme is proposed for contextual C-Ts that uses spliced speech utterances within each minibatch to minimize the synchronization overhead while preserving the sequential order of cross-utterance speech contexts. Experiments are conducted on four benchmark speech datasets across three languages: the English GigaSpeech and Mandarin Wenetspeech corpora used in contextual C-T models pre-training; and the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets used in domain fine-tuning. The best performing contextual C-T systems consistently outperform their respective baselines using no cross-utterance speech contexts in pre-training and fine-tuning stages with statistically significant average word error rate (WER) or character error rate (CER) reductions up to 0.9%, 1.1%, 0.51%, and 0.98% absolute (6.0%, 5.4%, 2.0%, and 3.4% relative) on the four tasks respectively. Their performance competitiveness against Wav2vec2.0-Conformer, XLSR-128, and Whisper models highlights the potential benefit of incorporating cross-utterance speech contexts into current speech foundation models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_10456 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems Cui, Mingyu Geng, Mengzhe Deng, Jiajun Deng, Chengxi Kang, Jiawen Hu, Shujie Li, Guinan Wang, Tianzi Li, Zhaoqing Chen, Xie Liu, Xunying Audio and Speech Processing This paper investigates four types of cross-utterance speech contexts modeling approaches for streaming and non-streaming Conformer-Transformer (C-T) ASR systems: i) input audio feature concatenation; ii) cross-utterance Encoder embedding concatenation; iii) cross-utterance Encoder embedding pooling projection; or iv) a novel chunk-based approach applied to C-T models for the first time. An efficient batch-training scheme is proposed for contextual C-Ts that uses spliced speech utterances within each minibatch to minimize the synchronization overhead while preserving the sequential order of cross-utterance speech contexts. Experiments are conducted on four benchmark speech datasets across three languages: the English GigaSpeech and Mandarin Wenetspeech corpora used in contextual C-T models pre-training; and the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets used in domain fine-tuning. The best performing contextual C-T systems consistently outperform their respective baselines using no cross-utterance speech contexts in pre-training and fine-tuning stages with statistically significant average word error rate (WER) or character error rate (CER) reductions up to 0.9%, 1.1%, 0.51%, and 0.98% absolute (6.0%, 5.4%, 2.0%, and 3.4% relative) on the four tasks respectively. Their performance competitiveness against Wav2vec2.0-Conformer, XLSR-128, and Whisper models highlights the potential benefit of incorporating cross-utterance speech contexts into current speech foundation models. |
| title | Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2508.10456 |