Cross-utterance ASR Rescoring with Graph-based Label Propagation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tankasala, Srinath, Chen, Long, Stolcke, Andreas, Raju, Anirudh, Deng, Qianli, Chandak, Chander, Khare, Aparna, Maas, Roland, Ravichandran, Venkatesh
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909098147053568
author Tankasala, Srinath
Chen, Long
Stolcke, Andreas
Raju, Anirudh
Deng, Qianli
Chandak, Chander
Khare, Aparna
Maas, Roland
Ravichandran, Venkatesh
author_facet Tankasala, Srinath
Chen, Long
Stolcke, Andreas
Raju, Anirudh
Deng, Qianli
Chandak, Chander
Khare, Aparna
Maas, Roland
Ravichandran, Venkatesh
contents We propose a novel approach for ASR N-best hypothesis rescoring with graph-based label propagation by leveraging cross-utterance acoustic similarity. In contrast to conventional neural language model (LM) based ASR rescoring/reranking models, our approach focuses on acoustic information and conducts the rescoring collaboratively among utterances, instead of individually. Experiments on the VCTK dataset demonstrate that our approach consistently improves ASR performance, as well as fairness across speaker groups with different accents. Our approach provides a low-cost solution for mitigating the majoritarian bias of ASR systems, without the need to train new domain- or accent-specific models.
format Preprint
id arxiv_https___arxiv_org_abs_2303_15132
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Cross-utterance ASR Rescoring with Graph-based Label Propagation
Tankasala, Srinath
Chen, Long
Stolcke, Andreas
Raju, Anirudh
Deng, Qianli
Chandak, Chander
Khare, Aparna
Maas, Roland
Ravichandran, Venkatesh
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
We propose a novel approach for ASR N-best hypothesis rescoring with graph-based label propagation by leveraging cross-utterance acoustic similarity. In contrast to conventional neural language model (LM) based ASR rescoring/reranking models, our approach focuses on acoustic information and conducts the rescoring collaboratively among utterances, instead of individually. Experiments on the VCTK dataset demonstrate that our approach consistently improves ASR performance, as well as fairness across speaker groups with different accents. Our approach provides a low-cost solution for mitigating the majoritarian bias of ASR systems, without the need to train new domain- or accent-specific models.
title Cross-utterance ASR Rescoring with Graph-based Label Propagation
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2303.15132