kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866910317433323520 |
|---|---|
| author | Zhou, Jiaming Zhao, Shiwan Liu, Yaqi Zeng, Wenjia Chen, Yong Qin, Yong |
| author_facet | Zhou, Jiaming Zhao, Shiwan Liu, Yaqi Zeng, Wenjia Chen, Yong Qin, Yong |
| contents | The success of retrieval-augmented language models in various natural language processing (NLP) tasks has been constrained in automatic speech recognition (ASR) applications due to challenges in constructing fine-grained audio-text datastores. This paper presents kNN-CTC, a novel approach that overcomes these challenges by leveraging Connectionist Temporal Classification (CTC) pseudo labels to establish frame-level audio-text key-value pairs, circumventing the need for precise ground truth alignments. We further introduce a skip-blank strategy, which strategically ignores CTC blank frames, to reduce datastore size. kNN-CTC incorporates a k-nearest neighbors retrieval mechanism into pre-trained CTC ASR systems, achieving significant improvements in performance. By incorporating a k-nearest neighbors retrieval mechanism into pre-trained CTC ASR systems and leveraging a fine-grained, pruned datastore, kNN-CTC consistently achieves substantial improvements in performance under various experimental settings. Our code is available at https://github.com/NKU-HLT/KNN-CTC. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2312_13560 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels Zhou, Jiaming Zhao, Shiwan Liu, Yaqi Zeng, Wenjia Chen, Yong Qin, Yong Sound Audio and Speech Processing The success of retrieval-augmented language models in various natural language processing (NLP) tasks has been constrained in automatic speech recognition (ASR) applications due to challenges in constructing fine-grained audio-text datastores. This paper presents kNN-CTC, a novel approach that overcomes these challenges by leveraging Connectionist Temporal Classification (CTC) pseudo labels to establish frame-level audio-text key-value pairs, circumventing the need for precise ground truth alignments. We further introduce a skip-blank strategy, which strategically ignores CTC blank frames, to reduce datastore size. kNN-CTC incorporates a k-nearest neighbors retrieval mechanism into pre-trained CTC ASR systems, achieving significant improvements in performance. By incorporating a k-nearest neighbors retrieval mechanism into pre-trained CTC ASR systems and leveraging a fine-grained, pruned datastore, kNN-CTC consistently achieves substantial improvements in performance under various experimental settings. Our code is available at https://github.com/NKU-HLT/KNN-CTC. |
| title | kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2312.13560 |