LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909314476670976 |
|---|---|
| author | Li, Shaojun Shang, Hengchao Wei, Daimeng Guo, Jiaxin Li, Zongyao He, Xianghui Zhang, Min Yang, Hao |
| author_facet | Li, Shaojun Shang, Hengchao Wei, Daimeng Guo, Jiaxin Li, Zongyao He, Xianghui Zhang, Min Yang, Hao |
| contents | Recent advancements in integrating speech information into large language models (LLMs) have significantly improved automatic speech recognition (ASR) accuracy. However, existing methods often constrained by the capabilities of the speech encoders under varied acoustic conditions, such as accents. To address this, we propose LA-RAG, a novel Retrieval-Augmented Generation (RAG) paradigm for LLM-based ASR. LA-RAG leverages fine-grained token-level speech datastores and a speech-to-speech retrieval mechanism to enhance ASR accuracy via LLM in-context learning (ICL) capabilities. Experiments on Mandarin and various Chinese dialect datasets demonstrate significant improvements in ASR accuracy compared to existing methods, validating the effectiveness of our approach, especially in handling accent variations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_08597 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation Li, Shaojun Shang, Hengchao Wei, Daimeng Guo, Jiaxin Li, Zongyao He, Xianghui Zhang, Min Yang, Hao Sound Computation and Language Audio and Speech Processing Recent advancements in integrating speech information into large language models (LLMs) have significantly improved automatic speech recognition (ASR) accuracy. However, existing methods often constrained by the capabilities of the speech encoders under varied acoustic conditions, such as accents. To address this, we propose LA-RAG, a novel Retrieval-Augmented Generation (RAG) paradigm for LLM-based ASR. LA-RAG leverages fine-grained token-level speech datastores and a speech-to-speech retrieval mechanism to enhance ASR accuracy via LLM in-context learning (ICL) capabilities. Experiments on Mandarin and various Chinese dialect datasets demonstrate significant improvements in ASR accuracy compared to existing methods, validating the effectiveness of our approach, especially in handling accent variations. |
| title | LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation |
| topic | Sound Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2409.08597 |