RAG-Boost: Retrieval-Augmented Generation Enhanced LLM-based Speech Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Pengcheng, Li, Sheng, Shinozaki, Takahiro |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
di: Saengthong, Phurich, et al.
Pubblicazione: (2025)
di: Saengthong, Phurich, et al.
Pubblicazione: (2025)
LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
di: Li, Shaojun, et al.
Pubblicazione: (2024)
di: Li, Shaojun, et al.
Pubblicazione: (2024)
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
di: Ko, Yuka, et al.
Pubblicazione: (2024)
di: Ko, Yuka, et al.
Pubblicazione: (2024)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
di: Min, Do June, et al.
Pubblicazione: (2024)
di: Min, Do June, et al.
Pubblicazione: (2024)
Retrieval-Augmented Speech Recognition Approach for Domain Challenges
di: Shen, Peng, et al.
Pubblicazione: (2025)
di: Shen, Peng, et al.
Pubblicazione: (2025)
Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
di: Komatsu, Ryota, et al.
Pubblicazione: (2024)
di: Komatsu, Ryota, et al.
Pubblicazione: (2024)
Data Augmentation for End-to-end Code-switching Speech Recognition
di: Du, Chenpeng, et al.
Pubblicazione: (2020)
di: Du, Chenpeng, et al.
Pubblicazione: (2020)
MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
di: Chien, Chung-Ming, et al.
Pubblicazione: (2026)
di: Chien, Chung-Ming, et al.
Pubblicazione: (2026)
Retrieval Augmented Generation based context discovery for ASR
di: Siskos, Dimitrios, et al.
Pubblicazione: (2025)
di: Siskos, Dimitrios, et al.
Pubblicazione: (2025)
Optimizing Contextual Speech Recognition Using Vector Quantization for Efficient Retrieval
di: Flemotomos, Nikolaos, et al.
Pubblicazione: (2024)
di: Flemotomos, Nikolaos, et al.
Pubblicazione: (2024)
Revisiting Interpolation Augmentation for Speech-to-Text Generation
di: Xu, Chen, et al.
Pubblicazione: (2024)
di: Xu, Chen, et al.
Pubblicazione: (2024)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)
SEAL: Speech Embedding Alignment Learning for Speech Large Language Model with Retrieval-Augmented Generation
di: Sun, Chunyu, et al.
Pubblicazione: (2025)
di: Sun, Chunyu, et al.
Pubblicazione: (2025)
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
di: Hu, Jiliang, et al.
Pubblicazione: (2025)
di: Hu, Jiliang, et al.
Pubblicazione: (2025)
Exploring Generative Error Correction for Dysarthric Speech Recognition
di: La Quatra, Moreno, et al.
Pubblicazione: (2025)
di: La Quatra, Moreno, et al.
Pubblicazione: (2025)
Unimodal Aggregation for CTC-based Speech Recognition
di: Fang, Ying, et al.
Pubblicazione: (2023)
di: Fang, Ying, et al.
Pubblicazione: (2023)
Streaming Speech-to-Confusion Network Speech Recognition
di: Filimonov, Denis, et al.
Pubblicazione: (2023)
di: Filimonov, Denis, et al.
Pubblicazione: (2023)
MTLM: Incorporating Bidirectional Text Information to Enhance Language Model Training in Speech Recognition Systems
di: Meng, Qingliang, et al.
Pubblicazione: (2025)
di: Meng, Qingliang, et al.
Pubblicazione: (2025)
CO-VADA: A Confidence-Oriented Voice Augmentation Debiasing Approach for Fair Speech Emotion Recognition
di: Tsai, Yun-Shao, et al.
Pubblicazione: (2025)
di: Tsai, Yun-Shao, et al.
Pubblicazione: (2025)
Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
di: Wang, Guansu, et al.
Pubblicazione: (2025)
di: Wang, Guansu, et al.
Pubblicazione: (2025)
SpeechColab Leaderboard: An Open-Source Platform for Automatic Speech Recognition Evaluation
di: Du, Jiayu, et al.
Pubblicazione: (2024)
di: Du, Jiayu, et al.
Pubblicazione: (2024)
PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech Recognition
di: Fu, Li, et al.
Pubblicazione: (2025)
di: Fu, Li, et al.
Pubblicazione: (2025)
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition
di: Noroozi, Vahid, et al.
Pubblicazione: (2023)
di: Noroozi, Vahid, et al.
Pubblicazione: (2023)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
di: Sudo, Yui, et al.
Pubblicazione: (2024)
di: Sudo, Yui, et al.
Pubblicazione: (2024)
Longer is (Not Necessarily) Stronger: Punctuated Long-Sequence Training for Enhanced Speech Recognition and Translation
di: Koluguri, Nithin Rao, et al.
Pubblicazione: (2024)
di: Koluguri, Nithin Rao, et al.
Pubblicazione: (2024)
Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model
di: Huang, Hukai, et al.
Pubblicazione: (2024)
di: Huang, Hukai, et al.
Pubblicazione: (2024)
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
di: Nozawa, Kento, et al.
Pubblicazione: (2024)
di: Nozawa, Kento, et al.
Pubblicazione: (2024)
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
di: Roll, Nathan, et al.
Pubblicazione: (2025)
di: Roll, Nathan, et al.
Pubblicazione: (2025)
Dialectal Coverage And Generalization in Arabic Speech Recognition
di: Djanibekov, Amirbek, et al.
Pubblicazione: (2024)
di: Djanibekov, Amirbek, et al.
Pubblicazione: (2024)
Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
di: Xu, Tianyi, et al.
Pubblicazione: (2025)
di: Xu, Tianyi, et al.
Pubblicazione: (2025)
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
di: Adila, Aulia, et al.
Pubblicazione: (2024)
di: Adila, Aulia, et al.
Pubblicazione: (2024)
Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition
di: Yang, Zhengdong, et al.
Pubblicazione: (2025)
di: Yang, Zhengdong, et al.
Pubblicazione: (2025)
Dolphin: A Large-Scale Automatic Speech Recognition Model for Eastern Languages
di: Meng, Yangyang, et al.
Pubblicazione: (2025)
di: Meng, Yangyang, et al.
Pubblicazione: (2025)
Triple X: A LLM-Based Multilingual Speech Recognition System for the INTERSPEECH2025 MLC-SLM Challenge
di: Gao, Miaomiao, et al.
Pubblicazione: (2025)
di: Gao, Miaomiao, et al.
Pubblicazione: (2025)
PhoWhisper: Automatic Speech Recognition for Vietnamese
di: Le, Thanh-Thien, et al.
Pubblicazione: (2024)
di: Le, Thanh-Thien, et al.
Pubblicazione: (2024)
Zero-resource Speech Translation and Recognition with LLMs
di: Mundnich, Karel, et al.
Pubblicazione: (2024)
di: Mundnich, Karel, et al.
Pubblicazione: (2024)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation
di: Wang, Shiyao, et al.
Pubblicazione: (2024)
di: Wang, Shiyao, et al.
Pubblicazione: (2024)
Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition
di: Moritz, Niko, et al.
Pubblicazione: (2024)
di: Moritz, Niko, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
di: Saengthong, Phurich, et al.
Pubblicazione: (2025) -
LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
di: Li, Shaojun, et al.
Pubblicazione: (2024) -
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
di: Ko, Yuka, et al.
Pubblicazione: (2024) -
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
di: Min, Do June, et al.
Pubblicazione: (2024) -
Retrieval-Augmented Speech Recognition Approach for Domain Challenges
di: Shen, Peng, et al.
Pubblicazione: (2025)