Speech Retrieval-Augmented Generation without Automatic Speech Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Min, Do June, Mundnich, Karel, Lapastora, Andy, Soltanmohammadi, Erfan, Ronanki, Srikanth, Han, Kyu
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913633963868160
author Min, Do June
Mundnich, Karel
Lapastora, Andy
Soltanmohammadi, Erfan
Ronanki, Srikanth
Han, Kyu
author_facet Min, Do June
Mundnich, Karel
Lapastora, Andy
Soltanmohammadi, Erfan
Ronanki, Srikanth
Han, Kyu
contents One common approach for question answering over speech data is to first transcribe speech using automatic speech recognition (ASR) and then employ text-based retrieval-augmented generation (RAG) on the transcriptions. While this cascaded pipeline has proven effective in many practical settings, ASR errors can propagate to the retrieval and generation steps. To overcome this limitation, we introduce SpeechRAG, a novel framework designed for open-question answering over spoken data. Our proposed approach fine-tunes a pre-trained speech encoder into a speech adapter fed into a frozen large language model (LLM)--based retrieval model. By aligning the embedding spaces of text and speech, our speech retriever directly retrieves audio passages from text-based queries, leveraging the retrieval capacity of the frozen text retriever. Our retrieval experiments on spoken question answering datasets show that direct speech retrieval does not degrade over the text-based baseline, and outperforms the cascaded systems using ASR. For generation, we use a speech language model (SLM) as a generator, conditioned on audio passages rather than transcripts. Without fine-tuning of the SLM, this approach outperforms cascaded text-based models when there is high WER in the transcripts.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16500
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Speech Retrieval-Augmented Generation without Automatic Speech Recognition
Min, Do June
Mundnich, Karel
Lapastora, Andy
Soltanmohammadi, Erfan
Ronanki, Srikanth
Han, Kyu
Audio and Speech Processing
Artificial Intelligence
Computation and Language
One common approach for question answering over speech data is to first transcribe speech using automatic speech recognition (ASR) and then employ text-based retrieval-augmented generation (RAG) on the transcriptions. While this cascaded pipeline has proven effective in many practical settings, ASR errors can propagate to the retrieval and generation steps. To overcome this limitation, we introduce SpeechRAG, a novel framework designed for open-question answering over spoken data. Our proposed approach fine-tunes a pre-trained speech encoder into a speech adapter fed into a frozen large language model (LLM)--based retrieval model. By aligning the embedding spaces of text and speech, our speech retriever directly retrieves audio passages from text-based queries, leveraging the retrieval capacity of the frozen text retriever. Our retrieval experiments on spoken question answering datasets show that direct speech retrieval does not degrade over the text-based baseline, and outperforms the cascaded systems using ASR. For generation, we use a speech language model (SLM) as a generator, conditioned on audio passages rather than transcripts. Without fine-tuning of the SLM, this approach outperforms cascaded text-based models when there is high WER in the transcripts.
title Speech Retrieval-Augmented Generation without Automatic Speech Recognition
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2412.16500