Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Pengchao, Ma, Ziyang, Chen, Wenxi, Li, Yao, Wang, Sheng, Yu, Kai, Chen, Xie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909897228025856
author Feng, Pengchao
Ma, Ziyang
Chen, Wenxi
Li, Yao
Wang, Sheng
Yu, Kai
Chen, Xie
author_facet Feng, Pengchao
Ma, Ziyang
Chen, Wenxi
Li, Yao
Wang, Sheng
Yu, Kai
Chen, Xie
contents End-to-end speech-to-speech (S2S) dialogue systems have recently garnered increasing research attention for their lower latency and more natural integration of nonverbal cues such as emotion and speaker identity. However, these systems face key challenges, particularly in incorporating external knowledge, a capability commonly addressed by Retrieval-Augmented Generation (RAG) in text-based large language models (LLMs). The core difficulty lies in the modality gap between input speech and retrieved textual knowledge, which hinders effective integration of information. To address this issue, we propose a novel end-to-end RAG framework that directly retrieves relevant textual knowledge from speech queries. Experimental results demonstrate that our method significantly improves the performance of end-to-end S2S dialogue systems while achieving higher retrieval efficiency. Although the overall performance still lags behind the SOTA cascaded models, our framework offers a promising direction for enhancing knowledge integration in end-to-end S2S systems. Our code and dataset are released.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00028
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
Feng, Pengchao
Ma, Ziyang
Chen, Wenxi
Li, Yao
Wang, Sheng
Yu, Kai
Chen, Xie
Computation and Language
Artificial Intelligence
Information Retrieval
End-to-end speech-to-speech (S2S) dialogue systems have recently garnered increasing research attention for their lower latency and more natural integration of nonverbal cues such as emotion and speaker identity. However, these systems face key challenges, particularly in incorporating external knowledge, a capability commonly addressed by Retrieval-Augmented Generation (RAG) in text-based large language models (LLMs). The core difficulty lies in the modality gap between input speech and retrieved textual knowledge, which hinders effective integration of information. To address this issue, we propose a novel end-to-end RAG framework that directly retrieves relevant textual knowledge from speech queries. Experimental results demonstrate that our method significantly improves the performance of end-to-end S2S dialogue systems while achieving higher retrieval efficiency. Although the overall performance still lags behind the SOTA cascaded models, our framework offers a promising direction for enhancing knowledge integration in end-to-end S2S systems. Our code and dataset are released.
title Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2505.00028