Chain-of-Thought Training for Open E2E Spoken Dialogue Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arora, Siddhant, Tian, Jinchuan, Futami, Hayato, Jung, Jee-weon, Shi, Jiatong, Kashiwagi, Yosuke, Tsunoo, Emiru, Watanabe, Shinji
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918041949831168
author Arora, Siddhant
Tian, Jinchuan
Futami, Hayato
Jung, Jee-weon
Shi, Jiatong
Kashiwagi, Yosuke
Tsunoo, Emiru
Watanabe, Shinji
author_facet Arora, Siddhant
Tian, Jinchuan
Futami, Hayato
Jung, Jee-weon
Shi, Jiatong
Kashiwagi, Yosuke
Tsunoo, Emiru
Watanabe, Shinji
contents Unlike traditional cascaded pipelines, end-to-end (E2E) spoken dialogue systems preserve full differentiability and capture non-phonemic information, making them well-suited for modeling spoken interactions. However, existing E2E approaches often require large-scale training data and generates responses lacking semantic coherence. We propose a simple yet effective strategy leveraging a chain-of-thought (CoT) formulation, ensuring that training on conversational data remains closely aligned with the multimodal language model (LM)'s pre-training on speech recognition~(ASR), text-to-speech synthesis (TTS), and text LM tasks. Our method achieves over 1.5 ROUGE-1 improvement over the baseline, successfully training spoken dialogue systems on publicly available human-human conversation datasets, while being compute-efficient enough to train on just 300 hours of public human-human conversation data, such as the Switchboard. We will publicly release our models and training code.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00722
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
Arora, Siddhant
Tian, Jinchuan
Futami, Hayato
Jung, Jee-weon
Shi, Jiatong
Kashiwagi, Yosuke
Tsunoo, Emiru
Watanabe, Shinji
Computation and Language
Sound
Audio and Speech Processing
Unlike traditional cascaded pipelines, end-to-end (E2E) spoken dialogue systems preserve full differentiability and capture non-phonemic information, making them well-suited for modeling spoken interactions. However, existing E2E approaches often require large-scale training data and generates responses lacking semantic coherence. We propose a simple yet effective strategy leveraging a chain-of-thought (CoT) formulation, ensuring that training on conversational data remains closely aligned with the multimodal language model (LM)'s pre-training on speech recognition~(ASR), text-to-speech synthesis (TTS), and text LM tasks. Our method achieves over 1.5 ROUGE-1 improvement over the baseline, successfully training spoken dialogue systems on publicly available human-human conversation datasets, while being compute-efficient enough to train on just 300 hours of public human-human conversation data, such as the Switchboard. We will publicly release our models and training code.
title Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.00722