Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918041949831168 |
|---|---|
| author | Arora, Siddhant Tian, Jinchuan Futami, Hayato Jung, Jee-weon Shi, Jiatong Kashiwagi, Yosuke Tsunoo, Emiru Watanabe, Shinji |
| author_facet | Arora, Siddhant Tian, Jinchuan Futami, Hayato Jung, Jee-weon Shi, Jiatong Kashiwagi, Yosuke Tsunoo, Emiru Watanabe, Shinji |
| contents | Unlike traditional cascaded pipelines, end-to-end (E2E) spoken dialogue systems preserve full differentiability and capture non-phonemic information, making them well-suited for modeling spoken interactions. However, existing E2E approaches often require large-scale training data and generates responses lacking semantic coherence. We propose a simple yet effective strategy leveraging a chain-of-thought (CoT) formulation, ensuring that training on conversational data remains closely aligned with the multimodal language model (LM)'s pre-training on speech recognition~(ASR), text-to-speech synthesis (TTS), and text LM tasks. Our method achieves over 1.5 ROUGE-1 improvement over the baseline, successfully training spoken dialogue systems on publicly available human-human conversation datasets, while being compute-efficient enough to train on just 300 hours of public human-human conversation data, such as the Switchboard. We will publicly release our models and training code. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_00722 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Arora, Siddhant Tian, Jinchuan Futami, Hayato Jung, Jee-weon Shi, Jiatong Kashiwagi, Yosuke Tsunoo, Emiru Watanabe, Shinji Computation and Language Sound Audio and Speech Processing Unlike traditional cascaded pipelines, end-to-end (E2E) spoken dialogue systems preserve full differentiability and capture non-phonemic information, making them well-suited for modeling spoken interactions. However, existing E2E approaches often require large-scale training data and generates responses lacking semantic coherence. We propose a simple yet effective strategy leveraging a chain-of-thought (CoT) formulation, ensuring that training on conversational data remains closely aligned with the multimodal language model (LM)'s pre-training on speech recognition~(ASR), text-to-speech synthesis (TTS), and text LM tasks. Our method achieves over 1.5 ROUGE-1 improvement over the baseline, successfully training spoken dialogue systems on publicly available human-human conversation datasets, while being compute-efficient enough to train on just 300 hours of public human-human conversation data, such as the Switchboard. We will publicly release our models and training code. |
| title | Chain-of-Thought Training for Open E2E Spoken Dialogue Systems |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.00722 |