OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Qinglin, Cheng, Luyao, Deng, Chong, Chen, Qian, Wang, Wen, Zheng, Siqi, Liu, Jiaqing, Yu, Hai, Tan, Chaohong, Du, Zhihao, Zhang, Shiliang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909447388921856
author Zhang, Qinglin
Cheng, Luyao
Deng, Chong
Chen, Qian
Wang, Wen
Zheng, Siqi
Liu, Jiaqing
Yu, Hai
Tan, Chaohong
Du, Zhihao
Zhang, Shiliang
author_facet Zhang, Qinglin
Cheng, Luyao
Deng, Chong
Chen, Qian
Wang, Wen
Zheng, Siqi
Liu, Jiaqing
Yu, Hai
Tan, Chaohong
Du, Zhihao
Zhang, Shiliang
contents Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and natural interactions in full-duplex dialogue systems remains a significant challenge, especially considering human conversation dynamics such as interruptions, backchannels, and overlapping speech. In this paper, we introduce a novel End-to-End GPT-based model OmniFlatten for full-duplex conversation, capable of effectively modeling the complex behaviors inherent to natural conversations with low latency. To achieve full-duplex conversation capabilities, we propose a multi-stage post-training scheme that progressively adapts a text large language model (LLM) backbone into a speech-text dialogue LLM, capable of generating text and speech in real time, without modifying the architecture of the backbone LLM. The training process comprises three stages: modality alignment, half-duplex dialogue learning, and full-duplex dialogue learning. In all training stages, we standardize the data using a flattening operation, which enables unifying the training methods and the GPT backbone across different modalities and tasks. Our approach offers a simple modeling technique and a promising research direction for developing efficient and natural end-to-end full-duplex spoken dialogue systems. Audio samples of dialogues generated by OmniFlatten can be found at this web site (https://omniflatten.github.io/).
format Preprint
id arxiv_https___arxiv_org_abs_2410_17799
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
Zhang, Qinglin
Cheng, Luyao
Deng, Chong
Chen, Qian
Wang, Wen
Zheng, Siqi
Liu, Jiaqing
Yu, Hai
Tan, Chaohong
Du, Zhihao
Zhang, Shiliang
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and natural interactions in full-duplex dialogue systems remains a significant challenge, especially considering human conversation dynamics such as interruptions, backchannels, and overlapping speech. In this paper, we introduce a novel End-to-End GPT-based model OmniFlatten for full-duplex conversation, capable of effectively modeling the complex behaviors inherent to natural conversations with low latency. To achieve full-duplex conversation capabilities, we propose a multi-stage post-training scheme that progressively adapts a text large language model (LLM) backbone into a speech-text dialogue LLM, capable of generating text and speech in real time, without modifying the architecture of the backbone LLM. The training process comprises three stages: modality alignment, half-duplex dialogue learning, and full-duplex dialogue learning. In all training stages, we standardize the data using a flattening operation, which enables unifying the training methods and the GPT backbone across different modalities and tasks. Our approach offers a simple modeling technique and a promising research direction for developing efficient and natural end-to-end full-duplex spoken dialogue systems. Audio samples of dialogues generated by OmniFlatten can be found at this web site (https://omniflatten.github.io/).
title OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.17799