SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Wenxi, Ma, Ziyang, Yan, Ruiqi, Liang, Yuzhe, Li, Xiquan, Xu, Ruiyang, Niu, Zhikang, Zhu, Yanqiao, Yang, Yifan, Liu, Zhanxun, Yu, Kai, Hu, Yuxuan, Li, Jinyu, Lu, Yan, Liu, Shujie, Chen, Xie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910756066295808
author Chen, Wenxi
Ma, Ziyang
Yan, Ruiqi
Liang, Yuzhe
Li, Xiquan
Xu, Ruiyang
Niu, Zhikang
Zhu, Yanqiao
Yang, Yifan
Liu, Zhanxun
Yu, Kai
Hu, Yuxuan
Li, Jinyu
Lu, Yan
Liu, Shujie
Chen, Xie
author_facet Chen, Wenxi
Ma, Ziyang
Yan, Ruiqi
Liang, Yuzhe
Li, Xiquan
Xu, Ruiyang
Niu, Zhikang
Zhu, Yanqiao
Yang, Yifan
Liu, Zhanxun
Yu, Kai
Hu, Yuxuan
Li, Jinyu
Lu, Yan
Liu, Shujie
Chen, Xie
contents Recent advancements highlight the potential of end-to-end real-time spoken dialogue systems, showcasing their low latency and high quality. In this paper, we introduce SLAM-Omni, a timbre-controllable, end-to-end voice interaction system with single-stage training. SLAM-Omni achieves zero-shot timbre control by modeling spoken language with semantic tokens and decoupling speaker information to a vocoder. By predicting grouped speech semantic tokens at each step, our method significantly reduces the sequence length of audio tokens, accelerating both training and inference. Additionally, we propose historical text prompting to compress dialogue history, facilitating efficient multi-round interactions. Comprehensive evaluations reveal that SLAM-Omni outperforms prior models of similar scale, requiring only 15 hours of training on 4 GPUs with limited data. Notably, it is the first spoken dialogue system to achieve competitive performance with a single-stage training approach, eliminating the need for pre-training on TTS or ASR tasks. Further experiments validate its multilingual and multi-turn dialogue capabilities on larger datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15649
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
Chen, Wenxi
Ma, Ziyang
Yan, Ruiqi
Liang, Yuzhe
Li, Xiquan
Xu, Ruiyang
Niu, Zhikang
Zhu, Yanqiao
Yang, Yifan
Liu, Zhanxun
Yu, Kai
Hu, Yuxuan
Li, Jinyu
Lu, Yan
Liu, Shujie
Chen, Xie
Audio and Speech Processing
Recent advancements highlight the potential of end-to-end real-time spoken dialogue systems, showcasing their low latency and high quality. In this paper, we introduce SLAM-Omni, a timbre-controllable, end-to-end voice interaction system with single-stage training. SLAM-Omni achieves zero-shot timbre control by modeling spoken language with semantic tokens and decoupling speaker information to a vocoder. By predicting grouped speech semantic tokens at each step, our method significantly reduces the sequence length of audio tokens, accelerating both training and inference. Additionally, we propose historical text prompting to compress dialogue history, facilitating efficient multi-round interactions. Comprehensive evaluations reveal that SLAM-Omni outperforms prior models of similar scale, requiring only 15 hours of training on 4 GPUs with limited data. Notably, it is the first spoken dialogue system to achieve competitive performance with a single-stage training approach, eliminating the need for pre-training on TTS or ASR tasks. Further experiments validate its multilingual and multi-turn dialogue capabilities on larger datasets.
title SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
topic Audio and Speech Processing
url https://arxiv.org/abs/2412.15649