GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915047042711552 |
|---|---|
| author | Zeng, Aohan Du, Zhengxiao Liu, Mingdao Wang, Kedong Jiang, Shengmin Zhao, Lei Dong, Yuxiao Tang, Jie |
| author_facet | Zeng, Aohan Du, Zhengxiao Liu, Mingdao Wang, Kedong Jiang, Shengmin Zhao, Lei Dong, Yuxiao Tang, Jie |
| contents | We introduce GLM-4-Voice, an intelligent and human-like end-to-end spoken chatbot. It supports both Chinese and English, engages in real-time voice conversations, and varies vocal nuances such as emotion, intonation, speech rate, and dialect according to user instructions. GLM-4-Voice uses an ultra-low bitrate (175bps), single-codebook speech tokenizer with 12.5Hz frame rate derived from an automatic speech recognition (ASR) model by incorporating a vector-quantized bottleneck into the encoder. To efficiently transfer knowledge from text to speech modalities, we synthesize speech-text interleaved data from existing text pre-training corpora using a text-to-token model. We continue pre-training from the pre-trained text language model GLM-4-9B with a combination of unsupervised speech data, interleaved speech-text data, and supervised speech-text data, scaling up to 1 trillion tokens, achieving state-of-the-art performance in both speech language modeling and spoken question answering. We then fine-tune the pre-trained model with high-quality conversational speech data, achieving superior performance compared to existing baselines in both conversational ability and speech quality. The open models can be accessed through https://github.com/THUDM/GLM-4-Voice and https://huggingface.co/THUDM/glm-4-voice-9b. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_02612 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Zeng, Aohan Du, Zhengxiao Liu, Mingdao Wang, Kedong Jiang, Shengmin Zhao, Lei Dong, Yuxiao Tang, Jie Computation and Language Sound Audio and Speech Processing We introduce GLM-4-Voice, an intelligent and human-like end-to-end spoken chatbot. It supports both Chinese and English, engages in real-time voice conversations, and varies vocal nuances such as emotion, intonation, speech rate, and dialect according to user instructions. GLM-4-Voice uses an ultra-low bitrate (175bps), single-codebook speech tokenizer with 12.5Hz frame rate derived from an automatic speech recognition (ASR) model by incorporating a vector-quantized bottleneck into the encoder. To efficiently transfer knowledge from text to speech modalities, we synthesize speech-text interleaved data from existing text pre-training corpora using a text-to-token model. We continue pre-training from the pre-trained text language model GLM-4-9B with a combination of unsupervised speech data, interleaved speech-text data, and supervised speech-text data, scaling up to 1 trillion tokens, achieving state-of-the-art performance in both speech language modeling and spoken question answering. We then fine-tune the pre-trained model with high-quality conversational speech data, achieving superior performance compared to existing baselines in both conversational ability and speech quality. The open models can be accessed through https://github.com/THUDM/GLM-4-Voice and https://huggingface.co/THUDM/glm-4-voice-9b. |
| title | GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2412.02612 |