GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Aohan, Du, Zhengxiao, Liu, Mingdao, Wang, Kedong, Jiang, Shengmin, Zhao, Lei, Dong, Yuxiao, Tang, Jie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915047042711552
author Zeng, Aohan
Du, Zhengxiao
Liu, Mingdao
Wang, Kedong
Jiang, Shengmin
Zhao, Lei
Dong, Yuxiao
Tang, Jie
author_facet Zeng, Aohan
Du, Zhengxiao
Liu, Mingdao
Wang, Kedong
Jiang, Shengmin
Zhao, Lei
Dong, Yuxiao
Tang, Jie
contents We introduce GLM-4-Voice, an intelligent and human-like end-to-end spoken chatbot. It supports both Chinese and English, engages in real-time voice conversations, and varies vocal nuances such as emotion, intonation, speech rate, and dialect according to user instructions. GLM-4-Voice uses an ultra-low bitrate (175bps), single-codebook speech tokenizer with 12.5Hz frame rate derived from an automatic speech recognition (ASR) model by incorporating a vector-quantized bottleneck into the encoder. To efficiently transfer knowledge from text to speech modalities, we synthesize speech-text interleaved data from existing text pre-training corpora using a text-to-token model. We continue pre-training from the pre-trained text language model GLM-4-9B with a combination of unsupervised speech data, interleaved speech-text data, and supervised speech-text data, scaling up to 1 trillion tokens, achieving state-of-the-art performance in both speech language modeling and spoken question answering. We then fine-tune the pre-trained model with high-quality conversational speech data, achieving superior performance compared to existing baselines in both conversational ability and speech quality. The open models can be accessed through https://github.com/THUDM/GLM-4-Voice and https://huggingface.co/THUDM/glm-4-voice-9b.
format Preprint
id arxiv_https___arxiv_org_abs_2412_02612
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Zeng, Aohan
Du, Zhengxiao
Liu, Mingdao
Wang, Kedong
Jiang, Shengmin
Zhao, Lei
Dong, Yuxiao
Tang, Jie
Computation and Language
Sound
Audio and Speech Processing
We introduce GLM-4-Voice, an intelligent and human-like end-to-end spoken chatbot. It supports both Chinese and English, engages in real-time voice conversations, and varies vocal nuances such as emotion, intonation, speech rate, and dialect according to user instructions. GLM-4-Voice uses an ultra-low bitrate (175bps), single-codebook speech tokenizer with 12.5Hz frame rate derived from an automatic speech recognition (ASR) model by incorporating a vector-quantized bottleneck into the encoder. To efficiently transfer knowledge from text to speech modalities, we synthesize speech-text interleaved data from existing text pre-training corpora using a text-to-token model. We continue pre-training from the pre-trained text language model GLM-4-9B with a combination of unsupervised speech data, interleaved speech-text data, and supervised speech-text data, scaling up to 1 trillion tokens, achieving state-of-the-art performance in both speech language modeling and spoken question answering. We then fine-tune the pre-trained model with high-quality conversational speech data, achieving superior performance compared to existing baselines in both conversational ability and speech quality. The open models can be accessed through https://github.com/THUDM/GLM-4-Voice and https://huggingface.co/THUDM/glm-4-voice-9b.
title GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.02612