JoyTTS: LLM-based Spoken Chatbot With Voice Cloning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Fangru, Zhao, Jun, Wang, Guoxin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916824260542464
author Zhou, Fangru
Zhao, Jun
Wang, Guoxin
author_facet Zhou, Fangru
Zhao, Jun
Wang, Guoxin
contents JoyTTS is an end-to-end spoken chatbot that combines large language models (LLM) with text-to-speech (TTS) technology, featuring voice cloning capabilities. This project is built upon the open-source MiniCPM-o and CosyVoice2 models and trained on 2000 hours of conversational data. We have also provided the complete training code to facilitate further development and optimization by the community. On the testing machine seed-tts-zh, it achieves a SS (speaker similarity) score of 0.73 and a WER (Word Error Rate) of 5.09. The code and models, along with training and inference scripts, are available at https://github.com/jdh-algo/JoyTTS.git.
format Preprint
id arxiv_https___arxiv_org_abs_2507_02380
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JoyTTS: LLM-based Spoken Chatbot With Voice Cloning
Zhou, Fangru
Zhao, Jun
Wang, Guoxin
Sound
Computation and Language
Audio and Speech Processing
JoyTTS is an end-to-end spoken chatbot that combines large language models (LLM) with text-to-speech (TTS) technology, featuring voice cloning capabilities. This project is built upon the open-source MiniCPM-o and CosyVoice2 models and trained on 2000 hours of conversational data. We have also provided the complete training code to facilitate further development and optimization by the community. On the testing machine seed-tts-zh, it achieves a SS (speaker similarity) score of 0.73 and a WER (Word Error Rate) of 5.09. The code and models, along with training and inference scripts, are available at https://github.com/jdh-algo/JoyTTS.git.
title JoyTTS: LLM-based Spoken Chatbot With Voice Cloning
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2507.02380