VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Jiacheng, Gao, Heting, Xie, Liufei, Yang, Zhenchuan, Li, Lijiang, Chen, Yiting, Zhang, Bin, Chen, Meng, Fu, Chaoyu, Zhao, Weifeng, Zhou, Wenjiang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909023368904704
author Xu, Jiacheng
Gao, Heting
Xie, Liufei
Yang, Zhenchuan
Li, Lijiang
Chen, Yiting
Zhang, Bin
Chen, Meng
Fu, Chaoyu
Zhao, Weifeng
Zhou, Wenjiang
author_facet Xu, Jiacheng
Gao, Heting
Xie, Liufei
Yang, Zhenchuan
Li, Lijiang
Chen, Yiting
Zhang, Bin
Chen, Meng
Fu, Chaoyu
Zhao, Weifeng
Zhou, Wenjiang
contents Human speech conveys expressiveness beyond linguistic content, including personality, mood, or performance elements, such as a comforting tone or humming a song, which we formalize as role-playing and singing. We present VITA-QinYu, the first expressive end-to-end (E2E) spoken language model (SLM) that goes beyond natural conversation to support both role-playing and singing generation. VITA-QinYu adopts a hybrid speech-text paradigm that extends interleaved text-audio modeling with multi-codebook audio tokens, a design enabling richer paralinguistic representation while preserving a clear separation between modalities to avoid interference. We further develop a comprehensive data generation pipeline to synthesize a total of 15.8K hours of natural conversation, role-playing, and singing data for training. VITA-QinYu demonstrates superior expressiveness, outperforming peer SLMs by 7 percentage points on objective role-playing benchmarks, and surpassing peer models by 0.13 points on a 5-point MOS scale for singing. Simultaneously, it achieves state-of-the-art conversational accuracy and fluency, exceeding prior SLMs by 1.38 and 4.98 percentage points on the C3 and URO benchmarks, respectively. We open-source our code and models and provide an easy-to-use demo with full-stack support for streaming and full-duplex interaction.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06765
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing
Xu, Jiacheng
Gao, Heting
Xie, Liufei
Yang, Zhenchuan
Li, Lijiang
Chen, Yiting
Zhang, Bin
Chen, Meng
Fu, Chaoyu
Zhao, Weifeng
Zhou, Wenjiang
Computation and Language
Artificial Intelligence
Human speech conveys expressiveness beyond linguistic content, including personality, mood, or performance elements, such as a comforting tone or humming a song, which we formalize as role-playing and singing. We present VITA-QinYu, the first expressive end-to-end (E2E) spoken language model (SLM) that goes beyond natural conversation to support both role-playing and singing generation. VITA-QinYu adopts a hybrid speech-text paradigm that extends interleaved text-audio modeling with multi-codebook audio tokens, a design enabling richer paralinguistic representation while preserving a clear separation between modalities to avoid interference. We further develop a comprehensive data generation pipeline to synthesize a total of 15.8K hours of natural conversation, role-playing, and singing data for training. VITA-QinYu demonstrates superior expressiveness, outperforming peer SLMs by 7 percentage points on objective role-playing benchmarks, and surpassing peer models by 0.13 points on a 5-point MOS scale for singing. Simultaneously, it achieves state-of-the-art conversational accuracy and fluency, exceeding prior SLMs by 1.38 and 4.98 percentage points on the C3 and URO benchmarks, respectively. We open-source our code and models and provide an easy-to-use demo with full-stack support for streaming and full-duplex interaction.
title VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.06765