VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Jiacheng, Gao, Heting, Xie, Liufei, Yang, Zhenchuan, Li, Lijiang, Chen, Yiting, Zhang, Bin, Chen, Meng, Fu, Chaoyu, Zhao, Weifeng, Zhou, Wenjiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
by: Dai, Ziqi, et al.
Published: (2025)
by: Dai, Ziqi, et al.
Published: (2025)
VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
by: Long, Zuwei, et al.
Published: (2025)
by: Long, Zuwei, et al.
Published: (2025)
Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling
by: Geng, Chen, et al.
Published: (2026)
by: Geng, Chen, et al.
Published: (2026)
RoleCraft-GLM: Advancing Personalized Role-Playing in Large Language Models
by: Tao, Meiling, et al.
Published: (2023)
by: Tao, Meiling, et al.
Published: (2023)
LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation
by: Wang, Qi, et al.
Published: (2026)
by: Wang, Qi, et al.
Published: (2026)
SingingSDS: A Singing-Capable Spoken Dialogue System for Conversational Roleplay Applications
by: Han, Jionghao, et al.
Published: (2025)
by: Han, Jionghao, et al.
Published: (2025)
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
by: Fu, Chaoyou, et al.
Published: (2025)
by: Fu, Chaoyou, et al.
Published: (2025)
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
by: Li, Lijiang, et al.
Published: (2026)
by: Li, Lijiang, et al.
Published: (2026)
CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness Condition
by: Wang, Jianzong, et al.
Published: (2024)
by: Wang, Jianzong, et al.
Published: (2024)
R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice Conversion
by: Zheng, Junjie, et al.
Published: (2025)
by: Zheng, Junjie, et al.
Published: (2025)
iSchools and Non-iSchools in the USA: An Examination of Their Master's Programs
by: Chu, Heting
Published: (2012)
by: Chu, Heting
Published: (2012)
Hyperlinks: How Well Do They Represent the Intellectual Content of Digital Collections?
by: Chu, Heting
Published: (1997)
by: Chu, Heting
Published: (1997)
VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
by: Dong, Shaoqi, et al.
Published: (2025)
by: Dong, Shaoqi, et al.
Published: (2025)
StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks
by: Huang, Jingyue, et al.
Published: (2025)
by: Huang, Jingyue, et al.
Published: (2025)
YuLan: An Open-source Large Language Model
by: Zhu, Yutao, et al.
Published: (2024)
by: Zhu, Yutao, et al.
Published: (2024)
Yizhan Qin
by: YizhanandQin
Published: (2026)
by: YizhanandQin
Published: (2026)
VITA: Vision-to-Action Flow Matching Policy
by: Gao, Dechen, et al.
Published: (2025)
by: Gao, Dechen, et al.
Published: (2025)
The Literacy-Enhancing Potential of Singing versus Spoken Language in Public Library Storytimes: A Text Analytics Approach
by: Soohyung Joo, et al.
Published: (2024)
by: Soohyung Joo, et al.
Published: (2024)
From Persona to Personalization: A Survey on Role-Playing Language Agents
by: Chen, Jiangjie, et al.
Published: (2024)
by: Chen, Jiangjie, et al.
Published: (2024)
Gravitational Wave Astronomy With TianQin
by: Li, En-Kun, et al.
Published: (2024)
by: Li, En-Kun, et al.
Published: (2024)
The Academic Library Meets Web 2.0: Applications and Implications
by: Xu, Chen, et al.
Published: (2009)
by: Xu, Chen, et al.
Published: (2009)
VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting
by: Liu, Xiaoyu, et al.
Published: (2025)
by: Liu, Xiaoyu, et al.
Published: (2025)
VITA: Towards Open-Source Interactive Omni Multimodal LLM
by: Fu, Chaoyou, et al.
Published: (2024)
by: Fu, Chaoyou, et al.
Published: (2024)
On payload architecture and pointing control strategies for TianQin
by: Fang, Yuzhou, et al.
Published: (2024)
by: Fang, Yuzhou, et al.
Published: (2024)
Progress of the TianQin project
by: Luo, Jun, et al.
Published: (2025)
by: Luo, Jun, et al.
Published: (2025)
WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training
by: Chen, Yifu, et al.
Published: (2026)
by: Chen, Yifu, et al.
Published: (2026)
The Imperial Qín Dynasty
by: Schneider, Marcel
Published: (2025)
by: Schneider, Marcel
Published: (2025)
EmoNews: A Spoken Dialogue System for Expressive News Conversations
by: Matsuura, Ryuki, et al.
Published: (2025)
by: Matsuura, Ryuki, et al.
Published: (2025)
Capturing Minds, Not Just Words: Enhancing Role-Playing Language Models with Personality-Indicative Data
by: Ran, Yiting, et al.
Published: (2024)
by: Ran, Yiting, et al.
Published: (2024)
Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing
by: Wang, Miao, et al.
Published: (2026)
by: Wang, Miao, et al.
Published: (2026)
A Hybrid Input based Deep Reinforcement Learning for Lane Change Decision-Making of Autonomous Vehicle
by: Gao, Ziteng, et al.
Published: (2025)
by: Gao, Ziteng, et al.
Published: (2025)
YuLan-OneSim: Towards the Next Generation of Social Simulator with Large Language Models
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
by: Xue, Hongfei, et al.
Published: (2023)
by: Xue, Hongfei, et al.
Published: (2023)
Scene Graph Generation with Role-Playing Large Language Models
by: Chen, Guikun, et al.
Published: (2024)
by: Chen, Guikun, et al.
Published: (2024)
VITA - Vocational Innovation through Teaching with AI
by: Ravotto, Pierfranco
Published: (2025)
by: Ravotto, Pierfranco
Published: (2025)
Innovations in Scholarly Electronic Journals: The Challenge from Nontraditional STM Publishers (SIG PUB, SIG STI).
by: Panagopoulos, Beata, et al.
Published: (2000)
by: Panagopoulos, Beata, et al.
Published: (2000)
YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models
by: Lin, Junyu, et al.
Published: (2026)
by: Lin, Junyu, et al.
Published: (2026)
Glycolysis Plays a Critical and Dual Role in Periodontitis
by: Hongyu Ming, et al.
Published: (2025)
by: Hongyu Ming, et al.
Published: (2025)
Fundamental Physics and Cosmology with TianQin
by: Luo, Jun, et al.
Published: (2025)
by: Luo, Jun, et al.
Published: (2025)
Spoken Language Intelligence of Large Language Models for Language Learning
by: Peng, Linkai, et al.
Published: (2023)
by: Peng, Linkai, et al.
Published: (2023)
Similar Items
-
Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
by: Dai, Ziqi, et al.
Published: (2025) -
VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
by: Long, Zuwei, et al.
Published: (2025) -
Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling
by: Geng, Chen, et al.
Published: (2026) -
RoleCraft-GLM: Advancing Personalized Role-Playing in Large Language Models
by: Tao, Meiling, et al.
Published: (2023) -
LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation
by: Wang, Qi, et al.
Published: (2026)