OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Haonan, Luo, Run, Liu, Xiong, Wu, Yuchuan, Lin, Ting-En, Zeng, Pengpeng, Qu, Qiang, Fang, Feiteng, Yang, Min, Gao, Lianli, Song, Jingkuan, Huang, Fei, Li, Yongbin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915326759796736
author Zhang, Haonan
Luo, Run
Liu, Xiong
Wu, Yuchuan
Lin, Ting-En
Zeng, Pengpeng
Qu, Qiang
Fang, Feiteng
Yang, Min
Gao, Lianli
Song, Jingkuan
Huang, Fei
Li, Yongbin
author_facet Zhang, Haonan
Luo, Run
Liu, Xiong
Wu, Yuchuan
Lin, Ting-En
Zeng, Pengpeng
Qu, Qiang
Fang, Feiteng
Yang, Min
Gao, Lianli
Song, Jingkuan
Huang, Fei
Li, Yongbin
contents Role-Playing Agents (RPAs), benefiting from large language models, is an emerging interactive AI system that simulates roles or characters with diverse personalities. However, existing methods primarily focus on mimicking dialogues among roles in textual form, neglecting the role's voice traits (e.g., voice style and emotions) as playing a crucial effect in interaction, which tends to be more immersive experiences in realistic scenarios. Towards this goal, we propose OmniCharacter, a first seamless speech-language personality interaction model to achieve immersive RPAs with low latency. Specifically, OmniCharacter enables agents to consistently exhibit role-specific personality traits and vocal traits throughout the interaction, enabling a mixture of speech and language responses. To align the model with speech-language scenarios, we construct a dataset named OmniCharacter-10K, which involves more distinctive characters (20), richly contextualized multi-round dialogue (10K), and dynamic speech response (135K). Experimental results showcase that our method yields better responses in terms of both content and style compared to existing RPAs and mainstream speech-language models, with a response latency as low as 289ms. Code and dataset are available at https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/OmniCharacter.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20277
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
Zhang, Haonan
Luo, Run
Liu, Xiong
Wu, Yuchuan
Lin, Ting-En
Zeng, Pengpeng
Qu, Qiang
Fang, Feiteng
Yang, Min
Gao, Lianli
Song, Jingkuan
Huang, Fei
Li, Yongbin
Computation and Language
Computer Vision and Pattern Recognition
Role-Playing Agents (RPAs), benefiting from large language models, is an emerging interactive AI system that simulates roles or characters with diverse personalities. However, existing methods primarily focus on mimicking dialogues among roles in textual form, neglecting the role's voice traits (e.g., voice style and emotions) as playing a crucial effect in interaction, which tends to be more immersive experiences in realistic scenarios. Towards this goal, we propose OmniCharacter, a first seamless speech-language personality interaction model to achieve immersive RPAs with low latency. Specifically, OmniCharacter enables agents to consistently exhibit role-specific personality traits and vocal traits throughout the interaction, enabling a mixture of speech and language responses. To align the model with speech-language scenarios, we construct a dataset named OmniCharacter-10K, which involves more distinctive characters (20), richly contextualized multi-round dialogue (10K), and dynamic speech response (135K). Experimental results showcase that our method yields better responses in terms of both content and style compared to existing RPAs and mainstream speech-language models, with a response latency as low as 289ms. Code and dataset are available at https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/OmniCharacter.
title OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.20277